University of Southampton — September 2026

Conformal defects in neural network field theories

symmetry breaking  ·  solvable nn-dCFTs

Benjamin Suzzoni
Department of Mathematical Sciences, UNIST
arXiv:2512.07946

Outline

1
Why NN? Why defects CFTs?
physics for ML for physics · conformal symmetry ubiquitous
2
CFT and dCFT data
conformal primaries · OPE · crossing · ambient and defect channels
3
Neural-network CFTs
UATs · parameter averages · symmetry in architecture and distribution
4
Neural-network defect CFTs
general construction · monomial and reciprocal examples
5
What we learn
results · limits · directions opened by the defect viewpoint

Chapter 1

Why NN? Why defect CFTs?

a lot can be found by combining the two

Strong interplay between ML and physics


ML → physics
  • Find hidden symmetries in data
  • Compute numerical approximations
  • Formulate exact conjectures
  • Assist theorem proving
physics → ML

Neural Network Field Theories

Then symmetry, correlation functions and defects become tools for organizing families of random networks.

A defect is a controlled way to break symmetry—and probe what survives.

Conformal symmetry and defect conformal symmetry


Why CFTs?

Symmetry controls dynamics

  • The conformal group strongly constrains correlators.
  • CFTs sit at RG fixed points; exactly marginal deformations can connect families of them.
  • They arise naturally in critical systems and D-brane constructions.
Why dCFTs?

Extended operators complete the picture

  • Boundaries, interfaces and defects carry their own degrees of freedom.
  • Higher-form and generalized symmetries act on extended operators.
  • Ambient spectrum + defect spectrum + couplings characterize the deformed theory.

The same organization can be transferred to neural-network ensembles.

Chapter 2

CFT and dCFT data

we can bootstrap from building blocks

A CFT is encoded by dimensions and OPE coefficients

correlators reveal the independent CFT data
Φᵢ Φⱼ two operator insertions
conformal primary
\(\mathcal O_i:\)( Δᵢ, Jᵢ)

\(\mathcal O_i(\rho x)=\rho^{-\Delta_i}\mathcal O_i(x)\)

symmetry fixes the scalar two-point function

\[ \langle\Phi_i(x_1)\Phi_j(x_2)\rangle =\frac{\delta_{ij}}{|x_{12}|^{2\Delta_i}} \]

Φᵢ Φⱼ Φₖ λᵢⱼₖ
the scalar three-point function defines the OPE coefficient
\(\langle\Phi_i(x_1)\Phi_j(x_2)\Phi_k(x_3)\rangle=\) λᵢⱼₖ \(|x_{12}|^{\Delta_i+\Delta_j-\Delta_k}|x_{23}|^{\Delta_j+\Delta_k-\Delta_i}\)
\(|x_{13}|^{\Delta_i+\Delta_k-\Delta_j}\)

Scalar correlators shown; spinning tensor structures are suppressed.

CFT data={ Δᵢ, Jᵢ, λᵢⱼₖ}
Δi Ji λijk

OPE data obeys the crossing equations

\[ \phi_1(x)\phi_2(y)\sim\sum_{\mathcal O} f_{\phi_1\phi_2\mathcal O}\, P(x-y,\partial_y)\mathcal O(y) \]

s-channel 𝒪
=
t-channel 𝒪′

The same four-point function must be reconstructed in either channel.

A defect breaks the symmetry subgroup

symmetry acts differently once the defect is present
\(SO(1,D+1)\)
\(SO(1,p+1)\)
×
\(SO(q)\)


\(q=D-p\)

schematic 2D slice

An ambient operator expands into defect primaries

bring \(\mathcal O_i\) to the defect defect orthogonal approach 𝒪ᵢ ambient Ôₐ Σₐ defect primaries
bulk-to-defect OPE (BOE) · transverse-scalar sector \(s_a=0\)

\[ \begin{aligned} \mathcal O_i(x_\parallel,x_\perp) &=\sum_a b_{ia}|x_\perp|^{\widehat\Delta_a-\Delta_i}\\[-1pt] &\quad\times\mathcal C_{ia}(x_\perp^2\partial_\parallel^2) \widehat{\mathcal O}_a(x_\parallel) \end{aligned} \]

ambient
\(\Delta_i,J_i,\lambda_{ijk}\)
defect
\(\widehat\Delta_a,s_a,\widehat\lambda_{abc}\)
couplings
\(a_i,b_{ia}\)
full dCFT data

Crossing: two OPE channels reconstruct one correlator

Two ambient insertions can reach the defect identity in two orders.
ambient OPE → identity term of the BOE
defect 𝒪ᵢ 𝒪ⱼ 𝒪ₖ λᵢⱼᵏ take 𝒪ₖ to the defect 𝟙̂ λᵢⱼᵏ aₖ
two bulk-to-defect OPEs → defect OPE
defect 𝒪ᵢ 𝒪ⱼ Ôₐ Ôᵦ bᵢₐ bⱼᵦ defect OPE 𝟙̂ δₐᵦ

\[ \boxed{\displaystyle \sum_k \lambda_{ij}{}^{k}a_k\,\mathcal G_k^{\mathrm{amb}}(\chi,\psi) =\sum_a b_{ia}b_{ja}\,\widehat{\mathcal G}_a^{\mathrm{def}}(\chi,\psi)} \]

\(a_k=b_{k\widehat{\mathbf 1}},\quad r_\ell=\lVert X_\ell\rVert_\circ,\quad \chi=-\dfrac{2X_1\bullet X_2}{r_1r_2},\quad \cos\psi=\dfrac{X_1\circ X_2}{r_1r_2}\)

Chapter 3

Neural-network CFTs

from function approximation to parameter-space field theory

A neural network is a composable map

Neural network:

\(\Phi_{\Theta}:\mathbb R^D\rightarrow\mathbb R\)


Operations: composition, addition and scalar multiplication.


Its functional form is the architecture; its weights and biases are the parameters.

one hidden layer x Φ(x) basis functions φᵢ weights Θᵢ

\[ \qquad\qquad\qquad\qquad\qquad\qquad\qquad\qquad\qquad\qquad\Phi_{\Theta}(x)=\sum_i \Theta_i \varphi_i(x) \]

The Universal Approximation Theorem (UAT)

one width parameter \(N\) controls both pictures
\(N=1\)   width parameter; displayed nodes are schematic
coarse approximation   schematic error
UAT: expressivity/density in a function space.
NNGP: a suitably normalized iid large-width limit can become Gaussian.

A functional integral through the UAT


Network at initialization: weights \(\Theta\) follow a distribution \(P\).


function space

\[ \int\mathcal D\phi\,\mu[\phi]\,\mathcal F[\phi] \]

integrate over field configurations

pushforward
through \(\Phi_\Theta\)
parameter space

\[ \lim_{N\to\infty}\int_{\Theta_N} d\Theta\,P_N(\Theta)\, \mathcal F[\Phi_\Theta] \]

ordinary integrals over network parameters


Key Idea: replace the functional integral with a statistical average.

But: conformal correlators can already emerge at finite architecture—the UAT limit is not required to define an nn-CFT.

What is a NN-CFT?[Halverson, et al; 2024]


Field-theory language Neural-network language
\(\phi(\lambda x)=\lambda^{-\Delta}\phi(x)\) \(\Phi_\Theta(\lambda X)=\lambda^{-\Delta}\Phi_\Theta(X)\)
\(\displaystyle Z[J]=\int\mathcal D\phi\,e^{-S[\phi]}e^{\int J\phi}\) \(\displaystyle Z[J]=\int d\Theta\,P(\Theta)e^{\int J\Phi_\Theta}\)
\(\langle\phi(x)\phi(y)\rangle\) \(\mathbb E_{P(\Theta)}[\Phi_\Theta(X)\Phi_\Theta(Y)]\)

Takeaway: conformal structure is a property of the architecture–distribution pair, not of the \(N\to\infty\) limit alone.

How to encode symmetry?


field theory

\[ Z[J]=\int\mathcal D\phi\,e^{-S[\phi]} e^{\int d^D x\,J(x)\phi(x)} \]

parameter ensemble

\[ Z[J]=\int d\Theta\,P(\Theta) e^{\int d^D x\,J(x)\Phi_\Theta(x)} \]


parameter ensemble
\(P(\Theta)\)
correlators

\[ \mathbb E_{P(\Theta)} [\Phi_{\Theta,1}\cdots\Phi_{\Theta,n}] \]

inherit the symmetries respected by both \(P(\Theta)\) and \(\Phi_\Theta\).


Key Idea: Symmetries of \(S[\phi]\) are encoded in the probability measure \(P(\Theta)\).

Embedding space gives a linear action

\(\mathbb R^{D+2}\) \(\mathbb R^{1,D+1}\)
  Wick rotation
\(\mathrm{NC}=\{X\in\mathbb R^{1,D+1}\mid X^2=0\}\) projective identification \(X\sim\lambda X\)
distribution

\(SO(D+2)\)-invariant \(P(\theta)\)

homogeneous architecture

\(\Phi_\Delta(X)=(X\cdot\theta)^{-\Delta}\)

Toy model: the monomial nn-CFT \(\Delta_n=-n\)

Gaussian parameter ensemble
Euclidean average → Wick rotate

\[ P_E(\Theta)=\frac{e^{-\Theta^2/(2\sigma^2)}}{(2\pi\sigma^2)^{(D+2)/2}} \]

normalized homogeneous primaries

\[ \mathcal O_n(X)=\frac{(\Theta\!\cdot\!X)^n}{\sigma^n\sqrt{n!}}, \qquad \Delta_n=-n,\quad n>0 \]

one-point function

\[ \langle\mathcal{O}_n(X)\rangle= \begin{cases}0,&n\ \mathrm{odd},\\\frac{(n-1)!!}{\sqrt{n!}}(X^2)^{n/2},&n\ \mathrm{even},\end{cases} \]

On the Poincaré section: \(\langle\mathcal O_n(X)\rangle=0\)
orthogonality on the null cone

\[ \boxed{\; \langle\mathcal{O}_n(X_1)\mathcal{O}_m(X_2)\rangle =\delta_{nm}(X_1\!\cdot\!X_2)^n\;}, \qquad X_1^2=X_2^2=0 \]

On the Poincaré section: \(\langle\mathcal{O}_n(X_1)\mathcal{O}_m(X_2)\rangle = \delta_{nm}|x_{12}|^{2n}\)
mixed four-point correlator · \((-n,-m,-n,-m)\)

\[ G_{nmnm}(x_i) =|x_{12}|^{n+m}|x_{34}|^{n+m} \left(\frac{|x_{24}|}{|x_{13}|}\right)^{m-n} g_{nm}(u,v) \]

conformal-block decomposition

\[ g_{nm}(u,v)= \sum_{\mathcal X\in\mathcal O_n\times\mathcal O_m} \lambda_{nm\mathcal X}^{\,2}\, g_{\Delta_{\mathcal X},J_{\mathcal X}}^{(m-n,m-n)}(u,v) \]

\(u=\dfrac{x_{12}^2x_{34}^2}{x_{13}^2x_{24}^2}\),  \(v=\dfrac{x_{14}^2x_{23}^2}{x_{13}^2x_{24}^2}\); \(\Delta_{12}=\Delta_{34}=m-n\).

paper example in \(D=4\) · \(n=1,\ m=2\)
\(\mathcal O_1=\Phi/\sigma,\quad \mathcal O_2=\Phi^2/(\sqrt2\,\sigma^2),\quad \mu_6=15\sigma^6\)  ⇒  divide by \(2\sigma^6\)
normalized reduced correlator

\[ \boxed{\; g_{12}(u,v)= 2u^{-1/2}+2v\,u^{-3/2}+u^{-3/2}\;} \]

normalized three-point data → the same \(g_{12}\)

\[ \boxed{\; g_{12}(u,v)= 3\,g_{-3,0}^{(1,1)}(u,v) +\frac52\,g_{-1,0}^{(1,1)}(u,v)\;} \]

\(\widehat{\mathcal O}_{0,3}=\dfrac{(\Theta\cdot X)^3}{\sigma^3\sqrt6}\), \(\lambda_{12,\widehat{\mathcal O}_{0,3}}=\sqrt3\); \(\widehat{\mathcal O}_{1,1}=\dfrac{(\Theta\cdot\Theta)(\Theta\cdot X)}{\sigma^3\sqrt{80}}\), \(\lambda_{12,\widehat{\mathcal O}_{1,1}}=\sqrt{\frac52}\).
Crossing: \(g_{12}(u,v)=\left(\frac vu\right)^{3/2}g_{12}(v,u)\).

Chapter 4

Neural-network defect CFTs

break the symmetry, choose an invariant distribution, pick an architecture

Ingredients for nn-dCFTs


1 · preserved symmetry

\[ SO(1,D+1)\rightarrow SO(1,p+1)\times SO(q) \]

with \(q=D-p\)

2 · invariant distribution

\[ P(g\theta)=P(\theta),\qquad g\in G_{\rm defect} \]

3 · covariant architectures

ambient   \(\Phi_\Delta(X)\)
defect    \(\widehat\varphi_{\widehat\Delta,s}(X)\)

\[ Z[\widehat J,J]= \int d\theta\,P(\theta)\, e^{\int d^p x_\parallel\,\widehat J(x_\parallel)\widehat\varphi(x_\parallel)} e^{\int d^D x\,J(x)\Phi(x)} \]

Example 1: a monomial nn-dCFT


1

Factorized distribution

\(\widehat P(\hat\theta)\widetilde P(\tilde\theta)\): centered Gaussians.

second moments \(\widehat\mu_2\) and \(\widetilde\mu_2\)

2

Defect network

\[ \widehat\varphi_{\widehat n}(X) =(X\bullet\hat\theta)^{\widehat n} \]

3

Ambient network

\[ \Phi_n(X)= (X\bullet\hat\theta+X\circ\tilde\theta)^n \]

4

Scaling weights

\(\widehat\Delta=-\widehat n\), \(\Delta=-n\)

negative dimensions; \(n,\widehat n\in\mathbb N\)


Convention: hats are tangent/\(\bullet\); tildes are normal/\(\circ\).

Example 1: Finite defect-network expansion


exact OPE-like decomposition

\[ \Phi_n(X)= \sum_{d=0}^{n}\binom nd \underbrace{(X\circ\tilde\theta)^{n-d}}_{\text{normal degree }n-d} \underbrace{(X\bullet\hat\theta)^d}_{\text{tangent degree }d} \]

Each term is a smaller network with a definite scaling dimensions.

For \(n=4\): exactly five network factors ⟾ no infinite defect tower.

example \(n=4\)

Example 1: One-point functions


selection by the defect
defect insertion ambient insertion
defect

\[ \mathbb E[\widehat\varphi_{\widehat n}(X)] \overset{\mathrm{P.S.}}{=}0,\qquad \widehat n>0 \]

The identity \(\widehat n=0\) has expectation value 1.

ambient

\[ \mathbb E[\Phi_n(X)]\overset{\mathrm{P.S.}}{=} \begin{cases} \displaystyle \frac{\Gamma(n+1)} {2^{n/2}\Gamma(\frac n2+1)} \big[(\widetilde\mu_2-\widehat\mu_2)X\circ X\big]^{n/2}, & n\in2\mathbb Z,\\[4pt] 0,&\text{otherwise.} \end{cases} \]


The variance difference measures the defect deformation; it vanishes when the two sectors match.

Example 1: Two-point functions


defect–defect
defect–ambient
ambient–defect
ambient–ambient
defect–defect

\[ \mathbb E[\widehat\varphi_{\widehat n_1}(X_1) \widehat\varphi_{\widehat n_2}(X_2)] =\delta_{\widehat n_1,\widehat n_2} \Gamma(\widehat n_1+1) (\widehat\mu_2X_1\bullet X_2)^{\widehat n_1} \]

defect–ambient

\[ \begin{aligned} &\mathbb E[\widehat\varphi_{\widehat n_1}(X_1)\Phi_{n_2}(X_2)] =\frac{\Gamma(n_2+1)} {2^{(n_2-\widehat n_1)/2} \Gamma(\frac{n_2-\widehat n_1}{2}+1)} (\widehat\mu_2X_1\bullet X_2)^{\widehat n_1}\\ &\hspace{2.3cm}\times [(\widetilde\mu_2-\widehat\mu_2)X_2\circ X_2]^{ (n_2-\widehat n_1)/2} \end{aligned} \]

nonzero for \(n_2\geq\widehat n_1\) and \(n_2-\widehat n_1\in2\mathbb Z\)

ambient–ambient

\[ \begin{aligned} \mathbb E[\Phi_{n_1}(X_1)\Phi_{n_2}(X_2)] ={}&\frac{\Gamma(n_2+1)} {\Gamma(\frac{n_2-n_1}{2}+1)} \left(\frac{\alpha_{22}}{2}\right)^{\frac{n_2-n_1}{2}} \alpha_{12}^{n_1}\\ &\times{}_2F_1\!\left( \frac{1-n_1}{2},-\frac{n_1}{2}; \frac{n_2-n_1}{2}+1; \frac{\alpha_{11}\alpha_{22}}{\alpha_{12}^2} \right) \end{aligned} \]

for \(n_1\leq n_2\) with equal parity; exchange \(1\leftrightarrow2\) otherwise

Example 1: exact scalar one-/two-point data

unit-normalized basis

\[ \rho=\frac{\widetilde\mu_2-\widehat\mu_2}{\widehat\mu_2},\qquad \mathcal O_n=\frac{\Phi_n}{\sqrt{n!\,\widehat\mu_2^{\,n}}},\qquad \widehat{\mathcal O}_m=\frac{\widehat\varphi_m}{\sqrt{m!\,\widehat\mu_2^{\,m}}}. \]


scalar spectrum

\[ (\Delta_n,J_n)=(-n,0),\qquad (\widehat\Delta_m,s_m)=(-m,0) \]


For a fixed \(\mathcal O_n\), only

\[ m=n,n-2,n-4,\ldots\geq0, \qquad r_{nm}\equiv\frac{n-m}{2}\in\mathbb N \]

one-point and scalar BOE coefficients

\[ \boxed{\displaystyle a_n=b_{n0}=\delta_{n\,\mathrm{even}} \frac{\sqrt{n!}}{2^{n/2}(n/2)!}\,\rho^{n/2}} \]


\[ \boxed{\displaystyle b_{nm}=\delta_{r_{nm}\in\mathbb N} \frac{\sqrt{n!/m!}}{2^{r_{nm}}r_{nm}!}\,\rho^{r_{nm}}} \]


\[ \langle\mathcal O_n(X)\widehat{\mathcal O}_m(Y)\rangle =b_{nm}(X\circ X)^{r_{nm}}(X\bullet Y)^m, \qquad \langle\widehat{\mathcal O}_m\rangle=\delta_{m0}. \]

Ambient–ambient check: \(b_{n_1m}b_{n_2m}=\dfrac{\sqrt{n_1!n_2!}}{m!\,2^{r_{1m}+r_{2m}}r_{1m}!r_{2m}!}\rho^{r_{1m}+r_{2m}}\), with \(r_{im}=(n_i-m)/2\). For a nondegenerate scalar family this fixes BOE products—not individual three-point coefficients.

Example 2: a reciprocal nn-dCFT


1

Same factorized distribution

\(\widehat P(\hat\theta)\widetilde P(\tilde\theta)\), with moments \(\widehat\mu_2,\widetilde\mu_2\).

2

Defect network

\[ \widehat\varphi_{\widehat\Delta}(X) =(X\bullet\hat\theta)^{-\widehat\Delta} \]

3

Ambient network

\[ \Phi_\Delta(X)= (X\bullet\hat\theta+X\circ\tilde\theta)^{-\Delta} \]

4

Positive dimensions

\(\widehat\Delta,\Delta>0\)

positive does not by itself imply unitarity

\[ \Phi_\Delta(X)= \sum_{n=0}^{\infty}\frac{(\Delta)_n}{n!}(-1)^n\, \widetilde\varphi_{-n}(X)\, \widehat\varphi_{\Delta+n}(X) \]

Example 2: One-point functions


spectrum starts at \(\widehat\Delta=\Delta>0\)
defect

\[ \mathbb E[\widehat\varphi_{\widehat\Delta}(X)]=0 \]

ambient

\[ \mathbb E[\Phi_\Delta(X)]=0 \]

With no \(\widehat\Delta=0\) term, there is no coupling to the defect identity.

Example 2:Two-point functions


defect–defect

\[ \begin{aligned} &\mathbb E[ \widehat\varphi_{\widehat\Delta_1}(X_1) \widehat\varphi_{\widehat\Delta_2}(X_2)]\\ &\quad= \frac{\sec(\pi\widehat\Delta_1)} {\Gamma(\widehat\Delta_1)} (\widehat\mu_2X_1\bullet X_2)^{-\widehat\Delta_1} \delta_{\widehat\Delta_1,\widehat\Delta_2}. \end{aligned} \]

For integer \(\widehat\Delta\), \(\sec(\pi\widehat\Delta)=(-1)^{\widehat\Delta}\).

ambient–defect

\[ \begin{aligned} &\mathbb E[ \Phi_{\Delta_1}(X_1) \widehat\varphi_{\widehat\Delta_2}(X_2)] = \frac{ 2^{(\Delta_1-\widehat\Delta_2)/2} \sec(\pi\widehat\Delta_2)} {\Gamma(\Delta_1) \Gamma(\frac{\widehat\Delta_2-\Delta_1}{2}+1)} \\[-1pt] &\qquad\times (\widetilde\mu_2X_1\circ X_1)^{ (\widehat\Delta_2-\Delta_1)/2} (\widehat\mu_2X_1\bullet X_2)^{ -\widehat\Delta_2}. \end{aligned} \]

nonzero for \(\widehat\Delta_2\geq\Delta_1\)


The mixed result has exactly the kinematic form required by defect conformal symmetry.

Example 2: exact scalar one-/two-point data

scalar spectrum and normalization

\[ (\Delta,J)=(\Delta,0),\qquad (\widehat\Delta_k,s)=(\Delta+2k,0), \quad k=0,1,\ldots \]


\[ c_{\hat\Delta}=\frac{\sec(\pi\hat\Delta)}{\Gamma(\hat\Delta)\widehat\mu_2^{\,\hat\Delta}},\qquad \mathcal O_{\hat\Delta}=c_{\hat\Delta}^{-1/2}\varphi_{\hat\Delta} \]

\[ a_\Delta=b_{\Delta\mathbf1}=0, \qquad \langle\widehat{\mathcal O}_{\widehat\Delta}\rangle=0 \]


No coupling of bulk primary to defect identity.

unit-normalized scalar BOE tower

\[ \boxed{\displaystyle b_{\Delta\hat\Delta_k}=\frac1{k!} \left(\frac{\widetilde\mu_2}{2\widehat\mu_2}\right)^k \sqrt{\frac{\Gamma(\Delta+2k)}{\Gamma(\Delta)}}}, \qquad k\geq0 \]

Leading restriction: \(b_0=1\)

The variance ratio controls every higher scalar coupling.

The exact ambient–ambient hypergeometric kernel resums this tower together with the transverse-spin sectors.

Meromorphic and generically nonunitary: \(c_\hat\Delta\) has poles at \(\hat\Delta=(2\ell+1)/2\) and changes sign across them. Outside the bare convergence window these are analytically continued, renormalized data.

Beyond Guassian: Deforming \(P\)

keep every primary architecture · deform only \(P\)

\[ P\propto\exp\!\left[ -\frac{u}{2\widehat\sigma} -\frac{v}{2\widetilde\sigma} +\sum_{a,b}\lambda_{ab}u^av^b \right], \quad u=\widehat\theta^2,\ v=\widetilde\theta^2 \]


\(\widehat\sigma=1.10\)
\(\widetilde\sigma=0.85\)
λ₂₀ = 0
λ₀₂ = 0
λ₁₁ = 0

Only the invariant radial moments \(M_{a,b}=\mathbb E_P[u^av^b]\) change. Spectra and selection rules stay fixed.

Null-projected angular contractions are analytic. The full Gaussian reduction agrees through degree six to \(2.7\times10^{-15}\).

Gaussian reference: the exact coefficients at \(\lambda_{ab}=0\) provide the starting point.

Chapter 5

What was achieved?

a solvable laboratory · a new toolset for dCFT data

A brief recap


completed here
  • Extended neural-network conformal field theories to conformal defects.
  • Built monomial and reciprocal nn-dCFT toy models.
  • Computed ambient, defect and mixed scalar correlators.
  • Made explicit the dCFT data of new theories.
next questions

Key takeaways


1

UATs give a well-defined approximation of the path integral.

2

Finite neural architectures can already define CFT and dCFT data. A UAT limit is not required.

3

Large-width Gaussian limits generally yield generalized-free theories

4

Can combine these nn-CFTs (nn-dCFTs) to generate infinitely many new theories


Thank you


Questions are welcome — b.suzzoni@benterre.com

arXiv:2512.07946