A neural network accelerator based on complex phase interference and a chip implementation method thereof
Patent Information
- Application Number
- CN202611280542.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-22
- Publication Date
- 2026-09-25
AI Technical Summary
本发明旨在解决现有Transformer加速器中计算复杂度过高存在存储墙瓶颈、位置编码外推性差限制上下文窗口、
计算复杂度降低:干涉计算采用物理模拟并行计算,计算复杂度从数字域的O(N²·d)降为O(N²)模拟干涉,且N×N干涉单元全并行工作,延时仅1~5ns;
Smart Images

Figure CN122819338A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of artificial intelligence hardware accelerators, analog computing chips, and photonic computing technologies, specifically to a neural network accelerator based on complex phase interference and its chip implementation method. This invention can be applied to cloud-based AI inference acceleration, edge computing devices, photonic computing chips, and neuromorphic computing chips, and is particularly suitable for Transformer model acceleration scenarios requiring high energy efficiency, low latency, and long sequence processing capabilities. Background Technology
[0002] Since its introduction in 2017, the Transformer architecture has become the mainstream model architecture in fields such as natural language processing, computer vision, and multimodal learning. Large-scale language models such as GPT, BERT, and LLaMA are all based on the Transformer architecture, with self-attention as its core mechanism. Attention(Q,K,V) = softmax(QK^T / √d_k)·V However, the existing Transformer architecture and its hardware implementation have the following fundamental problems: (i) The computational complexity is too high, and there is a "memory wall" bottleneck. Self-attention mechanisms require calculating attention scores between all token pairs, with a computational complexity of O(N²·d), where N is the sequence length and d is the feature dimension. When the sequence length N exceeds 10k, the computational complexity increases quadratically. Existing digital accelerators such as GPUs and TPUs use multiply-accumulate (MAC) arrays to implement matrix multiplication, with a single MAC operation consuming approximately 0.1–1 pJ, resulting in high power consumption for large-scale computations. Furthermore, data transfer accounts for over 90% of energy consumption, highlighting a severe "memory wall" bottleneck. (ii) Poor extrapolation of positional encoding limits the context window. Existing Transformers require additional sine / cosine position encoding or rotation position encoding (RoPE) to inject the token's position information. However, these position encoding methods experience a sharp performance drop when the sequence length during inference exceeds the training length, making arbitrary length extrapolation impossible. (iii) Softmax normalization is unstable and requires a large number of normalization layers. The Softmax operation in self-attention mechanisms involves exponential calculations, which can easily lead to numerical overflow, requiring additional normalization layers such as LayerNorm for compensation. These normalization layers increase hardware overhead and computational latency, and make hyperparameter tuning difficult during training. (iv) Real-valued vector addition destroys modulus information The residual connection uses vector addition to update the token state, which destroys the magnitude information of the token vector, requiring an additional normalization layer to recover it. Existing hardware accelerators lack effective utilization of complex phase information. (v) Existing photonic computing schemes do not utilize phase information. Existing optical neural networks mostly use Mach-Zehnder interferometer (MZI) grids to perform matrix multiplication, but they still use traditional real number weight matrices, do not utilize the phase of light waves as an information carrier, and lack a deep integration with physical geometric theory. To address the above problems, this invention provides a neural network accelerator based on complex phase interference and its chip implementation method. Summary of the Invention
[0003] (a) Purpose of the invention This invention aims to address the issues of excessive computational complexity, memory wall bottleneck, poor positional encoding extrapolation, and context window limitations in existing Transformer accelerators. The problems include the instability of Softmax normalization requiring numerous normalization layers, the destruction of mode length information by real-valued vector addition, and the failure of existing photonic computing schemes to utilize phase information. This paper presents a neural network accelerator based on complex phase interference and its chip implementation method. (II) Technical Solution This invention provides a neural network accelerator based on complex phase interference, characterized in that it comprises: An input encoding module is used to convert the input token sequence into a phase angle θ_i on the complex unit circle; An interferometric calculation array is used to calculate the interference intensity α_ij = cos(θ_i -θ_j) between any two phase angles; A phase update module is used to update the state of each phase angle according to the interference intensity and perform complex multiplication rotation operation; An output readout module is used to convert the updated phase distribution into an output result; A digital control unit is used to control the timing and parameter configuration of each module; All phase angle information is transmitted and processed within the chip in the form of analog signals, without the need to convert it into digital signals for matrix multiplication operations. (1) The input encoding module is used to convert the input token sequence into a phase angle θ_i on the complex unit circle; The transformation obtains the semantic phase θ_sem by looking up a table. (i),通过时间步长累积获取位置相位θ_pos (i); The total phase angle θ_i = θ_sem^(i) + θ_pos^(i) (modulo 2π); The output complex signal z_i = e^(i·θ_i) satisfies |z_i|² ≡ 1; The input encoding module includes at least one of a DAC or a phase modulator; No additional location encoding module is required. (2) Interferometric computation array Used to calculate the interference intensity α_ij = max(cos(θ_i - θ_j), 0) between any two phase angles; Only positive interference intensity values are retained, while negative interference intensity values are set to zero; The interference calculation array can be implemented in at least one of the following three ways: (2a) CMOS analog circuit implementation: It contains N×N Gilbert multiplier units, each of which receives two differential voltage inputs V_θ_i and V_θ_j; Each Gilbert multiplier unit contains a double-balanced differential pair and a tail current source I_tail; Output current I_out = I_tail · cos(θ_i - θ_j); The output current of all units in each row is summed directly through the metal wire; Includes a positive half-wave rectifier to achieve max(I, 0); (2b) Photonic integration implementation method: It contains N optical waveguides, each carrying phase φ_i; It includes a 1×N optical splitter, which splits each optical signal into N paths; It contains N×N Mach-Zehnder interferometer units; Each Mach-Zehnder interferometer unit contains two 3dB couplers and a phase delay arm ΔL; Output light intensity I_out = I_0 / 2 · [1 + cos(φ_i - φ_j)]; It includes a germanium-silicon balanced detector that converts light intensity signals into current signals; (2c) Superconducting microwave implementation method: It contains N microwave resonators, each carrying a phase φ_i; Interference is achieved using a SQUID coupler array; Includes a Josephson parametric amplifier for signal readout. (3) Phase update module Used to update the state of each phase angle based on the interference intensity; The updated formula is: θ_i^(t+1) = θ_i^(t) + η · Σ_j α_ij · θ_j^(t) (modulo 2π) Where η is the adjustable coupling coefficient; The update is a pure rotation operation, keeping the complex modulus constant at 1, and no normalization layer is required; It contains an array of accumulators, each accumulator containing an integration capacitor C_i; It includes a phase rotation actuator, wherein the actuator is at least one of a voltage-controlled oscillator or a phase modulator; It includes control logic, including clock input, reset input, and timing switches; The coupling coefficient η can be configured independently at different layers. (4) The output readout module is used to convert the updated phase distribution into an output result; The output probability is determined by calculating the number of implicit branches m_i corresponding to each phase: m_i = floor(M · (|sin θ_i| + 1) / 2) P_i = m_i / Σ_k m_k Where M is the maximum number of branches; The output action S_i = θ_i · ħ; It includes at least one of an ADC or a detector. (5) The digital control unit is used to control the timing and parameter configuration of each module; It includes a timing controller, parameter registers, and external interfaces; The external interface is at least one of SPI, I²C or AXI interfaces. (6) Multi-layer stacked architecture It includes a direct cascade of an L-layer interferometric calculation array and a phase update module; Direct interlayer cascading eliminates the need for residual connections; The value of L ranges from 4 to 12; Each layer contains an independent coupling coefficient η_l. Core mathematical framework: In this invention, the neural network state is represented by the phase angle θ_i on the complex unit circle, where: z_i = e^(i·θ_i),|z_i|² ≡ 1 The representation satisfies the postulate that the modulus is always 1 in the geometrically unified field theory. The state update is a pure rotation operation, mathematically equivalent to moving along a geodesic on the fiber bundle. Mathematical definition of interference computation: In this invention, the interference strength between any two tokens is defined as: α_ij = max(cos(θ_i - θ_j),0), The interference intensity directly corresponds to the visibility of the physical interference fringes, and does not require Softmax normalization. Loss function: This invention uses the principle of minimum action to define the loss function: ℒ = |Σ_i θ_i^(L) - Θ_target|² + λ · Σ_{i=2}^N [1 - cos(θ_i - θ_{i-1})] The first term is the alignment loss, which makes the total action approach the target phase; the second term is the smoothing loss, which constrains the phase continuity of adjacent tokens; and λ is the regularization coefficient. Training methods: Forward propagation: Calculate the L-layer interference-update sequentially, keeping all phases within the range [0, 2π); Backpropagation: Automatic differentiation is used, and all operations are differentiable in the complex field; Parameter update: Gradient descent updates the initial phase θ_i^(0) and coupling coefficient η_l. (III) Beneficial Effects Reduced computational complexity: Interference calculation adopts physical simulation parallel computing, reducing the computational complexity from O(N²·d) in the digital domain to O(N²) simulation interference, and the N×N interference units work in parallel with a delay of only 1 to 5 ns; Eliminating positional encoding: The phase angle θ_i contains both semantic and positional information, eliminating the need for an additional positional encoding module and enabling extrapolation of sequences of arbitrary length; No normalization layer required: Complex multiplication rotation keeps the modulus constant at 1, completely eliminating the need for LayerNorm and BatchNorm, reducing hardware overhead and computational latency; High energy efficiency: By using analog computing instead of a digital MAC array, the energy efficiency can reach 0.25~10 TOPS / W, which is significantly better than existing GPU solutions; Physical interpretability: All calculations have a clear physical correspondence, and the decision-making process is interpretable and traceable. Detailed Implementation
[0004] Example 1: CMOS Analog Chip Implementation Step 1: Input the code. The input token sequence is converted into an analog phase voltage V_θ_i using a DAC. Each token corresponds to an 8-bit digital value, and the semantic phase θ_sem is obtained by looking up a table. (i),通过时间计数器累积得到位置相位θ_pos(i) The sum of the two is used to output a differential voltage pair (V_p, V_n) by the DAC, where V_p - V_n ∝ θ_i. The DAC resolution is 8 to 12 bits, and the update rate is ≥100 MSPS. Step 2: Interference calculation. An N×N Gilbert multiplier array (N=256) was fabricated using a 28nm CMOS process. Each cell contains a double-balanced differential pair and a 100μA tail current source. The input is two differential voltage pairs (V_θ_i, V_θ_j), and the output current I_out = 100μA · cos(θ_i - θ_j). The output currents of each row of 256 cells are directly summed via metal lines, and after passing through a positive half-wave rectifier, the output is I_i = Σ_j max(cos(θ_i - θ_j), 0). The computation delay per layer is approximately 5ns, and the power consumption is approximately 50mW. Step 3: Phase update. The summing current I_i of each channel is input to the integrating capacitor C_i (C=1pF), the integration time T_int=5ns, and the output voltage V_i ∝ I_i · T_int / C_i = η · Σ_j α_ij · θ_j. V_i is input to the voltage-controlled oscillator (VCO), and the VCO outputs a phase increment Δθ_i ∝ V_i, which is fed back to the input via a delay line, completing θ_i^(t+1) = θ_i^(t) + Δθ_i (mod 2π). The VCO's center frequency is approximately 2GHz, and the tuning range is ±500MHz. Step 4: Multi-layer stacking. The single-layer structure described in steps 1-3 is directly cascaded into L=8 layers without the need for inter-layer buffering or normalization. The coupling coefficient η_l of each layer is independently configured through a digital control unit (8-bit register). The total delay is approximately 40ns (8 layers × 5ns), and the total power consumption is approximately 400mW. Step 5: Read the output. The 8th layer output phase distribution θ_i^(8) is the input / output readout module. The number of branches for each token is calculated as m_i = floor(16 · (|sinθ_i| + 1) / 2), and then converted to digital output (8-bit parallel) via an ADC. The ADC sampling rate is ≥200MSPS, and the resolution is 8-bit. Step 6: Training methods. Initialize all initial phases θ_i^(0) to random values (uniformly distributed [0, 2π)), and initialize the coupling coefficients of each layer η_l=0.1. Forward propagation: 8 layers of interference-update are calculated sequentially to obtain the output phase θ_i^(8). Calculate the loss: ℒ = |Σ_i θ_i^(8) - Θ_target|² + 0.01 · Σ_{i=2}^N [1 -cos(θ_i - θ_{i-1})]. Backpropagation: The gradients ∂ℒ / ∂θ_i^(0) and ∂ℒ / ∂η_l are calculated using automatic differentiation in the complex field. Parameter update: Gradient descent, learning rate 0.001, iterations 10k to 100k until convergence. Performance test results: Test item: Single-layer latency, existing GPU solution: ~100μs (including data transfer), this embodiment (CMOS): 5ns Test item: Energy efficiency ratio; Existing GPU solution: ~0.01 TOPS / W; This embodiment (CMOS): ~0.25 TOPS / W Test item: Normalization layer; Existing GPU solution: requires LayerNorm; This embodiment (CMOS): requires no normalization layer. Test item: Position encoding; Existing GPU solutions: require RoPE or Sinusoidal; This embodiment (CMOS): Phase self-contained position. Test item: Extrapolation capability; Existing GPU solution: Limited; This embodiment (CMOS): Arbitrary length extrapolation Example 2: Implementation of Photonic Integrated Chip Step 1: Input the code. A silicon-based optoelectronic platform is employed, with a silicon nanowire waveguide cross-section of 220nm × 500nm. The input token sequence is converted into a phase signal φ_i in the optical waveguide via a micro-ring modulator array (Q≈10000). A laser source (wavelength 1550nm, power 100mW) is input to each micro-ring modulator via a 1×N splitter. Each modulator is driven by an electrical signal to change its resonant wavelength, thereby modulating the optical phase. Step 2: Interference calculation. N optical waveguide signals are split by a 1×N splitter and then enter the Mach-Zehnder interferometer grid. Each Mach-Zehnder interferometer unit contains two 3dB couplers and a phase delay arm ΔL. The phase difference between the two inputs is Δφ = φ_i - φ_j, and the output light intensity is: I_+ = I_0 / 2 · [1 + cos(Δφ)], I_- = I_0 / 2 · [1 - cos(Δφ)] The differential current output by the balanced detector is I_out = I_+ - I_- = I_0 · cos(Δφ). All detectors output in parallel, with a single-layer delay of approximately 1 ns. Step 3: Phase update. The current signal output by the detector is amplified by a transimpedance amplifier and then drives a phase modulator. The phase modulator applies an increment Δφ_i ∝ Σ_j α_ij · φ_j to the optical phase. The updated phase is then fed back to the next layer or the next cycle via an optical delay line. Step 4: Multi-layer stacking and reading. The eight-layer Mach-Zehnder interferometer grid is directly cascaded with interlayer optical waveguides, eliminating the need for photoelectric-to-electro-optical conversion. The output terminal reads out the phase distribution via a coherent detector array, which is then converted into a digital output by an ADC. Performance test results: Test item: Single-layer latency, existing GPU solution: ~100μs (including data transfer), this embodiment (photon): 1ns Test project: Supports scale N; Existing GPU solution: Limited by video memory; This example (Photon): 1024 (scalable) Test item: Energy efficiency ratio, existing GPU solution: ~0.01 TOPS / W, this embodiment (Photon): ~10 TOPS / W (estimated) Test item: Operating temperature; Existing GPU solution: Room temperature; This embodiment (Photon): Room temperature Example 3: Digital Verification Prototype Implementation Step 1: Digital phase representation. In the FPGA, the phase angle θ_i ∈ [0, 2π] is represented in fixed-point numbers (16-bit, 12-bit decimal places). All operations reproduce the interferometric calculation and phase update logic in the digital domain for algorithm verification and model training. Step 2: Digital implementation of interferometric calculations. The CORDIC algorithm is used to calculate cos(θ_i - θ_j), with a parallelism of N=64, a clock frequency of 200MHz, and a delay of approximately 1280 clock cycles (6.4μs) per layer. Step 3: Training and Validation. The training algorithm was run on an FPGA prototype to verify the model's convergence and accuracy. The trained initial phase θ_i^(0) and coupling coefficient η_l can be derived and embedded into the ASIC chip. Step 4: Data export. After training converges, the configuration parameters (initial phase θ_i^(0), coupling coefficients of each layer η_l, number of output branches M) are exported as a configuration file for the initialization configuration of the ASIC chip. Attached Figure Description
[0005] Figure 1 A block diagram of the overall architecture of a neural network accelerator based on complex phase interference provided by the present invention; Figure 2 for Figure 1 Circuit schematic diagram of a CMOS analog computing array; Figure 3 for Figure 1 Schematic diagram of the optical path of the interferometric computing array (photonic integration embodiment); Figure 4 This is the circuit schematic for the phase update module; Figure 5 This is a schematic diagram of the multi-layer stacked architecture of the accelerator of the present invention. Terminology Explanation
[0006] Complex phase interference: Information is encoded using phase angles on complex unit circles, and similarity is calculated through physical interference; Interference calculation array: A hardware array used to calculate the interference intensity between any two phase angles; Gilbert multiplier: A unit circuit in CMOS analog circuitry that implements analog multiplication operations; Mach-Zehnder interferometer: The basic unit for realizing optical wave interference in photonic chips; Phase update: A complex multiplication rotation of the phase angle based on the interference intensity; The modulus is always 1: complex signals always remain on the unit circle and do not require normalization; The principle of least action: The principle that the actual path in a physical system causes the action to reach an extreme value.
Claims
1. A neural network accelerator based on complex phase interference, characterized in that, include: An input encoding module is used to convert an input token sequence into a phase angle θ_i on a complex unit circle. The input encoding module includes at least one of a DAC or a phase modulator and outputs a complex signal z_i = e^(i·θ_i) and satisfies |z_i|² ≡ 1. No additional position encoding module is required. An interferometric calculation array is used to calculate the interference intensity α_ij = max(cos(θ_i -θ_j), 0) between any two phase angles; A phase update module, connected to the interference calculation array, is used to update the state of each phase angle according to the interference intensity and perform a complex multiplication rotation operation θ_i^(t+1) = θ_i^(t) + η·Σ_j α_ij·θ_j^(t) (mod 2π), where η is an adjustable coupling coefficient; An output readout module is connected to the phase update module and is used to convert the updated phase distribution into an output result; A digital control unit is connected to the above modules and is used to control the timing and parameter configuration of each module.
2. The accelerator according to claim 1, characterized in that: The interference calculation array is implemented using a CMOS analog circuit and contains N×N Gilbert multiplier units. Each unit contains a double-balanced differential pair and a tail current source I_tail, with an output current I_out = I_tail · cos(θ_i - θ_j). The output currents of all units in each row are directly summed via metal lines, and a positive half-wave rectifier is included to achieve max(I, 0).
3. The accelerator according to claim 1, characterized in that: The interferometric computing array is implemented using photonic integration and includes N optical waveguides, 1×N optical splitters, N×N Mach-Zehnder interferometer units, and germanium-silicon balanced detectors. Each Mach-Zehnder interferometer unit includes two 3dB couplers and a phase delay arm ΔL, with an output light intensity I_out = I_0 / 2 · [1 + cos(φ_i - φ_j)].
4. The accelerator according to claim 1, characterized in that: The phase update module includes an accumulator array and a phase rotation actuator; the accumulator array includes an integrating capacitor C_i, and the phase rotation actuator is at least one of a voltage-controlled oscillator or a phase modulator.
5. The accelerator according to claim 1, characterized in that: The output readout module determines the output probability by calculating the implicit branch number m_i corresponding to each phase: m_i = floor(M·(|sin θ_i| + 1) / 2), P_i = m_i / Σ_km_k, where M is the maximum branch number; the output action S_i = θ_i·ħ.
6. The accelerator according to claim 1, characterized in that: The accelerator comprises a direct cascade of an L-layer interferometric computation array and a phase update module. The direct cascade between layers does not require residual connections. The value of L ranges from 4 to 12, and each layer contains an independent coupling coefficient η_l.
7. A neural network acceleration method based on complex phase interference, characterized in that, Includes the following steps: (1) The input token sequence is converted into a phase angle θ_i on the complex unit circle by the input encoding module, and the complex signal z_i = e^(i·θ_i) is output, satisfying |z_i|² ≡ 1; (2) Calculate the interference intensity α_ij = max(cos(θ_i -θ_j), 0) between any two phase angles using an interferometric calculation array; (3) The phase update module updates the state of each phase angle according to the interference intensity: θ_i^(t+1) = θ_i^(t) +η·Σ_j α_ij·θ_j^(t) (modulo 2π); (4) Repeat steps (2) and (3) L times, where L is 4 to 12; (5) The updated phase distribution is converted into an output result through the output readout module; All phase angle information is transmitted and processed within the chip in the form of analog signals, without the need to convert it into digital signals for matrix multiplication operations.
8. The method according to claim 7, characterized in that: The interference calculation described in step (2) is implemented in the CMOS analog circuit through a Gilbert multiplier array. The output current of all cells in each row is directly summed through the metal line, and the single-layer calculation delay is ≤5ns.
9. The method according to claim 7, characterized in that: The interference calculation described in step (2) is implemented in the photonic chip through a Mach-Zehnder interferometer grid, with N-way optical waveguides interfering in parallel, and a single-layer calculation delay ≤1ns.
10. The method according to claim 7, characterized in that: The training process uses the minimum action loss function: ℒ = |Σ_i θ_i^(L) - Θ_target|² + λ·Σ_{i=2}^N [1 - cos(θ_i - θ_{i-1})].