Self-attention network system based on optical computing acceleration

The self-attention network system accelerated by optical computing utilizes the optical domain operation module for matrix operations and signal processing, solving the problems of high computational latency, high energy consumption, and insufficient parallel performance of self-attention networks. It achieves efficient and accurate calculation of relevance parameters, improving the real-time performance and scalability of text classification.

CN120806013BActive Publication Date: 2025-11-21INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511250135.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-03
Publication Date
2025-11-21
Estimated Expiration
2045-09-03

AI Technical Summary

Technical Problem

Existing self-attention networks suffer from high computational latency, high energy consumption, and insufficient parallel performance when calculating relevance parameters, which affects the efficiency and accuracy of large-scale text classification tasks.

Method used

The self-attention network system, accelerated by optical computing, converts text feature vectors into optical signals through an input conversion module. It then utilizes an optical domain computation unit for adaptation and processing, an optical signal processing module for adaptation, an optical domain computation module for matrix operations, a signal processing module for wavelength matching, phase adjustment, and hybrid interference, and an output module for detecting correlation parameters.

Benefits of technology

By leveraging the high parallelism and low latency of photonic computing, the computational efficiency and accuracy of correlation parameters are improved, enhancing the applicability of self-attention networks in large-scale text classification tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120806013B_ABST
    Figure CN120806013B_ABST
Patent Text Reader

Abstract

The application discloses a self-attention network system based on optical computing acceleration, and relates to the technical field of data processing.The input conversion module converts a text feature vector into an optical signal, the optical domain operation module performs adaptive processing on a parameter matrix of the self-attention network and completes optical domain matrix operation, the signal processing module performs wavelength matching, phase adjustment and hybrid interference on the optical signal, and the output module generates a correlation parameter.The architecture replaces a traditional electronic processor to complete large-scale matrix operation by virtue of the natural high parallelism and low delay characteristics of photonic computing, thereby reducing delay and energy consumption in the data processing process.The architecture solves the problems of high calculation delay, high energy consumption and insufficient parallelism of the traditional electronic processor when running the self-attention network due to large-scale matrix operation, and achieves the technical effects of improving the calculation efficiency and accuracy of the correlation parameter in the text classification process and enhancing the real-time performance of the self-attention network in a large-scale text data scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and in particular to a self-attention network system based on optical computing acceleration. Background Technology

[0002] In natural language processing tasks such as text classification, related techniques generate relevance parameters that represent the degree of association between elements in a sequence through self-attention networks, achieving a deep understanding of the semantic structure of the text. These relevance parameters are the core of constructing attention weights, directly affecting the model's ability to focus on key information and thus determining classification accuracy. In these techniques, the calculation of relevance parameters relies on a large number of matrix multiplication operations, which must be completed using an electronic processor.

[0003] Related techniques have significant limitations in generating relevance parameters. In classic self-attention networks, the calculation of relevance parameters involves global correlation operations between sequence elements. As the data scale grows exponentially, the amount of matrix operations increases dramatically, leading to high computational latency and high energy consumption for electronic processors, making it difficult to generate accurate relevance parameters in real time. Furthermore, traditional electronic processors are limited by Moore's Law, resulting in insufficient performance for parallel processing of large-scale matrix operations. This directly affects the computational efficiency and accuracy of relevance parameters, restricting the application of self-attention networks in large-scale text classification tasks. Summary of the Invention

[0004] This application provides a self-attention network system based on optical computing acceleration, which at least solves the problems of high computational latency, high energy consumption and insufficient parallel performance in related technologies when calculating the correlation parameters of self-attention networks, thus affecting computational efficiency and accuracy.

[0005] This application provides a self-attention network system based on optical computing acceleration, including: an input conversion module, an optical domain operation module, a signal processing module, and an output module;

[0006] The input conversion module is used to receive input vectors and convert them into optical signals. The input vectors are feature vectors obtained by processing text sequences.

[0007] The optical domain computation module includes an adaptation unit and a computation unit. The adaptation unit is used to adapt the parameter matrix in the self-attention network; the computation unit is used to perform optical domain matrix operations based on the adapted parameter matrix and the optical signal of the input vector to generate an intermediate optical signal.

[0008] The signal processing module is used to perform wavelength matching, phase adjustment, and mixed interference on the intermediate optical signal and another input vector optical signal processed by the input conversion module to generate a mixed optical signal containing sum and difference information.

[0009] The output module is used to detect mixed optical signals and output correlation parameters that characterize the correlation between sequence elements. The correlation parameters are used to construct a self-attention mechanism to classify text.

[0010] This application converts text feature vectors into optical signals, then uses an optical domain computation module to perform optical domain matrix operations on the parameter matrix and the input vector, and generates correlation parameters through a signal processing module. By utilizing the high parallelism and low latency characteristics of photonic computing to replace the computation process of traditional electronic processors, this application can solve the technical problems of high computational latency, high energy consumption, and insufficient parallel performance of electronic processors when calculating correlation parameters. This achieves the effect of improving the efficiency and accuracy of correlation parameter calculation and enhancing the applicability of self-attention networks in large-scale text classification tasks. Attached Figure Description

[0011] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 A schematic diagram of the structure of a self-attention network system based on optical computing acceleration provided in an embodiment of this application;

[0013] Figure 2 A schematic diagram of a self-attention network system based on optical computing acceleration provided in an embodiment of this application;

[0014] Figure 3 A silicon-based MZI unit structure diagram provided in the embodiments of this application;

[0015] Figure 4 This is a schematic diagram of a modulator cascaded into a network array provided in an embodiment of this application;

[0016] Figure 5 A flowchart of a method for constructing a self-attention network with optical computing acceleration provided in an embodiment of this application;

[0017] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0018] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, other embodiments obtained by those of ordinary skill in the art without creative effort are all within the protection scope of this application.

[0019] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0020] This application addresses the problems of low computational efficiency, high energy consumption, and limited parallel processing capabilities caused by large-scale matrix operations in self-attention networks for text classification tasks. It constructs an input conversion module to convert text feature vectors into optical signals, and an optical domain computation module to adapt the self-attention network parameter matrix and complete optical domain matrix operations. Combined with wavelength matching, phase adjustment, and hybrid interference operations in the signal processing module, the optical signal is refined. Finally, the output module generates correlation parameters representing the correlation degree of sequence elements. By leveraging the high parallelism and low latency of photonic computing, matrix operations traditionally performed by electronic processors are migrated to the optical domain, achieving high efficiency and low energy consumption in the calculation of correlation parameters in the self-attention mechanism, thereby improving the real-time performance and scalability of text classification.

[0021] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0022] Figure 1 A schematic diagram of a self-attention network system based on optical computing acceleration provided in this application embodiment is shown below. Figure 1 As shown, an embodiment of this application provides a self-attention network system 10 based on optical computing acceleration, specifically including: an input conversion module 101, an optical domain operation module 102, a signal processing module 103, and an output module 104;

[0023] The input conversion module 101 is used to receive the input vector and convert it into an optical signal. The input vector is a feature vector obtained by processing the text sequence.

[0024] Specifically, the electrical signal of the text feature vector is converted into an optical signal by a modulator, and the wave-particle duality of light is used to realize the conversion of the signal carrier; an interface is established from electrical domain information to optical domain operation, providing a suitable signal form for subsequent optical domain processing; the physical limitations of electronic signals in parallel transmission and processing are broken, and the high bandwidth characteristics of optical signals support high-density data parallel transmission, laying the foundation for large-scale matrix operations.

[0025] The optical domain computation module 102 includes an adaptation unit and a computation unit. The adaptation unit is used to adapt the parameter matrix in the self-attention network. The computation unit is used to perform optical domain matrix operations based on the adapted parameter matrix and the optical signal of the input vector to generate an intermediate optical signal.

[0026] Specifically, the adaptation unit performs normalization preprocessing on the parameter matrix of the self-attention network and constrains its element values ​​to the [0,1] interval through linear transformation to match the dynamic range of the optical modulator, solve the nonlinear response problem of the optical modulator, ensure that the optical domain multiplication of the parameter matrix and the input vector is physically feasible, improve the numerical stability of optical domain operations, and avoid computational errors caused by modulator saturation. The signal processing module is used to perform wavelength matching, phase adjustment, and mixed interference on the intermediate optical signal and another input vector optical signal processed by the input conversion module to generate a mixed optical signal containing sum and difference information.

[0027] The computing unit is based on the optical domain matrix multiplication architecture. It maps the adapted parameter matrix to the computing modulation network and realizes matrix-vector multiplication through step-by-step modulation and optical field superposition. It replaces the von Neumann architecture of traditional electronic processors and directly completes linear transformation with the spatial parallelism of light waves. The computation time complexity is reduced from O(n²) in electronic devices to O(1) in the optical domain, which significantly improves the computing efficiency. At the same time, the dissipation-free characteristics of optical signals can effectively reduce energy consumption.

[0028] The signal processing module 103 is used to perform wavelength matching, phase adjustment and mixed interference on the intermediate optical signal and another input vector optical signal processed by the input conversion module to generate a mixed optical signal containing sum and difference information.

[0029] Specifically, by using wavelength division multiplexing (WDM) technology, channels with the same wavelength are allocated to the intermediate optical signal and the input vector optical signal, enabling synchronous transmission of multidimensional data. This ensures that the two optical signals are strictly aligned in space and spectrum, providing coherent conditions for subsequent interference, eliminating cross-channel crosstalk, and ensuring the accuracy of the interference results.

[0030] By applying precise phase shifts using thermo-optic or electro-optic phase shifters, the contrast of interference fringes is optimized, the phase relationship of light waves is controlled, and the interference results can effectively distinguish between sums and differences, thereby improving the signal-to-noise ratio of mixed optical signals and enhancing the detectability of correlation parameters.

[0031] Mode coupling of two optical signals is achieved through a directional coupler, generating orthogonal components containing sums and differences. The results of optical domain matrix operations are transformed into separable physical quantities, facilitating subsequent detection. Interference effects are used to amplify differences in weak signals, improving the resolution of correlation parameters.

[0032] The output module 104 is used to detect the mixed optical signal and output a correlation parameter that represents the correlation between sequence elements. The correlation parameter is used to construct a self-attention mechanism to classify the text.

[0033] Specifically, by using a balanced photodetector to detect the two orthogonal components of the mixed optical signal, suppressing common-mode noise through differential amplification, converting the optical domain interference result into an electrical signal, and analyzing the correlation parameter characterizing the correlation degree of sequence elements, a highly sensitive photo-to-electric conversion is achieved, ensuring the accuracy of the correlation parameter; the differential detection mechanism significantly reduces the impact of environmental noise.

[0034] The self-attention network architecture for accelerated optical computing provided in this invention achieves efficient conversion of electron-optical signals through an input conversion module, breaking through the data transport limitations of traditional computing paradigms; it completes large-scale linear transformations at the speed of light through parallel matrix multiplication in the optical domain operation module; it accurately extracts correlation parameters through wavelength matching, phase modulation, and interferometry in the signal processing module; and it achieves low-noise, high-resolution result analysis through high-sensitivity detection in the output module. By fully integrating the parallelism, low latency, and low power consumption of photonics, the computational flow of the self-attention mechanism is reconstructed, significantly improving the real-time performance and energy efficiency of text classification tasks.

[0035] In some alternative implementations, the input conversion module is a modulator array, which is composed of multiple cascaded modulator units, with each modulator unit corresponding to an element of the input vector.

[0036] Specifically, each modulator unit corresponds to an element of the input vector, and the amplitude or phase of the optical carrier is controlled in real time via an electrically driven signal. When the input vector is applied, each modulator unit independently modulates the light intensity of the corresponding optical path according to the input value, thereby encoding the discrete digital feature vector into a continuous optical signal. This process follows the physical laws of the electro-optic effect, that is, the applied electric field changes the refractive index of the waveguide material, thereby controlling the propagation characteristics of the optical signal. The digital feature vector in the electronic domain is converted into an analog signal in the optical domain, providing a physical carrier for subsequent optical domain matrix operations. This is achieved through a cascaded modulator array, such as... Figure 4 As shown, Figure 4 The diagram shows a modulator cascaded into a network array provided in this application embodiment. This enables one-time parallel loading of multidimensional data, avoiding the time overhead of element-by-element serial processing in traditional electronic computing. It provides the optical signal dynamic range that matches the parameter matrix for subsequent optical domain computing units, ensuring the feasibility of linear operations.

[0037] This invention achieves efficient electronic-optical signal conversion and synchronous loading of multidimensional data through a spatially parallel architecture of a modulator array; it ensures the accuracy of numerical mapping based on the linear characteristics of the electro-optic effect; and it significantly improves the data input efficiency of the self-attention mechanism by combining the low-loss transmission and high parallelism of optical signals. This provides a high-quality input foundation for optical domain matrix operations, overcoming the performance problems of data handling and serial processing in traditional electronic computing.

[0038] In some alternative implementations, the modulator array is a silicon-based Mach-Zehnder modulator array.

[0039] Specifically, a Mach-Zehnder interferometer (MZI) is an optical silicon-based device based on the principle of optical interference that can realize various functions, such as optical switches and modulators, and has wide applications in optical computing, optical communication, and other fields. A silicon-based MZI unit consists of two couplers and two phase shifters, such as... Figure 3 As shown, Figure 3 The silicon-based MZI cell structure diagram provided in the embodiments of this application is shown below, and the overall network architecture diagram is as follows. Figure 2 As shown, Figure 2 This diagram illustrates a self-attention network system based on optical computing acceleration, as provided in an embodiment of this application. The silicon-based MZI cell structure allows adjustment of the optical path difference to achieve interference with different phases. The input optical signal is split into two beams by the input coupler (also called a beam splitter), and each beam passes through two different phase shifters, resulting in a specific phase difference. Finally, these two beams are combined into a new optical signal by the output coupler (also called a beam combiner), forming an interference signal. Based on the phase difference, the amplitude and phase of the optical signal can be precisely controlled, thereby performing complex mathematical operations.

[0040] The input conversion module utilizes a silicon-based Mach-Zehnder modulator array as its core component. The silicon-based Mach-Zehnder modulator is based on the electro-optic effect of silicon photonics: when an external electrical signal is applied to the two arm-shaped waveguides of the modulator, the refractive index of the waveguides is altered through carrier depletion, thereby modulating the phase difference of the optical signal in the two arms. The values ​​of the input vector elements are mapped to electrical drive signals, which are applied to the corresponding modulator units, thus modulating the amplitude or phase of the optical carrier. The CMOS compatibility of the silicon-based platform allows for the large-scale integration of high-density modulator arrays, and each modulator unit can be independently controlled, achieving precise correspondence with the input vector elements.

[0041] This invention achieves efficient and low-loss conversion of electronic to optical signals by employing a silicon-based Mach-Zehnder modulator array, fully leveraging the advantages of silicon photonics in terms of high integration, low power consumption, and high linearity. It not only overcomes the limitations of traditional electronic interfaces in terms of bandwidth and latency but also provides high-quality input optical signals for subsequent optical domain matrix operations, significantly improving the computational efficiency and energy efficiency of the self-attention mechanism, thus laying the foundation for accelerating optical computation in text classification tasks.

[0042] This embodiment provides a detailed description of the process for adapting the parameter matrix in the self-attention network described in the above embodiments. The specific implementation of this process includes the following steps:

[0043] Step a1: Analyze the maximum and minimum values ​​of the elements in the parameter matrix.

[0044] Specifically, by analyzing the maximum and minimum values ​​of the parameter matrix, its numerical distribution range is quantified, providing a normalization benchmark for subsequent linear transformations. The numerical span of the parameter matrix is ​​clarified, providing a mathematical basis for defining the mapping relationship between the input and output domains of the linear transformation. This ensures that the transformed value range is strictly limited to the target interval, avoiding signal truncation or nonlinear distortion caused by parameter values ​​exceeding the physical dynamic range of the optical modulator. At the same time, the relative proportional relationship between parameters is preserved, providing a foundation for high-precision optical domain calculations.

[0045] Step a2: Perform linear transformations on the elements in the matrix according to the maximum and minimum values, so that the transformed elements are in the range of 0 to 1, and obtain the adapted parameter matrix.

[0046] Specifically, an affine transformation is performed on each element through linear normalization, mapping the original parameters to the [0,1] interval. This transformation preserves order, meaning the relative sizes of the elements in the original matrix remain unchanged after the transformation. This forces the numerical range of the parameter matrix to be constrained to the linear operating region of the optical modulator, matching it with the physically adjustable range of the optical signal (such as extinction ratio and modulation depth). This eliminates computational errors caused by numerical overflow, achieving dynamic range adaptation between the parameter matrix and the optical domain computation unit. This maximizes the utilization of the linear response characteristics of the optical modulator, improving the numerical accuracy and stability of optical domain matrix multiplication.

[0047] This invention analyzes the extreme values ​​of the parameter matrix to obtain its numerical distribution characteristics, and then uses a linear normalization transformation to map the parameters to the linear operating range of the optical modulator, achieving precise adaptation of the dynamic range of the parameter matrix and the optical domain computation unit. This process preserves the relative weighting relationships between parameters and avoids nonlinear distortion of the optical signal caused by numerical overflow, significantly improving the numerical accuracy and system robustness of optical domain matrix operations, and providing a key guarantee for efficient optical domain computation using a self-attention mechanism.

[0048] In some alternative implementations, the linear transformation formula is:

[0049]

[0050] Among them, W qk The weight matrix; max(W) qk ) represents the maximum value of the elements of the weight matrix; min(W) qk ) represents the minimum value of the elements in the weight matrix; (W) qk ) it The weight matrix W qk The element located in the i-th row and t-th column, where i is the row index and t is the column index, i=1,2,…d, t=1,2,…d.

[0051] Specifically, the linear transformation formula is based on the Min-Max normalization principle. By analyzing the global extrema (maximum and minimum values) of the weight matrix, it constructs an affine mapping from the original numerical space to the target interval [0,1]. Mathematically, it eliminates the absolute scale differences of the parameter values ​​by subtracting the minimum value through translation and scaling by dividing by the range, retaining only the relative proportional relationships. This process is a deterministic linear transformation, possessing order preservation, meaning that the relative sizes of the elements in the original matrix remain unchanged after the transformation.

[0052] By forcibly constraining the values ​​of the weight matrix elements to the linear operating range [0,1] of the optical modulator, the problem of signal truncation or nonlinear distortion caused by parameter values ​​exceeding the physical limits of optoelectronic devices is solved. By normalizing and compressing the dynamic range of parameter values, the noise amplification effect of optical signals caused by large numerical fluctuations is suppressed, the robustness of optical domain matrix multiplication is improved, and the parameter matrix is ​​made to perfectly match the input characteristics of optical domain operation units (such as silicon-based modulator arrays). This eliminates the need for additional complex dynamic gain control modules and simplifies the system implementation complexity.

[0053] This invention achieves precise adaptation of the dynamic range of the parameter matrix and the optical domain computation unit by mapping the weight matrix to the [0,1] interval. It preserves the relative weight relationships between parameters, avoids nonlinear distortion of the optical signal due to numerical overflow, significantly improves the numerical accuracy and system robustness of optical domain matrix operations, and provides a guarantee for efficient optical domain computation using a self-attention mechanism.

[0054] This embodiment details the process described above, which involves performing optical domain matrix operations on the optical signal based on the adapted parameter matrix and the input vector to generate an intermediate optical signal. The specific implementation of this process includes the following steps:

[0055] Step b1: Map the adapted parameter matrix to the computational modulation network. The computational modulation network corresponds to the modulator unit of the modulator array, and the computational unit in the computational modulation network corresponds to the element in the parameter matrix.

[0056] Specifically, by matching the elements of the parameter matrix with the independent modulation units in the computational modulation network to form a physical correspondence, the spatial distribution of the parameters is ensured to be consistent with the propagation path of the optical signal. This establishes a direct association between the parameter matrix and optical domain computational resources, enabling the parameter elements to independently control the signal characteristics of the corresponding optical path. This provides a physical carrier for subsequent distributed parallel computing, realizes the hardware deployment of the parameter matrix, eliminates storage access latency in traditional electronic computing, and provides a scalable physical foundation for large-scale parallel matrix operations.

[0057] Step b2: Input the optical signal of the input vector into the operational modulation network, so that the optical signal passes through each operational unit in row and column order.

[0058] Specifically, the propagation direction of the optical path and the arrangement order of the modulation units are determined by the cascaded optical waveguide architecture. The optical signal is forced to pass through each modulation unit in sequence according to the row and column priority of matrix operations. This ensures that the multiplication operation between the input vector and the parameter matrix conforms to the mathematically defined row and column order, avoids timing errors caused by parallel processing, and guarantees the correctness of the computational logic. The element-wise traversal of the input vector is completed at the speed of light, significantly shortening the computation cycle. At the same time, the low-loss characteristics of the waveguide structure maintain signal integrity.

[0059] Step b3: The arithmetic unit modulates the power of the optical signal according to the values ​​of the corresponding parameter matrix elements.

[0060] Specifically, by applying an external electrical / thermal excitation signal, the attenuation coefficient or phase offset of the modulation unit for the optical signal is dynamically adjusted to achieve a linear mapping between the optical signal power and the parameter value. The parameter value is converted into a physically controllable optical signal modulation depth, completing the core calculation step of parameter weighting. The high linearity of the optical modulation characteristics enables accurate weight allocation, avoiding quantization errors in electronic digital calculations. At the same time, the high dynamic range of the optical signal is used to improve the calculation accuracy. The modulated optical signal is transmitted and superimposed at the output of the operational modulation network to generate an intermediate optical signal corresponding to the multiplication operation of the parameter matrix and the input vector.

[0061] Step b4 involves transmitting and superimposing the modulated optical signal at the output of the computational modulation network to generate an intermediate optical signal corresponding to the multiplication operation of the parameter matrix and the input vector.

[0062] Specifically, the output optical signals of each modulation unit are spatially converged through optical couplers or waveguide cross structures to achieve the physical summation of row-column inner products in matrix multiplication. The single-point multiplications completed in a distributed parallel manner are counted as global vector outputs, completing the complete matrix operation process from parameter weighting to result aggregation. The parallel summation of massive amounts of data is completed at the speed of light without the need for additional computing resources, significantly reducing computational complexity and energy consumption, while avoiding the problem of bus contention in electronic systems.

[0063] This invention establishes a physical correspondence between parameters and optical path resources through precise mapping between parameter matrices and computational modulation networks; ensures the logical correctness of matrix operations by traversing the row and column order of optical signals; achieves high linearity and low error in weight allocation through parameter-driven optical power modulation; and completes efficient parallel summation calculations through spatial superposition of optical signals. By fully leveraging the high parallelism, low latency, and low power consumption advantages of photonic computing, it reconstructs the core computational paradigm of the self-attention mechanism, significantly improving the real-time performance and energy efficiency of text classification tasks.

[0064] In some alternative implementations, the modulated optical signal is:

[0065]

[0066] in, The modulated optical signal, W is the optical signal with input vector. qk This is the weight matrix.

[0067] Specifically, the input vector is linearly transformed using a weight matrix to generate a modulated optical signal. Based on the continuous power control of the input optical signal by each modulation unit according to the values ​​of the weight matrix elements, the inner product operation of the parameter matrix and the input vector is transformed into real-time modulation of the optical intensity signal. Through the spatial layout of the optical waveguide network and the phase / amplitude control of the modulator, the optical simulation of matrix multiplication is achieved.

[0068] Large-scale matrix multiplication is performed at the speed of light, avoiding the overhead of memory access and instruction relay in electronic computing, with single-step operation latency approaching the light propagation time; modulation energy is applied only to the optical path corresponding to non-zero weights, significantly reducing ineffective power consumption compared to the all-device power supply mode of electronic processors; the high dynamic range and low crosstalk characteristics of optical signals suppress the cumulative effects of quantization errors and thermal noise in traditional electronic computing, maintaining numerical accuracy; the multi-channel parallel capability of optical waveguides supports seamless expansion to higher dimensions, providing linearly increasing computing resources for long-sequence text processing.

[0069] This invention achieves light-speed-level acceleration of the core computation of the self-attention mechanism by completely migrating the linear transformation of the weight matrix and input vector to the optical domain. This process retains the mathematical rigor of matrix multiplication while fully utilizing the high parallelism and low energy consumption advantages of photonic computing, effectively reducing the generation delay and power consumption of intermediate optical signals.

[0070] In some optional implementations, the signal processing module includes a wavelength division multiplexing unit, a phase shifter, and a directional coupler; the specific implementation of wavelength matching, phase adjustment, and hybrid interference between the intermediate optical signal and another input vector optical signal processed by the input conversion module includes the following steps:

[0071] In step c1, the wavelength division multiplexing unit performs wavelength matching between the intermediate optical signal and another input vector optical signal processed by the input conversion module, and allocates channels with the same wavelength.

[0072] Specifically, based on wavelength division multiplexing, the intermediate optical signal and the elements to be correlated in the input vector optical signal are mapped to the same wavelength channel through wavelength selectors and channel allocators, establishing a spatial-spectral correspondence between the two optical signals. This ensures that the element pairs whose correlation needs to be calculated are in the same wavelength channel, providing a physical carrier for subsequent interference, eliminating crosstalk between different wavelengths, ensuring that interference only occurs on the target element pairs, and improving the accuracy of correlation calculation. At the same time, wavelength multiplexing supports parallel processing of multidimensional data, expanding the system throughput.

[0073] In step c2, the phase shifter adjusts the phase of the two optical signals by a preset phase offset.

[0074] Specifically, precise phase shifts (such as π / 2) are introduced through electro-optic or thermo-optic phase shifters. The propagation constant of the light wave is controlled by the electro- / thermal refractive index change of the material, achieving artificial intervention in the phase. This calibrates the phase relationship between the two optical signals, optimizes the contrast and orthogonality of the interference fringes, and effectively separates the sum and difference signals, significantly improving the signal-to-noise ratio of the interference results and suppressing measurement errors caused by initial phase randomness. Dynamic phase compensation adapts to process deviations and environmental fluctuations, enhancing system robustness. For example, for an input optical signal y, after passing through a phase shifter, its output is... ,in For the phase of the moving optical signal, It is the imaginary unit.

[0075] Step c3: The directional coupler couples the two phase-adjusted optical signals to generate a mixed optical signal including the sum and difference values.

[0076] Specifically, based on the evanescent field coupling effect, when two optical signals propagate at close range in a directional coupler, their mode fields exchange energy, resulting in an interference distribution at the output port. By controlling the coupling length and refractive index difference, a specific power distribution and phase reversal can be achieved. Two optical signals that have undergone wavelength matching and phase adjustment are coherently superimposed to generate a mixed optical signal containing both sum and difference values. The dot product operation required for the self-attention mechanism is directly implemented at the physical layer, eliminating the need for complex algorithms in the digital domain. The orthogonal components of the mixed optical signal naturally carry correlation information, simplifying subsequent photoelectric conversion and parameter extraction processes.

[0077] A directional coupler consists of two waveguides close to each other, allowing energy to be transferred between them. Its transfer matrix is ​​as follows. ,in This is the projection coefficient. For a 50:50 directional coupler, .

[0078] This invention achieves wavelength-space mapping of multidimensional data through wavelength division multiplexing (WDM) units, establishing a strict element-level correspondence; optimizes the controllability and stability of interference conditions through precise phase modulation of phase shifters; and transforms the optical domain matrix operation results into directly detectable physical quantities through coherent coupling of directional couplers. By fully integrating the high parallelism, low loss, and reconfigurability of photonics, the core computation of the self-attention mechanism is completed at a single physical layer, significantly improving the energy efficiency and real-time performance of text classification tasks.

[0079] In some alternative implementations, the wavelength division multiplexing unit includes:

[0080] A wavelength selector is used to provide an independent wavelength with the same number of dimensions as the input vector, and the independent wavelength is matched to the channel.

[0081] Specifically, based on wavelength-selective devices such as grating diffraction or arrayed waveguides, an independent set of wavelengths with the same dimension as the input vector is generated, with each wavelength corresponding to a unique channel. By adjusting device parameters (such as temperature or current), the center wavelength and channel spacing can be precisely controlled, assigning a unique wavelength label to each input vector element, and constructing an orthogonal optical frequency domain resource pool. This avoids timing conflicts and address contention in traditional electronic buses, enabling fine-grained partitioning of optical domain resources, with an upper limit of tens to hundreds of independent channels, providing physical layer guarantees for large-scale parallel computing; the low crosstalk characteristics between wavelengths improve system reliability.

[0082] A channel allocator is used to assign elements corresponding to positions in an intermediate optical signal and another input vector optical signal to the same wavelength channel.

[0083] Specifically, by integrating an optical switch matrix or a thermo-optical tunable router, a mapping relationship between intermediate optical signals and input vector optical signal elements is dynamically established based on control signals. Low insertion loss optical path redirection is achieved through waveguide cross-structures, forcibly associating element pairs requiring correlation calculation and constraining them to the same wavelength channel. This ensures that subsequent interference operations only apply to the target element combination, eliminating spurious interference between unrelated elements and improving the specificity of correlation calculation. The dynamic reconfigurability feature supports online adjustment of the mapping strategy to adapt to different task requirements.

[0084] A waveguide array, containing the same number of waveguides as the wavelength channels, is used to carry optical signals of different wavelengths for parallel transmission.

[0085] Specifically, by optimizing waveguide width, height, and cladding refractive index difference, efficient confinement and low bending loss transmission of specific wavelengths are achieved, providing physically isolated transmission paths for multi-wavelength signals and suppressing crosstalk between channels; the parallel waveguide architecture supports the simultaneous transmission and processing of multiple wavelength signals, ensuring the amplitude and phase stability of each wavelength signal and maintaining a high signal-to-noise ratio; the compact planar optical path design significantly reduces the system size and improves integration and scalability.

[0086] This invention constructs an orthogonal optical frequency domain resource pool through a wavelength selector to achieve wavelength-based encoding of input vector elements; establishes element-level mapping relationships through a channel allocator to ensure precise targeting of correlation calculations; and provides a low-loss parallel transmission channel through a waveguide array to ensure synchronous processing of multidimensional data. Leveraging the high density, low power consumption, and high parallelism advantages of photonic integration technology, it solves the resource contention and latency problems of large-scale matrix operations in self-attention mechanisms at the physical layer, providing hardware support for text classification tasks.

[0087] This embodiment details the process by which the wavelength division multiplexing unit in the above embodiment performs wavelength matching between the intermediate optical signal and another input vector optical signal processed by the input conversion module, and allocates channels with the same wavelength. The specific implementation of this process includes the following steps:

[0088] Step d1: Configure the wavelength channels of the wavelength division multiplexing unit. The number of wavelength channels is the same as the dimension of the input vector, and each channel corresponds to a different wavelength.

[0089] Specifically, based on wavelength division multiplexing (WDM) spectrum resource management, an equal number of independent wavelength channels are dynamically configured according to the dimension of the input vector. Precise wavelength allocation is achieved through tunable lasers or fixed filter arrays, pre-allocating dedicated optical frequency domain resources for each element of the input vector, establishing orthogonal optical signal transmission channels, avoiding timing conflicts and address contention issues in traditional electronic buses, realizing fine-grained partitioning of optical domain resources, and scalable to a large number of independent channels to meet the processing needs of high-dimensional text feature vectors; the low crosstalk characteristics between wavelengths ensure the stability of multi-channel parallel transmission.

[0090] Step d2 involves identifying the elements corresponding to the positions in the intermediate optical signal and the other input vector optical signal, and allocating the elements at the corresponding positions to the same wavelength channel, so that the elements of the two signals can share the channel for parallel transmission.

[0091] Specifically, the element positions of the intermediate optical signal and the input vector optical signal are identified step by step by control signals. Based on the optical switch matrix or thermo-optical tunable router, the element pairs for which correlation needs to be calculated are dynamically routed to the same wavelength channel. The element pairs that need to interact in the self-attention mechanism are forcibly associated and constrained to the same wavelength channel to ensure that subsequent interference operations only act on the target element combination, eliminate false interference between unrelated elements, and improve the specificity of correlation calculation. The dynamic routing capability supports online adjustment of mapping strategies to adapt to different task requirements and changes in data distribution.

[0092] This invention achieves efficient management and precise mapping of optical domain resources through the pre-configuration and dynamic allocation of wavelength channels. By establishing an orthogonal optical frequency domain resource pool, it provides physical layer guarantees for large-scale parallel computing; and through a dynamic routing mechanism, it ensures the targeting of correlation calculations and suppresses irrelevant interference.

[0093] This embodiment provides a detailed description of the process by which the phase shifter in the above embodiment adjusts the phase of two optical signals by a preset phase offset. The specific implementation of this process includes the following steps:

[0094] Step e1: Obtain the interference requirements of the directional coupler and set the preset phase offset according to the interference requirements.

[0095] Specifically, by analyzing the relationship between the light intensity distribution at the output port of the directional coupler and the phase difference of the input optical signal, the required phase difference value for achieving the target interference effect, such as the maximum extinction ratio or a specific splitting ratio, is determined. This provides a clear calibration benchmark for subsequent phase adjustment, ensuring that the phase relationship between the two optical signals meets the operating conditions of the interferometer. By calculating the preset ideal phase difference, suboptimal interference results caused by blind adjustment are avoided, thus improving the initial alignment efficiency of the system.

[0096] Step e2 involves associating the phase shifter with the transmission waveguides of the two optical signals, and changing the effective refractive index of the waveguides by applying voltage or controlling the temperature.

[0097] Specifically, by utilizing the electro-optic effect (electric field-induced refractive index change) or the thermo-optic effect (temperature gradient-induced refractive index distribution change), the propagation constant of the waveguide is dynamically adjusted by an external electrical signal or heat source to achieve precise phase control of the optical signal. A physical coupling is established between the phase shifter and the optical signal transmission path, converting the electrical / thermal excitation signal into a phase change in the optical signal, achieving subwavelength-level precision phase adjustment capability, and supporting dynamic correction of phase drift caused by process deviations and environmental fluctuations.

[0098] Step e3: Monitor the phase difference between the two optical signals to obtain the real-time phase difference.

[0099] Specifically, by using interferometer self-sensing technology or an external phase detector, the phase difference information is extracted in real time by analyzing the light intensity waveform or beat frequency signal after the interference of two optical signals. This provides real-time data support for subsequent error determination, and captures the instantaneous changes in phase difference with a high sampling rate, ensuring the timeliness and accuracy of feedback control.

[0100] Step e4: Compare the real-time phase difference with the preset phase offset.

[0101] Specifically, the real-time phase difference is compared with a preset value using a digital signal processor or related comparator circuit to generate an error signal. This determines whether the current phase difference meets the system performance requirements, quantifies the direction and magnitude of the phase deviation, and provides a basis for subsequent adjustment strategies.

[0102] Step e5: When the deviation between the real-time phase difference and the preset phase offset is within the preset tolerance range, maintain the current driving parameters of the phase shifter and maintain a stable phase difference.

[0103] Specifically, when the error signal amplitude is less than the preset threshold, the driving parameters of the phase shifter, such as the voltage / temperature setpoint, are locked, active adjustment is stopped, system oscillations caused by over-adjustment are suppressed, the phase difference is kept stable within the target range, control energy consumption is reduced, device life is extended, and the stability of the interference results is ensured.

[0104] Step e6: When the deviation between the real-time phase difference and the preset phase offset is not within the preset tolerance range, repeat the step of changing the effective refractive index of the waveguide by applying voltage or controlling temperature until the real-time phase difference is within the preset tolerance range.

[0105] Specifically, the driving parameters of the phase shifter are dynamically adjusted according to the error signal, driving the system to converge toward the target phase difference. Steady-state error is eliminated through iterative adjustment, forcing the system to return to the preset operating point, significantly improving the robustness of phase control and adapting to long-term drift caused by environmental disturbances and device aging.

[0106] This invention achieves precise adjustment of the phase difference between two optical signals by a phase shifter through phase demand modeling and closed-loop feedback control; establishes a low-delay, high-linearity phase control link through efficient energy conversion of electro-optic / thermo-optic effects; and ensures stability and reliability in complex environments through real-time monitoring and dynamic correction mechanisms.

[0107] This embodiment details the process by which the directional coupler in the above embodiment couples two phase-adjusted optical signals to generate a mixed optical signal including a sum and a difference. The specific implementation of this process includes the following steps:

[0108] Step m1: Input the two phase-adjusted optical signals into the two input terminals of the directional coupler respectively.

[0109] Specifically, the directional coupler has a dual-port input. By injecting the two optical signals to be interfered into adjacent waveguides, they interact according to the evanescent field coupling excitation mode, providing initial conditions for subsequent energy exchange and interference. This ensures that the two signals are highly coincident in space and time, achieving collimated input of the optical signals and reducing coupling loss caused by alignment errors.

[0110] Step m2, based on the mode coupling between waveguides, enables the two optical signals to exchange energy and interfere within the coupler.

[0111] Specifically, when the distance between two waveguides is less than several times the wavelength, the evanescent fields of the guided modes overlap, triggering mode coupling effects. By controlling the coupling length and the refractive index difference, partial or complete energy transfer can be achieved. By driving the amplitude and phase of the two optical signals to coherently superimpose, physically meaningful sum and difference components are generated, enabling optical domain simulation of the inner product operation of the self-attention mechanism within a single device.

[0112] In step m3, the interference results are obtained from the two outputs of the coupler. The interference result at the first output is the sum of the energies of the two signals, and the interference result at the second output is the difference between the energies of the two signals.

[0113] Specifically, one end of the directional coupler outputs a sum signal that is in phase and superimposed, while the other end outputs a difference signal that is out of phase and cancels out. By decomposing the optical domain interference result into independently detectable physical quantities, the complete information required for the self-attention weight calculation is retained. The real and imaginary parts of matrix multiplication are directly separated through the physical layer, eliminating the need for complex number operations in the digital domain.

[0114] Step m4: Integrate the interference results of the first output terminal and the second output terminal to obtain a mixed optical signal including the sum and difference values.

[0115] Specifically, the two output signals are combined into a composite optical signal by an optical beam combiner. The light intensity distribution carries information about the sum and difference, providing a multi-dimensional optical signal with complete correlation information for subsequent photoelectric detection. This simplifies the subsequent signal processing and reduces the requirements for the bandwidth and resolution of the analog-to-digital converter.

[0116] This invention implements inner product operations of the self-attention mechanism directly at the physical layer through the mode coupling effect of the directional coupler; it synchronously acquires sum and difference signals through orthogonal separation of dual outputs; and it eliminates the problems of memory access and bus contention in traditional electronic computing through optical domain integration design.

[0117] In some alternative implementations, the mixed optical signal is:

[0118]

[0119] in, The signal is optical, and j is the imaginary unit. This is the intermediate optical signal. To and Intermediate optical signals of the same wavelength, Input optical signals at the same location with the same wavelength and The element at the i-th position.

[0120] Specifically, for the input optical signal and The same position with the same wavelength After passing through the aforementioned phase shifter and directional coupler, the output mixed optical signal is z. i The optical domain matrix operation results are transformed into directly detectable physical quantities, where the sum component represents the energy superposition of two signals, and the difference component reflects their degree of difference. Together, they constitute the correlation measure required by the self-attention mechanism. Through imaginary units, sum and difference information are transmitted synchronously in a single physical channel, avoiding the resource overhead of multi-channel separation in traditional schemes. While amplifying the effective signal using interference effects, the influence of incoherent stray light is suppressed, improving the signal-to-noise ratio. The complex domain linear transformation is completed at the speed of light, replacing the multi-step numerical calculations of electronic processors, significantly reducing latency and energy consumption.

[0121] This invention achieves efficient computation of correlation parameters in the self-attention mechanism directly in the optical domain by combining complex-domain matrix transformation and coherent interference. It possesses the advantages of high parallelism and low power consumption of photonic computing, and further enhances text classification accuracy and robustness through orthogonal component separation and precise phase control.

[0122] This embodiment details the process of detecting mixed optical signals and outputting correlation parameters characterizing the correlation between sequence elements as described in the above embodiments. The output module includes a balanced photodetector, and the specific implementation of this process includes the following steps:

[0123] Step n1: Detect the optical power of the two outputs in the mixed optical signal through the two detection ends of the balanced photodetector.

[0124] Specifically, two orthogonal components of the mixed optical signal are simultaneously detected using two symmetrically arranged photodiodes. A differential amplifier circuit suppresses common-mode noise, such as ambient light interference, retaining only the optical power difference information between the two signals. The optical domain interference result is converted into an electrical signal, and the optical power values ​​of the two signals are extracted, providing raw data for subsequent correlation calculations. Differential detection significantly improves the signal-to-noise ratio and eliminates the influence of light source intensity fluctuations and background noise. Dual-end synchronous detection ensures the integrity of the sum / difference signals, avoiding information loss caused by single-end detection.

[0125] Step n2 involves analyzing the power of the two optical paths to obtain the power difference. The power difference is positively correlated with the semantic association between corresponding element pairs in the text sequence.

[0126] Specifically, by mapping the power difference of optical interference to the relevance parameters required by the self-attention mechanism, a quantitative index of semantic correlation between text elements is established. The semantic similarity between elements is directly reflected by the linear relationship of optical power difference, and key features can be extracted without complex algorithms.

[0127] Step n3: Output the power difference as a correlation parameter.

[0128] Specifically, by using the power difference as a self-attention weight parameter, it is directly input into subsequent neural network layers for weighted summation and classification, completing the final conversion from light signal to correlation parameters. This provides an interpretable numerical basis for the self-attention mechanism, enabling the extraction and output of correlation parameters at the speed of light, thus avoiding memory access delays in electronic processors.

[0129] This invention achieves efficient conversion of optical domain interferometry results to correlation parameters through differential detection using a balanced photodetector.

[0130] In some alternative implementations, the power difference is:

[0131]

[0132] in, This is the power difference. Encoding optical signals, To and Encoding of light signals of the same wavelength.

[0133] Specifically, the optical simulation of matrix multiplication is achieved through the physical phenomenon of optical interference, transforming the inner product operation of the parameter matrix and the input vector into a measurable difference in optical power; matrix multiplication in traditional electronic computing is transferred to the optical domain, utilizing the parallelism and low latency characteristics of photonics to achieve linear algebraic operations with time complexity; the core calculation can be completed solely through the interference and detection of optical signals, avoiding the large amount of unnecessary power consumption of electronic processors; and the continuous change of optical power directly reflects the dot product result, avoiding the accuracy loss caused by digital quantization.

[0134] The embodiments of the present invention can complete the inner product calculation of the entire vector space through a single optical interference operation, which far exceeds the serial execution speed of electronic processors; the high dynamic range of optical signals supports parameter matrices with large numerical spans, avoiding the overflow problem in electronic computing, eliminating the need for additional analog-to-digital conversion or caching mechanisms, and the optical signals are directly mapped to the final parameters, simplifying the computation link.

[0135] In some optional implementations, the method further includes: dividing the modulator array according to a preset function to obtain a first array and a second array; and assigning optical signals of different wavelength ranges to the modulators in the first array and the second array respectively.

[0136] Specifically, the modulator array can be functionally divided into coarse-grained and fine-grained subarrays. Long-wavelength optical signals are allocated to the coarse-grained subarrays to process sentence-level features, while short-wavelength optical signals are allocated to the fine-grained subarrays to process word-level details. Spatial partitioning can be achieved through left-right partitioning or top-bottom stacking. The two arrays process features of different granularities in parallel, with wavelength allocation matching their functions. Simultaneously, it enables fine-grained modeling of both long-sequence global semantics and local keywords, addressing the accuracy loss caused by single-granularity approaches and achieving multi-scale feature fusion while maintaining optical domain parallelism.

[0137] Through the above description of the embodiments, those skilled in the art can clearly understand that the process according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.

[0138] Figure 5 A flowchart illustrating a method for constructing a self-attention network with accelerated optical computing provided in an embodiment of this application. Figure 5 As shown, embodiments of this application also provide a method for constructing a self-attention network with optical computing acceleration, the process including the following steps:

[0139] Step S501: Receive the input vector and convert it into an optical signal. The input vector is the feature vector obtained by processing the text sequence.

[0140] Step S502: Adapt the parameter matrix in the self-attention network and perform optical domain matrix operations based on the adapted parameter matrix and the optical signal of the input vector to generate an intermediate optical signal.

[0141] Step S503: Perform wavelength matching, phase adjustment and mixed interference on the intermediate optical signal and another input vector optical signal processed by the input conversion module to generate a mixed optical signal containing sum and difference information.

[0142] Step S504: Detect the mixed optical signal and output the correlation parameter, which represents the correlation between sequence elements. The correlation parameter is used to construct a self-attention mechanism to classify the text.

[0143] The self-attention network construction method for optical computing acceleration provided in this invention achieves efficient cross-domain encoding by converting text feature vectors into optical signals. It performs optical domain matrix operations based on the adapted parameter matrix to leverage the high parallelism of photonics. It uses wavelength matching and phase modulation techniques to accurately align multidimensional data, directly extracts the correlation degree of sequence elements using hybrid interference effects, and finally outputs correlation parameters quickly by analyzing the interference pattern through photoelectric conversion. The method in this invention reconstructs the optical domain computing of the self-attention mechanism, significantly improves the energy efficiency and real-time performance of large-scale matrix operations, and solves the problem of insufficient computing power in traditional electronic computing for long sequence processing.

[0144] In some optional implementations, the parameter matrix in the self-attention network is adapted, including:

[0145] Analyze the maximum and minimum values ​​of the elements in the parameter matrix;

[0146] Perform linear transformations on the elements of the matrix based on the maximum and minimum values, so that the transformed elements are in the range of 0 to 1, and obtain the adapted parameter matrix.

[0147] In some optional implementations, optical domain matrix operations are performed based on the adapted parameter matrix and the optical signal of the input vector to generate an intermediate optical signal, including:

[0148] The adapted parameter matrix is ​​mapped to the operational modulation network. The operational modulation network corresponds to the modulator unit of the modulator array, and the operational unit in the operational modulation network corresponds to the element in the parameter matrix.

[0149] The optical signal of the input vector is input into the operational modulation network, so that the optical signal passes through each operational unit in row and column order;

[0150] The processing unit modulates the power of the optical signal according to the values ​​of the corresponding parameter matrix elements;

[0151] The modulated optical signal is transmitted and superimposed at the output of the computational modulation network to generate an intermediate optical signal corresponding to the multiplication operation of the parameter matrix and the input vector.

[0152] In some optional implementations, wavelength matching, phase adjustment, and hybrid interference are performed on the intermediate optical signal and another input vector optical signal processed by the input conversion module, including:

[0153] By performing wavelength matching between the intermediate optical signal and another input vector optical signal processed by the input conversion module, channels with the same wavelength are allocated.

[0154] The phase of the two optical signals is adjusted by a preset phase offset.

[0155] The two phase-adjusted optical signals are coupled to generate a mixed optical signal that includes the sum and difference values.

[0156] In some alternative implementations, channels with the same wavelength are allocated by wavelength matching between the intermediate optical signal and another input vector optical signal processed by the input conversion module, including:

[0157] The wavelength channels of the wavelength division multiplexing unit are configured. The number of wavelength channels is the same as the dimension of the input vector, and each channel corresponds to a different wavelength.

[0158] By identifying the elements corresponding to the positions in the intermediate optical signal and the other input vector optical signal, the elements at the corresponding positions are assigned the same wavelength channel, enabling the elements of the two signals to share the channel for parallel transmission.

[0159] In some optional implementations, phase adjustment of the two optical signals is performed by a preset phase offset, including:

[0160] Obtain the interference requirements of the directional coupler and set a preset phase offset based on the interference requirements;

[0161] The phase shifter is associated with the transmission waveguides of the two optical signals, and the effective refractive index of the waveguide is changed by applying voltage or controlling temperature.

[0162] The phase difference between the two optical signals is monitored to obtain the real-time phase difference;

[0163] The real-time phase difference is compared with the preset phase offset.

[0164] When the deviation between the real-time phase difference and the preset phase offset is within the preset tolerance range, the current driving parameters of the phase shifter are maintained to keep a stable phase difference.

[0165] When the deviation between the real-time phase difference and the preset phase offset is not within the preset tolerance range, the step of changing the effective refractive index of the waveguide by applying voltage or controlling temperature is repeated until the real-time phase difference is within the preset tolerance range.

[0166] In some optional implementations, the two phase-adjusted optical signals are coupled to generate a mixed optical signal including a sum and a difference, including:

[0167] The two phase-adjusted optical signals are input to the two input terminals of the directional coupler, respectively.

[0168] The mode coupling between waveguides enables energy exchange and interference between the two optical signals within the coupler;

[0169] Interference results are obtained from the two outputs of the coupler. The interference result at the first output is the sum of the energies of the two signals, and the interference result at the second output is the difference between the energies of the two signals.

[0170] The interference results from the first and second output terminals are integrated to obtain a mixed optical signal that includes both the sum and the difference.

[0171] In some optional implementations, the mixed optical signal is detected, and a correlation parameter characterizing the correlation between sequence elements is output, including:

[0172] Detect the optical power of the two outputs in the mixed optical signal;

[0173] The power difference between the two optical paths was obtained by analyzing the power of the two paths. The power difference was positively correlated with the semantic association between corresponding element pairs in the text sequence.

[0174] The power difference is output as a correlation parameter.

[0175] For a description of the features in the embodiment of the optical computing-accelerated self-attention network construction method, please refer to the relevant description of the embodiment of the optical computing-accelerated self-attention network system, which will not be repeated here.

[0176] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 6 As shown, the electronic device 60 provided in this embodiment includes at least one processor 601 and a memory 602. Optionally, the electronic device 60 further includes a communication component 603. The processor 601, memory 602, and communication component 603 are connected via a bus.

[0177] In the specific implementation process, at least one processor 601 executes computer execution instructions stored in memory 602, causing at least one processor 601 to execute the above-described embodiment of the optical computing accelerated self-attention network construction method.

[0178] The specific implementation process of processor 601 can be found in the above method embodiments, and its implementation principle and technical effect are similar. It will not be repeated here.

[0179] In the above embodiments, it should be understood that the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the application can be directly manifested as being executed by a hardware processor, or executed by a combination of hardware and software modules within the processor.

[0180] The memory may include random access memory (RAM) and may also include non-volatile memory (NVM), such as at least one disk storage device.

[0181] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, the buses shown in the accompanying drawings are not limited to a single bus or a single type of bus.

[0182] Embodiments of this application also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above embodiments of the optical computing accelerated self-attention network construction method.

[0183] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.

[0184] Embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above embodiments of the optical computing accelerated self-attention network construction method.

[0185] Embodiments of this application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in any of the above embodiments of the optical computing accelerated self-attention network construction method.

[0186] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0187] The foregoing has provided a detailed description of a self-attention network system based on optical computing acceleration provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only intended to aid in understanding the method and core ideas of this application. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from its principles, and these improvements and modifications also fall within the protection scope of the claims of this application.

Claims

1. A self-attention network system based on optical computing acceleration, characterized in that, include: Input conversion module, optical domain computation module, signal processing module, output module; The input conversion module is used to receive an input vector and convert it into an optical signal. The input vector is a feature vector obtained by processing a text sequence. The optical domain computation module includes an adaptation unit and a computation unit. The adaptation unit is used to adapt the parameter matrix in the self-attention network. The computation unit is used to perform optical domain matrix operations based on the adapted parameter matrix and the optical signal of the input vector to generate an intermediate optical signal. The signal processing module is used to perform wavelength matching, phase adjustment, and mixed interference on the intermediate optical signal and another input vector optical signal processed by the input conversion module to generate a mixed optical signal containing sum and difference information. The output module is used to detect the mixed optical signal and output a correlation parameter that characterizes the correlation between sequence elements. The correlation parameter is used to construct a self-attention mechanism to classify the text. The input conversion module is a modulator array, which is composed of multiple cascaded modulator units, and the modulator units correspond to the elements of the input vector. The process of performing optical domain matrix operations based on the adapted parameter matrix and the optical signal of the input vector to generate an intermediate optical signal includes: The adapted parameter matrix is ​​mapped to the computational modulation network, which corresponds to the modulator unit of the modulator array, and the computational unit in the computational modulation network corresponds to the element in the parameter matrix. The optical signal of the input vector is input into the operational modulation network, so that the optical signal passes through each operational unit in row and column order; The processing unit modulates the power of the optical signal according to the values ​​of the corresponding parameter matrix elements; The modulated optical signal is transmitted and superimposed at the output of the operational modulation network to generate an intermediate optical signal corresponding to the multiplication operation of the parameter matrix and the input vector. The signal processing module includes a wavelength division multiplexing unit, a phase shifter, and a directional coupler; The process of performing wavelength matching, phase adjustment, and hybrid interference between the intermediate optical signal and another input vector optical signal processed by the input conversion module includes: The wavelength division multiplexing unit allocates channels with the same wavelength by performing wavelength matching between the intermediate optical signal and another input vector optical signal processed by the input conversion module. The phase shifter adjusts the phase of the two optical signals by a preset phase offset; The directional coupler couples the two phase-adjusted optical signals to generate a mixed optical signal including the sum and difference values.

2. The self-attention network system based on optical computing acceleration according to claim 1, characterized in that, The modulator array is a silicon-based Mach-Zehnder modulator array.

3. The self-attention network system based on optical computing acceleration according to claim 1, characterized in that, The adaptation process for the parameter matrix in the self-attention network includes: Analyze the maximum and minimum values ​​of the elements in the parameter matrix; Based on the maximum and minimum values, perform linear transformations on the elements in the matrix respectively, so that the transformed elements are in the range of 0 to 1, and obtain the adapted parameter matrix.

4. The self-attention network system based on optical computing acceleration according to claim 3, characterized in that, The linear transformation formula is: Among them, W qk The weight matrix; max(W) qk ) represents the maximum value of the elements of the weight matrix; min(W) qk ) represents the minimum value of the elements in the weight matrix; (W) qk ) it The weight matrix W qk The element located in the i-th row and t-th column, where i is the row index and t is the column index, i=1,2,…d, t=1,2,…d.

5. The self-attention network system based on optical computing acceleration according to claim 1, characterized in that, The modulated optical signal is: in, The modulated optical signal, W is the optical signal with input vector. qk This is the weight matrix.

6. The self-attention network system based on optical computing acceleration according to claim 1, characterized in that, The wavelength division multiplexing unit includes: A wavelength selector is used to provide an independent wavelength with the same number of dimensions as the input vector, and the independent wavelength is matched to the channel. A channel allocator is used to assign elements corresponding to positions in an intermediate optical signal and another input vector optical signal to the same wavelength channel. A waveguide array, containing the same number of waveguides as the wavelength channels, is used to carry optical signals of different wavelengths for parallel transmission.

7. The self-attention network system based on optical computing acceleration according to claim 1, characterized in that, The wavelength division multiplexing unit allocates channels with the same wavelength by performing wavelength matching between the intermediate optical signal and another input vector optical signal processed by the input conversion module, including: The wavelength channels of the wavelength division multiplexing unit are configured such that the number of wavelength channels is the same as the dimension of the input vector, and each channel corresponds to a different wavelength. By identifying the elements corresponding to the positions in the intermediate optical signal and the other input vector optical signal, the elements at the corresponding positions are assigned the same wavelength channel, so that the elements of the two signals can share the channel for parallel transmission.

8. The self-attention network system based on optical computing acceleration according to claim 1, characterized in that, The phase shifter adjusts the phase of the two optical signals by a preset phase offset, including: Obtain the interference requirements of the directional coupler, and set a preset phase offset based on the interference requirements; The phase shifter is associated with the transmission waveguides of the two optical signals, and the effective refractive index of the waveguides is changed by applying voltage or controlling temperature. The phase difference between the two optical signals is monitored to obtain the real-time phase difference; The real-time phase difference is compared with a preset phase offset. When the deviation between the real-time phase difference and the preset phase offset is within the preset tolerance range, the current driving parameters of the phase shifter are maintained to keep a stable phase difference. When the deviation between the real-time phase difference and the preset phase offset is not within the preset tolerance range, the step of changing the effective refractive index of the waveguide by applying voltage or controlling temperature is repeated until the real-time phase difference is within the preset tolerance range.

9. The self-attention network system based on optical computing acceleration according to claim 1, characterized in that, The directional coupler couples the two phase-adjusted optical signals to generate a mixed optical signal including a sum and a difference, comprising: The two phase-adjusted optical signals are input to the two input terminals of the directional coupler, respectively. The mode coupling between waveguides enables energy exchange and interference between the two optical signals within the coupler; Interference results are obtained from the two outputs of the coupler respectively. The interference result at the first output is the sum of the energies of the two signals, and the interference result at the second output is the difference between the energies of the two signals. The interference results from the first output terminal and the second output terminal are integrated to obtain a mixed optical signal that includes both sum and difference values.

10. The self-attention network system based on optical computing acceleration according to claim 9, characterized in that, The mixed optical signal is: in, The signal is optical, and j is the imaginary unit. This is the intermediate optical signal. To and Intermediate optical signals of the same wavelength, Input optical signals at the same location with the same wavelength and The element at the i-th position.

11. The self-attention network system based on optical computing acceleration according to claim 9, characterized in that, The output module includes a balanced photodetector, which detects the mixed optical signal and outputs correlation parameters characterizing the correlation between sequence elements, including: The optical power of the two outputs in the mixed optical signal is detected by the two detection ends of a balanced photodetector. The power difference between the two optical paths is obtained by analyzing the power of the two paths. The power difference is positively correlated with the semantic association between corresponding element pairs in the text sequence. The power difference is output as a correlation parameter.

12. The self-attention network system based on optical computing acceleration according to claim 11, characterized in that, The power difference is: in, This is the power difference. Encoding optical signals, To and Encoding of light signals of the same wavelength.

Citation Information

Patent Citations

  • Tensor point product acceleration architecture and method of self-attention mechanism

    CN117610682A

  • Optical computing device, computing method and computing system for attention mechanism

    CN118261220A