Intelligent optical computing on-chip element learning training method, framework and system
By introducing pre-trained non-reconstructible diffraction modules and adaptable reconstructible Machzendel interferometer arrays into the photoelectric computing system, the problem of frequent retraining of the optical computing system in different data domains is solved, and learning efficiency and generalization capabilities are improved.
Patent Information
- Application Number
- CN202510423251.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-04-07
AI Technical Summary
Existing optical computing systems require frequent retraining when facing different data domains, poor generalization ability and low learning efficiency.
A method of on-chip learning training for intelligent optical computing is proposed. By introducing a pre-trained non-reconstructible diffraction module and an adaptable reconstructible Machzendel interferometer array in the photoelectric computing system, combining the high data throughput capability of diffraction and the reconfigurability of the interference network.
It significantly reduces training costs, improves learning efficiency, can quickly adapt to various data fields, reduces parameter adjustments and shortens the overall training time.
Smart Images

Figure CN119940486A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of optical computing technology, and in particular to an on-chip meta-learning training method, architecture and system for intelligent optical computing. Background Art
[0002] With the rapid development of artificial intelligence and scientific computing, the complexity and scale of computing needs are also increasing. However, the existing electronic computing technology is limited by Moore's Law, and its performance is gradually approaching saturation, making it difficult to effectively cope with the increasingly stringent requirements of large-scale complex algorithms on computing power and power consumption. Light has natural advantages such as high throughput and low latency in the propagation process. Optical computing technology that uses photons instead of electrons as computing carriers is seen as the key to breaking the existing computing bottleneck. Summary of the invention
[0003] The present disclosure aims to solve one of the technical problems in the related art at least to some extent.
[0004] To this end, the first objective of the present disclosure is to propose an on-chip meta-learning training method for intelligent optical computing to improve the training efficiency of optoelectronic computing systems.
[0005] The second objective of the present disclosure is to provide an optoelectronic hybrid chip architecture.
[0006] The third objective of the present disclosure is to provide a photoelectric computing system.
[0007] To achieve the above-mentioned purpose, the first aspect of the present disclosure proposes an on-chip meta-learning training method for intelligent optical computing, which is applied to an optoelectronic computing system, wherein the optoelectronic computing system includes at least one optoelectronic hybrid chip architecture, wherein the optoelectronic hybrid chip architecture includes a diffraction module and a Mach-Zehnder interferometer array, and the method includes: Pre-training the diffraction neural network in the diffraction module based on the target data domain, and solidifying the pre-trained diffraction neural network to obtain a non-reconfigurable diffraction module; In response to receiving the target light computing task corresponding to the target data domain, acquiring a target data set corresponding to the target light computing task in the target data domain; The phase of each Mach-Zehnder interferometer in the Mach-Zehnder interferometer array is trained based on the target data set to obtain a trained Mach-Zehnder interferometer array, so that the optoelectronic hybrid chip architecture performs the target optical computing task according to the non-reconfigurable diffraction module and the trained Mach-Zehnder interferometer array.
[0008] Optionally, the training the phase of each Mach-Zehnder interferometer in the Mach-Zehnder interferometer array based on the target data set includes: Dividing the target data set into a support set and a query set; Initialize the phase of each Mach-Zehnder interferometer in the Mach-Zehnder interferometer array using the support set to obtain an initialized Mach-Zehnder interferometer array; The query set is used to train the phase of each Mach-Zehnder interferometer in the initialized Mach-Zehnder interferometer array.
[0009] Optionally, the training the phase of each Mach-Zehnder interferometer in the Mach-Zehnder interferometer array based on the target data set includes: Performing one-dimensional processing on the target data set to obtain a one-dimensional target data set; Inputting the one-dimensionalized target data set into the non-reconstructible diffraction module for feature extraction to obtain a target feature data set; The phase of each Mach-Zehnder interferometer in the Mach-Zehnder interferometer array is trained based on the target feature data set.
[0010] Optionally, the training the phase of each Mach-Zehnder interferometer in the Mach-Zehnder interferometer array based on the target data set includes: The phase of each Mach-Zehnder interferometer in the Mach-Zehnder interferometer array is trained on the target data set using a gradient descent algorithm.
[0011] Optionally, the training the phase of each Mach-Zehnder interferometer in the Mach-Zehnder interferometer array based on the target data set includes: Based on the target data set, the input voltage of each Mach-Zehnder interferometer in the Mach-Zehnder interferometer array is trained to achieve the training of the phase of each Mach-Zehnder interferometer in the Mach-Zehnder interferometer array.
[0012] Optionally, the solidifying the pre-trained diffractive neural network includes: The pre-trained diffractive neural network is solidified in the diffractive module by using photolithography technology.
[0013] Optionally, the optoelectronic computing system includes a plurality of optoelectronic hybrid chip architectures, and the method further includes: Connecting the diffraction neural networks in the plurality of optoelectronic hybrid chip architectures to obtain a gradient synthesis network; The gradient synthesis network is pre-trained based on the target data domain, and the pre-trained gradient synthesis network is solidified.
[0014] Optionally, the diffractive neural networks in the plurality of optoelectronic hybrid chip architectures are connected using at least one of the following connection methods: Series connection; Connect in parallel.
[0015] To achieve the above-mentioned purpose, a second embodiment of the present disclosure proposes an optoelectronic hybrid chip architecture, including: A diffraction module, used for solidifying a pre-trained diffraction neural network, and when receiving an optical input signal, controlling the pre-trained diffraction neural network to extract features from the optical input signal to obtain an optical feature signal, wherein the pre-trained diffraction neural network is obtained by pre-training the diffraction neural network based on a target data domain; The Mach-Zehnder interferometer array is used to perform training based on the target optical computing task corresponding to the target data domain to obtain a trained Mach-Zehnder interferometer array, and when receiving the optical characteristic signal, control the trained Mach-Zehnder interferometer array to perform optical computing on the optical characteristic signal to obtain an optical output signal.
[0016] To achieve the above-mentioned purpose, a third aspect of the present disclosure provides an optoelectronic computing system, comprising: at least one optoelectronic hybrid chip architecture as shown in the aforementioned second aspect.
[0017] In summary, the method, architecture and system provided by the present invention, by introducing a design concept of combining a pre-trained non-reconfigurable diffraction module with an adaptable reconfigurable MZI array, combines the high data throughput capability of diffraction and the reconfigurability of the interference network, combines the high data throughput capability of photons and the flexibility of electronics, provides hardware advantages and improves training efficiency, and can solve the problems of current optical computing systems that require frequent retraining when facing different data domains, have poor generalization capabilities and low learning efficiency.
[0018] Additional aspects and advantages of the present disclosure will be given in part in the following description and in part will be obvious from the following description or learned through practice of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The above and / or additional aspects and advantages of the present disclosure will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which: Figure 1 A training flow chart of an existing optical computing architecture provided by an embodiment of the present disclosure; Figure 2 A schematic diagram of a flow chart of a meta-learning training method on an intelligent optical computing chip provided by an embodiment of the present disclosure; Figure 3A schematic diagram of a flow chart of a meta-learning training method on an intelligent optical computing chip provided by another embodiment of the present disclosure; Figure 4 An experimental flow chart of an in-task training provided by an embodiment of the present disclosure; Figure 5 A schematic diagram of a process of MZI training provided by an embodiment of the present disclosure; Figure 6 A flowchart of a meta-learning training method on an intelligent optical computing chip provided by yet another embodiment of the present disclosure. DETAILED DESCRIPTION
[0020] Embodiments of the present disclosure are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present disclosure, and should not be construed as limiting the present disclosure.
[0021] In the field of optical computing, a series of research and development on optical neural network processors have been made public. These studies mainly focus on the technology of using photons instead of electrons for information processing, aiming to build a high-speed and energy-efficient artificial neural network system through the characteristics of optical propagation. For example, some studies have shown how to use integrated photonics technology to implement complex matrix multiplication operations, which is one of the basic operations for building deep neural networks; others have explored the implementation of different types of nonlinear activation functions in the optical domain to enhance the model's expressiveness.
[0022] However, at the current stage, most reported optical computing architectures face a common technical challenge: they usually require training the entire network structure from scratch for a specific task or dataset, such as Figure 1 As shown. This means that every time the application scenario is changed, the long process of data collection, model design and optimization must be repeated. This method is not only time-consuming and labor-intensive, but it is also difficult to ensure that the newly trained model can be well generalized to unseen data, especially when faced with a variety of practical problems. In addition, due to the lack of an effective transfer learning mechanism, the flexibility of photon-based artificial intelligence (AI) solutions in practical applications is greatly reduced, limiting the possibility of its wider adoption. Therefore, improving the adaptability of optical computing systems to different tasks and accelerating the learning speed has become one of the key issues that need to be addressed.
[0023] The present disclosure is described in detail below with reference to specific embodiments.
[0024] In the first embodiment, if Figure 2 As shown, Figure 2 A flow chart of an on-chip meta-learning training method for intelligent optical computing provided in one embodiment of the present disclosure is provided. The method can be applied to an optoelectronic computing system, wherein the optoelectronic computing system includes at least one optoelectronic hybrid chip architecture, wherein the optoelectronic hybrid chip architecture includes a diffraction module and a Mach–Zehnder interferometer (MZI) array.
[0025] For example, the meta-learning training method on the intelligent optical computing chip includes the following steps: S101, pre-training the diffraction neural network in the diffraction module based on the target data domain, and solidifying the pre-trained diffraction neural network to obtain a non-reconfigurable diffraction module; According to some embodiments, the target data domain refers to a data domain where the optical computing system is subsequently used to perform an optical computing task.
[0026] In some embodiments, the pre-trained diffractive neural network can be solidified in the diffractive module using photolithography technology.
[0027] It should be noted that after solidifying the pre-trained diffractive neural network, the parameters in the obtained non-reconfigurable diffractive module cannot be changed again and are fixed parameters.
[0028] The diffraction module may also be referred to as a data processing unit (DPU).
[0029] Take a scenario as an example. Figure 3 This is a flow chart of a meta-learning training method for an intelligent optical computing chip provided by another embodiment of the present disclosure. Figure 3 As shown, a 4-layer diffractive neural network can be pre-trained on the first four categories of the MNIST and Fashion-MNIST datasets and solidified on the DPU, where the input image is resized to 4×4 and flattened into a vector to match the scale and shape of the input modulation.
[0030] S102, in response to receiving a target light computing task corresponding to a target data domain, obtaining a target data set corresponding to the target light computing task in the target data domain; According to some embodiments, the target optical computing task refers to an optical computing task that the optoelectronic hybrid chip architecture needs to execute.
[0031] S103, training the phase of each Mach-Zehnder interferometer in the Mach-Zehnder interferometer array based on the target data set to obtain a trained Mach-Zehnder interferometer array, so that the optoelectronic hybrid chip architecture performs the target optical computing task according to the non-reconfigurable diffraction module and the trained Mach-Zehnder interferometer array.
[0032] According to some embodiments, a gradient descent algorithm may be used to train the phase of each Mach-Zehnder interferometer in the Mach-Zehnder interferometer array on a target data set, such as Figure 3 shown.
[0033] In some embodiments, since the phase of a Mach-Zehnder interferometer is related to its input voltage, the input voltage of each Mach-Zehnder interferometer in the Mach-Zehnder interferometer array can be trained based on the target data set to train the phase of each Mach-Zehnder interferometer in the Mach-Zehnder interferometer array.
[0034] It should be noted that by using photolithography technology to solidify the diffraction module parameters of the DPU and configuring the MZI phase through voltage, high-speed neural network configuration and video rate inference capabilities can be achieved.
[0035] According to some embodiments, Figure 1 As shown, the cross-task training used in the prior art focuses on optimizing the fixed diffraction backbone, while the present disclosure has pre-trained the diffraction neural network in the diffraction module. Therefore, only the reconfigurable MZI array needs to be trained within the task subsequently. Compared with training the entire network from scratch, fewer parameter adjustments and shorter training time are required to achieve high performance in the target data domain, which significantly reduces the training cost and improves the learning efficiency. In-task training can be achieved using only small sample learning, which can not only quickly adapt to various data fields under the condition of a small number of samples, but also significantly reduce the required parameter adjustment amount and shorten the overall training time, thereby overcoming one of the main obstacles encountered by traditional optical neural networks in practical applications.
[0036] In some embodiments, in the process of training the phase of each Mach-Zehnder interferometer in the Mach-Zehnder interferometer array based on the target data set, the target data set can be divided into a support set and a query set; the support set is used to initialize the phase of each Mach-Zehnder interferometer in the Mach-Zehnder interferometer array to obtain the initialized Mach-Zehnder interferometer array; the query set is used to train the phase of each Mach-Zehnder interferometer in the initialized Mach-Zehnder interferometer array. Therefore, it is possible to quickly adapt to new data domains under limited data samples, and is particularly suitable for high-performance device deployment that needs to learn through computationally efficient local training data, and has a wide range of application scenarios.
[0037] In some embodiments, the ability of the optoelectronic computing system to quickly adapt the model to different data domains can be verified through a small number of sample learning experiments. Figure 4 An experimental flow chart of in-task training provided in an embodiment of the present disclosure. Figure 4 As shown, 5-way 1-shot and 20-way 1-shot experiments were conducted on the Omniglot dataset, and few-shot learning experiments were performed using the CIFAR-FS and Mini-ImageNet datasets. The adjustable classification weights were experimentally demonstrated by repeatedly using the MZI array. The experimental results showed its high efficiency in few-shot learning tasks and demonstrated efficient learning capabilities.
[0038] According to some embodiments, Figure 5 A schematic diagram of a MZI training process provided by an embodiment of the present disclosure. Figure 5 As shown, in the process of training the phase of each Mach-Zehnder interferometer in the Mach-Zehnder interferometer array based on the target data set, the target data set can be processed into one dimension to obtain a one-dimensional target data set; the one-dimensional target data set is input into the non-reconfigurable diffraction module for feature extraction to obtain a target feature data set; and the phase of each Mach-Zehnder interferometer in the Mach-Zehnder interferometer array is trained based on the target feature data set. This process is applicable to the process of initializing the phase of each Mach-Zehnder interferometer in the Mach-Zehnder interferometer array using the support set and training the phase of each Mach-Zehnder interferometer in the initialized Mach-Zehnder interferometer array using the query set.
[0039] In summary, the method provided in this embodiment, by introducing a design idea of combining a pre-trained non-reconfigurable diffraction module with an adaptable reconfigurable MZI array, combines the high data throughput of diffraction and the reconfigurability of the interference network, combines the high data throughput of photons and the flexibility of electronics, provides hardware advantages and improves training efficiency, and can solve the problems of frequent retraining, poor generalization ability and low learning efficiency of current optical computing systems when facing different data domains. Secondly, by adjusting the diffraction module parameters and MZI phase, the optoelectronic computing system can also be optimized to achieve higher computing speed and lower energy consumption.
[0040] This embodiment also provides another on-chip meta-learning training method for intelligent optical computing, which can be applied to an optoelectronic computing system including multiple optoelectronic hybrid chip architectures.
[0041] For example, the meta-learning training method on the intelligent optical computing chip may include the following steps: S201, connecting the diffraction neural networks in the multiple optoelectronic hybrid chip architectures to obtain a gradient synthesis network; In some embodiments, the diffractive neural networks in multiple optoelectronic hybrid chip architectures may be connected using at least one of the following connection methods: Series connection; Connect in parallel.
[0042] It should be noted that by connecting the diffractive neural networks in multiple optoelectronic hybrid chip architectures, multiple optoelectronic hybrid chip architectures can be combined into a large feature extraction backbone.
[0043] S202, pre-training the gradient synthesis network based on the target data domain, and solidifying the pre-trained gradient synthesis network.
[0044] It should be noted that by using synthetic gradient networks instead of electronic components for back-propagation, this method allows faster back-propagation and supports more complex data set processing, which can improve the versatility and adaptability of on-chip meta-learning, thereby expanding the application scope of optoelectronic computing systems.
[0045] Take a scenario as an example. Figure 6 A schematic diagram of a process flow of a meta-learning training method on an intelligent optical computing chip provided by another embodiment of the present disclosure. Figure 6 As shown, the diffraction neural networks in multiple optoelectronic hybrid chip architectures are connected in series and in parallel to obtain a gradient synthesis network. In addition, the classification parameters corresponding to the MZI array are initialized using the support set and optimized through inductive learning on the query set.
[0046] In summary, the method provided in this embodiment can effectively improve the performance and flexibility of the optoelectronic computing system through the optimized synthetic gradient network design, while reducing the training cost, meeting the needs of different application scenarios, and further enhancing the model's support capabilities for complex pattern recognition tasks, making the present invention an important step in promoting the development of optical computing technology.
[0047] In order to implement the above embodiments, the present disclosure also proposes an optoelectronic hybrid chip architecture.
[0048] For example, the optoelectronic hybrid chip architecture includes: A diffraction module is used to solidify the pre-trained diffraction neural network, and when receiving an optical input signal, control the pre-trained diffraction neural network to extract features of the optical input signal to obtain an optical feature signal, wherein the pre-trained diffraction neural network is obtained by pre-training the diffraction neural network based on a target data domain; The Mach-Zehnder interferometer array is used to perform training based on a target optical computing task corresponding to a target data domain to obtain a trained Mach-Zehnder interferometer array, and when receiving an optical characteristic signal, control the trained Mach-Zehnder interferometer array to perform optical computing on the optical characteristic signal to obtain an optical output signal.
[0049] According to some embodiments, the electrical layer in the optical-electrical hybrid chip architecture is located between cells and can be used for data rearrangement and / or nonlinear computing.
[0050] It should be noted that the aforementioned explanation of the embodiment of the meta-learning training method on the intelligent optical computing chip is also applicable to the optoelectronic hybrid chip architecture of this embodiment and will not be repeated here.
[0051] In summary, the optoelectronic hybrid chip architecture provided in the embodiments of the present disclosure integrates a pre-trained non-reconfigurable diffraction module and an adaptable reconfigurable MZI array by implementing a hybrid structure. It can fully utilize the efficient transmission characteristics of photons and has good adaptability and flexibility. It not only improves the computing efficiency of the optoelectronic hybrid chip architecture, but also enhances its adaptability in different application scenarios.
[0052] In order to implement the above embodiments, the present disclosure further proposes an optoelectronic computing system, including: an optoelectronic hybrid chip architecture provided by at least one of the above embodiments.
[0053] The collection, storage, use, processing, transmission, provision and disclosure of user personal information involved in this disclosure shall comply with the relevant laws and regulations and shall not violate public order and good morals.
[0054] It should be noted that personal information from users should be collected for legitimate and reasonable purposes and should not be shared or sold outside of these legitimate uses. In addition, such collection / sharing should be carried out after receiving the user's informed consent, including but not limited to notifying the user to read the user agreement / user notice and sign the agreement / authorization including authorization of relevant user information before the user uses the function. In addition, any necessary steps should be taken to protect and safeguard access to such personal information data and ensure that others who have access to personal information data comply with its privacy policy and procedures.
[0055] The present disclosure anticipates providing implementation schemes for users to selectively block the use or access of personal information data. That is, the present disclosure anticipates providing hardware and / or software to prevent or block access to such personal information data. Once the personal information data is no longer needed, risks can be minimized by limiting data collection and deleting the data. In addition, when applicable, such personal information is de-identified to protect the privacy of the user.
[0056] In the description of the aforementioned embodiments, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present disclosure. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, unless they contradict each other.
[0057] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of the features. In the description of the present disclosure, "plurality" means at least two, such as two, three, etc., unless otherwise clearly and specifically defined.
[0058] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment or portion of code that includes one or more executable instructions for implementing the steps of a custom logical function or process, and the scope of the preferred embodiments of the present disclosure includes additional implementations in which functions may not be performed in the order shown or discussed, including performing functions in a substantially simultaneous manner or in reverse order depending on the functions involved, which should be understood by technicians in the technical field to which the embodiments of the present disclosure belong.
[0059] The logic and / or steps represented in the flowchart or otherwise described herein, for example, may be considered as an ordered list of executable instructions for implementing logical functions, and may be embodied in any computer-readable medium for use by an instruction execution system, apparatus or device (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, apparatus or device and execute instructions), or in conjunction with such instruction execution system, apparatus or device. For purposes of this specification, "computer-readable medium" may be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, apparatus or device, or in conjunction with such instruction execution system, apparatus or device. More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection with one or more wires (electronic devices), a portable computer disk case (magnetic devices), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM) or flash memory, an optical fiber device, and a portable compact disc read-only memory (CDROM). In addition, the computer-readable medium may even be a paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, deciphering or, if necessary, processing in another suitable manner, and then stored in a computer memory.
[0060] It should be understood that the various parts of the present disclosure can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0061] A person skilled in the art may understand that all or part of the steps in the above-mentioned embodiment method may be completed by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, which, when executed, includes one or a combination of the steps of the method embodiment.
[0062] In addition, each functional unit in each embodiment of the present disclosure may be integrated into a processing module, or each unit may exist physically separately, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.
[0063] The storage medium mentioned above may be a read-only memory, a disk or an optical disk, etc. Although the embodiments of the present disclosure have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limitations of the present disclosure. A person of ordinary skill in the art may change, modify, replace and modify the above embodiments within the scope of the present disclosure.
Claims
1. A meta-learning training method for intelligent optical computing on-chip, characterized in that: Applied to an optoelectronic computing system, the optoelectronic computing system includes at least one optoelectronic hybrid chip architecture, the optoelectronic hybrid chip architecture includes a diffraction module and a Mach-Zehnder interferometer array, the method includes: Pre-training the diffraction neural network in the diffraction module based on the target data domain, and solidifying the pre-trained diffraction neural network to obtain a non-reconfigurable diffraction module; In response to receiving the target light computing task corresponding to the target data domain, acquiring a target data set corresponding to the target light computing task in the target data domain; The phase of each Mach-Zehnder interferometer in the Mach-Zehnder interferometer array is trained based on the target data set to obtain a trained Mach-Zehnder interferometer array, so that the optoelectronic hybrid chip architecture performs the target optical computing task according to the non-reconfigurable diffraction module and the trained Mach-Zehnder interferometer array.
2. The method according to claim 1, characterized in that The training of the phase of each Mach-Zehnder interferometer in the Mach-Zehnder interferometer array based on the target data set includes: Dividing the target data set into a support set and a query set; Initialize the phase of each Mach-Zehnder interferometer in the Mach-Zehnder interferometer array using the support set to obtain an initialized Mach-Zehnder interferometer array; The query set is used to train the phase of each Mach-Zehnder interferometer in the initialized Mach-Zehnder interferometer array.
3. The method according to claim 1, characterized in that: The training of the phase of each Mach-Zehnder interferometer in the Mach-Zehnder interferometer array based on the target data set includes: Performing one-dimensional processing on the target data set to obtain a one-dimensional target data set; Inputting the one-dimensionalized target data set into the non-reconstructible diffraction module for feature extraction to obtain a target feature data set; The phase of each Mach-Zehnder interferometer in the Mach-Zehnder interferometer array is trained based on the target feature data set.
4. The method according to claim 1, characterized in that: The training of the phase of each Mach-Zehnder interferometer in the Mach-Zehnder interferometer array based on the target data set includes: The phase of each Mach-Zehnder interferometer in the Mach-Zehnder interferometer array is trained on the target data set using a gradient descent algorithm.
5. The method according to claim 1, characterized in that The training of the phase of each Mach-Zehnder interferometer in the Mach-Zehnder interferometer array based on the target data set includes: Based on the target data set, the input voltage of each Mach-Zehnder interferometer in the Mach-Zehnder interferometer array is trained to achieve the training of the phase of each Mach-Zehnder interferometer in the Mach-Zehnder interferometer array.
6. The method according to claim 1, characterized in that The solidified pre-trained diffractive neural network comprises: The pre-trained diffractive neural network is solidified in the diffractive module by using photolithography technology.
7. The method according to claim 1, characterized in that The optoelectronic computing system includes a plurality of optoelectronic hybrid chip architectures, and the method further includes: Connecting the diffraction neural networks in the plurality of optoelectronic hybrid chip architectures to obtain a gradient synthesis network; The gradient synthesis network is pre-trained based on the target data domain, and the pre-trained gradient synthesis network is solidified.
8. The method according to claim 7, characterized in that The diffractive neural networks in the plurality of optoelectronic hybrid chip architectures are connected using at least one of the following connection methods: Series connection; Connect in parallel.
9. An optoelectronic hybrid chip architecture, characterized in that: include: A diffraction module, used for solidifying a pre-trained diffraction neural network, and when receiving an optical input signal, controlling the pre-trained diffraction neural network to extract features from the optical input signal to obtain an optical feature signal, wherein the pre-trained diffraction neural network is obtained by pre-training the diffraction neural network based on a target data domain; The Mach-Zehnder interferometer array is used to perform training based on the target optical computing task corresponding to the target data domain to obtain a trained Mach-Zehnder interferometer array, and when receiving the optical characteristic signal, control the trained Mach-Zehnder interferometer array to perform optical computing on the optical characteristic signal to obtain an optical output signal.
10. An optoelectronic computing system, characterized in that: include: At least one optoelectronic hybrid chip architecture as claimed in claim 9.
Citation Information
Patent Citations
Reconfigurable nonvolatile integrated three-dimensional optical diffraction neural network chip
CN115018041A
Optical computing device and method
CN115481361A
Large-scale distributed photoelectric intelligent computing architecture and chip system
CN117349225A
Robust optical neural network training method and apparatus, electronic device, and medium
WO2024198504A1