Intelligent optical computing attention mechanism model, method and system

Through the intelligent optical computing attention mechanism model, the matrix operation and weighted calculation are performed using the optical field intermodulation propagation sub-model, which solves the high power consumption and delay problems of electronic computing architecture in large-scale attention mechanism operations, and realizes efficient data processing.

CN119886212BActive Publication Date: 2025-09-02TSINGHUA UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510379734.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-09-02
Estimated Expiration
2045-03-28

AI Technical Summary

Technical Problem

Existing electronic computing architectures face problems of high power consumption and computational delay when dealing with large-scale attention mechanism operations.

Method used

The intelligent optical computing attention mechanism model is adopted, and the light field intermodulation propagation sub-model is used for matrix operations and weighted calculations, and high-speed parallel processing is performed through photons to realize the calculation in the attention mechanism.

Benefits of technology

The inference process of the model is greatly accelerated, energy consumption is reduced, and the ability to process large-scale data is improved, which solves the problems of high power consumption and computing latency of existing electronic computing architectures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119886212B_ABST
    Figure CN119886212B_ABST
Patent Text Reader

Abstract

The present disclosure relates to the field of optical computing technology, and in particular to an intelligent optical computing attention mechanism model, method, and system. The model includes multiple light field intermodulation propagation sub-models. In the light field intermodulation propagation sub-model, the spatial computing module encodes the acquired first sub-vector and second sub-vector into the phase and / or amplitude of the spatial light field, and calculates the first sub-vector and second sub-vector through optical propagation to obtain the target first sub-vector; the temporal computing module maps the acquired first sub-vector and second sub-vector to the phase and / or amplitude of the time-series optical signal, and calculates the first sub-vector and second sub-vector through optical propagation to obtain the target second sub-vector; and the output module performs a first fusion of the target first sub-vector of the spatial computing module and the target second sub-vector of the temporal computing module to obtain the target sub-vector. The present disclosure utilizes the high-speed parallel processing advantages of photonic computing to greatly accelerate the inference process of the model and reduce energy consumption.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of optical computing technology, and in particular to an intelligent optical computing attention mechanism model, method, and system. Background Art

[0002] With the development of deep learning and large models, the computational complexity and scale of the model training process are also increasing. Among them, attention mechanism models are key components for improving model performance, and are widely used in the Transformer architecture. However, existing electronic computing architectures may face high power consumption and computational latency when processing large-scale attention mechanism operations. Optical computing, which uses photons for information transmission and processing, offers high bandwidth, low latency, and parallel processing capabilities. Based on this, optical computing technology, which uses photons instead of electrons as computing carriers, is seen as the key to breaking through existing computing bottlenecks. Summary of the Invention

[0003] The present disclosure aims to solve one of the technical problems in the related art at least to a certain extent.

[0004] To this end, the first purpose of the present disclosure is to propose an intelligent optical computing attention mechanism model. Through the light field intermodulation propagator model, the high-speed parallel processing advantages of photonic computing are utilized to realize the matrix operations and weighted calculations in the attention mechanism, greatly accelerate the reasoning process of the model, reduce energy consumption, and improve the ability to process large-scale data.

[0005] The second objective of the present disclosure is to provide a question-answering processing method.

[0006] To achieve the above objectives, the first embodiment of the present disclosure proposes an intelligent optical computing attention mechanism model, which embeds the attention mechanism model into a large model. The attention mechanism model includes multiple light field intermodulation propagation sub-models, wherein each light field intermodulation propagation sub-model is applied to each sub-head in each phrase window, the phrase window is a window divided by word vectors, and the sub-head is a sub-head divided by feature vectors in each window. The light field intermodulation propagation sub-model includes a spatial calculation module, a temporal calculation module and an output module, wherein,

[0007] The spatial calculation module encodes the acquired first sub-vector and second sub-vector into the phase and / or amplitude of the spatial light field, and calculates the first sub-vector and the second sub-vector through light propagation to obtain a target first sub-vector;

[0008] The time calculation module maps the acquired first sub-vector and second sub-vector to the phase and / or amplitude of the time-series optical signal, and calculates the first sub-vector and the second sub-vector through optical propagation to obtain a target second sub-vector;

[0009] The output module performs a first fusion on the target first subvector of the spatial calculation module and the target second subvector of the temporal calculation module to obtain a target subvector.

[0010] Optionally, after the attention mechanism model obtains the target sub-vector in each light field intermodulation propagation sub-model, the target sub-vector in each light field intermodulation propagation sub-model is secondly fused to obtain a target vector, wherein the target vector is the output of the attention mechanism model.

[0011] Optionally, parallel computing of the multiple spatial computing modules is achieved through spatial multiplexing (SMUX).

[0012] Optionally, wavelength multiplexing (WMUX) is used to implement parallel calculation of the multiple time calculation modules.

[0013] Optionally, the intelligent optical computing attention mechanism model is used for question-answering processing.

[0014] To achieve the above objectives, a second embodiment of the present disclosure proposes a question-answering processing method, including:

[0015] Get the data to answer the questions you need;

[0016] Input the problem data into the large model, and determine the first vector and the second vector of the input attention mechanism model in the large model;

[0017] The first vector and the second vector are calculated by the attention mechanism model to obtain a target vector, and the target vector is transmitted to a subsequent model for calculation to obtain a target answer to the question data.

[0018] Optionally, the attention mechanism model includes multiple light field intermodulation propagation sub-models; and calculating the first vector and the second vector using the attention mechanism model to obtain a target vector includes:

[0019] The attention mechanism model divides the first vector and the second vector into multiple first sub-vectors and multiple second sub-vectors;

[0020] Inputting the first sub-vector and the second sub-vector into corresponding light field intermodulation propagation sub-models to obtain a target sub-vector of each light field intermodulation propagation sub-model;

[0021] The target sub-vectors of each light field intermodulation propagation sub-model are fused to obtain a target vector.

[0022] Optionally, the light field intermodulation propagation sub-model includes a spatial calculation module, a temporal calculation module, and an output module; and inputting the first sub-vector and the second sub-vector into the corresponding light field intermodulation propagation sub-model to obtain a target sub-vector for each light field intermodulation propagation sub-model includes:

[0023] encoding the first sub-vector and the second sub-vector into the phase and / or amplitude of the spatial light field by the spatial calculation module, and calculating the first sub-vector and the second sub-vector by light propagation to obtain a target first sub-vector;

[0024] Mapping the first sub-vector and the second sub-vector to the phase and / or amplitude of the time-series optical signal by the time calculation module, and calculating the first sub-vector and the second sub-vector by light propagation to obtain a target second sub-vector;

[0025] The output module fuses the target first subvector of the spatial calculation module and the target second subvector of the temporal calculation module to obtain a target subvector.

[0026] Optionally, fusing the target subvectors of each light field intermodulation propagation submodel to obtain the target vector includes: performing weighted fusion on the target subvectors of each light field intermodulation propagation submodel to obtain the target vector.

[0027] Another object of the present invention is to provide an electronic device, comprising:

[0028] at least one processor; and

[0029] a memory communicatively connected to the at least one processor; wherein,

[0030] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the methods described in the above two aspects.

[0031] Another object of the present invention is to provide a computer storage medium, wherein the computer storage medium stores computer-executable instructions; after the computer-executable instructions are executed by a processor, the computer executes the methods described in the above two aspects.

[0032] In summary, the intelligent optical computing attention mechanism model, method and system provided by the present disclosure, through the light field intermodulation propagator model, utilizes the high-speed parallel processing advantages of photonic computing to realize the matrix operations and weighted calculations in the attention mechanism, greatly accelerates the model's reasoning process, reduces energy consumption, and improves the ability to process large-scale data.

[0033] Additional aspects and advantages of the present disclosure will be given in part in the following description and in part will be obvious from the following description, or will be learned through practice of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] The above and / or additional aspects and advantages of the present disclosure will become apparent and readily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:

[0035] Figure 1 A schematic diagram of the structure of an intelligent optical computing attention mechanism model provided by an embodiment of the present disclosure;

[0036] Figure 2 This is a schematic diagram of the structure of a multi-head multi-window distributed architecture proposed in an embodiment of the present disclosure;

[0037] Figure 3 A flowchart of a question-answering processing method proposed in an embodiment of the present disclosure;

[0038] Figure 4 A structural diagram of a question-answering processing system provided in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0039] The following describes in detail embodiments of the present disclosure, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present disclosure, and should not be construed as limiting the present disclosure.

[0040] The present disclosure is described in detail below with reference to specific embodiments.

[0041] Figure 1 This is a schematic diagram of the structure of an intelligent optical computing attention mechanism model provided by an embodiment of the present disclosure. Figure 1 As shown, the attention mechanism model includes multiple light field intermodulation propagation sub-models, wherein the light field intermodulation propagation sub-model includes a spatial calculation module 101, a time calculation module 102 and an output module 103, wherein,

[0042] The spatial calculation module 101 encodes the obtained first sub-vector and second sub-vector into the phase and / or amplitude of the spatial light field, and calculates the first sub-vector and second sub-vector through light propagation to obtain a target first sub-vector;

[0043] The time calculation module 102 maps the obtained first sub-vector and second sub-vector to the phase and / or amplitude of the time-series optical signal, and calculates the first sub-vector and second sub-vector through optical propagation to obtain a target second sub-vector;

[0044] The output module 103 performs a first fusion on the target first subvector of the spatial calculation module and the target second subvector of the temporal calculation module to obtain a target subvector.

[0045] In one embodiment of the present disclosure, the above-mentioned large model can be applied to a variety of scenarios, such as intelligent question-answering processing scenarios.

[0046] In one embodiment of the present disclosure, the aforementioned light field intermodulation propagation model can perform continuous optical computations in spatial and temporal dimensions. To address the inherent dimensional mismatch between spatial and temporal light fields, a spatial multiplexing and spectral multiplexing computational model is proposed to match the spatial and temporal optical computational dimensions. All spatial and temporal computational operations are performed in the optical simulation domain. This allows the speed of spatial and temporal optical computation to be unconstrained by the transmission and read / write capabilities of electronic memory, enabling high-speed reasoning for complex visual intelligence tasks.

[0047] In one embodiment of the present disclosure, the light field intermodulation propagation submodel includes a spatial computation module, a temporal computation module, and an output module. The spatial computation module encodes the acquired first and second subvectors into the phase and / or amplitude of the spatial light field and calculates the first and second subvectors via optical propagation to obtain a target first subvector. The spatial computation module can perform numerous operations at the speed of light and fully extract spatial information from each time frame. Furthermore, the temporal computation module maps the acquired first and second subvectors into the phase and / or amplitude of a time-series optical signal and calculates the first and second subvectors via optical propagation to obtain a target second subvector.

[0048] Among them, in one embodiment of the present disclosure, parallel computing of multiple spatial computing modules can be achieved through spatial multiplexing SMUX and parallel computing of multiple time computing modules can be achieved through wavelength multiplexing WMUX to solve the problem of most information loss caused by coupling spatial content to a single-mode timing channel.

[0049] Furthermore, in one embodiment of the present disclosure, after the attention mechanism model obtains the target sub-vector in each light field intermodulation propagation sub-model, the target sub-vector in each light field intermodulation propagation sub-model can be secondly fused to obtain a target vector, wherein the target vector is the output of the attention mechanism model.

[0050] Furthermore, in one embodiment of the present disclosure, each of the aforementioned light field intermodulation propagation sub-models is applied to each sub-head within each phrase window, where a phrase window is a window divided by a word vector, and a sub-head is a sub-head divided by a feature vector within each window. Specifically, in one embodiment of the present disclosure, through a multi-head, multi-window distributed architecture, large feature dimension network operations and large matrix operations are implemented using parallel optical computing units, achieving high-efficiency and high-speed large-model inference and training.

[0051] In one embodiment of the present disclosure, Figure 2 This is a structural diagram of a multi-head multi-window distributed architecture proposed in an embodiment of the present disclosure. Figure 2 As shown, the multi-head, multi-window distributed architecture can be composed of two parts: a distributed multi-window computing architecture and a distributed multi-head computing architecture. Specifically, the input of a large model network consisting of a certain number of word vectors can be evenly decomposed into multiple phrase windows along the word vector dimension, resulting in a distributed multi-window computing architecture. Furthermore, a self-attention mechanism is applied to each phrase window, collecting correlation information between word vectors within the window. Simultaneously, information exchange between windows is leveraged to achieve global radiation of the self-attention mechanism, building contextual connections.

[0052] Furthermore, in one embodiment of the present disclosure, in each of the above-mentioned windows, the features can be evenly decomposed into several sub-heads containing partial information along the word feature dimension, thereby obtaining a distributed multi-head computing architecture. Specifically, in one embodiment of the present disclosure, after obtaining the multi-head, multi-window distributed architecture through the above steps, the above-mentioned light field intermodulation propagation sub-model can be applied to each sub-head in each phrase window, thereby utilizing efficient parallel optical computing to implement a feature processing network with large feature dimensions, supporting high-speed computation of feature maps in the Transformer unit.

[0053] The intelligent optical computing attention mechanism model of the disclosed embodiment, through the light field intermodulation propagator model, utilizes the high-speed parallel processing advantages of photonic computing to realize the matrix operations and weighted calculations in the attention mechanism, greatly accelerates the reasoning process of the model, reduces energy consumption, and improves the ability to process large-scale data.

[0054] In order to implement the above embodiment, Figure 3 The present disclosure also proposes a question-answering processing method, which may include the following steps:

[0055] Step 301, obtaining question data that needs to be answered;

[0056] Step 302: Input the question data into the large model and determine the first vector and the second vector of the input attention mechanism model in the large model;

[0057] In step 303, the first vector and the second vector are calculated by the attention mechanism model to obtain a target vector, and the target vector is transmitted to a subsequent model for calculation to obtain a target answer to the question data.

[0058] In one embodiment of the present disclosure, an attention mechanism model includes multiple light field intermodulation propagation sub-models; and a method for calculating a first vector and a second vector using the attention mechanism model to obtain a target vector may include the following steps:

[0059] Step 1: The attention mechanism model divides the first vector and the second vector into multiple first sub-vectors and multiple second sub-vectors;

[0060] Step 2: Input the first sub-vector and the second sub-vector into the corresponding light field intermodulation propagation sub-model to obtain the target sub-vector of each light field intermodulation propagation sub-model;

[0061] Step 3: Fuse the target sub-vectors of each light field intermodulation propagation sub-model to obtain the target vector.

[0062] In one embodiment of the present disclosure, the light field intermodulation propagation sub-model includes a spatial calculation module, a temporal calculation module, and an output module. The method of inputting the first sub-vector and the second sub-vector into the corresponding light field intermodulation propagation sub-model to obtain the target sub-vector of each light field intermodulation propagation sub-model may include the following steps:

[0063] Step a: encoding the first sub-vector and the second sub-vector into the phase and / or amplitude of the spatial light field through a spatial calculation module, and calculating the first sub-vector and the second sub-vector through light propagation to obtain a target first sub-vector;

[0064] Step b: mapping the first subvector and the second subvector to the phase and / or amplitude of the time-series optical signal through a time calculation module, and calculating the first subvector and the second subvector through optical propagation to obtain a target second subvector;

[0065] Step c: fusing the target first subvector of the spatial calculation module and the target second subvector of the temporal calculation module through the output module to obtain a target subvector.

[0066] In one embodiment of the present disclosure, the method for fusing the target first subvector of the spatial calculation module and the target second subvector of the temporal calculation module via the output module to obtain the target subvector may include: performing a weighted fusion of the target first subvector of the spatial calculation module and the target second subvector of the temporal calculation module via the output module to obtain the target subvector. In one embodiment of the present disclosure, the weights in the weighted fusion may be set based on manual experience.

[0067] In the embodiment of the present disclosure, the above-mentioned question-answering processing method, through multiple light field intermodulation propagation sub-models, utilizes the high-speed parallel processing advantages of photonic computing to realize matrix operations and weighted calculations in the attention mechanism, greatly accelerates the reasoning process of the model, reduces energy consumption, and improves the ability to process large-scale data.

[0068] In order to implement the above embodiment, Figure 4 The present disclosure further proposes a question-answering processing system, which may include:

[0069] Acquisition module 401, used to obtain question data that needs to be answered;

[0070] A determination module 402 is configured to input the question data into the large model and determine a first vector and a second vector of the input attention mechanism model in the large model;

[0071] The calculation module 403 is used to calculate the first vector and the second vector through the attention mechanism model to obtain a target vector, and transmit the target vector to a subsequent model for calculation to obtain a target answer to the question data.

[0072] In one embodiment of the present disclosure, the attention mechanism model includes multiple light field intermodulation propagation sub-models; the above-mentioned calculation module is specifically used to:

[0073] The attention mechanism model divides the first vector and the second vector into multiple first sub-vectors and multiple second sub-vectors;

[0074] Inputting the first sub-vector and the second sub-vector into the corresponding light field intermodulation propagation sub-model to obtain a target sub-vector of each light field intermodulation propagation sub-model;

[0075] The target sub-vectors of each light field intermodulation propagation sub-model are fused to obtain the target vector.

[0076] In one embodiment of the present disclosure, the light field intermodulation propagation sub-model includes a spatial calculation module, a temporal calculation module, and an output module; the above-mentioned calculation modules are specifically used to:

[0077] The first sub-vector and the second sub-vector are encoded into the phase and / or amplitude of the spatial light field by a spatial calculation module, and the first sub-vector and the second sub-vector are calculated by light propagation to obtain a target first sub-vector;

[0078] Mapping the first sub-vector and the second sub-vector to the phase and / or amplitude of the time-series optical signal through a time calculation module, and calculating the first sub-vector and the second sub-vector through optical propagation to obtain a target second sub-vector;

[0079] The target subvector is obtained by fusing the target first subvector of the spatial calculation module and the target second subvector of the temporal calculation module through the output module.

[0080] In one embodiment of the present disclosure, the calculation module is further configured to perform weighted fusion of the target first subvector of the spatial calculation module and the target second subvector of the temporal calculation module through the output module to obtain a target subvector.

[0081] In the disclosed embodiment, the question-answering processing system uses multiple light field intermodulation propagation sub-models and the high-speed parallel processing advantages of photonic computing to implement matrix operations and weighted calculations in the attention mechanism, greatly accelerating the model's reasoning process, reducing energy consumption, and improving the ability to process large-scale data.

[0082] The collection, storage, use, processing, transmission, provision and disclosure of user personal information involved in this disclosure are in compliance with relevant laws and regulations and do not violate public order and good morals.

[0083] It is important to note that personal information collected from users should be used for legitimate and reasonable purposes and should not be shared or sold beyond these legitimate uses. Furthermore, such collection / sharing should be conducted only after receiving the user's informed consent, including but not limited to notifying the user to read the user agreement / user notice and sign an agreement / authorization that includes the relevant user information before using the feature. Furthermore, any necessary steps must be taken to safeguard and secure access to such personal information and ensure that others with access to personal information comply with its privacy policy and procedures.

[0084] This disclosure contemplates providing implementations that allow users to selectively block the use or access of personal information data. Specifically, this disclosure contemplates providing hardware and / or software to prevent or block access to such personal information data. Risks can be minimized by limiting data collection and deleting data once it is no longer needed. Furthermore, where applicable, such personal information can be de-identified to protect user privacy.

[0085] The acquisition, transmission, storage, use, and processing of data in the technical solution disclosed herein are in compliance with the relevant provisions of national laws and regulations.

[0086] It should be noted that in the embodiments of the present disclosure, certain software, components, models and other existing solutions in the industry may be mentioned. They should be regarded as exemplary and their purpose is only to illustrate the feasibility of implementing the technical solution of this application, but it does not mean that the applicant has or will necessarily use the solution.

[0087] In the descriptions of the aforementioned embodiments, the reference terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" mean that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present disclosure. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and features of different embodiments or examples without contradiction.

[0088] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one such feature. Throughout the present disclosure, "plurality" means at least two, such as two, three, etc., unless otherwise specifically defined.

[0089] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment or portion of code comprising one or more executable instructions for implementing the steps of a custom logical function or process, and the scope of the preferred embodiments of the present disclosure includes additional implementations in which functions may be performed out of the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present disclosure belong.

[0090] The logic and / or steps represented in a flowchart or otherwise described herein, for example, can be considered a sequenced list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" is any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (not exhaustive) of computer-readable media include: an electrical connection with one or more wires (electronic devices), a portable computer disk cartridge (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and programmable read-only memory (EPROM or flash memory), fiber optic devices, and a portable compact disc read-only memory (CDROM). Furthermore, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting or otherwise processing it in a suitable manner if necessary, and then storing it in a computer memory.

[0091] It should be understood that various parts of the present disclosure can be implemented using hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used: a discrete logic circuit having logic gate circuits for implementing logical functions on data signals, an application-specific integrated circuit having suitable combinational logic gate circuits, a programmable gate array (PGA), a field-programmable gate array (FPGA), etc.

[0092] Those skilled in the art will appreciate that all or part of the steps in the method for implementing the above-mentioned embodiment can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.

[0093] In addition, the functional units in the various embodiments of the present disclosure may be integrated into a single processing module, each unit may exist physically separately, or two or more units may be integrated into a single module. The aforementioned integrated modules may be implemented in the form of hardware or in the form of software functional modules. If the integrated modules are implemented in the form of software functional modules and sold or used as independent products, they may also be stored in a computer-readable storage medium.

[0094] The storage medium mentioned above may be a read-only memory, a magnetic disk, or an optical disk, etc. Although the embodiments of the present disclosure have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present disclosure. A person of ordinary skill in the art may make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present disclosure.

Claims

1. An intelligent optical computing attention mechanism model, characterized by: The attention mechanism model is embedded in the large model. The attention mechanism model includes multiple light field intermodulation propagation sub-models, wherein each light field intermodulation propagation sub-model is applied to each sub-head in each phrase window, the phrase window is a window divided by word vectors, and the sub-head is a sub-head divided by feature vectors in each window. The light field intermodulation propagation sub-model includes a spatial calculation module, a temporal calculation module and an output module, wherein, The spatial calculation module encodes the obtained first sub-vector and second sub-vector into the phase and / or amplitude of the spatial light field, and calculates the first sub-vector and the second sub-vector through light propagation to obtain a target first sub-vector; The time calculation module maps the acquired first sub-vector and second sub-vector to the phase and / or amplitude of the time-series optical signal, and calculates the first sub-vector and the second sub-vector through optical propagation to obtain a target second sub-vector; The output module performs a first fusion of the target first subvector of the spatial calculation module and the target second subvector of the temporal calculation module to obtain a target subvector; After the attention mechanism model obtains the target sub-vector in each light field intermodulation propagation sub-model, the target sub-vector in each light field intermodulation propagation sub-model is secondly fused to obtain a target vector, wherein the target vector is the output of the attention mechanism model.

2. The model according to claim 1, characterized in that The parallel computing of the multiple spatial computing modules is achieved through spatial multiplexing SMUX.

3. The model according to claim 1, characterized in that The parallel calculation of the multiple time calculation modules is achieved through wavelength multiplexing WMUX.

4. The intelligent optical computing attention mechanism model according to any one of claims 1 to 3, characterized in that: Used for question-answering processing.

5. A question-answering processing method, characterized in that: include: Get the data to answer the questions you need; Input the problem data into the large model, and determine the first vector and the second vector of the input attention mechanism model in the large model; Calculating the first vector and the second vector using the attention mechanism model to obtain a target vector, and transmitting the target vector to a subsequent model for calculation to obtain a target answer to the question data; The attention mechanism model includes multiple light field intermodulation propagation sub-models, each of which includes a spatial calculation module, a temporal calculation module, and an output module. The attention mechanism model divides the first vector and the second vector into multiple first sub-vectors and multiple second sub-vectors; encoding the first sub-vector and the second sub-vector into the phase and / or amplitude of the spatial light field by the spatial calculation module, and calculating the first sub-vector and the second sub-vector by light propagation to obtain a target first sub-vector; Mapping the first sub-vector and the second sub-vector to the phase and / or amplitude of the time-series optical signal by the time calculation module, and calculating the first sub-vector and the second sub-vector by light propagation to obtain a target second sub-vector; The output module fuses the target first subvector of the spatial calculation module and the target second subvector of the temporal calculation module to obtain a target subvector; The target sub-vectors of each light field intermodulation propagation sub-model are fused to obtain a target vector.

6. The method according to claim 5, characterized in that The fusing the target first subvector of the spatial calculation module and the target second subvector of the temporal calculation module through the output module to obtain the target subvector includes: performing weighted fusion of the target first subvector of the spatial calculation module and the target second subvector of the temporal calculation module through the output module to obtain the target subvector.

7. A computer storage medium, wherein: The computer storage medium stores computer-executable instructions; after the computer-executable instructions are executed by the processor, the method according to any one of claims 5 to 6 can be implemented.

Citation Information

Patent Citations

  • Optical diffraction neural network online training method and system

    CN110929864A

  • Neuromorphic intelligent optical computing architecture system and device

    CN116739064A