A knowledge distillation-based on-chip light neural network design method
By optimizing the parameters of the optical neural network through algorithm optimization and knowledge distillation framework, the problems of large area and high complexity of on-chip optical neural networks are solved, thereby improving their performance and adaptability in handling complex tasks.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHANGSHA SEMICON TECH & APPL INNOVATION RES INST
- Filing Date
- 2025-12-30
- Publication Date
- 2026-04-21
AI Technical Summary
On-chip optical neural networks suffer from problems such as large area, high structural complexity, and poor performance when processing complex tasks.
An on-chip optical neural network design method based on knowledge distillation is adopted. The parameters of the optical neural network are optimized by optimization algorithm, and the optical neural network is trained by combining the knowledge distillation framework to determine the optimal optical neural network parameters and construct an optical neural network that meets the preset process constraints.
It effectively reduces redundant structures, decreases the area and complexity of on-chip optical neural networks, improves performance and task adaptability when handling complex tasks, and achieves high efficiency and accuracy.
Smart Images

Figure CN121436067B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of optical computing technology, and in particular relates to an on-chip optical neural network design method based on knowledge distillation. Background Technology
[0002] With the rapid development of deep learning technology, the scale and complexity of neural network models are constantly increasing. Although large models have achieved groundbreaking results in many tasks such as image recognition, natural language processing, and scientific computing, their massive number of parameters and high computational costs have also brought many problems, such as high inference latency, high energy consumption, and strong dependence on hardware resources. These bottlenecks limit the deployment and application of models in resource-constrained scenarios, prompting academia and industry to develop new, efficient, and low-power computing architectures.
[0003] Optical Neural Networks (ONNs) are considered an important direction for solving this problem due to their unique physical properties. Unlike traditional electronic-based computing, ONNs utilize the physical processes of light propagation, interference, and diffraction to achieve the computational functions of neural networks. Because light itself possesses advantages such as high-speed propagation, low energy consumption, and inherent parallelism, ONNs can theoretically achieve high throughput, low latency, and extremely low energy consumption for forward inference operations. In particular, linear operations in optical systems can be efficiently implemented through structures such as phase modulators, optical Fourier transforms, and integrated interferometer arrays, making ONNs a natural candidate platform for matrix multiplication tasks.
[0004] In recent years, with advancements in optical device fabrication processes and the development of integrated optoelectronic technology, On-Chip Neural Networks (ONNs) have evolved from physical systems to highly integrated on-chip chips. ONNs primarily include methods based on diffraction principles, achieving parallel information processing through phase modulation layers and natural light diffraction; methods based on interference principles, constructing matrix operation units using Mach-Zehnder interferometer arrays; methods based on nonlinear optical effects, implementing activation functions through saturated absorption and the Kerr effect; methods based on two-dimensional metasurfaces, designing topologies on silicon photonics platforms using recurrent neural network (RNN) principles to achieve high integration; and hybrid optoelectronic methods, combining the advantages of high-speed optical parallel computing with precise electronic control. The core principle of these methods is to utilize the physical properties of light (such as wave properties, coherence, and nonlinear effects) to replace digital computation in traditional electronic neural networks, thereby achieving ultra-high-speed, low-power neural network computation. However, compared to traditional optical neural network devices, highly integrated optical neural network chips (i.e., on-chip optical neural networks) suffer from large area, high structural complexity, and poor performance when handling complex tasks. Summary of the Invention
[0005] This application provides an on-chip optical neural network design method based on knowledge distillation, which can solve the problems of large area, high structural complexity, and poor performance when processing complex tasks in on-chip optical neural networks.
[0006] This application provides an on-chip optical neural network design method based on knowledge distillation, including:
[0007] Based on preset process constraints, an initial population for the optimization algorithm is constructed; multiple individuals in the initial population correspond one-to-one with multiple sets of optical neural network parameters.
[0008] The initial population is iteratively optimized using an optimization algorithm to obtain the optimal optical neural network parameters that satisfy the preset process constraints;
[0009] An optical neural network is constructed based on the optimal optical neural network parameters, and the constructed optical neural network is used as an on-chip optical neural network.
[0010] The process of calculating the fitness value of an individual in the optimization algorithm is as follows: the optical neural network corresponding to the individual is trained using the knowledge distillation framework, and the fitness value of the individual is determined based on the training loss value at the end of training and the physical parameters of the optical neural network corresponding to the individual; the fitness value of the individual is inversely proportional to the training loss value.
[0011] Optionally, the optical neural network corresponding to the individual can be trained using a knowledge distillation framework, including:
[0012] An optical neural network is constructed based on the optical neural network parameters corresponding to the individual, and the optical neural network constructed based on the optical neural network parameters corresponding to the individual is used as the student model;
[0013] The target model to be learned is used as the teacher model, and knowledge distillation is performed on the student model based on the teacher model to complete the training of the optical neural network corresponding to the individual.
[0014] Optionally, the training loss values include soft target loss values, hard target loss values, manufacturing errors of the optical neural network, and environmental noise during the operation of the optical neural network;
[0015] The soft target loss is the loss function value between the predicted values output by the student model and the predicted values output by the teacher model, while the hard target loss is the loss function value between the predicted values output by the student model and the true values.
[0016] Optionally, the expression for the training loss value is:
[0017] ;
[0018] in, This represents the training loss value. The weighting coefficients represent the loss values of soft targets. This represents the loss value for soft targets. The weighting coefficients represent the hard target loss values. This represents the loss value for hard targets. This represents the constant that influences the error. This represents the average error obtained from the experiment. This indicates the number of basic units that make up an optical neural network. This represents the environmental noise during the operation of the optical neural network.
[0019] Optional, ambient noise during optical neural network operation The expression is:
[0020] ;
[0021] in, This indicates the preset weighting coefficient. This represents calculating the mathematical expectation of all noise distributions. This represents the input to the optical neural network. Represents the controllable vector parameters of an optical neural network. This represents the random disturbance measured during the process. Indicates in The output of the lower optical neural network, Indicates the controllable vector parameters The output of the lower optical neural network.
[0022] Optionally, the expression for the soft target loss value is:
[0023] ;
[0024] in, This represents the loss value for soft targets. This represents the predicted value output by the student model. This represents the predicted value output by the teacher model. Representing measurement and The distance function between the differences.
[0025] Optionally, the expression for the hard target loss value is:
[0026] ;
[0027] in, This represents the loss value for hard targets. This represents the predicted value output by the student model. Represents the true value. Representing measurement and The distance function between the differences.
[0028] Optionally, each set of optical neural network parameters includes: the number of layers in the optical neural network, the layer width, the interlayer spacing, the number of pixels per layer, the parameters of the metasurface on which the optical neural network is located, and the parameters of MZI and MRR in the optical neural network.
[0029] Optional, preset process constraints include:
[0030] ;
[0031] ;
[0032] ;
[0033] ;
[0034] in, Indicates the distance between adjacent processing points. Indicates the layer width. This indicates the number of pixels that need to be processed in each layer. Indicates the wavelength of light. Indicates the interlayer spacing. This indicates the limit dimensions of the machining process used.
[0035] Optionally, physical parameters include the number of layers, layer width, and interlayer spacing. The fitness value of an individual is determined based on the training loss value at the end of training and the physical parameters of the corresponding optical neural network, including:
[0036] The fitness value of this individual is calculated using the following formula. :
[0037] ;
[0038] in, , and All are preset weighting coefficients. This indicates the layer width of the optical neural network corresponding to that individual. This represents the interlayer spacing of the optical neural network corresponding to that individual. This indicates the number of layers in the optical neural network corresponding to that individual. This indicates the number of pixels that need to be processed in each layer. This represents the training loss value at the end of the training.
[0039] The above-mentioned solution in this application has the following beneficial effects:
[0040] In the embodiments of this application, the parameters of the optical neural network are optimized using an optimization algorithm to determine the optimal optical neural network parameters that meet preset process constraints. The optical neural network constructed based on the optimal optical neural network parameters is then used as the on-chip optical neural network, thereby effectively reducing redundant structures and achieving the effects of reducing the area and complexity of the on-chip optical neural network. Simultaneously, since the optical neural network is trained using a knowledge distillation framework during the execution of the optimization algorithm, the fitness value of each individual is determined based on the training loss value at the end of training and the physical parameters of the optical neural network. The smaller the training loss value, the larger the fitness value, thus enabling the final on-chip optical neural network to possess good inference performance and task adaptability, greatly improving the performance of the on-chip optical neural network in handling complex tasks, and achieving high efficiency and accuracy in handling complex tasks.
[0041] Other beneficial effects of this application will be described in detail in the following detailed description section. Attached Figure Description
[0042] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0043] Figure 1 A flowchart of an on-chip optical neural network design method based on knowledge distillation provided in this application embodiment;
[0044] Figure 2a This is a fitness distribution diagram of the optimal ONN structure searched using a genetic algorithm in the experiment;
[0045] Figure 2b This is a chip area distribution diagram for searching the optimal ONN structure using a genetic algorithm in the experiment;
[0046] Figure 3a This is a schematic diagram illustrating the effect of ONN on the fitting accuracy of the Iris dataset as the number of intermediate layers increases in the experiment.
[0047] Figure 3b This is a graph showing the trend of the fitting effect of the ONN with two intermediate layers in the experiment as the number of iterations increases during training.
[0048] Figure 4 This is a comparative analysis chart of the knowledge distillation effects in the experiment;
[0049] Figure 5a This is the final topology diagram of the ONN model optimized for the Iris dataset recognition task in the experiment;
[0050] Figure 5b This is the final topological structure diagram of the matrix multiplication operator implemented by the two-dimensional metasurface optical neural network in the experiment. Detailed Implementation
[0051] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.
[0052] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.
[0053] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0054] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."
[0055] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0056] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0057] To address the issues of large area, high structural complexity, and poor performance when handling complex tasks in current on-chip optical neural networks (ONNs), this application provides an ONN design method based on knowledge distillation. This method optimizes the parameters of the optical neural network using an optimization algorithm to determine the optimal parameters that satisfy preset process constraints. The ONN constructed based on these optimal parameters is then used as the on-chip optical neural network, effectively reducing redundant structures and achieving the effects of reducing the area and complexity of the ONN. Simultaneously, during the execution of the optimization algorithm, the optical neural network is trained using a knowledge distillation framework. The fitness value of each individual is determined based on the training loss value and the physical parameters of the optical neural network at the end of training. The smaller the training loss value, the larger the fitness value, resulting in a final ONN with good inference performance and task adaptability. This significantly improves the performance of the ONN when handling complex tasks, achieving both efficiency and accuracy in processing complex tasks.
[0058] The on-chip optical neural network design method based on knowledge distillation provided in this application will be illustrated below with reference to specific embodiments.
[0059] like Figure 1 As shown in the embodiments of this application, the on-chip optical neural network design method based on knowledge distillation includes the following steps:
[0060] Step 11: Based on the preset process constraints, construct the initial population of the optimization algorithm; multiple individuals in the initial population correspond one-to-one with multiple sets of optical neural network parameters.
[0061] In some embodiments of this application, process constraints can be set based on the physical characteristics and manufacturing process requirements of the optical computing technology used, so that the final on-chip optical neural network meets the requirements. In some optional embodiments, the above-mentioned preset process constraints include physical and process constraints such as sampling rate limits, Fresnel approximation conditions, minimum feature size, and etching accuracy.
[0062] Specifically, the aforementioned preset process constraints include:
[0063] ;
[0064] ;
[0065] ;
[0066] ;
[0067] in, This indicates the distance between adjacent processing points. Taking a two-dimensional metasurface as an example, a processing point refers to a grid point on the two-dimensional metasurface that needs to be etched. Indicates the layer width. This indicates the number of pixels that need to be processed in each layer. Indicates the wavelength of light. Indicates the interlayer spacing. This indicates the limit size of the processing technology used. The limit size of the processing technology refers to the smallest unit size that can be etched. Taking a two-dimensional metasurface as an example, the limit size of the processing technology here refers to the smallest size of the grid points that need to be etched. Generally, 28nm, 45nm, 90nm or 180nm processes can be selected.
[0068] In some embodiments of this application, each set of optical neural network parameters includes all the parameters required to construct the optical neural network. Each individual in the initial population corresponds to a set of optical neural network parameters, which include: the number of layers, layer width, interlayer spacing, number of pixels per layer, parameters of the metasurface containing the optical neural network, and parameters of the Mach-Zehnder interferometer (MZI) and microring resonator (MRR) within the optical neural network. The parameters of the metasurface containing the optical neural network include the surface area of the metasurface, the number of pixels within the optimized region of the metasurface, and the pixel size of the metasurface. The parameters of the MZI include the radius, coupling coefficient, phase arm length difference, and phase modulator-related parameters, etc., while the parameters of the MRR include the radius, coupling coefficient, phase arm length difference, and phase modulator-related parameters, etc.
[0069] It should be noted that for any two sets of optical neural network parameters, the values of corresponding parameters in the two sets can be completely different, or some parameter values can be the same. Of course, the optical neural networks constructed based on the optical neural network parameters corresponding to each individual are all different.
[0070] Step 12: Iteratively optimize the initial population using an optimization algorithm to obtain the optimal optical neural network parameters that satisfy the preset process constraints. The calculation process of the fitness value of an individual in the optimization algorithm is as follows: the optical neural network corresponding to the individual is trained using a knowledge distillation framework, and the fitness value of the individual is determined based on the training loss value at the end of training and the physical parameters of the optical neural network corresponding to the individual; the fitness value of the individual is inversely proportional to the training loss value.
[0071] In some embodiments of this application, the iterative optimization process described above is as follows: Based on the constructed initial population, the optical neural network (ONN) corresponding to each individual is trained. Each ONN is trained with the same number of training iterations. After training, the fitness value is determined based on the training loss value at the end of training (the larger the training loss value, the smaller the fitness value). Subsequently, the initial ONN is randomly mutated to generate a new structure. The above operation is repeated, and the two sets of structures obtained serve as the initial parents of the optimization algorithm. The optimization algorithm is used to perform crossover and mutation operations on the parents structure to generate new structures. The structures that meet the constraints in the new structures are selected as the next generation of ONN structures. The training and fitness evaluation process is repeated, and the optimization is iteratively performed until the maximum number of iterations of the optimization algorithm or the preset optimization time is reached. The structure with the highest fitness is the optimal ONN structure. It should be noted that the search space of each parameter in the optimization algorithm can be preset. The purpose of using the optimization algorithm to iteratively optimize the initial population is to obtain an optical neural network architecture with the most compact structure, the fewest number of pixels, and the most efficient connection method.
[0072] The optimization algorithm described above can be any heuristic algorithm / evolutionary algorithm. In this embodiment, a genetic algorithm is used as an example to illustrate the on-chip optical neural network design method.
[0073] It should be noted that the genetic algorithm in this application uses the logic of traditional genetic algorithms to perform operations. The difference is that in this application, the individuals are optical neural network parameters, and the fitness value is calculated based on the training loss value during knowledge distillation training.
[0074] In some embodiments of this application, the specific implementation steps for training the optical neural network corresponding to the individual using the knowledge distillation framework are as follows: steps 12.1 to 12.2.
[0075] Step 12.1: Construct an optical neural network based on the optical neural network parameters corresponding to the individual, and use the optical neural network constructed based on the optical neural network parameters corresponding to the individual as the student model.
[0076] In related technologies, there are various ways to construct optical neural networks, including simulating the characteristics of recurrent neural networks (RNNs) and constructing optical neural networks on two-dimensional metasurfaces.
[0077] ;
[0078] ;
[0079] In the above formula, This is the current state. This refers to the state at the previous moment. and These are the linear operators from the hidden layer to the output layer and from the input layer to the hidden layer, respectively. It is a sparse matrix. For the output of the optical neural network, This is the input for the optical neural network.
[0080] In this application, a two-dimensional metasurface optimization region is first constructed. By defining a minimum pixel array, the refractive index is precisely controlled. Each pixel can only select two discrete refractive index states: pure silicon (Si) refractive index and a specific target refractive index. This design directly corresponds to the selective etching process in chip (i.e., optical neural network) manufacturing. In the specific implementation, this method is based on the PyTorch deep learning framework (an open-source deep learning framework) and combined with the Lumerical FDTD optical simulation platform (a numerical simulation technique for simulating micro-nano photonic devices). The optimization region undergoes systematic topological optimization, and the optimal solution is continuously approximated through iterative calculations. Once the optimization conditions are met, a specific optical signal is injected into the input of the ONN. After processing by the on-chip network, the output light intensity is accurately read using a photodetector at the output, ultimately achieving high-precision data fitting of the target model.
[0081] Another approach is to utilize the diffraction properties of light to construct an optical neural network. It is understood that the embodiments of this application do not limit the construction method of the optical neural network; any structure capable of implementing an optical neural network can be used with the knowledge distillation system of the optical neural network described in this application.
[0082] In related technologies, the forward propagation process of optical neural networks is divided into two main stages: the optical propagation stage and the optical modulation stage. In the propagation stage, each pixel in the previous layer acts as a secondary wave source, and the generated light wave propagates forward in space and transmits the signal to the next layer. In the modulation stage, when the light wave propagates to the modulation layer, its phase or amplitude changes accordingly according to the modulator characteristics, thereby achieving precise optical field control.
[0083] The propagation process can be modeled using rigorous Rayleigh-Sommerfeld diffraction theory, which is particularly important in chip-scale optical implementations. This physical modeling method provides a more accurate description of light propagation compared to traditional far-field or small-angle approximations. In this framework, each layer is located in spatial coordinates Each pixel is considered a secondary wave source, and its propagation process to the next layer can be described by the following formula:
[0084] ;
[0085] in, Indicates the first The first in the layer The secondary wave generated by each pixel in spatial coordinates Contribution at that location. Parameters The wavelength of light is represented by λ, and L is the axial spacing between adjacent layers. This is the Euclidean distance between the pixel and the target point. It is the imaginary unit.
[0086] Finally, the detector located on the imaging plane will measure the intensity distribution of this output light field, which will serve as the inference output of the optical neural network. The output is the square of the complex amplitude, i.e., the light intensity distribution.
[0087] Step 12.2: Use the target model to be learned as the teacher model, and perform knowledge distillation on the student model based on the teacher model to complete the training of the optical neural network corresponding to the individual.
[0088] The architecture of optical neural networks (ONNs) is not limited to mimicking a specific type of model during design. Regardless of whether the teacher model is a convolutional neural network, a graph neural network, a multimodal hybrid architecture, or even operators like convolution and matrix multiplication, as long as its output is learnable, the ONN can extract knowledge from it through a distillation mechanism and achieve autonomous training by learning the teacher model's predicted output. Therefore, optical neural networks, as student models, possess natural versatility and cross-model adaptability within the distillation framework. This lays the foundation for building a general-purpose optical computing platform and realizing transfer learning between different architectural models.
[0089] In the knowledge distillation process, the student model not only learns the positive label information from the teacher model but also simultaneously acquires the negative label knowledge, thus achieving more comprehensive knowledge transfer. For classification tasks, this application introduces a distillation temperature parameter. As a control parameter for the softmax function (which is a normalized exponential function), the adjusted expression for the softmax function is: This parameter effectively controls the impact of negative labels on the student model's learning process, optimizing knowledge distillation performance. Among other things, and These represent the teacher model for each category. and categories The output value of logits (logits is the linear output of the last layer of the model). For temperature parameters The adjusted softmax probability (softmax is a mathematical function that converts model output into a probability distribution; its core function is to convert model output into probability values) is used as soft annotation information for student model learning.
[0090] Because ONNs possess advantages such as lightweight structure, low power consumption, and parallel processing, they can reduce computational resource consumption while maintaining inference efficiency. Therefore, the knowledge distillation training loss function designed in this application (i.e., the training loss function trained in step 12 above) can be compatible with its optical propagation characteristics and manufacturing process features, and promote performance improvement. The constructed knowledge distillation training loss function consists of multiple components, among which the basic term is the standard supervised loss term, used to measure the output accuracy of the student model (i.e., ONN) under given real label conditions. The supplementary term is used to measure the difference between the student model and the teacher model in predicting labels, and the error caused by the manufacturing process is also added to the loss function. The error of the manufacturing process mainly comes from the deviation caused by process noise at each etching point. The more pixels that need to be etched in the manufacturing area, the higher the error will be, or the more devices used to form an array of MZI or MRR devices, the greater the process deviation will be. In addition, optical chips are affected by phase noise, thermal noise, and coupling loss in actual operation. These errors will significantly reduce the performance of the model on the real chip. Therefore, this application adds a noise sensitivity parameter and uses noise robustness as the goal of minimizing during training, so that the accuracy of the output results is still guaranteed when the actual on-chip device is subjected to noise disturbance.
[0091] Specifically, in some embodiments of this application, the aforementioned knowledge distillation training loss function values include soft target loss values, hard target loss values, manufacturing errors of the optical neural network, and environmental noise during the operation of the optical neural network. That is, the aforementioned training loss values include soft target loss values, hard target loss values, manufacturing errors of the optical neural network, and environmental noise during the operation of the optical neural network.
[0092] The soft target loss is the loss function value between the predicted values output by the student model and the predicted values output by the teacher model, while the hard target loss is the loss function value between the predicted values output by the student model and the true values.
[0093] Specifically, the expression for the training loss value is as follows:
[0094] ;
[0095] in, This represents the training loss value. The weighting coefficients represent the loss values of soft targets. This represents the loss value for soft targets. The weighting coefficients represent the hard target loss values. This represents the loss value for hard targets. This represents the constant that influences the error. , and All settings can be customized according to the actual situation. This represents the average error obtained from the experiment. This average error is obtained by calculating the ratio of the etching unit error or the performance error of the basic device used in the actual chip fabrication process to the total number of processing units. This indicates the number of basic units that make up an optical neural network. Here, a basic unit refers to each diffraction unit in the optical neural network, each etch point in the metasurface, or the MRR / MZI device in the MRR / MZI network. This represents the environmental noise during the operation of the optical neural network.
[0096] Among them, the environmental noise during the operation of the optical neural network The expression is:
[0097] ;
[0098] in, This represents a preset weighting coefficient used to adjust the proportion of environmental noise in the total loss. This represents calculating the mathematical expectation of all noise distributions. This represents the input to the optical neural network, i.e., the input to the knowledge distillation model. This represents the controllable vector parameters of an optical neural network. These controllable vector parameters include realizable parameters such as phase, coupling ratio, gain, and pixel size. This represents the random disturbance measured during the process. Random disturbance refers to the average random error caused by the manufacturing process of each basic unit, calculated through simulation or actual fabrication. Indicates in The output of the lower optical neural network, Indicates the controllable vector parameters The output of the lower optical neural network.
[0099] The expression for the soft target loss value is:
[0100] ;
[0101] in, This represents the loss value for soft targets. This represents the predicted value output by the student model. This represents the predicted value output by the teacher model. Representing measurement and The distance function between the differences, for example, could be the mean squared error function.
[0102] The expression for the hard target loss value is:
[0103] ;
[0104] in, This represents the loss value for hard targets. This represents the predicted value output by the student model. Represents the true value. Representing measurement and The distance function between the differences, for example, could be the mean squared error function.
[0105] It is worth mentioning that, through the above-mentioned loss function construction method, this application achieves the optimization goal of balancing functional consistency and on-chip device robustness, effectively completes the accurate fitting of the optical student model to the task objective, provides a general loss model framework for optical neural network distillation, lays the foundation for introducing other distillation information, has good scalability and adaptability, and can be applied to different optical platforms and diverse application tasks.
[0106] In some embodiments of this application, the physical parameters of the optical neural network include the number of layers, layer width, and interlayer spacing. Correspondingly, the specific implementation of determining the fitness value of an individual in step 12 above based on the training loss value at the end of training and the physical parameters of the optical neural network corresponding to that individual is as follows:
[0107] The fitness value of this individual is calculated using the following formula. :
[0108] ;
[0109] in, , and All are preset weighting coefficients. This indicates the layer width of the optical neural network corresponding to that individual. This represents the interlayer spacing of the optical neural network corresponding to that individual. This indicates the number of layers in the optical neural network corresponding to that individual. This indicates the number of pixels that need to be processed in each layer. This represents the training loss value at the end of the training.
[0110] Step 13: Construct an optical neural network based on the optimal optical neural network parameters, and use the constructed optical neural network as an on-chip optical neural network.
[0111] The aforementioned optimal set of optical neural network parameters enables the optical neural network structure to achieve optimized optical neural network parameters, thereby allowing the on-chip optical neural network constructed based on this to effectively reduce redundant structures, achieve the effect of reducing the area and complexity of the on-chip optical neural network.
[0112] It is understandable that after obtaining the optimal optical neural network parameters through iterative optimization algorithms (such as genetic algorithms), an optical neural network can be constructed based on traditional optical neural network construction methods, thereby obtaining an optical neural network with small area, low power consumption, and suitable for on-chip use.
[0113] For example, when addressing the issue of high energy consumption in traffic flow prediction, the aforementioned teacher model can be a Long Short-Term Memory (LSTM) network. Based on this, an on-chip optical neural network for traffic flow prediction is designed through steps 11 to 13, thereby reducing energy consumption when using the on-chip optical neural network for traffic flow prediction.
[0114] Furthermore, because on-chip ONNs cannot physically incorporate activation layers between each hidden layer, their nonlinear fitting capability is fundamentally limited. To effectively address this technical deficiency, this application adopts the following technical solution:
[0115] First, a bias circuit (which can be a traditional bias circuit) is added to the back end of the photodetector at the on-chip ONN output port. This bias circuit can provide bias voltage or bias current, and achieve precise control of the bias voltage or current value through a series connection of a resistor power supply.
[0116] Secondly, a nonlinear activation function circuit composed of high-speed diodes is set at the rear end of the bias circuit. This activation circuit has a clear threshold characteristic: when the input voltage or current value is higher than a preset threshold, the diode conducts; when the voltage or current value is lower than the threshold, the diode turns off, thereby realizing an effective nonlinear activation function.
[0117] Furthermore, this application fully utilizes the inherent square nonlinearity of the output light intensity of the ONN, using the square value of the output light intensity as the input signal for the bias circuit and the nonlinear activation circuit. Through this optoelectronic fusion design, the nonlinear fitting capability of the ONN chip is further enhanced, thereby significantly improving the learning performance of the overall system.
[0118] The following is an exemplary description of the on-chip optical neural network design method based on knowledge distillation provided in the embodiments of this application, with reference to specific experiments.
[0119] In this experiment, a multilayer perceptron (MLP) was used as the teacher model, and the classic Iris dataset (a commonly used dataset for classification analysis) was selected as the validation target. This dataset has a structure with 4 input features and 3 output categories. In terms of learning operators, the 4x4 matrix multiplication operator was selected for validation.
[0120] For Iris dataset recognition, when implementing the on-chip ONN using a photodiffraction neural network, based on the characteristic structure of the Iris dataset, this system sets the number of input channels of the ONN to 4 and the number of output channels to 3. The physical parameter configuration of the 2D on-chip ONN is found using an optimal structure search method based on a genetic algorithm. The adjustable structural parameters of the model (i.e., the photodiffraction neural network) include: number of layers. Layer width Interlayer spacing Simultaneously, the model should satisfy physical constraints, constraint 1: the sampling interval must satisfy the sampling theorem for diffraction propagation. ,in, , Indicates the distance between adjacent processing points. This represents the number of pixels that need to be processed in each layer. Constraint 2: The Fresnel approximation holds. , Indicates the interlayer spacing. Indicates the wavelength of light. Constraint 3: Manufacturing process limitations. ,in This represents the limit dimensions of the processing technology used. Since the goal is to select the minimum structure that meets the design objectives, structural trimming and optimization can be performed based on existing ONN structures that can achieve the design objectives, resulting in the minimum ONN structure area and the area requiring processing while meeting design requirements. Based on specific design requirements, a weighted quality assessment is conducted on the reduction of design parameters during the optimization process, because among the three parameters mentioned above, the number of model layers... The number of layers has the greatest impact on the area of the on-chip chip. The highest weighting factor, followed by the interlayer spacing. Then comes the layer width. The initial optical diffraction neural network structure is set as follows: each layer consists of 250 pixels, the distance between adjacent layers is 20 mm, the physical length of each layer is 1 mm, and monochromatic light with a wavelength of 550 nm is used as the incident light source, with a total of 5 modulation propagation layers. In each iteration, three variables randomly undergo abrupt changes, depending on the number of layers. Since it is a discrete variable, its transformation range takes integer values, and the width of the remaining layers... and interlayer spacing Mutations are performed at a rate of 10%. In each genetic iteration, the two individuals with the highest fitness are selected as parents to continue the next round of genetic iterations. At the same time, the top 10% of the elite individuals with the highest fitness are retained from all offspring generated by the mutation to participate in the next generation of mutations.
[0121] Based on the requirements of on-chip ONN chips (i.e., on-chip ONN), the area to be processed should be as small as possible, the chip area should be as small as possible, and the accuracy of fitting the chip should be as high as possible. Therefore, the genetic fitness of the ONN model structure is obtained. Calculation formula:
[0122] ;
[0123] Where M is the number of pixels that need to be processed in each layer. , , These are weighting coefficients used to control the optimization direction of the final ONN chip structure. The ONN model is set to undergo 200 distillation iterations. In this example, the weight of model accuracy is increased, and a high weight is selected. The value is such that the final distilled ONN model fitting cross-entropy is less than 0.01. Repeat the iteration until the number of iterations or the iteration time meets the upper limit, or the fitness is satisfied. The change in value is less than the threshold. If the iteration termination condition is met, the optimal ONN structure is obtained. During the genetic algorithm iteration process, the fitness and chip area distribution of the ONN structure are shown in the following diagram. Figure 2a and Figure 2b As shown in the figure, the vertical axis represents fitness and chip area, respectively, and the horizontal axis represents generations. This illustrates that as the number of generations increases, the chip design area tends to decrease, while the fitness gradually increases.
[0124] The data processing flow in ONN is as follows:
[0125] Input encoding stage: The input encoding layer evenly divides the original input data according to the four feature channels and maps them to different spatial regions of the light field respectively;
[0126] Light propagation stage: The encoded light field is transmitted through the propagation layer of the ONN in space, during which the light wave undergoes natural diffraction.
[0127] Phase modulation stage: After the light field reaches the modulation layer, the system uses the phase information of the light as a trainable parameter for optimized modulation to achieve precise control of the light field;
[0128] Iterative propagation: The phase-modulated light field continues to repeat the above propagation-modulation process in subsequent propagation and modulation layers;
[0129] Output decoding stage: The final light field reaches the output display layer. This layer divides the detection area into three sub-regions according to the three output channels. The average light intensity of each region is calculated by the photodetector to obtain the prediction result of ONN.
[0130] To address the classification task characteristics of the Iris dataset, the system employs a light intensity comparison strategy: calculating the average light intensity of the three output regions and selecting the category corresponding to the region with the highest light intensity value as the final prediction result.
[0131] ONN employs a typical multi-layer structure, comprising an input layer, several intermediate layers, and an output layer. The intermediate layers consist of alternating propagation and modulation layers, and the number of layers can be flexibly adjusted according to specific task requirements. In this experiment, a configuration with two intermediate layers was chosen to balance network expressiveness with computational efficiency. To accurately simulate the propagation process of the light field and improve computational efficiency, the Angular Spectrum Method (ASM) was used to numerically calculate the diffraction effect of the light field during spatial propagation. This method converts spatial convolution into frequency multiplication through Fourier transform, significantly reducing computational complexity and making it particularly suitable for high-precision calculations of far-field diffraction. When discretizing the optical diffraction process, to ensure the accuracy of the numerical calculation, the system strictly adheres to the constraints of the optical sampling theorem. Its sampling parameters must satisfy specific phase function relationships to avoid aliasing and ensure the reliability of the simulation results. The sampling phase function satisfies the following relationship:
[0132] ;
[0133] in, This represents the phase delay function in the spatial frequency domain. The wavelength of light For the distance of transmission in the square, For phase propagation factor, The frequency is the spatial frequency along the X direction.
[0134] The formula for calculating light propagation using the angular spectrum method is as follows:
[0135] ;
[0136] in, This represents the initial light field distribution. Indicates the distance of propagation The subsequent light field distribution Indicates Fourier transform, This represents the inverse Fourier transform. It is a propagation operator.
[0137] In determining the structure of an optical neural network (ONN), the choice of the number of intermediate layers has a decisive impact on the knowledge distillation effect. Too few layers will result in insufficient network expressive power, making it unable to fully learn the knowledge features of the teacher model; while too many layers will cause a sharp increase in the number of model parameters, not only increasing training complexity but, more importantly, significantly increasing the area of the final ONN chip and raising hardware implementation costs. Therefore, determining a reasonable configuration of intermediate layers is crucial for the successful implementation of knowledge distillation tasks in ONN models.
[0138] like Figure 3a As shown, this experiment systematically verifies the impact of different numbers of intermediate layers on the prediction accuracy of the ONN model on the Iris dataset. Experimental results show that when there is only one intermediate layer, the model's learning ability is significantly insufficient, resulting in low prediction accuracy. As the number of intermediate layers gradually increases, the prediction performance of the ONN model significantly improves. When the number of intermediate layers reaches two or more, the model's learning ability is sufficient to achieve 100% prediction accuracy. Based on the above experimental results, this embodiment selects a two-layer intermediate layer configuration, which ensures both the model's learning effect and the cost-effectiveness of hardware implementation.
[0139] like Figure 3b As shown, the loss function change curve during the knowledge distillation training process verifies the effective learning process of the model. With the continuous increase of the number of training steps, the knowledge distillation loss function value of the ONN student model shows a stable decreasing trend, eventually converging to a value close to 0, indicating that the student model has successfully learned and internalized the knowledge representation of the teacher model.
[0140] One of the significant advantages of ONN compared to traditional electronic neural networks is its physical visualization capability, which allows the model's prediction process to be observed through an intuitive light field distribution.
[0141] To fully verify the validity of this application, the ONN model trained by knowledge distillation was applied to the complete Iris dataset for prediction, and the prediction results were compared and analyzed with those of the teacher model. Figure 4 As shown, the ONN student model achieves a prediction accuracy of 97% on the entire Iris dataset, fully demonstrating that the proposed method can achieve high-accuracy knowledge transfer and verifying the feasibility and effectiveness of this technical approach. Different symbols are used in the figure to distinguish the prediction errors of the two models: crosses indicate prediction errors from the MLP teacher model, and asterisks indicate prediction errors from the ONN student model. Experimental results show that the ONN student model achieves a prediction accuracy of 97% on the entire Iris dataset, while the MLP teacher model achieves 99% accuracy, verifying the effectiveness of the knowledge distillation method in optical neural networks and achieving successful knowledge transfer from the teacher model to the student model.
[0142] For Iris dataset recognition, another on-chip ONN network is implemented using a two-dimensional metasurface topology. Within the optimization region, the adjoint method of the inverse design algorithm is introduced for precise pixel optimization. In the initial design, the pixel size is set to 200 nm. Considering the actual fabrication process is 90 nm, the pixel size is moderately adjusted to bridge the gap between theoretical design and manufacturing, ensuring the stability and reliability of the chip's function after fabrication. Based on the preset chip area and pixel size, the number of pixels within the optimization region is accurately calculated. The study uses the input-output data of the teacher model as the optimization objective function, and solves for the gradient information of each pixel during the inverse design process using the adjoint method. Subsequently, the gradient descent optimization algorithm is applied to iteratively optimize the target region, continuously approaching the optimal solution. The complete device design optimization process is systematically divided into four key stages: initialization, grayscale, quantization, and manufacturing constraint. Among these, the initialization stage is crucial, with the main task being to establish an accurate simulation optimization model. An advanced scripting language is used to automatically generate the initial model and accurately determine the initial parameter values, laying a solid theoretical foundation for subsequent optimization.
[0143] In the grayscale stage, this application introduces a parameter mapping method. Specifically, a key parameter is set within each region cell, which can be linearly mapped to the dielectric constant of the physical model. This method allows for continuous and precise changes in the dielectric constant within the region, effectively responding to the iterative optimization process of the objective function. The core process of iterative optimization includes two key electromagnetic field simulations: forward simulation and adjoint simulation. In the forward simulation stage, the updated parameters are substituted into the model for a comprehensive simulation and accurate calculation of the current graphic performance index (FOM). If the preset convergence condition is not met, the process transitions to the adjoint simulation stage, placing the light source on a monitor to acquire complete adjoint field data. Based on the obtained simulation data, the gradient value of the objective function with respect to the design parameters is accurately calculated according to the pre-derived theoretical formula, and the parameters are updated accordingly. When the preset convergence condition is met, the iterative process terminates. To improve the feasibility of subsequent manufacturing, a filter is introduced during the optimization process to perform a smooth convolution operation on the pattern. This key technique aims to effectively eliminate potential spikes and small holes in the pattern, significantly improving the overall quality and consistency of the pattern. After the grayscale phase, the system enters the quantization phase, where the optimized areas undergo rigorous quantization to address the limitations of actual manufacturing processes. To ensure the layout design fully meets actual manufacturing requirements, strict manufacturing rule constraint checks are performed after quantization. By introducing a carefully designed penalty function into the optimization algorithm, the quality factor is reduced accordingly when the design layout violates specific manufacturing rules. This strategy effectively guides the optimization process to proactively avoid potential violations. Finally, the layout design, after multiple rounds of optimization and rigorous verification, will be accurately output, completing the entire two-dimensional metasurface on-chip optical neural network chip design.
[0144] like Figure 5a As shown, the final optimized topology of the two-dimensional metasurface on-chip optical neural network for detecting the Iris dataset is displayed. In the figure, black pixels represent silicon dioxide material and gray pixels represent silicon material. The process is uniformly 180nm, that is, the length of a pixel is 180nm.
[0145] To further verify the effectiveness of the proposed method for other teacher models, this application validates the effect of operator learning, with the overall process being consistent with the Iris dataset detection model construction process described above. Figure 5b This is the final topological structure diagram of the matrix multiplication operator implemented by a two-dimensional metasurface optical neural network, which exhibits high accuracy in predicting the target dataset. In the diagram, black pixels represent silicon dioxide, and gray pixels represent silicon.
[0146] In summary, the on-chip optical neural network design method of this application can improve the performance of on-chip optical neural networks when processing complex tasks, while reducing the area and complexity of the on-chip optical neural network.
[0147] The above description is the preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principles described in this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. A method for designing on-chip optical neural networks based on knowledge distillation, characterized in that, include: Based on preset process constraints, an initial population for the optimization algorithm is constructed. The initial population has a one-to-one correspondence with multiple sets of optical neural network parameters; The initial population is iteratively optimized using the optimization algorithm to obtain the optimal optical neural network parameters that satisfy the preset process constraints; An optical neural network is constructed based on the optimal optical neural network parameters, and the constructed optical neural network is used as an on-chip optical neural network. The optimization algorithm calculates the fitness value of an individual as follows: the optical neural network corresponding to the individual is trained using a knowledge distillation framework, and the fitness value of the individual is determined based on the training loss value at the end of training and the physical parameters of the optical neural network corresponding to the individual; the fitness value of the individual is inversely proportional to the training loss value. The training loss values include soft target loss values, hard target loss values, manufacturing errors of the optical neural network, and environmental noise during the operation of the optical neural network; The preset process constraints include: ; ; ; ; in, Indicates the distance between adjacent processing points. Indicates the layer width. This indicates the number of pixels that need to be processed in each layer. Indicates the wavelength of light. Indicates the interlayer spacing. Indicates the limit dimensions of the processing technology used; The physical parameters include the number of layers, layer width, and interlayer spacing. The determination of the fitness value of an individual based on the training loss value at the end of training and the physical parameters of the corresponding optical neural network includes: The fitness value of this individual is calculated using the following formula. : ; in, , and All are preset weighting coefficients. This indicates the layer width of the optical neural network corresponding to that individual. This represents the interlayer spacing of the optical neural network corresponding to that individual. This indicates the number of layers in the optical neural network corresponding to that individual. This indicates the number of pixels that need to be processed in each layer. This represents the training loss value at the end of the training.
2. The on-chip optical neural network design method according to claim 1, characterized in that, The process of training the optical neural network corresponding to the individual using the knowledge distillation framework includes: An optical neural network is constructed based on the optical neural network parameters corresponding to the individual, and the optical neural network constructed based on the optical neural network parameters corresponding to the individual is used as the student model; The target model to be learned is used as the teacher model, and knowledge distillation is performed on the student model based on the teacher model to complete the training of the optical neural network corresponding to the individual.
3. The on-chip optical neural network design method according to claim 2, characterized in that, The soft target loss is the loss function value between the predicted value output by the student model and the predicted value output by the teacher model, and the hard target loss is the loss function value between the predicted value output by the student model and the true value.
4. The on-chip optical neural network design method according to claim 3, characterized in that, The expression for the training loss value is: ; in, This represents the training loss value. The weighting coefficients represent the loss values of soft targets. This represents the loss value for soft targets. The weighting coefficients represent the hard target loss values. This represents the loss value for hard targets. This represents the constant that influences the error. This represents the average error obtained from the experiment. This indicates the number of basic units that make up an optical neural network. This represents the environmental noise during the operation of the optical neural network.
5. The on-chip optical neural network design method according to claim 4, characterized in that, Environmental noise during optical neural network operation The expression is: ; in, This indicates the preset weighting coefficient. This represents calculating the mathematical expectation of all noise distributions. This represents the input to the optical neural network. Represents the controllable vector parameters of an optical neural network. This represents the random disturbance measured during the process. Indicates in The output of the lower optical neural network, Indicates the controllable vector parameters The output of the lower optical neural network.
6. The on-chip optical neural network design method according to claim 4, characterized in that, The expression for the soft target loss value is: ; in, This represents the soft target loss value. This represents the predicted value output by the student model. This represents the predicted value output by the teacher model. Representing measurement and The distance function between the differences.
7. The on-chip optical neural network design method according to claim 4, characterized in that, The expression for the hard target loss value is: ; in, This represents the loss value of the hard target. This represents the predicted value output by the student model. Represents the actual value. Representing measurement and The distance function between the differences.
8. The on-chip optical neural network design method according to claim 1, characterized in that, Each set of optical neural network parameters includes: the number of layers, layer width, interlayer spacing, number of pixels per layer, parameters of the metasurface on which the optical neural network is located, and parameters of MZI and MRR in the optical neural network.
Citation Information
Patent Citations
Deep optical neural network training method and system based on firefly algorithm
CN116842988A
Rapid high-performance online training method for optical diffraction neural network device
CN117291256A