Photon convolutional neural network system
By introducing a dynamic convolution kernel module into the photon convolution neural network system, combining non-volatile and volatile modulation, dynamically adjusting the weight and size of the convolution kernel, the problem of insufficient feature extraction capabilities of the convolution kernel is solved, and stronger feature extraction and multi-scale feature recognition are achieved, and network performance is improved.
Patent Information
- Application Number
- CN202510746642.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-05
AI Technical Summary
In the existing photon convolutional neural network system, the convolution kernel feature extraction capability is limited, and it cannot adapt to the characteristics of different input data, and is limited by the number of non-volatile phase states of phase change materials, resulting in insufficient feature extraction capability.
The dynamic convolution kernel module is adopted, combining non-volatile modulation and volatile modulation, and the coefficient generation model and size selection model are used to dynamically adjust the weight and size of the photon synaptic device to achieve flexible adaptability to the input data.
The feature extraction capability of the photon convolutional neural network system is improved, the recognition and expression of multi-scale features are enhanced, and the robustness and computing efficiency of the network are improved.
Smart Images

Figure CN120258052A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of optical neural networks, and more specifically, relates to a photonic convolutional neural network system. Background Art
[0002] Convolutional Neural Network (CNN) has made remarkable achievements in image classification, natural language processing and other fields. However, in actual data processing, convolution operation, as the pre-operation of CNN, occupies most of the computing power of CNN operation. With the explosive growth of CNN data volume, the von Neumann bottleneck of traditional electronic chips has become increasingly prominent, making it increasingly difficult to meet the rapidly growing hardware system requirements of CNN. Optical neuromorphic computing system has become a disruptive high-performance CNN computing architecture by virtue of the advantages of high speed, parallelism, low crosstalk, low energy consumption and high interconnection bandwidth of photons, as well as the rapid development of integrated optoelectronics in recent years.
[0003] At present, traditional photonic computing chips mainly use silicon-based materials as the substrate to construct optical components, and realize active regulation of devices through carrier dispersion effect or thermo-optical effect. However, the refractive index regulation range of carrier dispersion effect is limited (~10 -3 ) and the phase modulation effect is not ideal; the response time of the thermo-optical effect is slow, usually in the order of milliseconds. In recent years, phase change materials such as GST have received widespread attention in the field of integrated photonics due to their excellent optical programmable properties. Phase change materials are integrated with silicon-based devices to build photonic convolution accelerators, and the non-volatility of phase change materials is used to store different convolution kernel weight values to achieve "in-memory computing" of convolution. Based on this, the existing photonic convolutional neural network system encodes the pre-trained convolution kernel weight mapping into the non-volatile phase state of the phase change material. When data is input, the input data is convolved with the fixedly stored convolution kernel; however, in this way, the convolution kernel is fixed and cannot be adaptively adjusted according to the characteristics of the input data. The ability to extract different input data features is limited, which limits the performance under complex tasks. At the same time, due to the limited number of non-volatile states of phase change materials, the value of the convolution kernel weight is limited, which in turn limits the expression of the convolution kernel and the feature extraction capability is limited. Summary of the invention
[0004] In response to the above defects or improvement needs of the prior art, the present invention provides a photonic convolutional neural network system to solve the technical problem of limited convolution kernel feature extraction capability in the existing photonic convolutional neural network system.
[0005] In order to achieve the above-mentioned object, the present invention provides a photonic convolutional neural network system, comprising: dynamic convolution kernel modules corresponding one-to-one to each convolution layer in a pre-trained convolutional neural network model; The dynamic convolution kernel module includes: a photon synaptic device array; pre-trained convolution kernel weights corresponding to a convolution layer are stored in the array; wherein, the photon synaptic devices in the array are photon synaptic devices based on phase change materials; the convolution kernel weights are encoded into corresponding non-volatile phase states of the photon synaptic devices through a non-volatile modulation method, so as to achieve storage; one photon synaptic device corresponds to storing one convolution kernel weight; the light transmittance of the photon synaptic devices in different phase states is different; The dynamic convolution kernel module is used to dynamically adjust the convolution kernel weights stored in the array before performing a convolution operation: receive the dynamic coefficients corresponding to the convolution kernel weights stored in the array, and apply a pulse signal corresponding to the corresponding dynamic coefficient to each photon synaptic device for volatile modulation; The dynamic convolution kernel module is further used to, when performing a convolution operation, receive a continuous signal light carrying data information to be currently convolved, the continuous signal light is attenuated under the action of the array, and output the attenuated signal light; the light intensity of the output signal light is the current convolution operation result; Wherein, the dynamic coefficients are generated by inputting the data to be convolved into a corresponding pre-trained coefficient generation model; one dynamic convolution kernel module corresponds to one coefficient generation model, and the coefficient generation model is a deep learning model.
[0006] Further preferably, the coefficient generation model includes: a plurality of cascaded fully connected layers, an activation layer arranged between adjacent two fully connected layers, and a pooling layer connected before the first fully connected layer.
[0007] Further preferably, the training method of the coefficient generation model corresponding to each dynamic convolution kernel module includes: Introduce the above coefficient generation model during the training process of the convolutional neural network model, and the coefficient generation models corresponding to each dynamic convolution kernel module respectively correspond one-to-one with each convolution layer in the convolutional neural network model; during the forward propagation process, before each convolution layer performs a convolution operation, the corresponding coefficient generation model generates the dynamic coefficients corresponding to the convolution kernel weights in this convolution layer based on the data input to this convolution layer, and dynamically adjusts the corresponding convolution kernel weights in this convolution layer based on the obtained dynamic coefficients; during the backpropagation process, the parameters in the convolutional neural network model and each coefficient generation model are adjusted simultaneously.
[0008] Further preferably, the weights in the same convolution kernel are stored in the same column of the array; The dynamic convolution kernel module is also used to dynamically adjust the sizes of the convolution kernel weights stored in the array before performing the convolution operation, and then dynamically adjust the sizes of the convolution kernels stored in the array: select a number of photon synaptic devices from each column of the array, keep the phase states of the selected photon synaptic devices unchanged, and modulate the phase states of the unselected photon synaptic devices to the fully crystalline state by means of a non-volatile modulation method, so as to adjust the size of the convolution kernel stored in this column to the corresponding optimal size; Wherein, a dynamic convolution kernel module also corresponds to a size selection model, and the size selection model is a deep learning model; The optimal sizes of the convolution kernels stored in the array are generated by inputting the data to be convolved into the corresponding size selection module; the size selection model is used to extract the features of the data to be convolved, and then map them to the optimal sizes of the convolution kernels stored in the array within the corresponding dynamic convolution kernel module.
[0009] Further preferably, the size selection model includes: a cascaded feature extraction module and a classifier; wherein, the feature extraction module includes one or more of: an edge feature extraction unit, a texture feature extraction unit, and a smoothness feature extraction unit.
[0010] Further preferably, the above classifier is: a fully connected conditional network; wherein, the fully connected conditional network includes: a plurality of cascaded fully connected layers, an activation layer arranged between adjacent two fully connected layers, and a softmax layer connected after the last fully connected layer.
[0011] Further preferably, the training methods of the coefficient generation model and the size selection model corresponding to each dynamic convolution kernel module include: Introduce the above coefficient generation model and size selection model during the training process of the convolutional neural network model. The coefficient generation models corresponding to each dynamic convolution kernel module are respectively in one-to-one correspondence with each convolutional layer in the convolutional neural network model; the size selection models corresponding to each dynamic convolution kernel module are respectively in one-to-one correspondence with each convolutional layer in the convolutional neural network model; During the forward propagation process, before each convolutional layer performs the convolution operation, the corresponding coefficient generation model generates the dynamic coefficients corresponding to the convolution kernel weights in this convolutional layer based on the data input to this convolutional layer, and dynamically adjusts the corresponding convolution kernel weights in this convolutional layer based on the obtained dynamic coefficients; the corresponding size selection model generates the optimal sizes of the convolution kernels in this convolutional layer based on the data input to this convolutional layer, so as to dynamically adjust the sizes of the convolution kernels in this convolutional layer; During the backpropagation process, the parameters in the convolutional neural network model, each coefficient generation model, and each size selection model are adjusted simultaneously.
[0012] Further preferably, the above-mentioned photonic convolutional neural network system further includes: a light source module, which is used to load the data to be convolved currently onto an optical signal by using an optical signal modulator to obtain a continuous signal light carrying the data information to be convolved currently.
[0013] Further preferably, the above-mentioned photonic convolutional neural network system further includes: a photoelectric conversion module, a non-linear activation module, and a fully connected layer module; The photoelectric conversion module is used to convert the signal light output by the dynamic convolution kernel module into an electrical signal and output it to the non-linear activation module; The non-linear activation module is used to perform non-linear activation processing on the electrical signal input by the photoelectric conversion module; The fully connected layer module is used to calculate the result of the non-linear activation processing output by the non-linear activation module for the last time to obtain the final output result.
[0014] Further preferably, the photonic synaptic device includes: a substrate, a waveguide layer, a phase change material layer, a heating layer, a covering layer, and an electrode layer acting on the heating layer, which are arranged in sequence from bottom to top.
[0015] Generally speaking, through the above technical solutions conceived by the present invention, the following beneficial effects can be achieved: 1. The present invention provides a photonic convolutional neural network system, which includes each dynamic convolution kernel module corresponding one by one to each convolution layer in a pre-trained convolutional neural network model; wherein, the dynamic convolution kernel module includes a photonic synaptic device based on a phase change material, and the convolution kernel weights of the corresponding convolution layer are non-volatilely stored in the photonic synaptic device; before the dynamic convolution kernel module performs a convolution operation, the weights of the corresponding photonic synaptic device inside are dynamically adjusted by an easy modulation method according to the dynamic coefficients generated by the coefficient generation model according to the content of the input data, which can flexibly adapt to the changes of the input data; at the same time, the present invention does not solely rely on the non-volatile phase state of the phase change material for encoding, but further introduces volatile modulation encoding on this basis, effectively alleviating the problem that the existing photonic convolutional neural network system is severely limited by the number of non-volatile phase states of the phase change material, and greatly enriching the encoding expression ability of the convolution kernels in the dynamic convolution kernel module; based on this, the convolution kernel in the photonic convolutional neural network system provided by the present invention has strong feature extraction ability.
[0016] 2. Further, in the photon convolution neural network system provided by the present invention, the coefficient generation model includes: first, a pooling layer is used to extract local and global features, reducing the computational complexity while retaining important features; then, multiple cascaded fully connected layers are used to further gradually extract, combine, and abstract these features, and an activation layer is introduced between two adjacent fully connected layers, enabling the coefficient generation model to capture complex patterns in the input data, making the coefficient generation model adaptable to inputs of different sizes, enhancing the generalization ability of the model, generating coefficients suitable for adjusting the convolution kernel, and further improving the feature extraction ability of the photon convolution neural network system.
[0017] 3. Further, in the photon convolution neural network system provided by the present invention, the coefficient generation model adopted is obtained through collaborative training with the convolution neural network model corresponding to the photon convolution neural network system, which can, while paying attention to the data to be convolved, fully match the features of the convolution kernel weights in the convolution neural network model, and further improve the feature extraction ability of the photon convolution neural network system.
[0018] 4. Further, in the photon convolution neural network system provided by the present invention, before the dynamic convolution kernel module performs convolution operations, the size of the convolution kernel stored in its internal photon synaptic device array is dynamically adjusted through a non-volatile modulation method according to the optimal convolution kernel size obtained by the size selection model based on the features of the input data, enabling the dynamic convolution kernel module to extract data information at different optimal levels, and further improving the ability of the photon convolution neural network system to recognize and express multi-scale features.
[0019] 5. Further, in the photon convolution neural network system provided by the present invention, the feature extraction module of the size selection model includes one or more of an edge feature extraction unit, a texture feature extraction unit, and a smoothness feature extraction unit. By extracting statistical features of the complexity and detail level of the input data, the usage frequency of convolution kernels of different sizes can be dynamically adjusted according to input data of different complexities and features, enabling the optimal size of the convolution kernel to be determined more accurately, and thus further improving the ability of the photon convolution neural network system to recognize and express multi-scale features.
[0020] 6. Further, in the photon convolution neural network system provided by the present invention, the size selection model adopted is obtained through collaborative training with the coefficient generation model and the convolution neural network model corresponding to the photon convolution neural network system, which can, while paying attention to the data to be convolved, fully match the features of the convolution kernel weights in the convolution neural network model, and further improve the feature extraction ability of the photon convolution neural network system.
[0021] 7. Further, in the photon convolution neural network system provided by the present invention, the photon synaptic device preferably adopts an electrically tunable photon synaptic device, which has a fast regulation speed and can achieve regulation at the nanosecond or microsecond level; has strong anti-light noise interference ability and good stability; has high regulation precision, and the optical parameters can be accurately controlled and adjusted by electrothermal elements. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 It is a schematic structural diagram of the photon convolution neural network system provided by an embodiment of the present invention.
[0023] Figure 2 It is a schematic structural diagram of the dynamic convolution module provided by an embodiment of the present invention.
[0024] Figure 3 It is a schematic structural diagram of the electrically tunable phase change photon synaptic device provided by an embodiment of the present invention.
[0025] Figure 4 It is a non-volatile modulation performance diagram of the phase change photon synaptic device provided by an embodiment of the present invention.
[0026] Figure 5 It is a volatile modulation performance diagram of the phase change photon synaptic device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0027] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0028] To achieve the above objective, the present invention provides a photon convolution neural network system, including: each dynamic convolution kernel module corresponding one by one to each convolution layer in a pre-trained convolution neural network model; The dynamic convolution kernel module includes: a photon synaptic device array; the photon synaptic devices in the array are photon synaptic devices based on phase change materials; the pre-trained convolution kernel weights in the corresponding convolution layer are stored in the array; among them, the convolution kernel weights are encoded into the corresponding non-volatile phase states of the photon synaptic devices through non-volatile modulation, so as to achieve storage; one photon synaptic device corresponds to storing one convolution kernel weight; the light transmittance of the photon synaptic devices in different phase states is different; The dynamic convolution kernel module is used to dynamically adjust the convolution kernel weights stored in the array before performing the convolution operation: receive the dynamic coefficients corresponding to the convolution kernel weights stored in the array, and apply pulse signals corresponding to the corresponding dynamic coefficients to each photon synaptic device for volatile modulation; The dynamic convolution kernel module is also used to receive a continuous signal light carrying data information of the current convolution operation to be performed during the convolution operation. The continuous signal light decays under the action of the array, and the decayed signal light is output; the intensity of the output signal light is the result of the current convolution operation. Among them, the dynamic coefficient is generated by inputting the data to be convolved into the corresponding pre-trained coefficient generation model; one dynamic convolution kernel module corresponds to one coefficient generation model, and the coefficient generation model is a deep learning model.
[0029] It should be noted that the data to be convolved depends on the specific tasks of the photon convolution neural network system (such as natural language processing tasks, image classification tasks, etc.), and can be speech data, image data, etc.
[0030] It should be noted that the input layer dimension of the coefficient generation model is the data dimension of the convolution operation to be performed by the corresponding dynamic convolution kernel module, and the output layer dimension is the number of convolution kernel weights stored in the array in the corresponding dynamic convolution kernel module or the number of convolution kernels stored in the array in the corresponding dynamic convolution kernel module. When the dimension of the coefficient generation model is the number of convolution kernel weights stored in the array in the corresponding dynamic convolution kernel module, the generation of the coefficient generation model is directly the dynamic coefficient corresponding to the convolution kernel weights stored in the array. When the dimension of the coefficient generation model is the number of convolution kernels stored in the array in the corresponding dynamic convolution kernel module, the generation of the coefficient generation model is directly the dynamic coefficients corresponding to the convolution kernels stored in the array; at this time, the coefficients of the convolution kernel weights belonging to the same convolution kernel are all the dynamic coefficients corresponding to this convolution kernel, and thus the dynamic coefficients corresponding to each convolution kernel stored in the array are obtained.
[0031] Through the above different settings, the dynamic convolution kernel module can perform volatile modulation on each convolution kernel weight stored in its array, or the dynamic convolution kernel module can perform volatile modulation on each convolution kernel stored in its array.
[0032] The dynamic convolution kernel module uses the dynamic coefficient generated according to the input data to be convolved to act on the convolution kernel weights pre-stored in the array. After volatile modulation of the convolution kernel weights, a new dynamic convolution kernel is formed. The dynamically generated convolution kernel has stronger feature detection ability, thus improving the feature extraction ability of the convolution network.
[0033] It should be noted that the photonic synaptic devices (also known as optical synaptic devices) in the above-mentioned photonic synaptic device array are photonic synaptic devices based on phase change materials, which can specifically be electrically tunable phase change photonic synaptic devices or optically tunable phase change photonic synaptic devices, and are not limited here. For example, an electrically tunable phase change photonic synaptic device includes: a substrate, a waveguide layer, a phase change material layer, a heating layer, a covering layer, and an electrode layer acting on the heating layer, which are arranged in sequence from bottom to top. Preferably, the heating layer of the phase change optical synaptic device completely covers the phase change material layer to ensure obvious volatile modulation effects and at the same time does not affect the long-term retention of the non-volatile state. The difference between the optically tunable phase change photonic synaptic device and the structure of the electrically tunable phase change photonic synaptic device is only that it does not include the electrode layer acting on the heating layer, and the other structures are the same. The optical modulation method of the optically tunable phase change photonic synaptic device is to input the modulation light from the waveguide to modulate the phase change material. The power of the modulation light here is higher than the power of the optical signal.
[0034] When the photonic synaptic device in the above-mentioned photonic synaptic device array is an electrically tunable phase change photonic synaptic device, the pulse signal used for volatile modulation by the dynamic convolution kernel module is an electrical pulse signal. When the photonic synaptic device in the above-mentioned photonic synaptic device array is an optically tunable phase change photonic synaptic device, the pulse signal used for volatile modulation by the dynamic convolution kernel module is an optical pulse signal.
[0035] For the pulse signal applied by the dynamic convolution kernel module during volatile modulation, its parameters (such as amplitude, pulse width, etc.) correspond to the corresponding dynamic coefficients, and there can be a linear correspondence relationship, an exponential correspondence relationship, a logarithmic correspondence relationship, any combination of the above various correspondence relationships, etc., which are not limited here.
[0036] It should be noted that the data to be convolved is divided into multiple data segments by means of a sliding window, and the window size is consistent with the convolution kernel size in the dynamic convolution kernel module; each time a data segment is loaded onto an optical signal and input into the array of the dynamic convolution kernel module. Taking the example that the weights in the same convolution kernel are stored in the same column of the array, if the convolution kernel size is denoted as m, then m optical signals are used, and each optical signal loads a value on the data segment; during the convolution operation, the m optical signals are input into the rows storing the convolution kernel weights in the array one by one to complete the convolution operation of the current data segment with each convolution kernel in the dynamic convolution kernel module (the output of each column is the convolution operation result of the data segment and the corresponding convolution kernel). The situation where the weights in the same convolution kernel are stored in the same row of the array is similar to the situation where the weights in the same convolution kernel are stored in the same column of the array. The difference is that the m optical signals are input into the columns storing the convolution kernel weights in the array one by one, and the output of each row is the convolution operation result of the data segment and the corresponding convolution kernel.
[0037] It should be noted that there are various coefficient generation models that can be adopted, such as CNN, Transformer, etc., which are not limited here. Preferably, in an alternative implementation, the coefficient generation model includes: a plurality of cascaded fully connected layers, an activation layer arranged between two adjacent fully connected layers, and a pooling layer (which can be a max pooling layer, an average pooling layer, a global max pooling layer, a global average pooling layer, etc., not limited here) connected before the first fully connected layer. Specifically, for the data to be subjected to convolution operation as input, the pooling layer, as a feature extraction layer, will reduce the dimension of the input feature space and extract features, and then further map the features to the dynamic coefficients corresponding to the weights of each convolution kernel stored in the array after passing through the fully connected layer and the non-linear layer.
[0038] It should be noted that there are various training methods for the above coefficient generation model. It can be trained independently or co-trained with a convolutional neural network model. To further improve the feature extraction ability of the photon convolutional neural network system, it is preferably co-trained with the convolutional neural network model. Preferably, in an alternative implementation, the training method of the coefficient generation model corresponding to each dynamic convolution kernel module includes: Introduce the above coefficient generation model during the training process of the convolutional neural network model. The coefficient generation models corresponding to each dynamic convolution kernel module respectively correspond one by one to each convolutional layer in the convolutional neural network model; during the forward propagation process, before each convolutional layer performs a convolution operation, the corresponding coefficient generation model generates the dynamic coefficients corresponding to the weights of each convolution kernel in this convolutional layer based on the data input to this convolutional layer, and dynamically adjusts the corresponding convolution kernel weights in this convolutional layer based on the obtained dynamic coefficients; during the backpropagation process, the parameters in the convolutional neural model and each coefficient generation model are adjusted simultaneously.
[0039] Preferably, when training the coefficient generation module, the parameters in the coefficient generation module are initialized with nearly uniform initial values at the beginning of training, so that more convolution kernels can be optimized simultaneously at the beginning of training to promote the learning of all convolution kernels.
[0040] In an alternative implementation, the weights in the same convolution kernel are stored in the same column of the array; The dynamic convolution kernel module is also used to dynamically adjust the size of each convolution kernel stored in the array after dynamically adjusting the convolution kernel weights stored in the array before performing the convolution operation: select several photon synaptic devices from each column of the array, keep the phase states of the selected photon synaptic devices unchanged, and modulate the phase states of the unselected photon synaptic devices to the fully crystalline state through a non-volatile modulation method to adjust the size of the convolution kernel stored in this column to the corresponding optimal size; Among them, a dynamic convolution kernel module also corresponds to a size selection model, which is a deep learning model. The dimension of its input layer is the data dimension for the convolution operation to be performed by the corresponding dynamic convolution kernel module, and the dimension of the output layer is the number of convolution kernels stored in the array within the corresponding dynamic convolution kernel module; The optimal size of each convolution kernel stored in the array is generated by inputting the data to be convolved into the corresponding size selection module; the size selection model is used to extract the features of the data to be convolved (such as one or more of edge features, texture features, and smoothness features), and then map them to the optimal sizes of the convolution kernels stored in the array within the corresponding dynamic convolution kernel module.
[0041] It should be noted that there are multiple types of size selection models that can be used, such as CNN, Transformer, etc., which are not limited here. Preferably, in an alternative implementation, the size selection model includes: a cascaded feature extraction module and a classifier; among them, the feature extraction module includes: one or more of an edge feature extraction unit, a texture feature extraction unit, and a smoothness feature extraction unit. The statistical features of the complexity and detail level of the input data are extracted by one or more of the edge feature extraction unit, the texture feature extraction unit, and the smoothness feature extraction unit. Among them, the edge feature extraction unit can use an edge detection operator to extract edge features. In an alternative implementation, the edge detection operator can select the Roberts operator based on the first derivative, the Prewitt operator based on the first derivative, the Sobel operator based on the first derivative, the Laplacian operator based on the second derivative, the Canny operator based on non-differentiation, etc., to detect the discontinuous edge information in the input data. The texture feature extraction unit can use a texture detection operator to extract texture features. In an alternative implementation, the texture detection operator can select the gray-level co-occurrence matrix, local binary pattern, Gabor filter, Fourier transform, wavelet transform, etc., to detect the texture information existing in the input data. The smoothness feature extraction unit can calculate by means of local variance or global variance, etc., to detect the local or global complexity information of the input data, thereby extracting the smoothness feature of the input data.
[0042] Based on the scale selection model, the usage frequency of convolution kernels of different sizes can be dynamically adjusted according to input data of different complexities and features.
[0043] It should be noted that there are various classifiers that can be used, such as fully connected neural networks, support vector machines, decision trees, etc., and no limitation is made here. In an alternative embodiment, the above classifier is: a fully connected conditional network; wherein, the fully connected conditional network includes: a plurality of cascaded fully connected layers, an activation layer provided between two adjacent fully connected layers, and a softmax layer connected after the last fully connected layer; wherein, the activation layer can be a ReLU activation layer, a Sigmoid activation layer, a tanh activation layer, etc., and no limitation is made here.
[0044] It should be noted that the data to be subjected to convolution operation is divided into multiple data segments by means of a sliding window, and the window size is the same as the convolution kernel size before dynamic scale adjustment in the dynamic convolution kernel module; each time a data segment is loaded onto an optical signal and input into the array of the dynamic convolution kernel module. Taking the example that the weights in the same convolution kernel are stored in the same column of the array, let the convolution kernel size before dynamic scale adjustment be m, then m optical signals are used, and each optical signal is loaded with a value on the data segment; during the convolution operation, the m optical signals are input into the rows storing the convolution kernel weights in the array one by one to complete the convolution operation of the current data segment with each convolution kernel in the dynamic convolution kernel module (the output of each column is the convolution operation result of the data segment and the corresponding convolution kernel). The situation where the weights in the same convolution kernel are stored in the same row of the array is similar to the situation where the weights in the same convolution kernel are stored in the same column of the array, except that the m optical signals are input into the columns storing the convolution kernel weights in the array one by one, and the output of each row is the convolution operation result of the data segment and the corresponding convolution kernel.
[0045] It should be noted that there are various training methods for the above size selection model, which can be trained independently or co-trained with the convolutional neural network model. To further improve the feature extraction ability of the photon convolutional neural network system, it is preferably co-trained with the convolutional neural network model. In an alternative embodiment, the training methods of the coefficient generation model and the size selection model corresponding to each dynamic convolution kernel module include: Introducing the above coefficient generation model and size selection model during the training process of the convolutional neural network model, the coefficient generation models corresponding to each dynamic convolution kernel module are respectively in one-to-one correspondence with each convolutional layer in the convolutional neural network model; the size selection models corresponding to each dynamic convolution kernel module are respectively in one-to-one correspondence with each convolutional layer in the convolutional neural network model; During the forward propagation process, before each convolutional layer performs the convolution operation, the corresponding coefficient generation model generates dynamic coefficients corresponding to the weights of each convolutional kernel in this convolutional layer based on the data input to this convolutional layer, and dynamically adjusts the corresponding convolutional kernel weights in this convolutional layer based on the obtained dynamic coefficients; the corresponding size selection model generates the optimal size of each convolutional kernel in this convolutional layer based on the data input to this convolutional layer, so as to dynamically adjust the sizes of each convolutional kernel in this convolutional layer. During the backpropagation process, the parameters in the convolutional neural network model, each coefficient generation model, and each size selection model are adjusted simultaneously.
[0046] In an alternative embodiment, the above-mentioned photonic convolutional neural network system further includes: a light source module, which is used to load the data to be currently convolved onto an optical signal by using an optical signal modulator to obtain a continuous signal light carrying the data information to be currently convolved.
[0047] It should be noted that the optical signal modulator can precisely control the optical signal based on different physical mechanisms (such as the electro-optic effect, acousto-optic effect, etc.). Its types include but are not limited to: visible light modulator, electro-optic Mach-Zehnder modulator, micro-ring resonator, acousto-optic modulator, etc. The present invention does not limit this.
[0048] In an alternative embodiment, the above-mentioned photonic convolutional neural network system further includes: a photoelectric conversion module, a non-linear activation module, and a fully connected layer module; The photoelectric conversion module is used to convert the signal light output by the dynamic convolutional kernel module into an electrical signal and output it to the non-linear activation module; The non-linear activation module is used to perform non-linear activation processing on the electrical signal input by the photoelectric conversion module; The fully connected layer module is used to calculate the non-linear activation processing result output by the non-linear activation module for the last time to obtain the final output result.
[0049] It should be noted that the working principle of the photoelectric conversion module is based on the photoelectric effect. When light irradiates a semiconductor material, photons hit the electrons in the semiconductor, causing them to undergo transitions. After transitioning to the conduction band, electron-hole pairs are formed. These carrier pairs form a directional current under an applied bias voltage. The types of photodetectors in the photoelectric conversion module include but are not limited to: photodiodes, phototransistors, avalanche photodiodes, etc. The present invention does not limit this.
[0050] The non-linear activation module and the fully connected layer module can be existing hardware modules or software modules implemented on a host computer, which is not limited here. The non-linear activation module can be a ReLU activation module, a Sigmoid activation module, a tanh activation module, etc., which is not limited here.
[0051] In summary, the photon dynamic convolutional neural network system based on optical phase change materials provided by the present invention combines non-volatile modulation and volatile modulation methods, and constructs a coefficient generation model, a size selection model, and a dynamic convolution module. It can not only give full play to the non-volatile advantage of constructing photon convolution kernels with phase change materials, but also effectively improve the flexibility of photon convolution kernels. It can adaptively generate photon convolution kernels in a non-linear manner according to the content of the input data, overcoming the drawback that traditional static convolution kernels cannot adapt to changing inputs, and realizing flexible input-adaptive dynamic convolution processing. In addition, when performing convolution operations, since the photon convolution kernel does not solely rely on the non-volatile phase state of the phase change material for encoding, but introduces volatile modulation encoding on this basis, greatly enriching the encoding expression ability of the photon convolution kernel, it can effectively alleviate the problem that traditional photon convolutional neural network systems are severely limited by the number of non-volatile phase states of phase change materials. The dynamic adjustment of the convolution kernel size also improves the ability of the photon convolutional neural network system to recognize and express multi-scale features, retaining data information at different levels to improve the robustness of the network. The photon convolutional neural network system proposed in this application is also very efficient in computing. It can improve the network complexity without increasing the network depth or width, which is beneficial to improving the comprehensive performance of the photon convolutional neural network system, especially the classification accuracy of the image dataset in the image classification task, and at the same time helps to promote the application of the photon convolutional neural network system to larger-scale and more complex task scenarios.
[0052] To further illustrate the photon convolutional neural network system provided by the present invention, a specific embodiment is described in detail below: As Figure 1 shown, the present embodiment provides a photon convolutional neural network system for implementing the functions of a convolutional neural network model; the convolutional neural network model includes: a plurality of cascaded convolutional layers, a non-linear activation layer provided between two adjacent convolutional layers, and a fully connected layer connected after the last convolutional layer.
[0053] The photon convolutional neural network system provided by the present embodiment includes: a computer module, a light source module, a dynamic convolution module, and a photoelectric conversion module; The computer module includes: a training module, an input control module, P coefficient generation models, P size selection models, a coefficient cache module, a size cache module, a non-linear activation module, and a fully connected layer module; P is the number of convolutional layers in the convolutional neural network model; one convolutional layer in the convolutional neural network model corresponds to one coefficient generation model and also corresponds to one size selection model. Denote the coefficient generation model corresponding to the i-th convolutional layer as the i-th coefficient generation model; denote the size selection model corresponding to the i-th convolutional layer as the i-th size selection model; i = 1, 2, 3,..., P.
[0054] The training module is used to perform end-to-end training on the convolutional neural network model, each coefficient generation model, and each size selection model using a training set for performing corresponding tasks. Specifically, in the process of forward propagation, before each convolutional layer performs a convolutional operation, the corresponding coefficient generation model generates dynamic coefficients corresponding to the weights of each convolutional kernel in the convolutional layer based on the data input to the convolutional layer, and dynamically adjusts the corresponding convolutional kernel weights in the convolutional layer based on the obtained dynamic coefficients; the corresponding size selection model generates the optimal size of each convolutional kernel in the convolutional layer based on the data input to the convolutional layer to dynamically adjust the sizes of the convolutional kernels in the convolutional layer; in the process of backpropagation, the parameters in the convolutional neural network model, each coefficient generation model, and each size selection model are adjusted simultaneously. After the training is completed, the convolutional kernel weights in each convolutional layer of the convolutional neural network model are stored, and the weight parameters in the last fully connected layer of the stored convolutional neural network model are loaded into the fully connected layer module.
[0055] There are also P dynamic convolution modules, which correspond one by one to each convolutional layer in the convolutional neural network model; the dynamic convolution module corresponding to the i-th convolutional layer is denoted as the i-th dynamic convolution module; there is a one-to-one correspondence between the i-th dynamic convolution module, the i-th coefficient generation model, and the i-th size selection model. The dynamic convolution kernel module includes: a photonic synaptic device array; the pre-trained convolutional kernel weights in the corresponding convolutional layer are stored in the array; among them, the convolutional kernel weights are encoded as the corresponding non-volatile phase states of the photonic synaptic devices through a non-volatile modulation method to achieve storage; one photonic synaptic device corresponds to storing one convolutional kernel weight; the light transmittance of the photonic synaptic devices in different phase states is different.
[0056] In this embodiment, the data to be convolved is taken as image data as an example, such as Figure 2 The structure diagram of the dynamic convolution module is shown. For a certain dynamic convolutional layer with weights [C1×C2×k×k], there are C1 fixed convolutional kernels (pre-trained convolutional kernels) of size [C2×k i ×k j . The coefficient generation model generates a set of dynamic coefficients according to the input features, and dynamically modulates the weights with different levels or degrees of volatile modulation according to the correspondence between the coefficients and the volatile modulation excitation to enhance or suppress specific input-related features. The size selection model generates a set of size parameters according to the input information, and isolates part of the convolutional kernel weights by encoding them into the fully crystalline state level of high light attenuation to change the convolutional kernel size and enhance the recognition of features at different scales. Among them, the correspondence between the dynamic coefficients and the volatile modulation excitation can be one or a combination of linear correspondence, exponential correspondence, and logarithmic correspondence.
[0057] During the network inference process, for each input image, a set of static convolution kernel coefficients is generated by the coefficient generation module. Through the excitation correspondence relationship, the dynamic convolution module storing static convolution kernels is regulated in a volatile modulation manner, superimposing the volatile modulation effect on the basis of non-volatile phase modulation to achieve dynamic adjustment of the convolution kernel weights for this image. For each input image, a set of static convolution kernel size parameters can be generated by the size selection module. Through the association relationship, the static convolution kernels of the dynamic convolution module are modulated in a non-volatile modulation manner, encoding the phase change optical devices that do not need to play a convolution role into a high light attenuation level to achieve dynamic adjustment of the convolution kernel size for this image. Among them, the processing time of each input image matches the time when the volatile modulation takes effect. After each input image completes the dynamic convolution operation, the volatile modulation part of the dynamic convolution module will quickly fail, the part encoded as the high light attenuation level will be reset to the initial non-volatile value, and the non-volatile value will remain unchanged until the next input image enters.
[0058] In this embodiment, the dynamic convolution module has multiple photon synaptic devices based on phase change materials, which correspond one by one to the convolution kernel weights of the corresponding convolution layer. Taking a phase change photon synaptic device based on electrical tuning as an example, its device structure is as Figure 3 shown, specifically including: a substrate, a waveguide layer, a phase change material layer, a heating layer, a covering layer, and an electrode layer acting on the heating layer, which are arranged in sequence from bottom to top. The electrode layer is connected to an external power supply and applies an electrical pulse signal to the heating layer to adjust the temperature of the heating layer. When light travels in the waveguide layer, it is controlled by both the waveguide layer and the phase change material layer. Among them, the size of the heating layer of the phase change optical synaptic device should ensure that the volatile modulation effect is obvious while not affecting the long-term retention of the non-volatile state.
[0059] The encoding methods of the photon synaptic devices in the dynamic convolution module include: non-volatile modulation method and volatile modulation method. Among them, the non-volatile modulation method is used to encode the static convolution kernels (pre-trained convolution kernels) of the dynamic convolution module. By controlling the temperature of the heating layer, the modulation of the phase change material layer is achieved, and the static convolution kernel parameters obtained by the computer module training are stored in the non-volatile phase state of the phase change material. The modulation of the light transmittance by its phase state is the non-volatile modulation generated for the waveguide layer; the volatile modulation method is used to encode the part adapted to the input of the dynamic convolution module. Through one of the thermo-optical effect or carrier dispersion effect, based on the static convolution kernel coefficients generated by the coefficient generation module, different degrees of volatile modulation are performed on the waveguide layer, and then volatile modulation is generated for the transmitted light of the waveguide layer.
[0060] The phase-change photon synaptic device of the dynamic convolution module (photon synaptic device based on phase-change material) has x programmable non-volatile phase levels for transmitting light, which are used to construct static convolution kernels; each non-volatile phase has y volatile modulation levels for dynamically modulating the static convolution kernels; in addition, it also has a fully crystalline state level with high light attenuation.
[0061] The light transmission transmittance of the fully crystalline state level with high light attenuation is less than -60 dB, and the light transmission transmittances of the x non-volatile phase levels and the y volatile modulation levels of each non-volatile phase should be greater than -40 dB. The formula for calculating the light transmission transmittance is: dB = 10lg(P1 / P0); where P0 and P1 are the light power magnitudes before and after passing through the phase-change photon synaptic device respectively. According to the above formula, when the light transmission transmittance is less than -60 dB, it can be considered that the light transmission is completely blocked; when the light transmission transmittances differ by 20 dB, it can be considered that there is an obvious light transmission difference.
[0062] The transmittances of the x non-volatile phase levels of the phase-change photon synaptic device and the y volatile modulation levels of each non-volatile phase should be normalized to [W min , W max , where the minimum weight W min corresponds to the lowest device transmittance, and the maximum weight W max corresponds to the highest device transmittance. The other device transmittance levels correspond to different weight magnitudes within this weight range, and then they are substituted into the computer module for training.
[0063] The selection result of the size selection module is associated with the fully crystalline state level with high light attenuation. The phase-change photon synaptic devices corresponding to the positions of the convolution kernels that do not need to play a convolution role are encoded as the fully crystalline state level with high light attenuation.
[0064] As Figure 4 and Figure 5 shown are respectively the non-volatile modulation performance diagram and the volatile modulation performance diagram of a phase-change photon synaptic device. It can be seen from the figure that under the action of different electrical modulation excitations, the phase-change photon synaptic device can achieve non-volatile regulation and volatile regulation, and is used in the dynamic convolution module of the photon convolutional neural network system. As Figure 4 shown, by applying a voltage pulse on the metal electrode and controlling the ratio of the crystalline phase to the amorphous phase of the phase-change material layer by modulating the amplitude and / or pulse width of the voltage pulse, non-volatile multi-state regulation of the phase-change photon synaptic device can be achieved. It should be noted that by regulating the amplitude and / or pulse width of the voltage pulse applied to the metal electrode, the applied energy can be precisely controlled, and the precise encoding of the non-volatile phase state of the phase-change material can be controlled. As Figure 5As shown, when the phase change material layer is in a certain non-volatile phase state, by applying a lower voltage pulse (below the energy threshold for inducing non-volatile phase change) to the metal electrode, modulating the amplitude and / or pulse width of the voltage pulse can control the degree of volatile modulation such as the thermo-optic effect in the waveguide layer, realizing different levels of volatile modulation at a certain non-volatile phase state level.
[0065] The following is the complete working process of the photon convolution neural network system: In the inference stage, the operation of inputting the original data into the convolution neural network model for calculation is realized; the specific process includes: Performing a convolution operation corresponding to the first convolution layer of the convolution neural network model: The input control module outputs the original data as the data to be currently convolved to the light source module, the first coefficient generation model, and the first size selection module respectively; After receiving the data to be convolved, the first coefficient generation model generates the dynamic coefficients corresponding to the convolution kernel weights stored in the array in the first dynamic convolution module, and outputs them to the first dynamic convolution module through the coefficient cache module; After receiving the data to be convolved, the first size selection module generates the optimal size of each convolution kernel stored in the array in the first dynamic convolution module, and outputs them to the first dynamic convolution module through the size cache module; The first dynamic convolution module is used to dynamically adjust the convolution kernel weights stored in the array before performing the convolution operation: receiving the dynamic coefficients corresponding to the convolution kernel weights stored in the array in the first dynamic convolution module input by the first coefficient generation module, and applying pulse signals corresponding to the corresponding dynamic coefficients to each photon synaptic device for volatile modulation; dynamically adjusting the size of each convolution kernel stored in the array: receiving the optimal size of each convolution kernel stored in the array in the first dynamic convolution module input by the first size selection module; selecting several photon synaptic devices from each column of the array, keeping the phase states of the selected photon synaptic devices unchanged, and modulating the phase states of the unselected photon synaptic devices to the fully crystalline state by non-volatile modulation to adjust the size of the convolution kernel stored in this column to the corresponding optimal size; The light source module is used to, after receiving the data to be convolved, load the data to be currently convolved onto the optical signal by using an optical signal modulator to obtain a continuous signal light carrying the data information to be currently convolved, and output it to the first dynamic convolution module; The first dynamic convolution module is also used to receive, during the convolution operation, the continuous signal light carrying the data information to be currently convolved, which is input by the light source module. The continuous signal light is attenuated under the action of the array, and the attenuated signal light is output and sent to the photoelectric conversion module; the light intensity of the output signal light is the current convolution operation result. The photoelectric conversion module is used to convert the signal light carrying the convolution operation result information output by the first dynamic convolution module into an electrical signal, obtain the convolution operation result of the current layer, and output it to the non-linear activation module. The non-linear activation module is used to perform non-linear activation processing on the convolution operation result input by the photoelectric conversion module, and output the result after non-linear processing to the input control module; among them, the non-linear activation module performs non-linear activation processing on the electrical signal transmitted by the photoelectric conversion module through any one of the following functions: Sigmoid function, tanh function, ReLU function, LeakyReLU function.
[0066] The input control module outputs the result after non-linear activation processing input by the non-linear activation module as the data to be currently convolved to the light source module, the second coefficient generation module, and the second size selection module respectively, to perform the convolution operation corresponding to the second convolution layer of the convolutional neural network model, and so on. After the convolution operation of the P-th convolution layer is completed, the non-linear activation module outputs its current result after non-linear activation processing to the fully connected layer module to obtain the final calculation result.
[0067] It should be noted that the above computer modules can be a computer, a server, a workstation, etc. The computer modules perform simulation through the simulation system installed on them, and use the pre-collected data set to train the convolutional neural network model, the coefficient generation module, and the size selection module. In some embodiments, the computer module can implement the above functions using the Python programming language, and control the modulation and weighting process of the entire photon convolutional neural network system and the operation logic of each module through Python scripts or programs. The Python programming language provides rich interfaces and libraries to interact with hardware devices. In addition to the Python programming language, other methods and tools can also be used to control the modulation and weighting of signals and the operation logic of each module, including but not limited to: MATLAB, Simulink, C / C++ programming, application-specific integrated circuit (ASIC), etc. The embodiments of the present application do not make any limitations in this regard.
[0068] To further illustrate the construction process of the photon convolutional neural network system provided by the present invention, the following will be described in detail with a specific embodiment: In this embodiment, the step process of the construction method of the photon convolution neural network system based on the optical phase change material mainly includes steps S1 to S5, and the following is a detailed introduction to each step.
[0069] Step S1: Through the vertical grating coupling test platform, obtain all non-volatile encodable phase levels of light that the phase change optical synapse device can transmit, all volatile modulation levels of each phase, and a fully crystalline level with high light attenuation. Map all device transmittance levels (except the fully crystalline level) into the weight range required by the convolution neural network model.
[0070] Step S2: Construct a convolution neural network model, a coefficient generation model, and the size selection model in the computer module; according to all available weight values obtained in step S1, input the training data set into the computer module for training. The parameters to be trained include the parameters of the coefficient generation model, the parameters of the size selection model, the correspondence between the coefficient generation result and the volatile modulation excitation, the static convolution kernel value of the dynamic convolution module, and other parameters of the convolution neural network model; adopt the quantization training method for training, use optimization algorithms such as gradient descent, and continuously adjust and optimize the weights in the network through the forward propagation and backward propagation processes to minimize the loss function; co-train all parameters, and the weight value of the static convolution kernel is selected from the non-volatile phase levels.
[0071] Step S3: After the training is completed, determine the number of phase change optical synapse devices in the dynamic convolution module according to S2, and use the phase change optical synapse device to realize the convolution operation weight encoding in the convolution neural network model. One phase change optical synapse device corresponds to one convolution kernel weight; load the static convolution kernel weight value obtained in step S2 into the phase change material unit of the dynamic convolution module in a non-volatile modulation manner such as electro-controlled heating phase change, and encode the weight value into the corresponding non-volatile phase of the corresponding phase change optical synapse device, and characterize the weight size by the light transmittance.
[0072] Step S4: Determine the number of optical signals and the corresponding hardware quantity of the light source module and the photoelectric conversion module, and then determine the number of multiplexers and balanced photodetectors; the light source module simultaneously outputs m optical signals encoded with input data information, where m depends on the scale of the convolution kernel. Specifically, when the convolution kernel size is [3×3], m = 9; the photoelectric conversion module simultaneously performs photoelectric conversion on p optical signals, where p depends on the number of convolution kernels; then determine the hardware quantity required for the light source module and the photoelectric conversion module.
[0073] Step S5: Deploy the corresponding light source module and photoelectric conversion module, map the neural network onto the chip, establish the volatile modulation relationship between the coefficient generation model and the dynamic convolution module, establish the high-light attenuation modulation relationship between the size selection model and the dynamic convolution module, and set the control logic of the computer module to form a photon convolution neural network system.
[0074] In summary, the present invention combines the non-volatile modulation and volatile modulation of the phase state of the phase change material to jointly control the transmission of light in the waveguide and construct a dynamic photon convolution kernel; through the cooperative modulation of the two modulation methods, the flexibility of the phase change photon convolution kernel is improved, enabling the photon convolution neural network system to dynamically adjust according to different inputs, which helps to promote the application of the photon convolution neural network system in larger-scale and more complex task scenarios and the improvement of performance.
[0075] The above embodiments only represent several implementation manners of the present invention, and their descriptions are relatively specific and detailed, but should not be construed as limiting the scope of the patent application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention.
[0076] It is easy for those skilled in the art to understand that the above is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. A photon convolutional neural network system, characterized in that Comprising: Each dynamic convolution kernel module corresponding one by one to each convolution layer in the pre-trained convolutional neural network model; The dynamic convolution kernel module includes: a photon synaptic device array; the array stores the pre-trained convolution kernel weights in the corresponding convolution layer; wherein, the photon synaptic devices in the array are photon synaptic devices based on phase change materials; the convolution kernel weights are encoded into the corresponding non-volatile phase states of the photon synaptic devices by a non-volatile modulation method, so as to achieve storage; one photon synaptic device corresponds to storing one convolution kernel weight; the light transmittance of the photon synaptic devices in different phase states is different; The dynamic convolution kernel module is used to dynamically adjust the convolution kernel weights stored in the array before performing the convolution operation: receive the dynamic coefficients corresponding to the convolution kernel weights stored in the array, and apply a pulse signal corresponding to the corresponding dynamic coefficient to each photon synaptic device for volatile modulation; The dynamic convolution kernel module is also used to, when performing the convolution operation, receive the continuous signal light carrying the data information to be convolved currently, the continuous signal light is attenuated under the action of the array, and output the attenuated signal light; the light intensity of the output signal light is the current convolution operation result; Wherein, the dynamic coefficients are generated by inputting the data to be convolved into the corresponding pre-trained coefficient generation model; one dynamic convolution kernel module corresponds to one coefficient generation model, and the coefficient generation model is a deep learning model.
2. The photon convolution neural network system according to claim 1, wherein The coefficient generation model includes: a plurality of cascaded fully connected layers, an activation layer arranged between adjacent two fully connected layers, and a pooling layer connected before the first fully connected layer.
3. The photon convolution neural network system according to claim 1, characterized in that, The training method of the coefficient generation model corresponding to each dynamic convolution kernel module includes: Introducing the above coefficient generation model during the training process of the convolutional neural network model, the coefficient generation models corresponding to each dynamic convolution kernel module respectively correspond one by one to each convolution layer in the convolutional neural network model; during the forward propagation process, before each convolution layer performs the convolution operation, the corresponding coefficient generation model generates the dynamic coefficients corresponding to the convolution kernel weights in this convolution layer based on the data input to this convolution layer, and dynamically adjusts the corresponding convolution kernel weights in this convolution layer based on the obtained dynamic coefficients; during the backward propagation process, the parameters in the convolutional neural network model and each coefficient generation model are adjusted simultaneously.
4. The photon convolution neural network system according to any one of claims 1-3, characterized in that The weights in the same convolution kernel are stored in the same column of the array; The dynamic convolution kernel module is also used to, before performing the convolution operation and after dynamically adjusting the convolution kernel weights stored in the array, dynamically adjust the sizes of the convolution kernels stored in the array: select several photon synaptic devices from each column of the array, keep the phase states of the selected photon synaptic devices unchanged, and modulate the phase states of the unselected photon synaptic devices to the fully crystalline state by a non-volatile modulation method, so as to adjust the size of the convolution kernel stored in this column to the corresponding optimal size; Wherein, one dynamic convolution kernel module also corresponds to one size selection model, and the size selection model is a deep learning model; The optimal size of each convolution kernel stored in the array is generated by inputting the data to be convolved into the corresponding size selection module; the size selection model is used to extract the features of the data to be convolved and map them to the optimal sizes of each convolution kernel stored in the array within the corresponding dynamic convolution kernel module.
5. The photon convolution neural network system according to claim 4, wherein The size selection model includes: a cascaded feature extraction module and a classifier; wherein, the feature extraction module includes one or more of: an edge feature extraction unit, a texture feature extraction unit, and a smoothness feature extraction unit, and is used to extract one or more of the edge features, texture features, and smoothness features of the data to be convolved.
6. The photon convolution neural network system according to claim 5, characterized in that, The classifier is: a fully connected conditional network; wherein, the fully connected conditional network includes: a plurality of cascaded fully connected layers, an activation layer arranged between adjacent two fully connected layers, and a softmax layer connected after the last fully connected layer.
7. The photon convolution neural network system according to claim 4, characterized in that, The training method of the coefficient generation model and the size selection model corresponding to each dynamic convolution kernel module includes: During the training process of the convolutional neural network model, the coefficient generation model and the size selection model are introduced. The coefficient generation models corresponding to each dynamic convolution kernel module are respectively in one-to-one correspondence with each convolutional layer in the convolutional neural network model; the size selection models corresponding to each dynamic convolution kernel module are respectively in one-to-one correspondence with each convolutional layer in the convolutional neural network model; During the forward propagation process, before each convolutional layer performs a convolution operation, the corresponding coefficient generation model generates the dynamic coefficients corresponding to the weights of each convolution kernel in this convolutional layer based on the data input to this convolutional layer, and dynamically adjusts the corresponding convolution kernel weights in this convolutional layer based on the obtained dynamic coefficients; the corresponding size selection model generates the optimal sizes of each convolution kernel in this convolutional layer based on the data input to this convolutional layer to dynamically adjust the sizes of each convolution kernel in this convolutional layer; During the backpropagation process, the parameters in the convolutional neural network model, each coefficient generation model, and each size selection model are adjusted simultaneously.
8. The photon convolution neural network system according to any one of claims 1-3, characterized in that, It further includes: A light source module, which is used to load the data to be convolved currently onto an optical signal by using an optical signal modulator to obtain a continuous signal light carrying the data information to be convolved currently.
9. The photon convolution neural network system according to any one of claims 1-3, characterized in that, It further includes: A photoelectric conversion module, a non-linear activation module, and a fully connected layer module; The photoelectric conversion module is used to convert the signal light output by the dynamic convolution kernel module into an electrical signal and output it to the non-linear activation module; The non-linear activation module is used to perform non-linear activation processing on the electrical signal input by the photoelectric conversion module; The fully connected layer module is used to calculate the non-linear activation processing result output by the non-linear activation module for the last time to obtain the final output result.
10. The photon convolution neural network system according to any one of claims 1-3, characterized in that, The photon synaptic device includes: a substrate, a waveguide layer, a phase change material layer, a heating layer, a covering layer, and an electrode layer acting on the heating layer arranged successively from bottom to top.
Citation Information
Patent Citations
Data processing method based on photonic neural network chip and related device or equipment
CN111008982A
Photon convolutional neural network accelerator based on micro-ring resonator and nonvolatile phase change material
CN113657580A
Photon convolution accelerator based on mode multiplexing
CN114819089A
Photon nerve synaptic device with double-micro-ring structure and convolution operation network model
CN116011538A
Photon pulse neural network implementation method based on MRR and phase change material
CN116029343A
Cited By
Adjustable optical convolution system and method based on semiconductor laser
CN120745707A