Image classification method and system based on photoelectric hybrid computing architecture
By adopting the photoelectric hybrid computing architecture in image classification, using optical diffraction network to extract image features and combining electronic neural networks for deep feature learning, the high computational volume and high energy consumption problems of electronic computing architecture in deep neural network computing are solved, and high-efficiency and low-energy image classification is achieved.
Patent Information
- Application Number
- CN202510146465.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-05-16
AI Technical Summary
The existing electronic computing architectures have problems such as high computing volume, high energy consumption and limited computing speed in deep neural network computing, which is difficult to meet the needs of large-scale data processing.
The image classification method based on the photoelectric hybrid computing architecture is adopted to perform optical diffraction calculation on the upsampled images through the optical diffraction network, extract optical features, and convert them into electrical signals to input into electronic neural network for feature extraction and classification.
Partial feature extraction is achieved through optical diffraction calculation, which reduces the computational burden of electronic computing, improves calculation efficiency, reduces energy consumption, and maintains high classification accuracy.
Smart Images

Figure CN120014357A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision and artificial intelligence technology, and specifically to an image classification method and system based on an optoelectronic hybrid computing architecture. Background Art
[0002] Image classification is a technology based on image feature recognition. It uses computer vision and machine learning methods to classify input images into different categories or labels. It has broad application prospects in many fields such as security monitoring, autonomous driving, and financial payment. With the development of deep learning, convolutional neural network (CNN) has become one of the mainstream image classification methods. Convolutional neural network extracts image features through multi-layer convolution operations and pooling operations, and uses fully connected layers for classification. It has adaptive learning ability, efficient feature extraction ability and strong generalization ability, which greatly improves the accuracy and computational efficiency of image classification.
[0003] In the development of convolutional neural networks, researchers have proposed a variety of improved structures to improve classification performance. For example, the residual network (ResNet) alleviates the gradient vanishing and degradation problems in deep network training by introducing residual learning and jump connections, allowing deeper networks to be trained stably and show excellent performance in multiple computer vision tasks. However, as the depth of the network increases, the computational complexity increases significantly, requiring a large amount of computing resources to be consumed during training and inference, resulting in longer computing time and increased energy consumption, which limits its application in resource-constrained scenarios.
[0004] At present, deep learning models mainly rely on electronic computing chips for calculations, but the performance of electronic chips is limited by the slowdown of Moore's Law. Traditional electronic computing architectures are gradually encountering bottlenecks in computing speed, power consumption and bandwidth, and it is difficult to meet the needs of large-scale data processing. Therefore, how to reduce the computing burden and improve computing efficiency while ensuring computing accuracy has become an urgent problem to be solved in the field of deep learning.
[0005] As a new type of computing hardware based on the principle of photonics, optical computing chips provide a new solution to break through the bottleneck of electronic computing. Optical computing chips use photons as information carriers and perform calculations through optical devices (such as optical modulators, photodetectors, light sources, etc.). Compared with traditional electronic computing methods, optical computing has the advantages of high-speed transmission, low energy consumption, high bandwidth and strong anti-interference ability. In recent years, optical computing methods based on diffractive neural networks (DNNs) have received widespread attention. Diffractive neural networks simulate the information processing process of biological neural networks through the interference and diffraction effects of light to achieve energy-free parallel computing. This method uses optical elements (such as lenses, diffraction gratings, etc.) to perform mathematical operations such as Fourier transform and convolution on the input light field, and finally converts the calculation results into electrical signals through photodetectors. However, all-optical computing systems still face many challenges in practical applications, such as lack of nonlinear computing capabilities and limited storage units, which make them have great limitations when used independently for complex deep learning tasks.
[0006] Based on this, researchers began to explore hybrid computing architectures that combine optical computing with electronic computing to fully utilize the high speed and low energy consumption advantages of optical computing, while combining the high precision and storage capacity of electronic computing to break through the performance bottleneck of traditional computing architectures. Optical-electrical hybrid computing architectures have gradually become an important direction for the development of computing hardware, and have shown good application potential in fields such as computer vision and pattern recognition. Summary of the invention
[0007] In view of the shortcomings of the prior art, the present invention provides an image classification method and system based on an optoelectronic hybrid computing architecture, which solves the problems of high computing volume, high energy consumption and limited computing speed in deep neural network calculations of the existing electronic computing architecture.
[0008] To achieve the above objectives, the present invention is implemented by the following technical solutions: an image classification method based on an optoelectronic hybrid computing architecture comprises the following steps: Upsampling the input image to meet the input requirements of the optical diffraction network; Use an optical diffraction network to perform optical diffraction calculation on the upsampled image to obtain optical features; Convert optical features into electrical signals and input them into an electronic neural network for feature extraction; Convolutional neural network is used in electronic neural network for feature extraction and output image classification results.
[0009] Preferably, the upsampling operation adopts an upsampling method based on "area" interpolation, and the upsampling method calculates the neighborhood average value of each pixel by performing regional interpolation on the input image, thereby generating an upsampled image that matches the size of the original image.
[0010] Preferably, the step of optical diffraction calculation includes: Perform Fourier transform on the upsampled image to convert the image from spatial domain to frequency domain representation; Applying the angular spectrum transfer function to the frequency domain information, modulating the phase and amplitude of the frequency domain signal through an element-by-element multiplication operation; The modulated frequency domain signal is inversely Fourier transformed to restore the frequency domain signal to the image features in the spatial domain.
[0011] Preferably, the angular spectrum transfer function is a preset angular spectrum filter function, which weights different amplitude and phase information according to the frequency domain characteristics of the input image, adjusts the propagation characteristics of the frequency domain signal, thereby enhancing or suppressing specific frequency components to obtain a modulated frequency domain signal.
[0012] Preferably, the electronic neural network uses a deep convolutional neural network ResNet50 for feature extraction.
[0013] Preferably, the convolutional neural network ResNet50 includes multiple residual modules, each residual module consists of at least two convolutional layers, and adds the input signal to the output signal through a jump connection.
[0014] Preferably, the method further comprises: calculating the loss based on the image classification result and performing model training.
[0015] Preferably, the step of calculating the loss adopts a cross entropy loss function, which calculates the loss value based on the difference between the predicted probability of the image classification result and the true label, wherein the predicted probability is obtained by performing softmax activation on the network output, and the cross entropy loss function takes the logarithm of the predicted probability of each category and multiplies it with the true label, and finally calculates the total loss of the entire data set.
[0016] Preferably, the loss is calculated based on the image classification result and the model training is performed, specifically comprising the following steps: Compare the predicted category output by the electronic neural network with the true label and calculate the cross entropy loss; Calculate the gradient based on the loss value through the back-propagation algorithm; The network parameters are updated using an optimization algorithm to minimize the loss function, thereby optimizing the image classification model.
[0017] The present invention also provides an image classification system based on an optoelectronic hybrid computing architecture, comprising: A data input module, used for receiving an input image and performing an upsampling operation on the input image to meet the input requirements of the optical diffraction network; An optical diffraction calculation module is used to perform optical diffraction calculation on the upsampled image to obtain optical features; Photoelectric conversion module, used to convert optical features into electrical signals and transmit them to the electronic neural network; The electronic neural network module is used to receive the electrical signal after photoelectric conversion, use the convolutional neural network to extract features, and output the image classification result; The loss calculation and optimization module is used to calculate the loss according to the image classification results and update the parameters of the electronic neural network through the optimization algorithm, thereby optimizing the image classification model.
[0018] The present invention provides an image classification method and system based on an optoelectronic hybrid computing architecture, which has the following beneficial effects: 1. The present invention realizes partial feature extraction through optical diffraction calculation, which reduces the computational burden of a large number of matrix operations compared to pure electronic computing architecture. Since optical computing has strong parallel processing capabilities and does not require additional energy consumption, the overall computing efficiency of the system is improved, while the power consumption of the electronic computing part is reduced, making this method have a great advantage in high-performance computing tasks.
[0019] 2. The optical diffraction calculation of the present invention can extract high-dimensional spatial information and combine with electronic neural networks for deep feature learning, so that the image classification task has both the high-resolution characteristics of optical calculation and the strong expression ability of deep learning.
[0020] 3. Traditional electronic computing is easily limited by storage bandwidth and computing resources when processing high-dimensional, large-scale data. This invention innovatively introduces optical computing as a preprocessing method, reducing the amount of computation in the electronic computing part, thereby breaking through the computing bottleneck of the traditional electronic computing architecture and providing a new solution for large-scale image classification tasks.
[0021] 4. The optoelectronic hybrid computing architecture of the present invention can be seamlessly connected to the existing deep learning framework. The electronic neural network part can adopt the standard convolutional neural network (CNN) architecture and be trained in combination with mainstream optimization algorithms. This compatibility allows the system to be quickly integrated into the existing AI reasoning platform and is suitable for different application scenarios, such as medical image analysis, target detection, and remote sensing image classification. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 It is a schematic diagram of the method flow of the present invention; Figure 2It is a framework diagram of the optoelectronic hybrid neural network computing architecture system of the present invention; Figure 3 The photoelectric hybrid system flow chart of the present invention; Figure 4 This is the ResNet50 structure diagram of the present invention. DETAILED DESCRIPTION
[0023] The following will be combined with the drawings in the specification of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0024] Please see attached Figure 1 -Attached Figure 4 The present invention provides an image classification method based on an optoelectronic hybrid computing architecture, which combines the advantages of optical diffraction calculation and electronic neural network. By using the optoelectronic hybrid computing architecture, the calculation speed of image classification is improved, the energy consumption is reduced, and a high classification accuracy is maintained.
[0025] In the optical part, optical diffraction calculation is used. In the electrical part, the residual network (ResNet) is used to implement the forward reasoning process and reverse gradient loss calculation process of the convolutional neural network. The optical-electrical hybrid neural network computing architecture system framework is as follows Figure 2 shown.
[0026] First, the image and coherent light source are input into the optical modulator as electrical signals and optical signals respectively. The optical modulator will change its physical properties according to the image information and encode the information of the input image in the phase and amplitude channels of the input light field. Secondly, the modulated light wave is propagated through the various diffraction layers of the optical diffraction neural network and receives phase and amplitude modulation from different neurons in each diffraction layer. The optical diffraction calculation is mainly composed of 4 identical diffraction layers, and a single diffraction layer has 2 Fourier lenses. and , respectively modulate the amplitude and phase of the light wave, which corresponds to Fourier change and inverse Fourier change in mathematics. Finally, the optical signal is converted into an electrical signal through a photodetector and a power amplifier and input into the electrical neural network to further extract image features and output classification results.
[0027] like Figure 1 As shown, the image classification method based on the optoelectronic hybrid computing architecture may include the following steps: S1, upsampling the input image; S2, using an optical diffraction network to perform optical diffraction calculation on the upsampled image to obtain optical features; S3, converting the optical features into electrical signals and inputting them into the electronic neural network for feature extraction; S4. Use convolutional neural network to extract features in electronic neural network and output image classification results.
[0028] The principles, technical contents and implementation details of each step of the present invention will be described in detail below.
[0029] For step S1, in this embodiment, the first step of the method is to upsample the input image so that it can meet the input requirements of the optical diffraction network. The input image and its label are represented as a data set ,in is a collection of images, is a set of labels. Specifically, for each image in the training dataset , all need to be resized through upsampling operations to ensure that they are suitable for the processing requirements of the subsequent optical diffraction network.
[0030] In this embodiment, the upsampling operation adopts an upsampling method based on "area" interpolation. This method calculates the neighborhood average value of each pixel through area interpolation, thereby generating an upsampled image that matches the size of the original image. The upsampled image is resized and can be directly used as the input of the optical diffraction network.
[0031] In terms of technical implementation, the upsampling operation includes the following steps: Get the input image .
[0032] Interpolate the image and use the "area" interpolation method to generate an upsampled image .
[0033] This operation performs regional interpolation on the pixel neighborhood of the input image, calculates the neighborhood average of each pixel, and generates an upsampled image that matches the size of the original image.
[0034] The mathematical formula for the upsampling operation is as follows:
[0035] in, The upsampled image has been resized to fit the requirements of the optical diffraction network.
[0036] The "Area" interpolation method is an upsampling method based on regional interpolation, which performs upsampling by weighted averaging each pixel and its neighboring pixels. Compared with simple copy extension (such as nearest neighbor interpolation), this method has a better smoothing effect, can better preserve the local features in the image, and avoid information loss or image blurring due to over-extension.
[0037] Specifically, regional interpolation calculates the weighted average of the neighboring pixels around each pixel in the input image to ensure that after the image is expanded, the new pixel value can reasonably reflect the image characteristics of the surrounding area. This method works better when the image size is increased, especially when it is greatly enlarged, and it can minimize image distortion during the interpolation process.
[0038] In an exemplary implementation, assuming that the input image The size of the target image is 128×128. After upsampling, the size of the target image becomes 256×256. During the upsampling process, the value of each pixel will be interpolated by the average value of the surrounding neighborhood pixels to generate the upsampled image. , whose size is 256 × 256. In this way, the image not only maintains the original structural characteristics, but also adapts to the input requirements of the optical diffraction network.
[0039] This upsampling step ensures that the input image can meet the size requirements of the optical diffraction calculation, providing good input data for the subsequent optical diffraction calculation process.
[0040] For step S2, in this embodiment, the second step of the method is to use the optical diffraction network to perform optical diffraction calculation on the upsampled image to extract the optical features of the image. The optical diffraction calculation utilizes the principle of the optical diffraction neural network, and through operations such as Fourier transform, angular spectrum transfer and inverse Fourier transform, the spatial features of the image are converted into frequency domain features, thereby extracting useful image information.
[0041] First, the upsampled image and the coherent light source are input into the optical modulator as electrical signals and optical signals respectively. The optical modulator changes its physical properties according to the image information and encodes the information of the input image in the phase and amplitude channels of the input light field. After this process, the optical modulator converts the input image features into the phase and amplitude information of the light wave to form a modulated light wave. Next, these modulated light waves will propagate through the various diffraction layers of the optical diffraction neural network and be further modulated in each layer.
[0042] Specifically, the optical diffraction network is mainly composed of four identical diffraction layers. In each diffraction layer, there are two Fourier lenses. and , which modulate the amplitude and phase of the light wave respectively. This modulation method corresponds to Fourier transform and inverse Fourier transform, so the spatial domain features of the image can be effectively extracted. The following are the detailed steps: 1. Fourier transform (FFT): for the upsampled image Perform Fourier transform. Fourier transform is a process of converting an image from the spatial domain to the frequency domain. It can convert the structural information of the image into frequency characteristics and reveal the spectral information of the image. The process is mathematically expressed as:
[0043] in, Represents the converted frequency domain information, which contains the frequency characteristics of the image.
[0044] This step converts the image from the spatial domain to the frequency domain, which allows the different frequency components of the image to be separated. In the frequency domain, high-frequency components usually represent the details of the image, while low-frequency components represent the overall structure of the image. Through Fourier transform, the frequency characteristics in the image can be clearly analyzed.
[0045] 2. Angular spectrum transfer (multiplication of angular spectrum): In the frequency domain, optical diffraction calculations are performed by applying the angular spectrum transfer function , for frequency domain signals Specifically, angular spectrum transfer is to perform the angular spectrum transfer operation by element-by-element multiplication. With frequency domain signal Multiply and adjust the phase and amplitude of the frequency domain signal. This process helps to enhance or suppress specific frequency components of the image and highlight the key information of the image. The mathematical expression is:
[0046] in, It is a frequency domain signal after angular spectrum transmission, with modulated amplitude and phase information.
[0047] The function of angular spectrum transfer is to modulate the frequency domain signal through the transfer function Different frequency components are weighted. In this way, the high-frequency information in the image can be enhanced to highlight the detailed features of the image, or unnecessary noise can be reduced by suppressing certain frequency components.
[0048] 3. Inverse Fourier Transform (IFFT): frequency domain signal after angular spectrum transfer It is input into the inverse Fourier transform module, and after the inverse Fourier transform, the frequency domain information is restored to the image features in the spatial domain. The inverse Fourier transform process is a mapping from the frequency domain to the spatial domain, and finally the spatial domain features of the image are obtained. , this feature can better represent the spatial structure of the image. Mathematically expressed as:
[0049] in, It is the restored spatial domain image feature, which provides spatial feature information related to the input image.
[0050] The inverse Fourier transform restores the frequency domain signal to the spatial domain signal, ensuring that the image features can eventually return to a form that can be used for subsequent processing. , we are able to obtain an efficient feature representation associated with the input image.
[0051] In an exemplary implementation, assuming that the input image The size is 256×256, and the frequency domain image after Fourier transformation It may contain components of different frequencies. Through the angular spectrum transfer operation, the high-frequency information (details) in the image can be amplified and the low-frequency part (overall shape) can be retained. Finally, through the inverse Fourier transform, the frequency domain information is restored to the new image features. , these features will be processed and classified in the subsequent electronic neural network.
[0052] The optical diffraction calculation step in this embodiment converts the spatial information of the image into frequency domain information through the three steps of Fourier transform, angular spectrum transfer and inverse Fourier transform, and extracts the effective features of the image by modulating and restoring the frequency domain signal. Through this optical diffraction calculation process, the key information of the image can be extracted, providing valuable feature data for subsequent electronic neural network processing.
[0053] For step S3, in this embodiment, after the optical diffraction calculation, the optical features need to be converted into electrical signals so that the subsequent electronic neural network can further extract features and classify them. The conversion process mainly relies on the cooperation of the photodetector and the power amplifier to achieve effective mapping of optical signals to electrical signals.
[0054] First, the light waves processed by the optical diffraction neural network carry the image feature information after multi-layer diffraction modulation, which exists in the form of light intensity distribution. In order to use these features for electronic neural network analysis, a photodetector (PD) is needed to convert them. The function of the photodetector is to receive the final diffracted light field and generate a corresponding electrical signal based on the intensity of the incident light. The working principle of the photodetector is based on the photoelectric effect, that is, when the incident photon acts on the photosensitive material, photoelectrons are generated, thereby forming a current signal.
[0055] In the photoelectric conversion process, the light intensity information of different pixel points will be mapped to different electrical signal values, forming a spatially corresponding electrical signal matrix. This conversion method can retain the features extracted by optical diffraction calculations and provide efficient input for subsequent electronic neural networks.
[0056] Secondly, the converted electrical signal is usually weak, and directly using it in an electronic neural network may cause the signal quality to deteriorate, so a power amplifier (PA) is needed to enhance the signal. The function of the power amplifier is to amplify the input weak electrical signal to ensure that the signal has a sufficient power level to adapt to the subsequent electronic neural network processing.
[0057] Through the above-mentioned photoelectric conversion process, this embodiment ensures that the features extracted by optical diffraction calculation can be effectively transmitted to the electronic neural network, while retaining key image information, providing a data basis for subsequent classification tasks.
[0058] For step S4, in this embodiment, after completing the conversion and amplification of the photoelectric signal, the obtained electrical signal is sent as input to the electronic neural network (Conv) to perform deep feature extraction and classification operations to achieve pattern recognition of the data after optical diffraction processing. This process includes feature extraction, loss calculation, and classification decision, ensuring that the information transmitted from the optical diffraction network can be further parsed and used for target classification tasks.
[0059] First, the data after diffraction processing Input into the convolutional neural network (Conv) to extract image features The specific process is as follows:
[0060] like Figure 4 As shown, in this step, the electronic neural network adopts the ResNet50 structure, which is mainly composed of multiple residual blocks, including convolutional layers, batch normalization layers (BatchNorm), ReLU activation functions and skip connections (Skip Connection), so that the network can effectively extract feature information at different levels.
[0061] For example, the input data After the first layer of 7×7 large convolution kernels, preliminary feature extraction is performed, and the computational complexity is reduced through the maximum pooling layer. Subsequently, the data enters the backbone network of ResNet50 and passes through multiple residual blocks for deep feature learning. Among them, the residual block is composed of multiple 1×1 and 3×3 convolution kernels to achieve efficient feature fusion.
[0062] After feature extraction is completed, the obtained features With the corresponding label The cross entropy loss (Loss) is calculated as follows:
[0063] in, Represents the total number of categories in the dataset, The loss function measures the deviation between the network prediction and the true label and is used to optimize the network parameters so that the model can continuously learn the optimal classification boundary during the training process.
[0064] Finally, in the classification layer, the network converts the extracted features into category probability distribution through the fully connected layer and normalizes them using the Softmax activation function to generate the final classification result. For the output layer, a 1000-dimensional fully connected layer can be used to adapt to the classification task of large-scale data sets.
[0065] Through the above steps, this embodiment can perform in-depth analysis on the features obtained by optical diffraction calculation based on the ResNet50 structure, and optimize the classification performance in combination with the cross entropy loss, thereby achieving accurate recognition and classification of the input image.
[0066] In general, the present invention first performs upsampling on the input image to adapt to the input requirements of the optical diffraction system; then, the image information is encoded into the phase and amplitude channels of the coherent light field using an optical modulator, and propagated and modulated through a multi-layer optical diffraction network to achieve preliminary feature extraction; then, after photoelectric conversion, the obtained electrical signal is input into the electronic neural network for further deep feature extraction, and the classification performance is optimized by combining the cross entropy loss, and finally the classification result is output. This scheme fully combines the high-speed parallelism of optical computing with the powerful learning ability of electronic neural networks, improving the efficiency and accuracy of image recognition.
[0067] In a preferred embodiment of the present invention, the method further comprises: S5. Calculate the loss based on the image classification results and perform model training In this step, the output of the electronic neural network is first processed to obtain the predicted probability distribution of each category. Specifically, the softmax function is used to normalize the original output of the network to make it conform to the probability distribution. The softmax calculation formula is as follows:
[0068] in, The output of the electronic neural network is The raw scores of the categories, is the total number of categories in the dataset, is the predicted probability of this category.
[0069] After obtaining the predicted probability, compare it with the true category label Compare and use the cross-entropy loss function (Cross-Entropy Loss) to calculate the classification error. The loss calculation formula is as follows:
[0070] in, One-hot encoding is used, and only the true category index position is set to 1, and the rest of the positions are set to 0. The cross entropy loss measures the accuracy of the model prediction by calculating the information entropy difference between the predicted distribution and the true distribution.
[0071] After calculating the loss value, the backpropagation algorithm is used to calculate the parameter gradients of each layer of the neural network. Specifically, the chain rule is used to calculate the gradient layer by layer, propagating from the output layer to the input layer to ensure that the impact of the loss on each network parameter can be effectively calculated.
[0072] After calculating the gradient, an optimization algorithm is used to update the parameters of the neural network to minimize the loss function and improve the classification ability of the model. The optimization algorithm can be selected from the following methods: Stochastic Gradient Descent (SGD): Each update calculates the gradient based on only one sample, which has strong randomness and can speed up convergence; Adam (Adaptive Moment Estimation): combines momentum and adaptive learning rate adjustment to improve convergence stability; RMSprop: Performs exponentially weighted averaging of historical gradients to reduce gradient oscillation and improve training convergence.
[0073] The update formula of the optimization algorithm (taking Adam as an example) is as follows:
[0074]
[0075]
[0076] in, and denote the first-order moment estimate and the second-order moment estimate of the gradient, respectively. is the learning rate, is a stabilizing factor to prevent the denominator from being zero.
[0077] During the entire training process, the model will undergo multiple rounds of iterative updates, which will gradually reduce the loss value and improve the classification accuracy. The training process usually includes the following stages: Initialization phase: set initial parameters, including learning rate, batch size, and optimization algorithm type; Forward propagation stage: input image data, perform feature extraction and classification prediction through optical diffraction neural network and electronic neural network; Loss calculation stage: Calculate the classification error of the current batch of samples based on cross entropy; Back propagation phase: calculate the gradient of loss to neural network parameters and update the weights; Model evaluation phase: During the training process, the model is regularly tested on the validation set to determine the convergence and generalization ability of the model.
[0078] This embodiment realizes an efficient end-to-end image classification solution by combining optical computing with electronic neural network optimization, and has high computing efficiency and strong classification capability.
[0079] The present invention also provides an image classification system based on an optoelectronic hybrid computing architecture, comprising: A data input module, used for receiving an input image and performing an upsampling operation on the input image to meet the input requirements of the optical diffraction network; An optical diffraction calculation module is used to perform optical diffraction calculation on the upsampled image to obtain optical features; Photoelectric conversion module, used to convert optical features into electrical signals and transmit them to the electronic neural network; The electronic neural network module is used to receive the electrical signal after photoelectric conversion, use the convolutional neural network to extract features, and output the image classification result; The loss calculation and optimization module is used to calculate the loss according to the image classification results and update the parameters of the electronic neural network through the optimization algorithm, thereby optimizing the image classification model.
[0080] In a specific embodiment of the present invention, an image classification system based on an optoelectronic hybrid computing architecture includes a plurality of modules working in collaboration to achieve efficient image classification tasks.
[0081] First, the data input module receives the original input image and preprocesses it, including upsampling operations, to ensure that the image resolution meets the input requirements of subsequent optical diffraction calculations.
[0082] Subsequently, the preprocessed image enters the optical diffraction calculation module, which processes the input image using the optical diffraction characteristics and extracts the spatial frequency features through diffraction calculation to achieve preliminary feature transformation.
[0083] The result of the optical diffraction calculation is converted by the photoelectric conversion module, which is responsible for mapping the optical features into electrical signals for subsequent electronic neural network processing. The conversion process may involve sampling and signal modulation of the optical signal by a photodetector (such as a CMOS sensor).
[0084] After that, the electronic neural network module receives the photoelectric converted data and uses a convolutional neural network (CNN) for feature extraction. This module uses deep learning technology to extract more discriminative information from the features obtained from optical calculations through multi-layer convolution operations and nonlinear activation functions, and finally outputs the classification results.
[0085] Finally, the loss calculation and optimization module is responsible for calculating the loss based on the classification results and adjusting the parameters of the electronic neural network through the backpropagation algorithm. Specifically, this module uses the cross entropy loss function to measure the error between the predicted result and the true label, and uses optimization algorithms (such as Adam, SGD, etc.) to update the network parameters to improve the classification accuracy.
[0086] This implementation improves the computational efficiency of image classification and reduces the computational burden of electronic computing, while maintaining a high classification accuracy, through the collaborative work of optical computing and electronic computing.
[0087] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. An image classification method based on an optoelectronic hybrid computing architecture, characterized in that: The following steps are involved: Upsampling the input image to meet the input requirements of the optical diffraction network; Use an optical diffraction network to perform optical diffraction calculation on the upsampled image to obtain optical features; Convert optical features into electrical signals and input them into an electronic neural network for feature extraction; Convolutional neural network is used in electronic neural network for feature extraction and output image classification results.
2. The image classification method based on the optoelectronic hybrid computing architecture according to claim 1, characterized in that: The upsampling operation adopts an upsampling method based on "area" interpolation. The upsampling method performs regional interpolation on the input image and calculates the neighborhood average value of each pixel point, thereby generating an upsampled image that matches the size of the original image.
3. The image classification method based on optoelectronic hybrid computing architecture according to claim 1, characterized in that: The steps of optical diffraction calculation include: Perform Fourier transform on the upsampled image to convert the image from spatial domain to frequency domain representation; Applying angular spectrum transfer function to the frequency domain information, modulating the phase and amplitude of the frequency domain signal through element-by-element multiplication operation; The modulated frequency domain signal is inversely Fourier transformed to restore the frequency domain signal to the image features in the spatial domain.
4. The image classification method based on the optoelectronic hybrid computing architecture according to claim 3 is characterized in that: The angular spectrum transfer function is a preset angular spectrum filter function, which weights different amplitude and phase information according to the frequency domain characteristics of the input image, adjusts the propagation characteristics of the frequency domain signal, thereby enhancing or suppressing specific frequency components to obtain a modulated frequency domain signal.
5. The image classification method based on optoelectronic hybrid computing architecture according to claim 1, characterized in that: The electronic neural network uses the deep convolutional neural network ResNet50 for feature extraction.
6. The image classification method based on the optoelectronic hybrid computing architecture according to claim 5, characterized in that: The convolutional neural network ResNet50 includes multiple residual modules, each of which is composed of at least two convolutional layers, and adds the input signal to the output signal through a jump connection.
7. The image classification method based on optoelectronic hybrid computing architecture according to claim 1, characterized in that: The method also includes: calculating the loss based on the image classification result and performing model training.
8. The image classification method based on the optoelectronic hybrid computing architecture according to claim 7, characterized in that: The step of calculating the loss adopts a cross entropy loss function, which calculates the loss value according to the difference between the predicted probability of the image classification result and the true label, wherein the predicted probability is obtained by performing softmax activation on the network output, and the cross entropy loss function takes the logarithm of the predicted probability of each category and multiplies it with the true label, and finally calculates the total loss of the entire data set.
9. The image classification method based on optoelectronic hybrid computing architecture according to claim 7, characterized in that: The calculation of loss based on the image classification result and the model training specifically include the following steps: Compare the predicted category output by the electronic neural network with the true label and calculate the cross entropy loss; Calculate the gradient based on the loss value through the back-propagation algorithm; The network parameters are updated using an optimization algorithm to minimize the loss function, thereby optimizing the image classification model.
10. An image classification system based on an optoelectronic hybrid computing architecture, used to execute the method according to any one of claims 1 to 9, characterized in that: include: A data input module, used for receiving an input image and performing an upsampling operation on the input image to meet the input requirements of the optical diffraction network; An optical diffraction calculation module is used to perform optical diffraction calculation on the upsampled image to obtain optical features; Photoelectric conversion module, used to convert optical features into electrical signals and transmit them to the electronic neural network; The electronic neural network module is used to receive the electrical signal after photoelectric conversion, use the convolutional neural network to extract features, and output the image classification result; The loss calculation and optimization module is used to calculate the loss according to the image classification results and update the parameters of the electronic neural network through the optimization algorithm, thereby optimizing the image classification model.
Citation Information
Cited By
Industrial defect detection method and system based on photoelectric hybrid deep learning architecture
CN122473570A