Image processing acceleration method, system, chip and equipment based on photoelectric fusion

By combining FPGA, optical chip and GPU in the image processing system, and using photoelectric fusion technology for image feature extraction and deep learning classification, the problems of slow image processing speed, large energy consumption and insufficient photoelectric computing collaborative processing in the prior art are solved, and the efficient and low-energy image processing acceleration effect is achieved.

CN120071068APending Publication Date: 2025-05-30PHOTON ARITHMETIC(BEIJING)TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510188169.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The prior art has problems such as high computational complexity, slow processing speed, large energy consumption, and insufficient collaborative processing between photoelectric and electronic computing when processing large-scale images.

Method used

Using an image processing acceleration method based on photoelectric fusion, the input image data is converted into analog signals through the FPGA chip and transmitted to the optical chip for feature extraction and high-dimensional mapping. Then, the optical signal is converted into electrical signals through the photodetector and transmitted to the GPU chip, and a deep learning algorithm is used for image classification and recognition.

Benefits of technology

It significantly improves the speed and accuracy of image processing, reduces energy consumption, solves the problems of high computational complexity and slow processing speed, and effectively coordinates photoelectric and electronic computing through photoelectric fusion technology to achieve efficient acceleration of large-scale image processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071068A_ABST
    Figure CN120071068A_ABST
Patent Text Reader

Abstract

The invention relates to the field of optical chips, photoelectric fusion calculation and image recognition, and discloses an image processing acceleration method, system, chip and device based on photoelectric fusion, and the method comprises the following steps: receiving input image data, and transmitting the input image data to an FPGA chip; the invention also discloses a system which comprises the components of an upper computer which is used for receiving input image data and transmitting the input image data to the FPGA chip; the FPGA chip is used for converting a received digital signal into an analog signal and carrying out feature extraction and high-dimensional mapping through the optical chip; the invention also discloses a chip which comprises an optical calculation module configured to perform image feature extraction and high-dimensional mapping. According to the invention, the photoelectric fusion technology is combined with optical calculation and electronic calculation, so that the image processing speed, accuracy and energy efficiency are improved; feature extraction is accelerated by adopting an echo state network, and the calculation complexity is reduced; and meanwhile, through cooperative work of the FPGA and the GPU, data processing is optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of optical chips, optoelectronic fusion computing, and image recognition, and specifically to an image processing acceleration method, system, chip, and device based on optoelectronic fusion. Background Art

[0002] As a core deep learning model in the field of artificial intelligence, neural networks are widely used in fields such as image recognition and object detection. It simulates the working principle of the human visual system and realizes the efficient processing and recognition of complex data through multi-level abstraction and feature extraction. With the continuous development of neural network structures, especially the introduction of convolutional neural networks (CNNs) and other deep learning models, the network can process more complex scenarios and larger-scale image data. However, the problems brought about in this process are becoming increasingly prominent - with the increase in the number of network layers, the number of parameters and the amount of calculations increase sharply, resulting in a complex training process and high energy consumption. Especially when dealing with large-scale image data processing, problems such as slow training speed and gradient explosion frequently occur.

[0003] Currently, the training and operation of neural networks rely on the computing power of electronic chips. However, the slowdown of Moore's Law has made the performance improvement speed of electronic chips much lower than the demand growth, bringing huge energy consumption and computing bottlenecks. Although traditional electronic chips (such as CPUs and GPUs) can effectively execute the computing tasks of neural networks, they often have problems such as slow computing speed, high energy consumption, and low efficiency when facing large-scale data sets and complex computing tasks, which limit the further application of the existing technology.

[0004] In the field of optical computing, optical chips have received extensive attention due to their characteristics of fast propagation speed, high bandwidth, and low energy consumption. Traditional optical chips usually perform convolution operations through MZI (Mach-Zehnder interferometer) arrays to extract image features. However, the traditional optical computing method using MZI arrays still has problems of high computing complexity and poor scalability when facing complex image recognition tasks. In addition, the problem of the combination of optical computing and electronic computing has not been effectively solved, which makes the potential of optical computing not fully exerted.

[0005] Therefore, the deficiencies of the existing technology are mainly reflected in the following aspects: First, the traditional electronic computing module cannot balance speed and energy efficiency when processing large-scale images, resulting in performance bottlenecks; second, the convolution computing method of traditional optical chips cannot be extended to complex image processing tasks, and the computing complexity is relatively high; third, the collaborative work of optical and electronic computing has not been fully optimized, and the computing bottleneck of optoelectronic fusion cannot be effectively solved.

[0006] To solve the above problems, the present invention proposes an image processing acceleration method, system, chip, and device based on optoelectronic fusion. Summary of the Invention

[0007] Aiming at the deficiencies of the prior art, the present invention provides an image processing acceleration method, system, chip and device based on optoelectronic fusion, which solves the problems of high computational complexity, slow processing speed, high energy consumption and insufficient collaborative processing of optoelectronic computing and electronic computing in traditional image processing technologies.

[0008] To achieve the above objectives, the present invention is realized through the following technical solutions: An image processing acceleration method based on optoelectronic fusion includes the following steps: Receive the input image data and transmit it to the FPGA chip; Convert the input digital signal into an analog signal through the digital-to-analog converter in the FPGA chip and transmit it to the optical chip for feature extraction and high-dimensional mapping; In the optical chip, perform feature extraction and high-dimensional mapping on the image signal through an optical reservoir computing network. The optical reservoir computing network processes the signal in parallel through optical components such as waveguide delay lines and directional couplers and completes image dimensionality increase; Convert the output optical signal of the optical chip into an electrical signal through a photodetector and convert the optical signal into a digital signal through an analog-to-digital converter, and transmit it to the GPU chip; Apply a deep learning algorithm in the GPU chip to perform image classification and recognition on the digital signal and output the classification result; Return the classification result to the host computer for display and storage.

[0009] Preferably, the steps of the optical reservoir computing network for performing feature extraction and high-dimensional mapping on the image signal specifically include: Perform parallel processing on the input image signal through waveguide delay lines and directional couplers, and the waveguide delay lines and directional couplers are responsible for transmitting and weighting the image signal; Execute a non-linear activation function in the optical reservoir computing network to further enhance the mapping effect. The activation function can be Sigmoid, ReLU.

[0010] Preferably, the photodetector generates a current through the photoelectric effect, outputs an electrical signal, and the digital signal is used to maintain the feature information of the image.

[0011] Preferably, the steps of the GPU chip applying a deep learning algorithm to perform image classification and recognition on the digital signal specifically include: Transmit the digital signal as an input to the GPU and perform forward propagation processing through a convolutional neural network; In the convolutional neural network, perform feature extraction and classification on the image through convolutional layers, pooling layers, and fully connected layers, and finally output the classification result through the Softmax function.

[0012] Preferably, the step of returning the classification result to the host computer for display and storage specifically includes: Return the classification result to the host computer through a high-speed interface to display the final result of image classification; The host computer saves the recognition result and can perform subsequent processing or transmit it to other systems.

[0013] Preferably, the GPU chip adopts the ResNet algorithm in the process of image classification and recognition and calculates through the following formula: ; Where, is the output image classification result, representing the category of the input image, is the high-dimensional feature vector obtained from the optical chip and is passed as input to the ResNet network for classification.

[0014] Preferably, in the training process of returning the classification result to the host computer for display and storage, the network parameters are optimized by defining a loss function and adopting the backpropagation algorithm. The calculation formula of the loss function is: ; Where, is the total number of categories, is the one-hot encoding of the true label, indicating whether the th category is the correct category, is the probability of the th category predicted by the model.

[0015] The present invention also provides an optoelectronic fusion-based image processing acceleration system for executing the optoelectronic fusion-based image processing acceleration method described above, including: A host computer for receiving input image data and transmitting it to the FPGA chip; An FPGA chip configured to convert the received digital signal into an analog signal and perform feature extraction and high-dimensional mapping through an optical chip; An optical chip configured to perform feature extraction and high-dimensional mapping on the input signal through an optical reservoir network. The optical reservoir network performs parallel processing through waveguide delay lines and directional couplers to generate high-dimensional features of the image; An analog-to-digital converter and a digital-to-analog converter for converting between optoelectronic signals and digital signals to ensure that the optical signals output by the optics can be received by the digital signal processing system for further analysis; A GPU chip configured to apply deep learning algorithms to perform image classification and recognition on digital signals and finally output the image classification result; The host computer display module is configured to receive and display the image classification results output by the GPU, and it can update the classification results in real time and provide feedback to the user.

[0016] The present invention also provides a chip, including: An optical computing module configured to perform image feature extraction and high-dimensional mapping, which uses an optical reservoir computing network for efficient parallel computing; A signal conversion module configured to convert digital signals into analog signals through a digital-to-analog converter and convert optical signals into digital signals through an analog-to-digital converter; A data transmission module configured to transmit digital signals from the optoelectronic module to the deep learning processing module to ensure the efficient flow and timely processing of signals.

[0017] The present invention also provides an optoelectronic fusion-based image processing acceleration device, including an optoelectronic fusion-based image processing acceleration system as described above and a chip as described above The present invention provides an optoelectronic fusion-based image processing acceleration method, system, chip and device. It has the following beneficial effects: 1. By adopting the optoelectronic fusion technology and combining the advantages of optical computing and electronic computing, the present invention significantly improves the speed and accuracy of image processing. Compared with traditional image processing schemes based on pure electronic computing, the optoelectronic fusion system can greatly reduce energy consumption and accelerate the computing process when processing large-scale data. This not only solves the computing bottleneck in electronic computing but also effectively avoids the problem of low system efficiency caused by high computational complexity in traditional methods.

[0018] 2. Through the application of the echo state network in the optical chip, the present invention realizes the acceleration of data feature extraction and high-dimensional mapping, improves the accuracy and speed of the system for processing images. Compared with the traditional method of using a convolutional neural network to extract features, the echo state network has a lower computational complexity, and its parallel processing in the optical domain greatly reduces the consumption of computing resources and system latency. This design significantly optimizes the processing speed and energy efficiency, meeting the requirements of high-speed and large-data image processing.

[0019] 3. The present invention combines FPGA and GPU to optimize the data processing flow and efficiently identify and classify images through the ResNet algorithm. Compared with the traditional method of using a single computing chip, this system realizes the rapid completion of image classification tasks through flexible hardware cooperation while maintaining a high accuracy rate, solving the performance bottleneck caused by insufficient computing speed and hardware adaptability in traditional methods and providing a more efficient solution for large-scale image recognition tasks. Description of the Drawings

[0020] Figure 1 This is the architecture diagram of the optoelectronic fusion hardware system of the present invention; Figure 2 This is the schematic diagram of the optoelectronic reservoir computing system of the present invention; Figure 3 This is the schematic diagram of the model network structure and training process of the present invention; Figure 4 This is the architecture diagram of the ResNet network of the present invention; Figure 5 This is the flowchart of the method of the present invention; Figure 6 This is the framework diagram of the system of the present invention; Figure 7 This is the schematic diagram of the structure of the chip of the present invention. Detailed implementation manners

[0021] Next, in combination with the accompanying drawings of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0022] Please refer to the attached Figure 1 - attached Figure 5 , the embodiment of the present invention provides an optoelectronic fusion-based image processing acceleration method, including the following steps: S1. Receive the input image data and transmit it to the FPGA chip; In this embodiment, the image processing system first receives the input image data transmitted from an external device (such as a camera, a sensor, or a computer system). The receiving module of the input image data is usually responsible for the host computer. As the central node for data reception and control, the host computer processes the image data and transmits it to the FPGA chip. This process is the starting step of the image processing task, ensuring that the system can effectively obtain the image data to be processed and prepare to enter the subsequent feature extraction and classification process.

[0023] In this embodiment, the input image data is generally represented in the form of a digital signal and is transmitted from the external device to the host computer through a standard interface such as Ethernet, USB, or PCIe protocol. The image data usually adopts the form of a two-dimensional matrix, where each matrix element represents a pixel in the image and contains color or brightness information, such as RGB or grayscale values, etc.

[0024] The input image data is usually transmitted in the following format: ; Wherein: represents the pixel value at a position in the image , which contains color or luminance information.

[0025] and represent the width and height of the image respectively.

[0026] Generally, the resolution of the image is determined according to the application requirements. The typical resolution can be 224×224, or it can be a higher resolution to meet the requirements of high-precision image processing. In some embodiments, the image data will first be preprocessed by image processing software, such as denoising, color adjustment, etc., to ensure the image quality.

[0027] The image data is transmitted from the host computer to the FPGA chip through a transmission interface. During this process, a high-speed data bus (such as PCIe3.0, USB3.0, etc.) is usually used to ensure that the data can be transmitted at high speed and without loss. To ensure the reliability and stability of data transmission, each transmission stage in the system may include data caching or buffering to reduce the risk of data loss.

[0028] In some embodiments, the transmission of image data can be optimized through compression algorithms, such as compression technologies in JPEG, PNG and other formats, especially when processing high-definition images. The compressed image data is sent to the FPGA through a transmission link, and then decompressed and decoded within the FPGA to restore the original image data. In this way, not only can the bandwidth be saved, but also the data transmission process can be accelerated.

[0029] For the convenience of subsequent processing, the input image data may be preprocessed during transmission, including denoising, grayscale conversion, size transformation, etc. of the image. Especially in image classification and recognition tasks, size normalization (for example, scaling the image to 224×224 size) is a common preprocessing step. For this reason, the image size normalization processing formula is: ; where: is the preprocessed image, is the normalized coordinate, usually mapped to a fixed size , such as 224×224, is the corresponding coordinate in the original image.

[0030] In this case, the image pixel point will be mapped to the new size through interpolation (such as bilinear interpolation).

[0031] After receiving the image data, the FPGA chip is responsible for preliminarily processing and formatting the digital image signal for transmission to the optical computing module in subsequent steps. The FPGA is responsible for processing and converting the image data to ensure that the input data is passed to the optical chip in the correct format for optical computing.

[0032] Generally, the FPGA chip will perform the following types of operations on the image at this stage: Data caching: Since the amount of image data is large, the FPGA usually uses a cache to temporarily store the input image data to ensure that no information is lost during processing.

[0033] Signal format conversion: The FPGA will format the input signal into a format suitable for processing by the optical chip as needed. For example, converting the input RGB image to a grayscale image or extracting specific channels in the image.

[0034] Example of the image data preprocessing formula: ; Where: is the pixel value in the input image, , , …, are the weight coefficients for interpolation calculation, , , …, are the coordinates of the interpolation points.

[0035] After receiving the preprocessed image data, the FPGA chip converts the image data into an analog signal through a digital-to-analog converter (DAC) and transmits the data to the optical chip through a high-speed transmission interface. At this time, the optical chip will use optical computing technologies (such as optical reservoir computing networks) to perform feature extraction and high-dimensional mapping on the image.

[0036] Generally, when the image data is transmitted to the optical chip, the transmission speed is required to reach several hundred megabytes per second or higher to ensure the real-time processing of the image. To ensure the efficiency of transmission, the system may adopt a hardware-optimized signal transmission channel to reduce latency and improve processing efficiency.

[0037] S2. Convert the input digital signal into an analog signal through the digital-to-analog converter in the FPGA chip and transmit it to the optical chip for feature extraction and high-dimensional mapping; In an image processing system, the input digital image signal needs to undergo certain conversions and processing before it can be effectively utilized by the optical computing unit. Specifically, step S2 describes the process of converting the digital signal into an analog signal through the digital-to-analog converter (DAC) in the FPGA after the input image data is transmitted to the FPGA chip, and then transmitting the analog signal to the optical chip for subsequent feature extraction and high-dimensional mapping. The implementation of this process is crucial because the optical chip can only process analog signals, and the signal quality will directly affect the effect of subsequent processing.

[0038] In this embodiment, the FPGA chip first receives the digital image data from the previous step. To achieve compatibility with the optical chip, the digital-to-analog converter (DAC) integrated in the FPGA chip converts these digital signals into analog signals. After this conversion, the signals are transmitted to the optical chip for further processing. It should be noted that this conversion process involves precise regulation of the image data and high-quality signal transmission to ensure that the feature information of the image can be accurately transmitted to the optical computing module.

[0039] Generally, the digital-to-analog converter (DAC) inside the FPGA chip is responsible for converting the digital signal into an analog signal compatible with the optical chip. The digital-to-analog converter performs the conversion through the following formula: ; Where: is the output analog voltage signal, representing the converted signal, is the -th bit value of the digital signal, usually represented in binary, is the corresponding weight coefficient, which determines the contribution of each digital bit to the final analog signal, is the bit width of the digital signal.

[0040] In the FPGA, the digital signal is the image data transmitted from the host computer, representing the pixel information of the image. Through the digital-to-analog converter, the digital signal is converted into the corresponding analog voltage signal , and these analog signals are then transmitted to the optical chip.

[0041] Specifically, the analog signal after digital-to-analog conversion will be transmitted to the optical chip through a high-speed transmission link. In this embodiment, the transmission method usually adopts a high-bandwidth data bus, such as LVDS (Low Voltage Differential Signaling) or other standard protocols suitable for analog signal transmission. During the transmission of the analog signal, the minimum signal distortion and delay are required to ensure that the optical chip can receive as accurate signal data as possible.

[0042] As an option, before the analog signal is transmitted to the optical chip, it may pass through a signal conditioning module, such as an amplifier or a filter, to ensure signal quality and eliminate possible noise interference. After the conditioned signal enters the optical chip, it will become the input for the optical computing module to process.

[0043] After the analog signal is output via the digital-to-analog converter, it usually has the following characteristics: Continuity: The analog signal is continuous, while the digital signal is discrete. Through the conversion of the digital-to-analog converter, the discrete information of the signal will become smooth.

[0044] Frequency characteristics: The analog signal usually carries richer frequency domain information and is particularly suitable for frequency domain processing in optical computing.

[0045] At this time, the analog signal is ready to enter the optical chip for feature extraction and high-dimensional mapping. The conversion accuracy of the signal, signal conditioning, and transmission stability directly affect the calculation results of the optical chip. Therefore, it is crucial to ensure that the signal is not lost during transmission.

[0046] In some embodiments, step S2 is not just a simple digital-to-analog conversion. It also includes the enhancement and optimization of image data. For example, by preprocessing the image signal (such as denoising, contrast enhancement, etc.), certain signal optimization can be achieved in the FPGA. These optimization processes help ensure that the input image signal has good quality, thereby guaranteeing the effect of subsequent optical computing.

[0047] Specifically, the analog signal after digital-to-analog conversion relies on the optical computing module (such as the optical reservoir network, ESN) in the optical chip for feature extraction and high-dimensional mapping. The optical computing module can process the frequency domain information and spatial information of the image data in parallel, which is crucial for subsequent image recognition and classification.

[0048] S3. In the optical chip, the optical reservoir network is used to perform feature extraction and high-dimensional mapping on the image signal. The optical reservoir network parallelly processes the signal through optical components such as waveguide delay lines and directional couplers and completes image dimension elevation; In this embodiment, the core part of the optical chip adopts the optical reservoir network (ESN). This network performs feature extraction and high-dimensional mapping of the input signal through nonlinear optical computing. Step S3 immediately follows step S2 above, where the digital signal is converted into an analog signal by the digital-to-analog converter (DAC) and transmitted to the optical chip. The optical reservoir network uses optical components such as waveguide delay lines and directional couplers to parallelly process these analog signals, performing high-dimensional mapping and feature extraction of the signal.

[0049] Specifically, the structure of the optical reservoir computing network is a non-linear dynamic system based on the principle of random projection, capable of efficiently processing high-dimensional input data. Each node represents a computing unit in the network and is connected by optical waveguides. Information transfer between nodes is carried out using optical components such as directional couplers and waveguide delay lines. This network can complete the tasks of feature extraction and high-dimensional mapping of image signals in the optical domain through an efficient parallel processing method.

[0050] In general, the input signal of the optical reservoir computing network becomes an analog signal after digital-to-analog conversion and is transmitted into the optical chip. The optical chip maps these signals to a high-dimensional space through its internal optical reservoir computing network and extracts the feature information therein. Specifically, the output of the optical reservoir computing network is a high-dimensional feature vector obtained after feature extraction and high-dimensional mapping. This process can be expressed by the following formula: ; where: is the output signal of the optical reservoir computing network, representing the high-dimensional feature vector, is the input signal, representing the analog signal obtained from DAC conversion, is the mapping function from optical output to electrical input, which can be a non-linear or linear function and determines the conversion method of optical signals to electrical signals, represents the processing process of the optical reservoir computing network on the input signal to complete the signal dimensionality increase and feature extraction.

[0051] The waveguide delay lines and directional couplers in the optical reservoir computing network play crucial roles. The waveguide delay lines are used to transmit optical signals in the network and generate a phase delay during each signal transmission, thus realizing signal weighting and feature extraction. The delay length of the waveguide delay line and the refractive index of the optical medium jointly determine the magnitude of the signal delay.

[0052] The directional coupler is used to distribute optical signals to different paths, allowing parallel processing of signals among multiple nodes in the network. The working principle of the directional coupler is similar to that of a shunt in an electric circuit, which distributes the input signal proportionally to multiple output ports. Through the combination of different directional couplers, the network can achieve complex signal processing operations, including weighting, summing, beam splitting, etc.

[0053] Specifically, the synergistic effect of the waveguide delay lines and directional couplers enables the optical reservoir computing network to process multiple dimensions from the input signal in parallel and transform these signals into richer and more abstract features. Spatial features of images such as texture, edges, colors, etc. are effectively extracted in this process.

[0054] The core task of the optical reservoir computing network is to enhance the information expression in the input image signal through high-dimensional mapping and feature extraction. In the high-dimensional space, the low-dimensional features of the image are transformed into high-dimensional feature vectors, thereby improving the distinguishability and recognition of the image signal. This high-dimensional mapping is accomplished through parallel computing of multiple nodes, where each node weights the input signal and passes it to subsequent nodes.

[0055] In the reservoir computing network, the activation values and connection weights of the nodes remain fixed, and only the output weights need to be trained. This simplifies the training process of the optical reservoir computing network, as only the output layer needs to be optimized, thus significantly reducing the computational complexity. Specifically, the output feature vector of the optical reservoir computing network represents the mapping of the image in the high-dimensional space. This feature vector includes various spatial information and complex features of the image, providing rich information for subsequent image recognition tasks.

[0056] High-dimensional mapping is one of the core advantages of the optical reservoir computing network. In optical computing, the mapping process of the signal is achieved through the non-linear behavior of optical components, which can capture the complex features of the input image in the high-dimensional space. In this way, optical computing can process high-dimensional data and extract useful information from the image without relying on traditional computing processes. Compared with traditional electronic computing, optical computing has higher parallelism and lower energy consumption, which can significantly improve the efficiency of image processing tasks.

[0057] After completing feature extraction and high-dimensional mapping, the output signal of the optical reservoir computing network is transmitted to a photodetector. The photodetector converts the optical signal into an electrical signal and further amplifies it through a transimpedance amplifier (TIA), finally converting the electrical signal into a voltage signal. This process ensures that the optical signal can be received and processed by the electronic computing part (such as a GPU chip).

[0058] With the help of the photodetector, the optical signal can be efficiently converted into an electrical signal, ensuring that the signal is not lost or distorted during the conversion process. Subsequently, the signal will enter the subsequent image classification and recognition stage, where it will be processed by a GPU chip using deep learning algorithms (such as ResNet).

[0059] S4. Convert the output optical signal of the optical chip into an electrical signal through a photodetector, and convert the optical signal into a digital signal through an analog-to-digital converter, and transmit it to the GPU chip; In this embodiment, after being processed by the optical reservoir network, the optical signal output by the optical chip needs to go through a series of conversion steps before it can be processed by subsequent electronic computing modules (such as GPU chips). The core task of step S4 is to convert the optical signal into an electrical signal and digitize the electrical signal through an analog-to-digital converter (ADC), so that the digital signal can be further transmitted to an electronic chip (such as a GPU) for classification and recognition tasks.

[0060] Generally, the output of the optical signal comes from the optical reservoir network in the optical chip. After being processed by this network, the signal exists in the form of light. In order to transmit the optical signal to the subsequent electronic processing unit, it is first necessary to convert the optical signal into an electrical signal through a photodetector (PD). The photodetector relies on the photoelectric effect. By capturing photons and converting them into free electrons, it generates a current proportional to the intensity of the incident light. The current signal output by the photodetector is closely related to the intensity of the input optical signal and can accurately reflect the characteristic information of the input signal.

[0061] In this embodiment, the photodetector converts the optical signal output by the optical chip into a current signal . The specific conversion relationship can be expressed by the following formula: ; Where: is the current signal output by the photodetector, representing the electrical signal after being converted by the photodetector, is the conversion efficiency of the photodetector, representing the efficiency of converting the optical signal into current. Usually, it is a constant and depends on the physical characteristics and working conditions of the detector, is the optical signal output from the optical reservoir network, representing the characteristic information of the image signal.

[0062] The efficiency of the photodetector directly affects the quality of the signal and the accuracy of subsequent processing. Therefore, selecting a high-efficiency photodetector is the key to ensuring the system performance.

[0063] The current signal output by the photodetector is usually relatively weak. Therefore, it needs to be amplified by a transimpedance amplifier (TIA) to improve the signal strength and quality. The main function of the transimpedance amplifier is to convert the current signal into a voltage signal, thereby enhancing the signal amplitude and making it suitable for subsequent analog-to-digital conversion and electronic processing.

[0064] The working principle of the transimpedance amplifier is based on Ohm's law. It converts the current into a voltage , and controls the gain through a feedback resistor. The formula is expressed as: ; Wherein: is the output voltage signal, representing the signal amplified by the transimpedance amplifier, is the feedback resistor of the transimpedance amplifier, which determines the amplification factor and signal strength, is the current signal from the photodetector.

[0065] In the design, selecting an appropriate feedback resistor can adjust the gain of the transimpedance amplifier so that the amplitude of the signal is suitable for subsequent analog-to-digital conversion processing.

[0066] The voltage signal passing through the transimpedance amplifier is sent to an analog-to-digital converter (ADC) for digitization. The role of the analog-to-digital converter is to convert the continuous voltage signal into a discrete digital signal for subsequent processing by an electronic computing unit (such as an FPGA or GPU). The analog-to-digital converter converts the voltage signal into a set of digital values through sampling and quantization, represented as binary numbers.

[0067] In this embodiment, a high-precision ADC chip (such as the AD9238 of Analog Devices, Inc.) is selected, which has a high sampling rate and low distortion, ensuring that no important detail information is lost during the signal conversion process. The conversion process of the ADC can be expressed by the following formula:

[0068] Wherein: is the output digital signal, representing the discrete signal after analog-to-digital conversion, is the voltage signal amplified by the transimpedance amplifier, ready for analog-to-digital conversion, represents the process of the analog-to-digital converter digitizing the voltage signal into a digital signal.

[0069] During the operation of the ADC, the voltage signal will be discretized into a series of digital values at a predetermined sampling frequency, and each value corresponds to a fixed time point. This process ensures the accurate representation of the analog signal in the digital domain.

[0070] The digital signal after analog-to-digital conversion will be transmitted to the FPGA chip through a high-speed bus (such as PCIe, Ethernet, etc.), and then transmitted to the GPU chip. In this embodiment, the FPGA chip acts as a data transfer relay station, which is responsible for receiving the digital signal output by the ADC and transmitting it to the GPU chip. The deep learning algorithm (such as ResNet) in the GPU chip uses these digital signals for image recognition and classification processing.

[0071] Specifically, the digital signal will be fed into a deep neural network in the GPU chip. The network extracts features and classifies images through multiple convolutional layers and fully connected layers, and finally outputs the result of image recognition.

[0072] S5. Apply a deep learning algorithm in the GPU chip to classify and recognize the digital signal, and output the classification result. In this embodiment, step S5 is a key link in the image processing flow. It involves inputting the digital signal obtained after conversion from the optical chip into the GPU chip, and performing image classification and recognition through a deep learning algorithm (such as ResNet). This process utilizes the powerful parallel processing ability of the GPU chip to extract deep features from the image data and output the final classification result. Step S5 is not only the last step of image processing, but also undertakes the task of combining the previously processed optical data with the deep learning algorithm to achieve accurate image classification.

[0073] Generally, due to its advantages in parallel computing, the GPU chip is widely used for the training and inference of large-scale neural networks. In the present invention, the GPU chip will receive the digital signal D(t) after analog-to-digital conversion, and use the ResNet deep learning network for image recognition and classification. ResNet effectively solves the degradation problem in the training of deep networks by constructing multiple residual blocks, while improving the stability and accuracy of training.

[0074] Specifically, the ResNet algorithm plays a crucial role in the image recognition task. It can effectively extract the spatial features of the image and perform multi-level processing through the residual learning method. Each residual block contains a skip connection, that is, directly adding the input signal to the output signal, thus avoiding the common gradient disappearance problem in deep networks.

[0075] Specifically, in the GPU chip, the input digital signal D(t) serves as the input feature map of ResNet for processing. This process includes multiple convolutional layers, activation layers, and pooling layers. Each convolutional layer is responsible for extracting spatial features from the image, while the activation layer is used to increase the non-linearity of the network. Through the action of these layers, the image features will be gradually strengthened and optimized.

[0076] The structure of each residual block ResBlock in the ResNet network generally contains two 3x3 convolutional layers. Its core purpose is to ensure the accurate transmission of the signal while extracting spatial features. The formula is as follows: ; Wherein: is the output of the ResNet network, representing the final classification result, is the high-dimensional feature vector output from the optical chip and serves as the input feature map, represents the weight matrix of the th convolutional layer, represents the convolutional operation of the th convolutional layer on the input feature map is the bias term of the convolutional layer, usually a learnable parameter, is the activation function, usually using the ReLU (Rectified Linear Unit) function to increase the non-linearity of the network.

[0077] The output of each residual block is directly added to the input through a skip connection to form a so-called "residual mapping", which helps the deep network better perform feature learning and reduce information loss during training.

[0078] During the training process of ResNet, the gradient descent method (such as Adam or SGD) is usually adopted to minimize the loss function, thereby optimizing the weight parameters of the network. The loss function usually uses cross-entropy loss, which is suitable for classification tasks. The cross-entropy loss function can be expressed as: ; Wherein: is the value of the loss function, representing the difference between the predicted value of the model and the true label, is the total number of categories in the classification task, is the one-hot encoding of the true label, indicating whether the category is the correct category, is the predicted probability of the model for the category and is usually output through the Softmax activation function.

[0079] By minimizing the loss function, the weights of the network are optimized, enabling the model to achieve higher accuracy in image recognition tasks.

[0080] Specifically, the trained ResNet network can effectively classify the input images. In the present invention, the high-dimensional feature vector processed by the optical reservoir computing network is input into the ResNet network. After a series of convolutional, pooling, and fully connected layer processes, the classification result y of the image, that is, the category label, is finally output. The classification result may be the object category of the image, such as "cat", "dog", or other specific categories.

[0081] To evaluate the system performance, a common metric, Accuracy, can be used, and its calculation formula is as follows: ; Where: is the number of true positives, representing the number of samples correctly predicted as the positive class, is the number of true negatives, representing the number of samples correctly predicted as the negative class, is the number of false positives, representing the number of negative samples wrongly predicted as the positive class, is the number of false negatives, representing the number of positive samples wrongly predicted as the negative class.

[0082] As an evaluation metric, Accuracy can directly reflect the accuracy and effectiveness of image recognition.

[0083] In step S5, the GPU chip processes the digital signals output from the optical chip by applying the ResNet deep learning algorithm to complete image classification and recognition. Through the residual block structure, ResNet can effectively learn image features and optimize the network weights through the training process to achieve efficient and accurate image recognition. Finally, the classification result is output and returned to the host computer for storage and display. This process realizes the seamless connection of image signals from optics to electronic computing, gives full play to the advantages of the optoelectronic fusion computing system, and achieves a significant acceleration effect in large-scale image processing and recognition tasks.

[0084] S6. Return the classification result to the host computer for display and storage.

[0085] In this embodiment, step S6 is the last step of the entire image processing acceleration system, mainly involving transmitting the image classification result output by the GPU chip to the host computer for display and storage. The host computer is not only responsible for displaying the processing result but also for long-term data storage and subsequent further analysis. The core task of this step is to effectively transmit the classification data output by the GPU chip to the host computer while ensuring the accuracy and timeliness of the data.

[0086] Generally, the classification result processed by the GPU chip is transmitted to the host computer through a high-speed data channel. Data transmission usually adopts communication protocols such as PCIe (Peripheral Component Interconnect) and Ethernet to ensure efficient and high-bandwidth data transmission. In this embodiment, the FPGA chip acts as a data transfer station, responsible for receiving the classification result output by the GPU chip and transmitting the data to the host computer through Ethernet. During the transmission process, the FPGA chip needs to pack and format the digital data according to the predetermined protocol to ensure that the host computer can correctly interpret these data.

[0087] Specifically, the classification result generated by the GPU chip It includes the class labels of the input images and their corresponding confidence levels or probability values. This information will be transmitted to the host computer through the FPGA. The functions of the host computer include: Receiving data: The host computer receives the transmitted image classification results through a high-speed network interface (such as 100G Ethernet).

[0088] Parsing data: The host computer parses the received data packets and extracts the classification labels of the images and their corresponding probability values.

[0089] Displaying results: Once the data parsing is completed, the interface of the host computer will display the classification results to the user through a graphical interface. Usually, the classification results will include the class name of the image, the prediction probability, and the possible relevant confidence levels. For the convenience of users to understand, the original display of the image may also be attached.

[0090] In addition to displaying the classification results, the host computer is also responsible for storing these classification data. Specifically, the classification results will be stored in a database for further data analysis, model optimization, or subsequent data retrieval. In some embodiments, the storage process may involve the following steps: Saving the image classification results: The classification labels and related confidence values will be stored as standardized database records to ensure the consistency of the data structure and the efficiency of queries.

[0091] Archiving the original images and classification results: The original images and their corresponding classification results will be archived together to ensure the integrity of the data. In this way, users can trace the classification process and classification labels of each image.

[0092] Generating reports: For the convenience of subsequent analysis, the system can generate reports based on the classification results, recording the predicted classes of each image and related statistical information.

[0093] To improve the user experience, in some embodiments, the host computer interface will allow users to interact with the classification results. For example, users can click on the image to view its classification label, classification probability, and even the detailed information of model prediction errors. Through these interaction functions, users can have a deeper understanding of the image processing process and adjust system parameters or conduct secondary analysis as needed.

[0094] Generally, the visualization of data is not limited to simple label display, but also includes the visualization of prediction results so that users can intuitively understand the prediction effect of the model. For example, the host computer can display each image together with its predicted class and use different colors or graphical marks to display the confidence level of the prediction results. This feedback mechanism can help users evaluate the classification performance of the system.

[0095] In addition, the system may also collect and analyze various data during processing in real time and generate a real-time monitoring report on the system performance. For example, the report can display the time taken for each processing step, GPU computing load, classification accuracy, etc., to help users comprehensively evaluate the overall performance of the image processing system.

[0096] In some embodiments, the host computer not only presents the classification result to the user but may also use it for further applications. For example, an image recognition system can provide the classification result as input to downstream applications such as video surveillance, autonomous driving systems, medical image analysis, etc. In these scenarios, an accurate classification result is the basis for the system to make decisions.

[0097] The image processing acceleration system based on optoelectronic fusion described below can be correspondingly referred to the image processing acceleration method based on optoelectronic fusion described above.

[0098] Please refer to the append Figure 6 , an image processing acceleration system based on optoelectronic fusion, for performing the above-mentioned image processing acceleration method based on optoelectronic fusion, includes: A host computer, configured to receive input image data and transmit it to the FPGA chip; An FPGA chip, configured to convert the received digital signal into an analog signal and perform feature extraction and high-dimensional mapping through an optical chip; An optical chip, configured to perform feature extraction and high-dimensional mapping on the input signal through an optical reservoir network, and the optical reservoir network performs parallel processing through waveguide delay lines and directional couplers to generate high-dimensional features of the image; An analog-to-digital converter and a digital-to-analog converter, configured to convert between optoelectronic signals and digital signals to ensure that the optical signals output by the optics can be received by the digital signal processing system and used for further analysis; A GPU chip, configured to apply deep learning algorithms to classify and recognize images of digital signals and finally output an image classification result; A host computer display module, configured to receive and display the image classification result output by the GPU, which can update the classification result in real time and provide feedback to the user.

[0099] The system of this embodiment can be used to execute the above method embodiment, and its principle and technical effect are similar, which will not be elaborated here.

[0100] Please refer to the append Figure 7 , the present invention also provides a chip, including: An optical computing module, configured to perform image feature extraction and high-dimensional mapping, and it performs efficient parallel computing using an optical reservoir network; A signal conversion module, configured to convert digital signals into analog signals through a digital-to-analog converter and convert optical signals into digital signals through an analog-to-digital converter; A data transmission module, configured to transmit digital signals from the optoelectronic module to the deep learning processing module to ensure the efficient flow and timely processing of signals.

[0101] The present invention also provides an image processing acceleration device based on optoelectronic fusion, including the above-mentioned image processing acceleration system based on optoelectronic fusion and a chip.

[0102] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. An image processing acceleration method based on photoelectric fusion, characterized in that: The following steps are involved: Receive input image data and transmit it to the FPGA chip; The input digital signal is converted into an analog signal through the digital-to-analog converter in the FPGA chip and transmitted to the optical chip for feature extraction and high-dimensional mapping; In the optical chip, the image signal is subjected to feature extraction and high-dimensional mapping through an optical reservoir network, and the optical reservoir network processes the signal in parallel and completes image dimensionality upgrading through optical components such as waveguide delay lines and directional couplers; The output optical signal of the optical chip is converted into an electrical signal through a photodetector, and the optical signal is converted into a digital signal through an analog-to-digital converter and transmitted to the GPU chip; Apply deep learning algorithms in GPU chips to classify and recognize digital signals and output classification results; The classification results are returned to the host computer for display and storage.

2. The image processing acceleration method based on photoelectric fusion according to claim 1 is characterized in that: The steps of extracting features and performing high-dimensional mapping on the image signal by the optical reservoir network specifically include: Processing input image signals in parallel through waveguide delay lines and directional couplers, which are responsible for transmitting and weighting image signals; A nonlinear activation function is executed in the optical reservoir network to further improve the mapping effect. The activation function may be Sigmoid or ReLU.

3. The image processing acceleration method based on photoelectric fusion according to claim 1 is characterized in that: The photoelectric detector generates current through the photoelectric effect and outputs an electrical signal, and the digital signal is used to maintain characteristic information of the image.

4. The image processing acceleration method based on photoelectric fusion according to claim 1 is characterized in that: The step of applying the deep learning algorithm to classify and recognize digital signals by the GPU chip specifically includes: The digital signal is transmitted as input to the GPU and processed by forward propagation through the convolutional neural network; In the convolutional neural network, the image features are extracted and classified through the convolution layer, pooling layer, and fully connected layer, and finally the classification result is output through the Softmax function.

5. The image processing acceleration method based on photoelectric fusion according to claim 1 is characterized in that: The step of returning the classification results to the host computer for display and storage specifically includes: The classification results are returned to the host computer through a high-speed interface to display the final results of image classification; The host computer saves the recognition results and can perform subsequent processing or transmit them to other systems.

6. The image processing acceleration method based on photoelectric fusion according to claim 1 is characterized in that: The GPU chip uses the ResNet algorithm in the image classification and recognition process, and calculates using the following formula: ; in, is the output image classification result, indicating the category of the input image. It is a high-dimensional feature vector obtained from the optical chip and passed as input to the ResNet network for classification.

7. The image processing acceleration method based on photoelectric fusion according to claim 1 is characterized in that: The classification results are returned to the host computer for display and storage. The training process defines a loss function and uses a back propagation algorithm to optimize network parameters. The calculation formula of the loss function is: ; in, is the total number of categories, is the one-hot encoding of the true label, indicating Is the category the correct category? The model predicts The probability of a class.

8. The image processing acceleration system based on optoelectronic fusion is characterized by: The method for accelerating image processing based on photoelectric fusion according to any one of claims 1 to 7 comprises: The host computer is used to receive the input image data and transmit it to the FPGA chip; An FPGA chip configured to convert a received digital signal into an analog signal and perform feature extraction and high-dimensional mapping through an optical chip; An optical chip configured to perform feature extraction and high-dimensional mapping of an input signal through an optical reservoir network, wherein the optical reservoir network performs parallel processing through a waveguide delay line and a directional coupler to generate high-dimensional features of an image; Analog-to-digital converters and digital-to-analog converters are used to convert between optical and digital signals, ensuring that the optical output light signal can be received by the digital signal processing system and used for further analysis; A GPU chip configured to apply a deep learning algorithm to perform image classification and recognition on digital signals, and finally output image classification results; The host computer display module is configured to receive and display the image classification results output by the GPU. It can update the classification results in real time and provide feedback to the user.

9. A chip, characterized in that: include: An optical computing module configured to perform image feature extraction and high-dimensional mapping, which uses an optical reservoir network for efficient parallel computing; A signal conversion module configured to convert a digital signal into an analog signal through a digital-to-analog converter, and to convert an optical signal into a digital signal through an analog-to-digital converter; The data transmission module is configured to transmit the digital signal from the optoelectronic module to the deep learning processing module to ensure efficient flow and timely processing of the signal.

10. Image processing acceleration device based on optoelectronic fusion, characterized in that: It comprises an image processing acceleration system based on optoelectronic fusion as described in claim 8 and a chip as described in claim 9.

Citation Information

Cited By

  • Image generation system

    CN121235897A

  • Image generation system

    CN121235897B