Device and method for processing images and audios based on MZI network time division multiplexing

By optimizing MZI network resources through time division multiplexing technology and nonlinear activation function processing modules, the problems of low phase adjustment efficiency and insufficient hardware resources in large-scale data processing are solved, efficient and accurate image and audio signal processing is achieved, and the scalability and computing power of the photonic computing system are improved.

CN120785458AInactive Publication Date: 2025-10-14BEIJING CORE TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510885424.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-10-14
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing MZI networks face problems such as low phase adjustment efficiency, insufficient hardware resources, low efficiency due to reliance on electronic modules for nonlinear processing, and poor signal stability when processing large-scale data, which limits the application of photonic computing in image and audio processing.

Method used

Time division multiplexing technology is used to load data into the MZI network in batches. Combined with the nonlinear activation function processing module, nonlinear activation processing is performed through the FPGA module to optimize MZI network resources and improve signal processing accuracy and speed.

Benefits of technology

It improves the flexibility and scalability of photonic computing systems in processing large-scale data, reduces hardware requirements, improves computing accuracy and efficiency, and is suitable for high-resolution image and long-term audio signal processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120785458A_ABST
    Figure CN120785458A_ABST
Patent Text Reader

Abstract

The invention discloses a device and a method for processing an image and an audio based on MZI network time division multiplexing. The method comprises the following steps: S1, dividing an image signal and an audio signal into an image data block and an audio data block; s2, inputting the image data blocks and the audio data blocks into an MZI network through a time division multiplexing controller; s3, the MZI network adjusts the phase shift, the image data block is processed by the MZI network to output an optical signal after matrix operation, and spectrum analysis is performed on the audio data block; s4, converting the optical signal into an electric signal through a photoelectric conversion module; and S5, performing nonlinear activation function processing on the electric signal, and transmitting the processed signal to a next layer of output module. And S6, outputting the signal subjected to nonlinear activation processing to external equipment, and outputting feature and distribution information, a frequency spectrum and a feature signal. And S7, the time division multiplexing controller iteratively optimizes the photon MZI network. The method is low in delay and high in precision, and image audio efficient processing and feature extraction are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of photonic computing and optical communication technology, and particularly relates to an apparatus and method for processing images and audio based on an MZI network time division multiplexing. BACKGROUND

[0002] With the rapid development of information technology, photonic computing as a new computing method is gradually entering people's field of vision. Photonic computing has many unique advantages compared to traditional electronic computing, especially in processing large-scale data and complex neural network tasks. Photonic computing can utilize the characteristics of light speed propagation to achieve faster computing speed, and due to the fact that photons do not accumulate electric charge during propagation, the energy efficiency of photonic computing is significantly improved compared to traditional electronic computing methods. Photonic computing is particularly suitable for application scenarios that require strong parallel computing capability and low energy consumption, such as image recognition, speech recognition, and big data processing. In particular, through the Mach-Zehnder interferometer (MZI) network, the introduction of such photonic computing devices greatly enhances its application in neural networks.

[0003] As a core device in photonic computing, the MZI photonic network controls the interference of optical signals by adjusting the phase of the interferometer, thereby realizing large-scale matrix operations. In recent years, MZI networks have been widely used in photonic neural networks for performing matrix operations, such as image feature extraction, audio signal spectrum analysis, and other tasks. Due to the advantages of photonic networks in parallel computing and large-scale data processing, MZI networks have important application value in many fields. However, despite the great potential of photonic computing, existing MZI networks still face a series of technical challenges.

[0004] Firstly, although MZI networks can realize matrix operations by adjusting the phase of light, how to accurately control the phase of each MZI when dealing with complex matrix transformations is still an important technical problem. The phase adjustment of each MZI is usually controlled by voltage, which means that a specific matrix operation is realized by adjusting the phase of multiple MZIs. For large-scale MZI networks, how to efficiently optimize the phase setting of each MZI to ensure that the output optical signal meets the expected mathematical model is a major bottleneck in current technology. Although existing optimization algorithms can solve some small-scale problems, they usually have slow speed and are prone to local optimization when facing large-scale MZI networks, which makes it difficult to realize efficient photonic network optimization in practical applications.

[0005] In addition, photonic computing also faces problems of physical interference and thermal noise in the application process. Due to the influence of these physical interferences on the operation of the photonic chip, the stability of the signal is poor, which further increases the difficulty of optimizing the setting of the MZI phase shifter. In the prior art, a relatively complex algorithm is usually adopted to compensate for these interferences, but these algorithms usually have a high requirement for computing resources and are difficult to implement in the case of limited hardware resources, resulting in that the efficiency of the overall system cannot be fully utilized.

[0006] In order to solve the above problems, the introduction of time division multiplexing technology provides an effective solution for the application of MZI network. Time division multiplexing technology can process large-scale data under limited physical hardware resources by dividing the input data into multiple batches and loading them into the MZI network one by one in time sequence. This technology can significantly reduce the demand for the number of MZI networks, and through time-sharing data loading, it can realize efficient processing of large-scale data such as images and audio. Through time division multiplexing, the system can cut large-scale data into multiple smaller data blocks, while ensuring efficient processing and minimizing hardware resource consumption.

[0007] However, even if the time division multiplexing technology can effectively improve the processing capability of the photonic network, how to process and output the data processed by the MZI network is still a problem to be solved. In the prior art, since the MZI network is mainly good at performing linear operations, nonlinear processing often needs to be realized through electronic modules. Traditional nonlinear activation function processing, such as ReLU and Sigmoid, is usually performed through electronic computing units. This process involves a large amount of electronic computing resources, and due to the interconnection mode between the photonic network and the electronic computing module, there is often a signal transmission delay, which limits the processing efficiency of the overall system.

[0008] Therefore, in the face of large-scale image and audio data processing, although the hardware resources can be optimized through time division multiplexing technology, there are still significant defects in the nonlinear processing of optical signals, signal conversion and output, etc. These defects result in that the efficiency and accuracy of the system in processing complex signals such as images and audio cannot reach the expected level, further limiting the widespread application of photonic computing in practical applications.

[0009] The present application is based on the processing method of MZI network time division multiplexing, combined with a nonlinear activation function processing module, aiming at overcoming the defects of the prior art. By introducing the FPGA module for efficient nonlinear activation processing, the present application can quickly complete the calculation of the nonlinear activation function on the basis of ensuring the efficient linear operation of the photonic network, further improving the calculation accuracy and efficiency of the overall system. This method can effectively improve the flexibility and scalability of the photonic computing system in processing large-scale data, and is especially suitable for image recognition, speech recognition, signal processing and other high-concurrency and large-scale data processing application scenarios. In addition, the introduction of time division multiplexing technology not only reduces the demand for MZI network physical hardware, but also further improves the computing capacity of the system, especially when processing high-resolution images and long-period audio signals, which shows significant advantages.

[0010] In summary, the present application optimizes the phase adjustment of the MZI network, combines the fast processing of the nonlinear activation function and the time division multiplexing technology, and successfully solves the defects of the prior art in large-scale data processing, nonlinear processing and hardware resource optimization, providing a new technical path for the wide application of photonic computing. SUMMARY

[0011] One object of the present application is to provide an image and audio signal processing method based on MZI network time division multiplexing. The present application fully combines the Mach-Zehnder interferometer (MZI) network, time division multiplexing technology and nonlinear activation function processing module, and describes in detail how to load large-scale image and audio data into the MZI network in batches for efficient processing through time division multiplexing technology, and execute the nonlinear activation function through the FPGA module. This method can effectively optimize the MZI network resources, improve the signal processing accuracy and speed, and has the advantages of high hardware resource utilization, fast calculation speed, high accuracy and strong system scalability.

[0012] According to an embodiment of the present application, a method for processing images and audio based on MZI network time division multiplexing comprises the following steps:

[0013] S1, dividing the image signal and the audio signal into image data blocks and audio data blocks;

[0014] S2, inputting the image data blocks and the audio data blocks into the MZI network through a time division multiplexing controller;

[0015] S3, adjusting the phase shift of the MZI network, the image data blocks are processed through the MZI network to output the matrix operation of the optical signal, and the MZI network processes the frequency spectrum analysis of the audio data blocks;

[0016] S4, using a detector to detect the optical signal output by the MZI network, and converting the optical signal into an electrical signal through an optoelectronic conversion module;

[0017] S5, inputting the electrical signal to the nonlinear activation module, performing nonlinear activation function processing, and transmitting the processed signal to the next layer output module.

[0018] S6, outputting the signal processed by the nonlinear activation to an external device, the image signal outputting the feature and distribution information of the image, and the audio signal outputting the processed spectrum and feature signal.

[0019] S7, the time division multiplexing controller iteratively optimizes the phase shift setting of the photonic MZI network and the parameters of the nonlinear activation function.

[0020] Optionally, the MZI network time division multiplexing is a processing method combining time division multiplexing technology and MZI network, which is used for processing image and audio signals. The MZI network time division multiplexing processes the signals by dividing them into data blocks and loading them into the MZI network in time.

[0021] Optionally, the image data block is divided into a 6x6 matrix pixel block by a high-resolution image; and the audio data block is divided into a plurality of time periods, each time period containing 1024 sampling points.

[0022] Optionally, the S2 specifically comprises:

[0023] S21, inputting the image data block and the audio data block into the time division multiplexing controller, and loading each data block into the MZI network in time sequence according to the set time sequence and batch;

[0024] S22, the time division multiplexing controller adjusts the loading time sequence of each data block through the time sequence module built in the controller, to ensure that the system processes only one data block at a time within the processing capacity of the MZI network;

[0025] S23, after loading one data block each time, the time division multiplexing controller waits for the processing of the current data block to be completed, and then loads the next data block, until all data blocks are processed;

[0026] Optionally, the S3 specifically comprises:

[0027] S31, converting the image data block and the audio data block into optical signals, for the image data block, converting the input high-resolution image data into optical signals, the optical signals carrying the spatial feature information of the image; for the audio data block, converting the long-period audio signal into optical signals, the optical signals carrying the time-domain feature information of the audio

[0028] S32, in the MZI photon network, the input image and audio optical signals pass through multiple MZIs, and the phase shifters of each MZI precisely control the phase of the optical signal by adjusting the phase shift amount;

[0029] S33, for the optical signal of the image data block, the MZI photon network performs linear matrix operation to extract the spatial features in the image data block, and finally outputs the optical signal containing the linear feature information of the image data block, and the linear matrix operation is realized by adjusting the phase matrix M of the multiple MZI units:

[0030] S o ut=M*I;

[0031] Wherein M is the phase adjustment matrix of each interferometer in the MZI network, I is the input image data matrix, S o ut is the output optical signal;

[0032] S34, for the optical signal of the audio data block, the MZI photon network performs frequency spectrum analysis to analyze the frequency characteristics of the audio signal, and the output optical signal contains the frequency characteristics of the audio data block, and the frequency spectrum analysis is realized by adjusting the phase matrix F of the MZI network:

[0033] S f req=F*A;

[0034] Wherein F is the MZI optical matrix for performing frequency spectrum analysis, A is the input audio data, S f req is the output frequency domain optical signal;

[0035] Optionally, the S4 specifically comprises:

[0036] S41, the detector detects the optical signal S o ut, S f req output from the MZI photon network and converts it into an electrical signal, and the detector generates a corresponding current signal I according to the intensity P of the optical signal:

[0037] I=ηP;

[0038] Wherein, I is the current signal, η is the response efficiency of the photodetector, and P is the intensity of the optical signal.

[0039] S42, the position of the detector is flexibly set according to actual needs, and only the light signal intensity at key points is measured to reduce unnecessary measurement complexity. The detector converts the light signal intensity into a current signal through the photoelectric effect to ensure accurate correspondence between the electronic signal and the light signal. In this process, the selection and configuration method of the detector will affect its detection efficiency, and thus affect the conversion effect of the final signal. Considering the case of multiple light sources or multiple signals, the detector may produce cross interference, so a suitable space and light beam selection algorithm needs to be designed to optimize the position selection and signal sampling of the detector.

[0040] Assuming that the light signal intensity detected by the detector at a specific position x0 is P(x0), the output current I(x0) can be represented as:

[0041] I(x0) = η·P(x0);

[0042] At this time, the light signal intensity P(x0) received by the detector will change due to changes in the optical path (such as the angle of the light source, the superposition of interference patterns, etc.). In order to avoid the influence of the difference in signal intensity at different positions on the measurement results, a multi-point signal sampling method based on an optimization algorithm is adopted, which reduces the influence of signal intensity fluctuations through the average value of multi-point data.

[0043] If the signal intensity changes with the detection position x, the relationship between the measured current and the signal intensity can be derived as:

[0044]

[0045] where P(x) represents the light signal intensity received by the detector at different positions x, x1 and x2 represent the starting and ending positions of the detector respectively, and the integral result represents the total light signal intensity in the position range.

[0046] S43, the photoelectric conversion module converts the current signal generated by the detector into a format suitable for further electronic processing. Through the photoelectric conversion module, the generated current signal I is converted into a voltage signal V, and the conversion relationship can be represented as:

[0047] V = R·I;

[0048] where V is the voltage signal, R is the conversion coefficient, and I is the current signal.

[0049] Optionally, the S5 specifically includes:

[0050] S51, the electrical signal from the photoelectric conversion module is transmitted to the nonlinear activation module. The module receives the electrical signal V o ut after detection and photoelectric conversion, and performs nonlinear transformation on the electrical signal according to a predetermined activation function. The form of the electrical signal is a digital signal, represented as Ei n = V o ut.

[0051] S52, in the non-linear activation module, a Sigmoid activation function is selected to perform non-linear processing on the electrical signal. The definition of the Sigmoid activation function is:

[0052]

[0053] This function compresses the input signal to the range of (0, 1), which is suitable for probability estimation and binary classification tasks.

[0054] S53, the Sigmoid activation function is executed at high speed through the FPGA module, and the Sigmoid activation function is applied to each input electrical signal E i n, to generate the output electrical signal E o ut, that is:

[0055]

[0056] Where E i n is the electrical signal from the photoelectric conversion module, and E o ut is the electrical signal output after processing by the Sigmoid activation function. The FPGA module uses its parallel processing capability and high-speed operation characteristics to speed up this process.

[0057] S54, this processing process is performed on each input electrical signal one by one, ensuring that all signals passing through the MZI network and the photoelectric conversion module are processed by the Sigmoid activation function, generating signals suitable for further analysis and output.

[0058] S55, the electrical signal E o ut after Sigmoid activation processing will be used as the input of the next layer network or output module for further feature extraction, classification or other task processing.

[0059] Optionally, the S6 specifically includes:

[0060] S61, the electrical signal E o ut after non-linear activation processing is transmitted from the non-linear activation module to the output module;

[0061] S62, for image signals, the output module converts the received electrical signal E o ut into image features and distribution information. The electrical signal E o ut contains spatial feature information of the image data block after MZI photon network matrix operation processing, and the output module converts it into an image feature matrix F i mg by image feature extraction algorithm, which is represented as:

[0062] F i mg=[f1,f2,...,f n ];

[0063] Among them, f i is the feature vector of the i-th image data block, n is the total number of image data blocks, and the image feature matrix F i mg uses the HOG feature descriptor format;

[0064] S63: For audio signals, the output module receives the electrical signal E o ut is converted into the audio spectrum and characteristic signal, the electrical signal E o ut contains the frequency characteristic information of the audio data block after the MZI photonic network spectrum analysis processing, and the output module extracts the spectrum matrix S through the spectrum analysis method. a udio, expressed as:

[0065] S a udio=FFT(E o ut);

[0066] Among them, FFT is the fast Fourier transform function, and the spectrum matrix S a udio represents the frequency domain feature distribution of the audio data block. The audio feature signal adopts the MFCC feature descriptor format. The MFCC coefficient matrix is ​​expressed as:

[0067] MFCC=[c1,c2,...,c m ];

[0068] Among them, c j is the MFCC coefficient vector of the audio data block in the jth time period, and m is the total number of audio data blocks;

[0069] S64, the output format of the image signal is the image feature matrix F i mg is transmitted to the image recognition system through a standardized digital interface, and the output format of the audio signal is the spectrum matrix S a udio and MFCC coefficient matrix MFCC are transmitted to the audio analysis system through a standardized digital interface;

[0070] S65, the output module outputs the processed image feature matrix F i mg and the audio feature matrix including the spectrum matrix S a The udio and MFCC coefficient matrix MFCC are output to external devices through the USB interface to ensure the compatibility of the signal with the external device interface, and complete the image and audio signal processing output based on MZI network time division multiplexing.

[0071] Optionally, the S7 specifically includes:

[0072] S71、the time division multiplexing controller obtains the image feature matrix F output by the output module i mg and the audio feature matrix includes a spectrum matrix S a udio and the MFCC coefficient matrix MFCC, calculates an error function E between the current processing result and the expected target t otal, which is expressed as:

[0073] E t otal=E i mg+E a udio;

[0074] Wherein, E i mg is the image processing error, E a udio is the audio processing error;

[0075] S72, the image processing error E i mg is calculated by comparing the output image feature matrix F i mg with the standard image feature matrix F t arget, which is expressed as:

[0076] E i mg=||F i mg-F t arget|| 2 ;

[0077] Wherein, F t arget is a preset standard image feature matrix, and ||·|| 2 represents the square of the two-norm;

[0078] S73, the audio processing error E a udio is calculated by comparing the output spectrum matrix S a udio and the MFCC coefficient matrix MFCC with the standard audio feature matrix S t arget and MFCC t arget, which is expressed as:

[0079] E a udio=||S a udio-S t arget|| 2 +||MFCC-MFCC t arget|| 2 ;

[0080] Wherein, S t arget is a preset standard spectrum matrix, and MFCCt Target is a preset standard MFCC coefficient matrix;

[0081] S74, the time division multiplexing controller optimizes the phase shift parameter of each MZI in the MZI photonic network according to the total error E t Total adopts the gradient descent algorithm to iteratively optimize the phase shift parameter of each MZI in the MZI photonic network The phase shift parameter update formula is represented as:

[0082]

[0083] Wherein, is the phase shift parameter of the kth MZI at the tth iteration, is the learning rate of the phase shift parameter, is the partial derivative of the total error with respect to the kth phase shift parameter;

[0084] S75, the time division multiplexing controller simultaneously optimizes the parameters of the Sigmoid activation function in the nonlinear activation module, including the weight parameter w j And the bias parameter b j , the parameter update formula is represented as:

[0085]

[0086] Wherein, w j (t) and b j (t) are the weight parameter and the bias parameter of the jth activation unit at the tth iteration, respectively, α w And α b are the learning rates of the weight parameter and the bias parameter, respectively;

[0087] S76, the time division multiplexing controller sets the iteration termination condition, when the total error E t Total is less than the preset threshold ε or reaches the maximum iteration number T m Max, stop optimization, the optimization termination condition is represented as:

[0088] E t Total<ε or t≥T m Max;

[0089] Wherein, ε is a preset error threshold, T m Max is the maximum iteration number, and t is the current iteration number;

[0090] S77, after completing the parameter optimization, the time division multiplexing controller applies the optimized phase shift parameter And the activation function parameters w j , b j To the next round of image data block and audio data block processing, realizing the adaptive optimization of the MZI network time division multiplexing system.

[0091] The device for processing images and audio based on MZI network time division multiplexing according to an embodiment of the application comprises the following modules:

[0092] The input signal module is configured to receive a high-resolution image signal and an audio signal, divide the high-resolution image signal into image data blocks of 6*6 matrix pixels, and divide a long-time period audio signal into audio data blocks each containing 1024 sampling points in each time period.

[0093] The time division multiplexing controller module is configured to sequentially load the image data blocks and the audio data blocks into the MZI network according to a set time sequence and batch.

[0094] The MZI network module is configured to convert the input image data blocks and audio data blocks into optical signals, adjust the phase shift amount through the phase shifters of multiple MZIs to accurately control the phase of the optical signals, perform linear matrix operation on the optical signals of the image data blocks to extract spatial features, and perform frequency spectrum analysis on the optical signals of the audio data blocks to extract frequency features.

[0095] The detector module is configured to detect the optical signals output by the MZI network, measure the intensity of the optical signals at specific positions, and convert the optical signals into electrical signals.

[0096] The photoelectric conversion module is configured to convert the electrical signals generated by the detector into digital signal format.

[0097] The nonlinear activation module is configured to perform nonlinear processing on the electrical signals through the FPGA module to execute a Sigmoid activation function at high speed.

[0098] The output module is configured to convert the electrical signals E o out after nonlinear activation into an image feature matrix F i mg and an audio feature matrix including a frequency spectrum matrix S a audio and an MFCC coefficient matrix MFCC, and output them to an external device.

[0099] The parameter optimization module is configured to calculate a total error E t otal, and iteratively optimize the phase shift parameters of the MZI network and the weight parameters w and the bias parameters b j of the nonlinear activation module using a gradient descent algorithm. j .

[0100] The present application has the following beneficial effects:

[0101] The present application introduces time division multiplexing technology and optoelectronic hybrid computing architecture, solves multiple technical bottlenecks faced by existing MZI networks when processing large-scale data, and has significant beneficial effects.

[0102] Firstly, the present application breaks through the limitation of the number of physical neurons. In the prior art, the computing power of the MZI network is directly limited by the number of physical MZI units, which leads to the inability to handle large-scale data sets beyond the hardware scale. Through the time division multiplexing technology, the present application loads the input data into the MZI network in batches for processing, not only effectively avoiding the bottleneck of hardware resources, but also enabling the processing of large-scale data without increasing the number of physical MZI units, greatly improving the processing capacity of the system.

[0103] Secondly, the present application significantly reduces the hardware complexity. In the prior art, as the amount of data increases, it is usually necessary to expand the physical scale of the MZI network to handle larger-scale data, thereby increasing the complexity and cost of hardware. Through the time division multiplexing technology, the present application processes data in batches, avoiding excessive dependence on the number of physical MZI units, effectively reducing the complexity of hardware design and manufacturing, and improving the scalability and maintainability of the system in terms of hardware.

[0104] In addition, the present application adopts an efficient optoelectronic hybrid computing architecture, solving the problem that the MZI network can only handle linear operations. In the prior art, nonlinear operations often require additional electronic modules, which increases the number of computing steps and the complexity of the system, affecting the computing efficiency. By combining the linear processing capability of the MZI network and the nonlinear activation processing capability of the electronic module, the present application realizes the organic combination of photonic computing and electronic computing, greatly improving the computing efficiency and task processing capacity of the system.

[0105] At the same time, the present application reduces the loss and noise in the process of optical signal transmission through time division multiplexing technology, significantly improving the signal processing accuracy and noise immunity of the system. In the prior art, the MZI network is inevitably affected by noise and optical loss during multi-stage phase shift adjustment and signal transmission, resulting in a decrease in signal processing accuracy. The present application reduces the interference in the signal transmission process by processing signals in batches, ensuring higher computing accuracy.

[0106] The present application also has flexible scalability and system optimization capability. In traditional MZI networks, the scalability of the system is poor, and a large amount of hardware needs to be added to cope with different sizes of data sets. Through time division multiplexing and optimization iteration technology, the present application can flexibly adjust the batch loading strategy and the parameters of the MZI network to adapt to different sizes of input data, with strong scalability and adjustability, further improving the flexibility and performance of the system.

[0107] In addition, the present application maintains the low energy consumption advantage of photonic computing and further reduces the overall power consumption of the system through the optoelectronic hybrid architecture. Although the MZI network usually leads to an increase in energy consumption when processing large-scale data, the present application effectively reduces the overall energy consumption by reducing the hardware expansion requirement and optimizing the signal processing procedure, ensuring efficient and low-power operation of the system.

[0108] Finally, the present application meets the demand for large-scale data processing. The existing MZI network is difficult to process large data sets beyond its physical scale, especially the processing of high-resolution images and long-period audio signals, which is often limited. The present application uses time division multiplexing technology to enable the system to efficiently process large-scale image, audio signal and other data sets, meeting the high demand for data processing capacity in the big data era and solving the bottleneck of the prior art in large data processing.

[0109] In summary, the present application combines time division multiplexing technology, optoelectronic hybrid architecture and optimized processing strategy innovatively, breaking through the limitations of the prior art, improving the processing capacity, precision, energy efficiency and scalability of the system, and having higher performance and wider application prospects. BRIEF DESCRIPTION OF DRAWINGS

[0110] The accompanying drawings are included to provide a further understanding of the present application, and constitute a part of the specification, together with the embodiments of the present application, to explain the present application, and do not constitute a limitation of the present application. In the drawings:

[0111] Figure 1 A flowchart of a method for processing images and audio based on MZI network time division multiplexing according to the present application;

[0112] Figure 2 A structural schematic diagram of an apparatus for processing images and audio based on MZI network time division multiplexing according to the present application. DETAILED DESCRIPTION

[0113] The present application will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams and only illustrate the basic structure of the present application in a schematic manner, and therefore only show the components related to the present application.

[0114] REFERENCE Figure 1 A method for processing images and audio based on MZI network time division multiplexing, comprising the following steps:

[0115] S1, dividing the image signal and the audio signal into image data blocks and audio data blocks;

[0116] S2, inputting the image data blocks and the audio data blocks into the MZI network through a time division multiplexing controller;

[0117] S3, the MZI network adjusts the phase shift, the image data block is processed by the MZI network to output the optical signal after matrix operation, and the MZI network processing performs spectrum analysis on the audio data block;

[0118] S4, using a detector to detect the optical signal output by the MZI network, and converting the optical signal into an electrical signal through an optical-electric conversion module;

[0119] S5, inputting the electrical signal into a nonlinear activation module, performing nonlinear activation function processing, and transmitting the processed signal to a next layer output module.

[0120] S6, outputting the signal processed by the nonlinear activation to an external device, the image signal outputs the feature and distribution information of the image, and the audio signal outputs the processed spectrum and feature signal.

[0121] S7, the time division multiplexing controller iteratively optimizes the phase shift setting of the photonic MZI network and the parameters of the nonlinear activation function.

[0122] In the embodiment, the MZI network time division multiplexing is a processing method combining time division multiplexing technology and MZI network, which is used for processing image and audio signals, and the MZI network time division multiplexing processes the signals by dividing them into data blocks and loading them into the MZI network for processing in time.

[0123] In the embodiment, the image data block is divided into 6x6 matrix pixel blocks by high-resolution image; the audio data block is divided into multiple time periods by long-period audio signal, and each time period contains 1024 sampling points.

[0124] In the embodiment, the S2 specifically comprises:

[0125] S21, inputting the image data block and the audio data block into the time division multiplexing controller, and loading each data block into the MZI network in time period order according to the set time sequence and batch by the time division multiplexing controller;

[0126] S22, the time division multiplexing controller adjusts the loading time sequence of each data block through the time sequence module built in the controller, so that the system only processes one data block at a time within the processing capacity of the MZI network;

[0127] S23, after loading one data block each time, the time division multiplexing controller waits for the processing of the current data block to be completed, and then loads the next data block, until all data blocks are processed;

[0128] The image data block and the audio data block are input in sequence through the time division multiplexing controller, and are loaded into the MZI network in time sequence one by one according to the set time sequence and batch. The time division multiplexing controller adjusts the loading time sequence of each data block by using the built-in time sequence module, ensures that only one data block is processed at a time, and avoids overloading of the system. After the processing of a data block is completed, the controller waits for the current data block to be processed before loading the next data block, until all data blocks are processed. In this way, the time division multiplexing controller can efficiently process large-scale data in batches, and fully exert the processing capacity of the MZI network.

[0129] In the embodiment, the S3 specifically comprises:

[0130] S31, the image data block and the audio data block are converted into optical signals respectively, for the image data block, the input high-resolution image data is converted into an optical signal, and the optical signal carries the spatial feature information of the image; for the audio data block, the long-period audio signal is converted into an optical signal, and the optical signal carries the time-domain feature information of the audio

[0131] S32, in the MZI photon network, the input image and audio optical signals pass through a plurality of MZIs, and the phase shifter of each MZI precisely controls the phase of the optical signal by adjusting the phase shift amount;

[0132] S33, for the optical signal of the image data block, the MZI photon network performs linear matrix operation to extract the spatial features in the image data block, and the finally output optical signal contains the linear feature information of the image data block, and the linear matrix operation is realized by adjusting the phase matrix M of the plurality of MZI units:

[0133] S o ut=M*I;

[0134] Wherein M is the phase adjustment matrix of each interferometer in the MZI network, I is the input image data matrix, S o ut is the output optical signal;

[0135] S34, for the optical signal of the audio data block, the MZI photon network performs frequency spectrum analysis to analyze the frequency characteristics of the audio signal, and the output optical signal contains the frequency characteristics of the audio data block, and the frequency spectrum analysis is realized by adjusting the phase matrix F of the MZI network:

[0136] S f req=F*A;

[0137] Wherein F is the MZI optical matrix for performing frequency spectrum analysis, A is the input audio data, S f req is the output frequency domain optical signal;

[0138] By converting image data blocks and audio data blocks into optical signals, the present application utilizes MZI photonic networks to process large-scale data. First, image data blocks are converted into optical signals carrying spatial feature information of images, while audio data blocks are converted into optical signals carrying time-domain feature information of audio. In the MZI photonic network, these optical signals are processed through multiple MZI units, and the phase shifters of each MZI precisely adjust the amount of phase shift to control the phase of the optical signal. For image signals, the MZI network performs linear matrix operations to extract the spatial features of the image, and this processing is achieved by adjusting the phase matrix, and finally outputs optical signals containing the linear features of the image. For audio signals, the MZI network performs frequency spectrum analysis to extract the frequency features of the audio signal, and this frequency spectrum analysis is achieved by adjusting the phase matrix in the MZI network, and finally outputs optical signals containing the frequency features of the audio. Through this series of processing, the present application can effectively extract the features of image and audio signals and convert them into optical signals, and utilize MZI photonic networks to efficiently perform data processing tasks.

[0139] In this embodiment, S4 specifically includes:

[0140] S41, the detector detects the optical signal S output from the MZI photonic network o ut, S f req is converted into an electrical signal, and the detector generates a corresponding current signal I according to the intensity P of the optical signal:

[0141] I = ηP;

[0142] where I is the current signal, η is the response efficiency of the photodetector, and P is the intensity of the optical signal.

[0143] S42, the position of the detector is flexibly set according to actual needs, and only the intensity of the optical signal at the key point is measured to reduce unnecessary measurement complexity. The detector converts the intensity of the optical signal into a current signal through the photoelectric effect to ensure the accurate correspondence between the electronic signal and the optical signal. In this process, the selection and configuration method of the detector will affect its detection efficiency, and thus affect the conversion effect of the final signal. Considering the case of multiple light sources or multiple signals, the detector may produce cross interference, so a suitable space and light beam selection algorithm needs to be designed to optimize the position selection and signal sampling of the detector.

[0144] Assuming that the intensity of the optical signal detected by the detector at a specific position x0 is P(x0), the output current I(x0) can be represented as:

[0145] I(x0) = η·P(x0);

[0146] At this time, the light signal intensity P(x0) received by the detector will change due to changes in the optical path (such as the angle of the light source, the superposition of interference patterns, and other factors). In order to avoid the influence of the difference in signal intensity at different positions on the measurement results, a multi-point signal sampling method based on an optimization algorithm is adopted, which reduces the influence of signal intensity fluctuations through the average value of multi-point data.

[0147] If the signal intensity changes with the detection position x, the relationship between the measured current and the signal intensity can be derived as follows:

[0148]

[0149] where P(x) represents the light signal intensity received by the detector at different positions x, x1 and x2 represent the starting and ending positions of the detector respectively, and the integral result represents the total light signal intensity in the position range.

[0150] S43, the photoelectric conversion module converts the current signal generated by the detector into a format suitable for further electronic processing. Through the photoelectric conversion module, the generated current signal I is converted into a voltage signal V, and the conversion relationship can be represented as:

[0151] V = R · I;

[0152] where V is the voltage signal, R is the conversion coefficient, and I is the current signal.

[0153] In this embodiment, S5 specifically includes:

[0154] S51, transmit the electrical signal from the photoelectric conversion module to the nonlinear activation module. The module receives the electrical signal V o ut that has been detected and photoelectrically converted, and performs nonlinear transformation on the electrical signal according to a predetermined activation function. The electrical signal is in the form of a digital signal, represented as E i n = V o ut.

[0155] S52, in the nonlinear activation module, select the Sigmoid activation function to perform nonlinear processing on the electrical signal. The definition of the Sigmoid activation function is:

[0156]

[0157] This function compresses the input signal to the range of (0, 1), which is suitable for probability estimation and binary classification tasks.

[0158] S53, execute the Sigmoid activation function at high speed through the FPGA module, apply the Sigmoid activation function to each input electrical signal E i n, and generate an output electrical signal E out, i.e.

[0159]

[0160] where E i n is the electrical signal from the photoelectric conversion module, E o ut is the electrical signal output after the Sigmoid activation function processing. The FPGA module uses its parallel processing capability and high-speed operation characteristics to accelerate this process.

[0161] S54, the processing process is executed for each input electrical signal, ensuring that all signals passing through the MZI network and the photoelectric conversion module are processed by the Sigmoid activation function, generating signals suitable for further analysis and output.

[0162] S55, the electrical signal E o ut after Sigmoid activation processing will be the input of the next layer network or output module, for further feature extraction, classification or other task processing.

[0163] In this embodiment, S6 specifically includes:

[0164] S61, the electrical signal E o ut after nonlinear activation processing is transmitted from the nonlinear activation module to the output module;

[0165] S62, for image signals, the output module converts the received electrical signal E o ut into image features and distribution information, and the electrical signal E o ut contains the spatial feature information of the image data block after the MZI photon network matrix operation processing, and the output module converts it into an image feature matrix F i mg by image feature extraction algorithm, which is represented as:

[0166] F i mg = [f1, f2,..., fn] ; n ];

[0167] where f i i is the feature vector of the i-th image data block, and n is the total number of image data blocks. The image feature matrix F i mg is described in the format of HOG feature descriptor;

[0168] S63, for audio signals, the output module converts the received electrical signal E o ut into audio spectrum and feature signals, and the electrical signal E o ut contains the frequency feature information of the audio data block after the MZI photon network spectrum analysis processing, and the output module extracts the spectrum matrix Sa udio, is expressed as:

[0169] S a udio=FFT(E o ut);

[0170] Wherein, FFT is the fast Fourier transform function, the frequency spectrum matrix S a udio represents the frequency domain feature distribution of the audio data block, the audio feature signal adopts the MFCC feature descriptor format, and the MFCC coefficient matrix is expressed as:

[0171] MFCC=[c1,c2,...,c m ];

[0172] Wherein, c j is the MFCC coefficient vector of the jth time period audio data block, and m is the total number of audio data blocks;

[0173] S64, the output format of the image signal is the image feature matrix F i mg, which is transmitted to the image recognition system through the standardized digital interface, and the output format of the audio signal is the frequency spectrum matrix S a udio and the MFCC coefficient matrix MFCC, which are transmitted to the audio analysis system through the standardized digital interface;

[0174] S65, the output module outputs the processed image feature matrix F i mg and the audio feature matrix including the frequency spectrum matrix S a udio and the MFCC coefficient matrix MFCC to the external device through the USB interface, ensuring the compatibility of the signal with the external device interface, and completing the image and audio signal processing output based on the MZI network time division multiplexing.

[0175] In the embodiment, the S7 specifically includes:

[0176] S71, the time division multiplexing controller obtains the image feature matrix F i mg and the audio feature matrix including the frequency spectrum matrix S a udio and the MFCC coefficient matrix MFCC output by the output module, calculates the error function E t otal between the current processing result and the expected target, and is expressed as:

[0177] E t otal=E i mg+E a udio;

[0178] Wherein, E i mg is the image processing error, and E a udio is the audio processing error;

[0179] S72, image processing error E i mgby comparing the output image feature matrix F i mgwith the standard image feature matrix F t arget, is calculated, denoted as:

[0180] E i mg=||F i mg-F t arget|| 2 ;

[0181] where F t argetis the preset standard image feature matrix, and ||·|| 2 represents the square of the two-norm;

[0182] S73, audio processing error E a udioby comparing the output spectrum matrix S a udioand the MFCC coefficient matrix MFCC with the standard audio feature matrix S t argetand MFCC t arget, is calculated, denoted as:

[0183] E a udio=||S a udio-S t arget|| 2 +||MFCC-MFCC t arget|| 2 ;

[0184] where S t argetis the preset standard spectrum matrix, and MFCC t argetis the preset standard MFCC coefficient matrix;

[0185] S74, the time division multiplexing controller adopts a gradient descent algorithm to iteratively optimize the phase shift parameters of each MZI in the MZI photonic network according to the total error E t otal The phase shift parameter update formula is represented as:

[0186]

[0187] where, is the phase shift parameter of the kth MZI at the tth iteration, is the learning rate of the phase shift parameter, is the partial derivative of the total error with respect to the kth phase shift parameter;

[0188] S75, the time division multiplexing controller simultaneously optimizes the parameters of the Sigmoid activation function in the nonlinear activation module, including the weight parameter w j and the bias parameter b j , and the parameter update formula is expressed as:

[0189]

[0190] where w j (t) and b j (t) are the weight parameter and the bias parameter of the jth activation unit at the tth iteration, α w and α b are the learning rates of the weight parameter and the bias parameter, respectively;

[0191] S76, the time division multiplexing controller sets an iteration termination condition, and stops optimization when the total error E t otal is less than a preset threshold ε or the maximum iteration number T m max is reached, and the optimization termination condition is expressed as:

[0192] E t otal<ε or t≥T m max;

[0193] where ε is a preset error threshold, T m max is the maximum iteration number, and t is the current iteration number;

[0194] S77, after completing the parameter optimization, the time division multiplexing controller applies the optimized phase shift parameters and the activation function parameters w j , b j to the processing of the next round of image data blocks and audio data blocks, to realize adaptive optimization of the MZI network time division multiplexing system.

[0195] Referring to Figure 2 , an apparatus for processing images and audio based on MZI network time division multiplexing includes the following modules:

[0196] An input signal module is configured to receive a high-resolution image signal and an audio signal, divide the high-resolution image signal into image data blocks of a 6x6 matrix of pixels, and divide a long-period audio signal into audio data blocks each containing 1024 sampling points in a time period;

[0197] A time division multiplexing controller module is configured to load the image data blocks and the audio data blocks into the MZI network one by one in time period order according to a set time sequence and batch;

[0198] The MZI network module is used for converting input image data blocks and audio data blocks into optical signals, adjusting the phase shift amount through a phase shifter of a plurality of MZIs to accurately control the phase of the optical signals, performing linear matrix operation on the optical signals of the image data blocks to extract spatial features, and performing frequency spectrum analysis on the optical signals of the audio data blocks to extract frequency features.

[0199] The detector module is used for detecting the optical signals output by the MZI network, measuring the intensity of the optical signals at a specific position, and converting the optical signals into electrical signals.

[0200] The photoelectric conversion module is used for converting the electrical signals generated by the detector into a digital signal format.

[0201] The nonlinear activation module is used for performing nonlinear processing on the electrical signals through the FPGA module to execute a Sigmoid activation function at a high speed.

[0202] The output module is used for converting the electrical signals E o after the nonlinear activation processing into an image feature matrix F i mg and an audio feature matrix including a frequency spectrum matrix S a audio and an MFCC coefficient matrix MFCC, and outputting the same to an external device.

[0203] The parameter optimization module is used for calculating a total error E t otal, and iteratively optimizing the phase shift parameters of the MZI network and the weight parameters w j and the bias parameters b j of the nonlinear activation module through a gradient descent algorithm.

[0204] Embodiment 1

[0205] In June 2024, in the "Intelligent Information Processing and Optical Computing Technology Laboratory", we carried out a systematic experiment on processing image and audio signals based on MZI network time division multiplexing in response to the demand of image recognition and voice control in intelligent classroom of colleges and universities. Background: traditional image recognition system has low processing efficiency and is difficult to simultaneously consider multimedia fusion processing; and voice command recognition system is often limited by delay and computing resources. Therefore, we use the technical scheme of MZI network combined with time division multiplexing and optoelectronic hybrid architecture proposed in the present application to build a complete test platform to process high-definition image frames and synchronous audio signals, and verify its actual performance.

[0206] ​The experimental site was selected in the B3 laboratory building 406 classroom of Nanjing University of Science and Technology, and two sets of high-definition monitoring cameras (resolution of 3840x2160) and an array pickup system with a sampling frequency of 44.1 kHz were installed. In a 50-minute classroom recording, a frame of image was intercepted every 1 second and its 6x6 pixel block was extracted to form an image data block, a total of 3000 image data blocks were extracted; at the same time, the continuously recorded voice signals were divided into audio data blocks according to every 1024 sampling points, a total of 1292 audio blocks were generated. All data entered the time division multiplexing controller through the input module and loaded into the MZI photon network according to the set timing.

[0207] After the image signal enters the MZI network, the photon matrix constructed by 32 MZI units performs linear transformation on each image data block, extracts spatial features, and outputs optical signals. The optical signals are photoelectrically converted by a high-sensitivity detector array, and the converted electrical signals enter the FPGA module to perform Sigmoid nonlinear activation. The output image features are extracted by the HOG algorithm to obtain an image feature matrix, which is transmitted to the image classification module. The audio data is also analyzed by the MZI network to convert it into frequency domain optical signals, and further extracted by the MFCC algorithm to extract audio features.

[0208] The system uses Intel Stratix 10 FPGA to execute the nonlinear activation function, with an execution rate of 1.68 billion Sigmoid function operations per second. The average processing delay of all image and audio data is 7.2 ms, which is much lower than the 53 ms of the traditional GPU fusion processing architecture. In terms of signal recognition accuracy, the image recognition accuracy is improved from 78.5% to 91.3%, and the audio instruction recognition accuracy is improved from 82.7% to 95.1%. In addition, in terms of power consumption, the average power consumption of the system is only 9.2 W when processing the same batch of data, which is about 41% lower than the traditional CPU+GPU hybrid architecture.

[0209] More importantly, through the parameter optimization module, we control the error function below 0.013. The system meets the convergence condition (ε=0.015) after the 43th iteration, and the MZI phase shifter parameters and the activation function bias parameters reach the optimal configuration. The experiment maintains stable operation of the system during different batches of data loading, without delay accumulation or signal loss.

[0210] Table 1 Experimental data statistics of the image and audio fusion processing system based on MZI network

[0211]

[0212] As can be seen from the table, the system of the application is significantly better than the traditional architecture in the two key processing dimensions of image recognition and audio analysis, the processing delay is reduced by more than 85%, the recognition accuracy is improved by more than 15%, and the system also has obvious advantages in energy consumption control, robustness, adaptive optimization and the like. This fully verifies the application value and engineering feasibility of the application in the actual teaching scene. The system can be subsequently expanded to complex scenes such as smart classrooms, industrial security, medical image-speech fusion analysis and the like.

[0213] The above embodiments verify that the application significantly solves the problems of resource bottleneck, response delay and system expansion difficulty faced by the traditional image and audio processing in the demand of multi-modal information processing in teaching scenes and the like, and embodies the adaptability and technical advantages of the application in the actual complex environment. In the future, the scheme has the real-time multi-modal processing capability in the smart medical and automatic driving scenes, and is especially suitable for the high-concurrency signal processing demand in the big data background.

[0214] The above is only the preferred specific implementation of the application, but the protection scope of the application is not limited to this, any person skilled in the art can make equivalent replacement or change according to the technical scheme and the inventive concept of the application within the technical range disclosed by the application, which should be covered in the protection scope of the application.

Claims

1. A method for processing images and audio using time division multiplexing based on an MZI network, characterized in that: The steps include: S1. Divide the image signal and the audio signal into image data blocks and audio data blocks; S2, inputting the image data block and the audio data block into the MZI network through the time division multiplexing controller; S3, the MZI network adjusts the phase shift, the image data block is processed by the MZI network to output the optical signal after matrix operation, and the MZI network performs spectrum analysis on the audio data block; S4, using a detector to detect the optical signal output by the MZI network, and converting the optical signal into an electrical signal through a photoelectric conversion module; S5. Input the electrical signal into the nonlinear activation module, perform nonlinear activation function processing, and pass the processed signal to the next layer output module. S6. Output the signal after nonlinear activation processing to an external device. The image signal outputs the characteristics and distribution information of the image, and the audio signal outputs the processed spectrum and characteristic signal. S7. The time-division multiplexing controller iteratively optimizes the phase shift settings of the photonic MZI network and the parameters of the nonlinear activation function.

2. The method for processing images and audio based on MZI network time division multiplexing according to claim 1, characterized in that: The MZI network time division multiplexing is a processing method that combines time division multiplexing technology and MZI network, and is used to process image and audio signals. The MZI network time division multiplexing divides the signal into data blocks and loads them into the MZI network for processing in a time-sharing manner.

3. The method for processing images and audio based on MZI network time division multiplexing according to claim 1, characterized in that: The image data block is obtained by dividing a high-resolution image into 6×6 matrix pixel blocks; the audio data block is obtained by dividing a long-term audio signal into multiple time periods, each time period containing 1024 sampling points.

4. The method for processing images and audio based on MZI network time division multiplexing according to claim 1, characterized in that: The S2 specifically includes: S21, input the image data block and the audio data block into the time division multiplexing controller, and the time division multiplexing controller loads each data block into the MZI network one by one in the time period sequence according to the set timing and batch; S22, the time division multiplexing controller adjusts the loading timing of each data block through the built-in timing module of the controller to ensure that the system processes only one data block at a time within the processing capacity of the MZI network; S23. Each time a data block is loaded, the time division multiplexing controller waits for the processing of the current data block to be completed, and then loads the next data block until all data blocks are processed.

5. The method for processing images and audio based on MZI network time division multiplexing according to claim 1, characterized in that: The S3 specifically includes: S31. Converting the image data block and the audio data block into optical signals: For the image data block, convert the input high-resolution image data into an optical signal, which carries the spatial feature information of the image; for the audio data block, convert the long-term audio signal into an optical signal, which carries the temporal feature information of the audio; S32. In the MZI network, the input image and audio optical signals pass through multiple MZIs, and the phase shifter of each MZI accurately controls the phase of the optical signal by adjusting the phase shift amount. S33 For the optical signal of the image data block, the MZI network performs linear matrix operations to extract the spatial features in the image data block, and finally outputs the optical signal S o ut contains the linear feature information of the image data block. S34 For the optical signal of the audio data block, the MZI network performs spectrum analysis to analyze the frequency characteristics of the audio signal, and the output optical signal S f req contains the frequency characteristics of the audio data block.

6. The method for processing images and audio based on MZI network time division multiplexing according to claim 1, characterized in that: The S4 specifically includes: S41, the detector detects the optical signal S output from the MZI network o ut、S f req is converted into an electrical signal; S42. The position of the detector is flexibly set according to actual needs, and the light signal intensity is measured only at key points. The light signal intensity detected by the detector at a specific position x0 is P(x0); S43, the photoelectric conversion module converts the electrical signal generated by the detector into a digital signal format E i n.

7. The method for processing images and audio based on MZI network time division multiplexing according to claim 1, characterized in that: The S5 specifically includes: S51. In the nonlinear activation module, select the Sigmoid activation function to perform nonlinear processing on the electrical signal. S53, execute the Sigmoid activation function at high speed through the FPGA module, and apply the Sigmoid activation function to each input electrical signal E i n, generates an output electrical signal E o ut, that is: Among them, E i n is the electrical signal from the photoelectric conversion module, E o ut is the electrical signal output after Sigmoid activation function processing. The FPGA module uses its parallel processing capability and high-speed computing characteristics to accelerate this process. S54, the processing process for each input electrical signal E i nExecute one by one to ensure that all signals passing through the MZI network and the photoelectric conversion module are processed by the Sigmoid activation function.

8. The method for processing images and audio based on MZI network time division multiplexing according to claim 1, characterized in that: The S6 specifically includes: S61, the electrical signal E after nonlinear activation processing o ut is transferred from the nonlinear activation module to the output module; S62: For the image signal, the output module receives the electrical signal E o ut is converted into image features and distribution information, and the electrical signal E o ut contains the spatial feature information of the image data block after the MZI network matrix operation processing, and the output module converts it into the image feature matrix F through the image feature extraction algorithm i mg; S63: For audio signals, the output module receives the electrical signal E o ut is converted into the audio spectrum and characteristic signal, the electrical signal E o ut contains the frequency characteristic information of the audio data block after the MZI network spectrum analysis processing, and the output module extracts the spectrum matrix S through the spectrum analysis method. a udio, audio feature signal uses MFCC feature descriptor format; S64, the output format of the image signal is the image feature matrix F i mg is transmitted to the image recognition system through a standardized digital interface, and the output format of the audio signal is the spectrum matrix S a udio and MFCC coefficient matrix MFCC are transmitted to the audio analysis system through a standardized digital interface; S65, the output module outputs the processed image feature matrix F i mg and the audio feature matrix including the spectrum matrix S a The udio and MFCC coefficient matrix MFCC are output to external devices through the USB interface.

9. The method for processing images and audio based on MZI network time division multiplexing according to claim 1, characterized in that: The S7 specifically includes: S71, the time division multiplexing controller obtains the image feature matrix F output by the output module i mg and the audio feature matrix including the spectrum matrix S a udio and MFCC coefficient matrix MFCC, calculate the error function E between the current processing result and the expected target t otal; S72, image processing error E i mg is calculated by comparing the output image feature matrix F i mg and standard image feature matrix F t The difference between arget is calculated; S73, audio processing error E a udio compares the output spectrum matrix S a udio and MFCC coefficient matrix MFCC and standard audio feature matrix S t arget and MFCC t The difference between arget is calculated S74, the time division multiplexing controller is based on the total error E t otal uses the gradient descent algorithm to iteratively optimize the phase shift parameter φ of each MZI in the MZI network k , the phase shift parameter update formula is expressed as: in, is the phase shift parameter of the k-th MZI at the t-th iteration, is the learning rate of the phase shift parameter, is the partial derivative of the total error with respect to the kth phase shift parameter; S75, the time division multiplexing controller simultaneously optimizes the parameters of the Sigmoid activation function in the nonlinear activation module, including the weight parameter w j and bias parameter b j ; S76, the time division multiplexing controller sets the iteration termination condition, when the total error E t otal is less than the preset threshold ε or reaches the maximum number of iterations T m Stop optimization when ax; S77, after completing the parameter optimization, the time division multiplexing controller will optimize the phase shift parameters and activation function parameter w j 、b j Applied to the next round of processing of image data blocks and audio data blocks.

10. A device for processing images and audios based on MZI network time division multiplexing, executing the method for processing images and audios based on MZI network time division multiplexing according to any one of claims 1 to 9, characterized in that: Includes the following modules: An input signal module is configured to receive a high-resolution image signal and an audio signal, divide the high-resolution image signal into image data blocks of 6×6 matrix pixel blocks, and divide the long-period audio signal into audio data blocks containing 1024 sampling points per time period; The time division multiplexing controller module is used to load the image data blocks and audio data blocks into the MZI network one by one in the time period sequence according to the set timing and batch; The MZI network module is used to convert the input image data blocks and audio data blocks into optical signals. The phase of the optical signal is precisely controlled by adjusting the phase shift amount through multiple MZI phase shifters. Linear matrix operations are performed on the optical signals of the image data blocks to extract spatial features, and spectrum analysis is performed on the optical signals of the audio data blocks to extract frequency features. The detector module is used to detect the optical signal output by the MZI network, measure the optical signal intensity at a specific location and convert the optical signal into an electrical signal; Photoelectric conversion module, used to convert the electrical signal generated by the detector into digital signal format; A nonlinear activation module is used to perform nonlinear processing on electrical signals by executing the Sigmoid activation function at high speed through the FPGA module; Output module, used to convert the electrical signal E after nonlinear activation processing o ut is converted to image feature matrix F i mg and the audio feature matrix including the spectrum matrix S a udio and MFCC coefficient matrix MFCC, output to external devices; Parameter optimization module, used to calculate the total error E t otal, uses the gradient descent algorithm to iteratively optimize the phase shift parameter φ of the MZI network k and the weight parameter w of the nonlinear activation module j , bias parameter b j .