Underwater robot imaging algorithm optimization method based on deep learning
Through the optimization method of underwater robot imaging algorithm based on deep learning, the surface CMOS array sensor, quantum convolution kernel, OAM modal power distribution and programmable metasurface are used to solve the problem of failure of traditional optical imaging algorithms in turbid waters, achieving efficient and accurate imaging, and adapting to different water turbidity and light conditions.
Patent Information
- Application Number
- CN202510242433.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-06-20
AI Technical Summary
Traditional optical imaging algorithms fail in turbid waters and cannot effectively deal with the scattering and absorption of light. The existing fusion technology adopts a fixed fusion ratio and cannot adapt to the dynamic environment, resulting in poor fusion effect.
The optimization method of underwater robot imaging algorithm based on deep learning is adopted, including the use of surface CMOS array sensors for dynamic capture of light field, designing quantum convolution kernels for joint coding of optical and acoustic data and cross-domain feature fusion, dynamically adjusting the power distribution of OAM modes, and designing programmable metasurfaces to predict wavefront distortion through feedforward neural networks and adjusting metasurface parameters in real time.
It improves the accuracy and robustness of imaging, and can dynamically adjust the weight of optical images and sonar data under different water turbidity and lighting conditions, reduce model errors, realize sub-wavelength-level wavefront regulation, and improve the adaptability and robustness of the imaging system.
Smart Images

Figure CN120180356A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of underwater robots, and more specifically, to an optimization method for underwater robot imaging algorithms based on deep learning. Background Art
[0002] Underwater robots have a wide range of application scenarios, covering multiple fields. Specific devices include special types such as tethered underwater robots, untethered underwater robots, and underwater gliders. These devices are usually equipped with devices such as sonar systems, cameras, lighting lamps, and robotic arms, and can provide real-time videos, sonar images, and can lift heavy objects through robotic arms;
[0003] In terms of imaging, underwater robots mainly use optical imaging technology and acoustic imaging technology. Optical imaging technology obtains underwater images through high-definition cameras, while acoustic imaging technology uses sonar systems for detection. Single-beam mechanical scanning sonar realizes imaging through the mechanical rotation of the beam and the rotating pan-tilt, and is suitable for navigation and collision avoidance in complex environments. The advantage of existing algorithms is that they can automatically extract image features, learn complex patterns, and are suitable for training on large-scale data sets;
[0004] However, in the actual use process, traditional optical imaging algorithms fail in turbid waters because they cannot effectively handle the problems of light scattering and absorption. When a single optical imaging algorithm processes images in turbid waters, it cannot restore the clarity and color authenticity of the images. Moreover, existing fusion technologies use fixed fusion ratios and cannot adapt to dynamic environments. Under different water turbidities and lighting conditions, the fixed fusion ratio cannot dynamically adjust the weights of optical images and sonar data, resulting in poor fusion effects. In cross-water tests, models lacking physical constraints have large errors and cannot meet the imaging requirements of different waters. Summary of the Invention
[0005] To solve the above problems, the present invention provides an optimization method for underwater robot imaging algorithms based on deep learning.
[0006] The present invention provides an optimization method for underwater robot imaging algorithms based on deep learning, including the following steps:
[0007] Step 1: First, use a curved CMOS array sensor to perform dynamic capture of the underwater light field, and adaptively adjust the focal length of the curved CMOS array sensor according to the change of the scene light intensity, and activate local pixels when the scene changes to collect the light field information of the environment where the underwater robot is located;
[0008] Step 2: Design a quantum convolution kernel, map the classical convolution operation to a quantum circuit, construct a hybrid computing pipeline, perform quantum state joint encoding and cross-domain feature fusion on optical and acoustic data to obtain the fused data;
[0009] Step 3: Dynamically adjust the power distribution of the OAM mode to maximize the channel capacity in turbid waters. When the water turbidity increases, automatically reduce the power proportion of the high-order mode, and instead enhance the transmission efficiency of the low-order mode. Generate a light beam using optical orbital angular momentum modulation to increase the capacity of cross-media communication;
[0010] Step 4: Design a programmable metasurface to predict wavefront distortion through a feedforward neural network, and combine a PID controller to adjust the metasurface parameters in real time to achieve sub-wavelength wavefront control.
[0011] Preferably, the specific working steps of the said Step 1 are as follows:
[0012] Transfer the curved surface array of optoelectronic sensors to flexible rubber, and install a dual liquid lens directly in front to obtain a complete curved surface CMOS array sensor, and then install it at the front end of the underwater robot for collecting the light field information of the environment where the underwater robot is located;
[0013] Obtain the ambient light intensity Lt of the environment where the underwater robot is located at the current time point through the optoelectronic sensor; According to the formula Ft = F0 + ΔF * sigmoid(0.5 * Lt), calculate and obtain the focal length Ft of the dual liquid lens adjusted according to the ambient light intensity Lt at the current time point, and adjust the focal length of the dual liquid lens according to the focal length Ft to make the focal length of the dual liquid lens reach Ft, where F0 is the initial focal length of the dual liquid lens under zero voltage, and the focal length change amount ΔF is the maximum change range of the focal length of the dual liquid lens.
[0014] Preferably, the specific working steps of the said Step 1 further include the following:
[0015] Establish a plane rectangular coordinate system, take the upper left corner of the curved surface CMOS array sensor as the coordinate origin, and obtain the coordinates (X, Y) of each optoelectronic sensor pixel;
[0016] According to the formula Obtain the light intensity change rate R of each optoelectronic sensor pixel in time in real time, where represents the illumination intensity of this pixel point at the position (X, Y) and time t, is the time difference between this calculation of the instantaneous change rate R and the previous calculation;
[0017] Preset a threshold θ for the light intensity change rate in advance. If the light intensity change rate R of the optoelectronic sensor pixel in time is greater than the threshold θ, then record the polarity of the light intensity change at the (X, Y) position at the moment when the image exceeds the preset threshold θ. The activated optoelectronic sensor pixels will transmit the data through the readout circuit; then convert the electrical signal into a digital signal through an analog-to-digital converter, and finally output the digital image signal data and transmit it.
[0018] Preferably, the specific working steps of step two are as follows:
[0019] First, obtain acoustic data through a sonar sensor installed on the underwater robot;
[0020] Then receive the digital image signal data output by step one;
[0021] Process the low-frequency features through a classical convolutional neural network to extract classical features;
[0022] Process the high-frequency features through a quantum convolutional neural network to extract quantum features;
[0023] According to the formula
[0024]
[0025] Convert the classical features from the time domain to the frequency domain to extract the global feature QFT(F), where n is the number of qubits and F(j) is the j-th component of the classical feature;
[0026] Perform a convolution operation on the quantum features through quantum convolution to extract the local feature QConv(F), and the calculation formula is:
[0027]
[0028] where H k is the quantum Hamiltonian and θ k is the trainable parameter;
[0029] Through quantum entanglement, manage the global feature QFT(F) and the local feature QConv(F) extracted by quantum convolution to obtain the cross-domain fusion feature F3. The specific steps are:
[0030]
[0031] where U is the quantum entanglement gate operation and Tr represents the partial trace operation on the auxiliary qubits.
[0032] Preferably, step two can also adaptively adjust the weights of the optical image and sonar data according to the change of the ambient light intensity. The specific steps are:
[0033] Obtain the ambient light intensity Lt of the environment where the underwater robot is located at the current time point through a photoelectric sensor. According to the formula Calculate the weight ω1 of the optical image, and then according to ω2 = 1 - ω1, obtain the weight ω2 of the sonar data. According to the weight ω1 of the optical image and the weight ω2 of the sonar data, dynamically adjust the weights of the optical image and sonar data to perform data fusion.
[0034] Preferably, the specific working steps of Step 3 are as follows:
[0035] Generate a beam carrying orbital angular momentum using a spiral phase plate, and its complex amplitude distribution is;
[0036]
[0037] where L is the topological charge number, W0 is the beam waist radius, and P0 is the transmission power; obtain the measured water turbidity τ and light attenuation coefficient c through the sonar of the underwater robot, and construct a cross-media channel measurement model C;
[0038]
[0039] where P m is the transmission power of the m-th mode, is the noise power;
[0040] Then, solve the optimal power allocation strategy PM0 according to the Lagrange multiplier method, and the specific steps are;
[0041]
[0042] where δ is the multiplier satisfying the total power constraint, and (x) + represents max(x, 0).
[0043] Preferably, the specific working steps of Step 3 further include the following:
[0044] During the sonar signal transmission period, stagger the transmission timings of the acoustic pulse and the optical pulse through time-interleaved multiplexing technology to avoid mutual interference;
[0045] Establish an acoustic-optic propagation time-delay difference model, and calculate the time-delay difference Δt according to the formula where C1 is the speed of sound in water and C2 is the speed of sound in light;
[0046] Obtain the sonar received signal Y1 and the received signal Y2 of the optical image;
[0047] According to the formula obtain the parameter that maximizes the joint probability of the sonar and optical received signals, where θ is the parameter to be estimated.
[0048] Preferably, the specific working steps of Step 3 further include the following:
[0049] The specific acquisition methods of the sonar received signal Y1 and the received signal Y2 of the optical image are as follows;
[0050] For the sonar received signal Y1, through convolution operation, the transmitted signal of the sonar is convolved with the channel impulse response to obtain the received signal Y1. For the received signal Y2 of the optical image, based on the acquisition of Y1, real part (Re) and conjugate (*) operations are taken.
[0051] Preferably, the specific steps of step four are as follows:
[0052] Construct a metasurface phase modulation model, and the phase response of its unit structure is;
[0053] where μ1 is the equivalent permittivity, B is the operating wavelength, representing the wavelength of the electromagnetic wave, and μ2 is the unit height, representing the physical height of the metasurface unit;
[0054] Then train a feedforward neural network to predict the wavefront distortion;
[0055] Finally, design a PID controller to adjust the voltage of the metasurface unit, and the control law is;
[0056]
[0057] where eij is the phase error, K P is 0.8, K i is 0.2, K d is 0.05.
[0058] Preferably, the specific steps of training the feedforward neural network to predict the wavefront distortion further include:
[0059] Input layer, which receives the original light field data of the curved surface CMOS array and the sonar point cloud data. The original light field data of the curved surface CMOS array contains the light intensity information of the underwater environment, while the sonar point cloud data provides the three-dimensional position information of underwater objects. These data are used as the input of the neural network to predict the wavefront phase error distribution;
[0060] The hidden layer consists of 3 layers of fully connected networks, with 512 neurons in each layer. The fully connected network means that each neuron is connected to every neuron in the previous layer, and can capture the complex non-linear relationships of the input data. The activation function of each layer is LeakyReLU;
[0061] The output layer predicts the wavefront phase error distribution.
[0062] Beneficial effects: By means of quantum Fourier transform and quantum convolution, data of different modalities are fused to improve the accuracy and robustness of imaging. It can dynamically adjust the weights of optical images and sonar data under different water turbidities and lighting conditions, improve the fusion effect, reduce model errors, can adjust the metasurface parameters in real time to achieve sub-wavelength wavefront control, improve the adaptability and robustness of the imaging system. Through metasurface dynamic wavefront correction and intelligent regulation, it can dynamically adjust the parameters of the imaging system under different water turbidities and lighting conditions, improve the imaging quality, and reduce model errors. Description of the Drawings
[0063] Figure 1 It is the flowchart of the system of the present invention. Detailed Embodiments
[0064] Application scenarios: In the actual use process, traditional optical imaging algorithms fail in turbid waters because they cannot effectively handle the problems of light scattering and absorption. When a single optical imaging algorithm processes images in turbid waters, it cannot restore the clarity and color authenticity of the images. Moreover, existing fusion technologies use fixed fusion ratios and cannot adapt to dynamic environments. Under different water turbidities and lighting conditions, the fixed fusion ratio cannot dynamically adjust the weights of optical images and sonar data, resulting in poor fusion effects. In cross-water tests, the model errors without physical constraints are relatively large, and it also cannot meet the imaging requirements of different waters.
[0065] As Figure 1 shown: The optimization method for the imaging algorithm of an underwater robot based on deep learning includes the following steps:
[0066] Step 1: First, a curved CMOS array sensor is used to dynamically capture the underwater light field, and the focal length of the curved CMOS array sensor is adaptively adjusted according to the change of the scene light intensity, and local pixels are activated when the scene changes for collecting the light field information of the environment where the underwater robot is located. It should be noted that it can better capture the light field information, reduce the influence of light scattering and absorption on imaging, only activate local pixels when the scene changes, reduce redundant data transmission, improve the imaging efficiency and quality, and can capture clearer images in turbid waters, improving the dynamic range and adaptability of imaging;
[0067] Step 2: Design a quantum convolution kernel, map the classical convolution operation to a quantum circuit, construct a hybrid computing pipeline, perform quantum state joint encoding and cross-domain feature fusion on optical and acoustic data to obtain the fused data. It should be noted that through quantum Fourier transform and quantum convolution, data of different modalities are fused to improve the accuracy and robustness of imaging. It can dynamically adjust the weights of optical images and sonar data under different water turbidities and lighting conditions, improve the fusion effect, and reduce model errors;
[0068] Step 3: Dynamically adjust the power distribution of the OAM mode to maximize the channel capacity in turbid waters. When the water turbidity increases, automatically reduce the power ratio of the high-order mode and instead enhance the transmission efficiency of the low-order mode. Generate a light beam using optical orbital angular momentum modulation to increase the capacity of cross-media communication. It should be noted that the metasurface parameters can be adjusted in real time to achieve sub-wavelength wavefront control, improve the adaptability and robustness of the imaging system. Through dynamic wavefront correction and intelligent control of the metasurface, the parameters of the imaging system can be dynamically adjusted under different water turbidities and lighting conditions to improve the imaging quality and reduce model errors.
[0069] Step 4: Design a programmable metasurface to predict wavefront distortion through a feedforward neural network and combine a PID controller to adjust the metasurface parameters in real time to achieve sub-wavelength wavefront control. It should be noted that through dynamic wavefront correction and intelligent control of the metasurface, the parameters of the imaging system can be dynamically adjusted under different water turbidities and lighting conditions to improve the imaging quality and reduce model errors.
[0070] Since traditional optical imaging algorithms fail in turbid waters, mainly because they cannot effectively handle the problems of light scattering and absorption. When a single optical imaging algorithm processes images in turbid waters, it cannot restore the clarity and color authenticity of the images. Existing fusion techniques use a fixed fusion ratio and cannot adapt to dynamic environments. Under different water turbidities and lighting conditions, the fixed fusion ratio cannot dynamically adjust the weights of optical images and sonar data, resulting in poor fusion effects. In cross-water tests, the model errors without physical constraints are large and cannot meet the imaging requirements of different waters.
[0071] Aiming at the failure problem of traditional optical imaging algorithms in turbid waters, this technical solution provides comprehensive solution steps. Each of the above steps optimizes the defects of the existing technology, improves the adaptability, robustness and imaging quality of the imaging system, and can achieve efficient and accurate imaging in complex underwater environments.
[0072] As an optional embodiment: The specific working steps of Step 1 are as follows:
[0073] Transfer the curved surface array of the photoelectric sensor onto the flexible rubber, and install a dual liquid lens directly in front to obtain a complete curved surface CMOS array sensor. Then install it at the front end of the underwater robot to collect the light field information of the environment where the underwater robot is located. It should be noted that the core function of the photoelectric sensor is to convert optical signals into electrical signals. The smallest photosensitive unit in the photoelectric sensor is a pixel, and each pixel contains a photosensitive diode. When light irradiates the photosensitive diode, corresponding charges will be generated, and these charges are then converted into electrical signals and processed by the readout circuit. The sensor consists of a pixel array, and each pixel is responsible for capturing, the filter, the metal wire arrangement, and the photosensitive diode. The electrical signal is processed by the analog signal processing unit and finally outputs a digital image signal;
[0074] It should also be noted that the radius of curvature of the curved surface array of the photoelectric sensor is 8 mm ± 0.5 mm, which can better simulate the optical characteristics of the human eye retina. The dual liquid lens is a lens composed of two liquids with different refractive indices, and the focal length is adjusted by changing the curvature of the interface between the two liquids;
[0075] Obtain the ambient light intensity Lt of the environment where the underwater robot is located at the current time point through the photoelectric sensor. It should be noted that the ambient light intensity L is obtained by converting the ambient light intensity into a voltage signal through a photodiode and a conversion circuit, and then converting it into a digital signal through an analog-to-digital converter for output;
[0076] According to the formula Ft = F0 + ΔF * sigmoid(0.5 * Lt), calculate and obtain the focal length Ft of the dual liquid lens adjusted according to the ambient light intensity Lt at the current time point, and adjust the focal length of the dual liquid lens according to the focal length Ft to make the focal length of the dual liquid lens reach Ft, where F0 is the initial focal length of the dual liquid lens under zero voltage, and the focal length change amount ΔF is the maximum change range of the focal length of the dual liquid lens. It should be noted that the initial focal length F0 and the maximum change range of the focal length of the dual liquid lens are usually determined by the design and manufacture of the lens and obtained through the information provided by the manufacturer;
[0077] It should also be noted that the focal length adjustment mechanism of the dual liquid lens can adaptively adjust the focal length according to the change of the ambient light intensity, so that the system can maintain good imaging performance under different lighting conditions. This dynamic adaptability enables the underwater robot to work stably in the complex and changeable underwater environment. The curved surface CMOS array sensor supports high dynamic range and can adapt to the sudden change of light in the deep-sea to shallow-sea transition area. This high dynamic range imaging ability enables the underwater robot to maintain the clarity and color authenticity of the image in the environment with drastic light changes.
[0078] As an optional embodiment: The specific working steps of the first step further include the following:
[0079] Establish a plane rectangular coordinate system, take the upper left corner of the curved surface CMOS array sensor as the coordinate origin, and obtain the coordinates (X, Y) of each photoelectric sensor pixel; it should be noted that although the curved surface CMOS array is a curved surface structure, during the manufacturing and use processes, the coordinates of each pixel can be determined by means of plane mapping. Specifically, the curved surface array can be laid out on a plane, and then the plane layout can be mapped onto the curved surface through a transfer technology;
[0080] According to the formula Obtain the light intensity change rate R of each photoelectric sensor pixel in time in real time, where represents the light intensity of this pixel point at position (X, Y) and time t, is the time difference between this calculation of the instantaneous change rate R and the previous calculation; it should be noted that by calculating the change rate R of the light intensity of each photoelectric sensor pixel in time, the dynamic change situation of the light intensity can be understood in real time. In some optical experiments, such as studying the flicker characteristics of a light source or the change of reflected light on the surface of an object, this change rate can help researchers accurately capture the instantaneous increase or decrease of the light intensity;
[0081] Set a threshold θ for the light intensity change rate in advance. If the light intensity change rate R of this photoelectric sensor pixel in time is greater than the threshold θ, then record the polarity of the light intensity change at this (X, Y) position at the moment when the image exceeds the preset threshold θ. The activated photoelectric sensor pixel transmits the data through the readout circuit; then the analog signal is converted into a digital signal by an analog-to-digital converter, and finally the digital image signal data is output and transmitted. It should be noted that traditional optical imaging systems usually use planar sensors and global shutters, resulting in motion blur underwater and a decline in imaging quality. Planar sensors cannot effectively capture light field information and are easily affected by light scattering and absorption. Traditional systems usually use lenses with fixed focal lengths and cannot adaptively adjust the focal length according to the change of ambient light intensity, resulting in a decline in imaging performance under different lighting conditions. This technical solution significantly reduces redundant data transmission through an event-driven imaging mechanism by only activating local pixels when the scene changes, not only improving the imaging efficiency but also reducing the power consumption of the system and extending the endurance time of the underwater robot.
[0082] As an optional embodiment: The specific working steps of the second step are as follows:
[0083] First, obtain acoustic data through a sonar sensor installed on the underwater robot;
[0084] Then receive the digital image signal data output in the first step;
[0085] Process the low-frequency features through a classical convolutional neural network to extract classical features. It should be noted that the low-frequency features are features with a frequency < 1 THz;
[0086] Process the high-frequency features through a quantum convolutional neural network to extract quantum features. It should be noted that the high-frequency features are features with a frequency > 1 THz;
[0087] According to the formula
[0088]
[0089] Convert the classical features from the time domain to the frequency domain to extract the global feature QFT(F), where n is the number of qubits and F(j) is the j-th component of the classical features. It should be noted that j and k are the summation indices, ranging from 0 to 2 n -1, and the complex exponential terms have different phases but the same modulus of 1. These complex numbers with different phases perform weighted superposition on different frequency components of the classical features. The low-frequency components usually represent the smooth parts of the image, while the high-frequency components represent the edges and details of the image. Through QFT, we can weight different frequency components, thereby highlighting important features during the fusion process;
[0090] It should also be noted that by converting the classical features from the time domain to the frequency domain, the frequency components of the signal can be analyzed more intuitively. The frequency domain representation can reveal the amplitude and phase information of different frequency components in the signal;
[0091] After converting the classical features from the time domain to the frequency domain, global features can be extracted. These global features can reflect the overall characteristics of the signal, rather than just local information;
[0092] Perform a convolution operation on the quantum features through quantum convolution to extract the local feature QConv(F), where the calculation formula is:
[0093]
[0094] where H k is the quantum Hamiltonian and θ k is the trainable parameter;
[0095] Through quantum entanglement, manage the global feature QFT(F) and the local feature QConv(F) extracted by quantum convolution to obtain the cross-domain fusion feature F3. The specific steps are as follows:
[0096]
[0097] Among them, U is the quantum entanglement gate operation, and Tr represents the partial trace operation on the auxiliary qubits. It should be noted that the quantum Hamiltonian is an operator that describes the energy of a quantum system. In quantum computing, the Hamiltonian is usually used to define the evolution and energy state of the system. The trainable parameters are adjustable parameters in the quantum circuit and are usually used to control the operation angles of quantum gates. The overall formula describes the evolution process of the quantum state in the quantum circuit. Through the quantum entanglement gate operation and the partial trace operation, the global features and local features can be fused.
[0098] As an alternative embodiment: Step 2 can also adaptively adjust the weights of the optical image and the sonar data according to the change of the ambient light intensity. The specific steps are as follows:
[0099] Obtain the ambient light intensity Lt of the environment where the underwater robot is located at the current time point through a photoelectric sensor. According to the formula Calculate the weight ω1 of the optical image, and then according to ω2 = 1 - ω1, obtain the weight ω2 of the sonar data. According to the weight ω1 of the optical image and the weight ω2 of the sonar data, dynamically adjust the weights of the optical image and the sonar data to perform data fusion. It should be noted that through quantum entanglement to achieve non-linear fusion of features can break through the limitations of the classical information superposition principle and significantly improve the correlation of cross-modal data.
[0100] Dynamically adjust the weights: Dynamically adjust the weights of the optical image and the sonar data according to the ambient light intensity, improve the fusion effect, reduce the model error, enable the system to adapt to different water turbidities and lighting conditions, and improve the dynamic range and adaptability of imaging.
[0101] As an alternative embodiment: The specific working steps of Step 3 are as follows:
[0102] Generate a beam carrying orbital angular momentum using a spiral phase plate, and its complex amplitude distribution is;
[0103]
[0104] Among them, L is the topological charge number, W0 is the beam waist radius, and P0 is the emission power; it should be noted that by calculating the complex amplitude distribution, the propagation characteristics of the beam in space can be understood, including the intensity distribution, phase distribution of the beam, and the focusing and diffusion characteristics of the beam;
[0105] Through the sonar of the underwater robot, obtain the measured water turbidity τ and light attenuation coefficient c, and construct a cross-media channel measurement model C;
[0106]
[0107] Where P m is the emission power of the m-th mode, is the noise power; it should be noted that by constructing the cross - medium channel capacity model C, the maximum information transmission rate of the underwater communication channel can be evaluated. This is crucial for designing an efficient underwater communication system because it can help determine the performance limit of the system under specific water conditions. When an underwater robot conducts data transmission, understanding the channel capacity can help select appropriate modulation and coding schemes to maximize data transmission efficiency;
[0108] By analyzing the transmission power and noise power of different modes, the parameter configuration of the communication system can be optimized. For example, the transmission power can be adjusted to adapt to different water conditions, thereby minimizing energy consumption while ensuring communication quality;
[0109] Then, the optimal power allocation strategy PM0 is solved according to the Lagrange multiplier method, and the specific steps are as follows;
[0110]
[0111] where δ is the multiplier satisfying the total power constraint, and (x) + represents max(x, 0). It should be noted that this scheme maximizes the channel capacity in turbid waters by dynamically adjusting the power allocation of OAM modes. When the water turbidity increases, the power ratio of high - order modes is automatically reduced, and instead, the transmission efficiency of low - order modes is enhanced, thus maintaining the communication rate within the turbidity range of 5 - 15 NTU.
[0112] As an optional embodiment: The specific working steps of step three further include the following:
[0113] During the sonar signal transmission period, the transmission timings of acoustic pulses and optical pulses are staggered through time - interleaved multiplexing technology to avoid mutual interference; it should be noted that to avoid the mutual interference between acoustic pulses and optical pulses, during the sonar signal transmission period, the transmission timings of acoustic pulses and optical pulses are staggered through time - interleaved multiplexing technology, and no specific parameters need to be calculated, mainly achieved through time scheduling;
[0114] Establish an acoustic - optical propagation time - delay difference model, and according to the formula calculate the time - delay difference Δt, where C1 is the sound speed in water and C2 is the sound speed in light; it should be noted that according to the known transmission distance d, sound speed, and light speed, directly substitute them into the formula to calculate the time - delay difference Δt;
[0115] Obtain the sonar received signal Y1 and the received signal Y2 of the optical image;
[0116] According to the formula obtain the parameter that maximizes the joint probability of the sonar and optical received signals Where θ is the parameter to be estimated. It should be noted that θ is an unknown parameter in the model and needs to be estimated through data. In sonar and optical signal processing, θ can be parameters such as the position, velocity, and shape of the target. Through the method of acousto-optic joint modulation, the observation data of sonar and optical image signals are fused, significantly improving the accuracy of target positioning. Through time-interleaved multiplexing technology and a matched filter bank, the mutual interference between acoustic pulses and optical pulses is effectively avoided, improving the anti-interference ability of the system. By maximum likelihood estimation to fuse acousto-optic observation data, the signal processing process is optimized, improving the overall performance of the system.
[0117] As an optional embodiment: The specific working steps of step three further include the following:
[0118] The specific acquisition methods of the sonar received signal Y1 and the received signal Y2 of the optical image are as follows;
[0119] For the sonar received signal Y1, through convolution operation, the transmitted signal of the sonar is convolved with the channel impulse response to obtain the received signal Y1. For the received signal Y2 of the optical image, based on the acquisition of Y1, the real part (Re) and conjugate (*) operations are taken.
[0120] As an optional embodiment: The specific steps of step four are as follows:
[0121] Construct a metasurface phase modulation model, and its phase response of the unit structure is;
[0122] Where μ1 is the equivalent permittivity, B is the operating wavelength, representing the wavelength of the electromagnetic wave, μ2 is the unit height, representing the physical height of the metasurface unit;
[0123] Then train a feedforward neural network to predict the wavefront aberration;
[0124] Finally, design a PID controller to adjust the voltage of the metasurface unit, and the control law is;
[0125]
[0126] Where eij is the phase error, K P is 0.8, K i is 0.2, K d is 0.05. It should be noted that the phase error needs to be obtained by acquiring the predicted phase and the measured phase. The predicted phase can be obtained through a theoretical model or a known reference signal, while the measured phase can be obtained through actual measurement;
[0127] It should be noted that under the condition of short response time, this programmable metasurface can correct most of the phase aberrations, greatly improving the imaging resolution in turbid waters compared with traditional methods.
[0128] As an alternative embodiment: The specific steps of training the feedforward neural network to predict wavefront distortion further include:
[0129] The input layer receives the original light field data of the curved surface CMOS array and the sonar point cloud data. The original light field data of the curved surface CMOS array contains the light intensity information of the underwater environment, while the sonar point cloud data provides the three-dimensional position information of underwater objects. These data serve as the input of the neural network and are used to predict the wavefront phase error distribution.
[0130] The hidden layer consists of 3 layers of fully connected networks, with 512 neurons in each layer. A fully connected network means that each neuron is connected to every neuron in the previous layer, capable of capturing the complex non-linear relationships in the input data. The activation function for each layer is LeakyReLU. It should be noted that LeakyReLU is an improved ReLU activation function, which has a small slope in the negative value part, avoiding the "dead zone" problem of ReLU and making it easier for the network to be optimized during training. The slope parameter α of LeakyReLU is 0.01, which means the slope in the negative value part is 0.01, capable of maintaining a certain gradient and avoiding neuron death.
[0131] The output layer predicts the wavefront phase error distribution. It should be noted that the output is a 64x64 matrix, and each element represents the wavefront phase error at the corresponding position. By predicting the wavefront distortion, it can provide a basis for subsequent metasurface phase regulation.
[0132] Working principle:
[0133] Through metasurface dynamic wavefront correction and intelligent regulation, it is possible to dynamically adjust the parameters of the imaging system under different water turbidities and lighting conditions, improve the imaging quality, and reduce model errors;
[0134] Because traditional optical imaging algorithms fail in turbid waters, mainly because they cannot effectively handle the problems of light scattering and absorption. When a single optical imaging algorithm processes images in turbid waters, it cannot restore the clarity and color authenticity of the images. Existing fusion technologies use a fixed fusion ratio and cannot adapt to dynamic environments. Under different water turbidities and lighting conditions, the fixed fusion ratio cannot dynamically adjust the weights of optical images and sonar data, resulting in poor fusion effects. In cross-water tests, models without physical constraints have large errors and cannot meet the imaging requirements of different waters;
[0135] Aiming at the failure problem of traditional optical imaging algorithms in turbid waters, this technical solution provides comprehensive solution steps. Each of the above steps is optimized for the defects of the existing technology, improving the adaptability, robustness, and imaging quality of the imaging system, and enabling efficient and accurate imaging in complex underwater environments.
[0136] The above are only the preferred embodiments of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions falling within the concept of the present invention belong to the protection scope of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements should also be regarded as within the protection scope of this template.
Claims
1. A method for optimizing underwater robot imaging algorithm based on deep learning, characterized in that: The following steps are involved: Step 1: First, a curved CMOS array sensor is used to capture underwater light field dynamics, and the focal length of the curved CMOS array sensor is adaptively adjusted according to the change of scene light intensity, and local pixels are activated when the scene changes to collect light field information of the environment where the underwater robot is located; Step 2: Design a quantum convolution kernel, map the classical convolution operation to a quantum circuit, build a hybrid computing pipeline, perform quantum state joint encoding and cross-domain feature fusion on the optical and acoustic data, and obtain the fused data; Step 3: Dynamically adjust the power allocation of OAM modes to maximize the channel capacity in turbid waters. When the turbidity of the water increases, the power share of high-order modes is automatically reduced, and the transmission efficiency of low-order modes is enhanced. Optical orbital angular momentum modulation is used to generate light beams to increase the capacity of cross-medium communication. Step 4: Design a programmable metasurface, predict the wavefront distortion through a feedforward neural network, and use a PID controller to adjust the metasurface parameters in real time to achieve subwavelength-level wavefront control.
2. According to the method for optimizing underwater robot imaging algorithm based on deep learning in claim 1, it is characterized in that: The specific working steps of step one are as follows: The photoelectric sensor curved array is transferred onto flexible rubber, and a double liquid lens is installed in front of it to obtain a complete curved CMOS array sensor, which is then installed at the front end of the underwater robot to collect light field information of the environment in which the underwater robot is located. The ambient light intensity Lt of the environment where the underwater robot is located at the current time point is obtained by the photoelectric sensor; according to the formula Ft=F0+ΔF*sigmoid(0.5*Lt), the focal length Ft of the double liquid lens adjusted according to the ambient light intensity Lt at the current time point is calculated and obtained, and the focal length of the double liquid lens is adjusted according to the focal length Ft to make the focal length of the double liquid lens reach Ft, wherein F0 is the initial focal length of the double liquid lens at zero voltage, and the focal length change ΔF is the maximum change range of the focal length of the double liquid lens.
3. The method for optimizing underwater robot imaging algorithm based on deep learning according to claim 2, characterized in that: The specific working steps of step one also include the following: Establish a plane rectangular coordinate system, take the upper left corner of the curved CMOS array sensor as the coordinate origin, and obtain the coordinates (X, Y) of each photoelectric sensor pixel; According to the formula The light intensity change rate R of each photoelectric sensor pixel over time is obtained in real time, where Indicates the light intensity of the pixel at position (X, Y) and time t, It is the time difference between the instantaneous rate of change R calculated this time and the last calculation; A threshold value θ of the light intensity change rate is set in advance. If the light intensity change rate R of the photoelectric sensor pixel over time is greater than the threshold θ, the polarity of the light intensity change at the (X, Y) position is recorded at the moment when the image exceeds the preset threshold θ, and the activated photoelectric sensor pixel transmits the data through the readout circuit; then the electrical signal is converted into a digital signal through the analog-to-digital converter, and finally the digital image signal data is output and transmitted.
4. The method for optimizing underwater robot imaging algorithm based on deep learning according to claim 3 is characterized in that: The specific working steps of step 2 are as follows: First, acoustic data is obtained through a sonar sensor installed on an underwater robot; Then receiving the digital image signal data outputted in step 1; Low-frequency features are processed through classic convolutional neural networks to extract classic features; The high-frequency features are processed by quantum convolutional neural network to extract quantum features; According to the formula The classical features are transformed from the time domain to the frequency domain, and the global features QFT(F) are extracted, where n is the number of quantum bits and F(j) is the jth component of the classical features; The quantum features are convolved by quantum convolution to extract the local features QConv(F), where the calculation formula is: Among them, H k is the quantum Hamiltonian, θ k is a trainable parameter; Through quantum entanglement, the global feature QFT(F) and the local feature QConv(F) extracted by quantum convolution are managed to obtain the cross-domain fusion feature F3. The specific steps are as follows: Where U is the quantum entanglement gate operation, and Tr represents the partial trace operation on the auxiliary quantum bit.
5. The method for optimizing underwater robot imaging algorithm based on deep learning according to claim 4, characterized in that: The step 2 can also adaptively adjust the weights of the optical image and the sonar data according to the change of the ambient light intensity. The specific steps are: The ambient light intensity Lt of the environment where the underwater robot is located at the current time point is obtained through the photoelectric sensor. According to the formula The weight ω1 of the optical image is calculated, and then the weight ω2 of the sonar data is obtained according to ω2=1-ω1. The weights of the optical image and the sonar data are dynamically adjusted according to the weight ω1 of the optical image and the weight ω2 of the sonar data to perform data fusion.
6. The method for optimizing underwater robot imaging algorithm based on deep learning according to claim 1, characterized in that: The specific working steps of step three are as follows: The spiral phase plate is used to generate a beam carrying orbital angular momentum, and its complex amplitude distribution is: Where L is the topological charge, W0 is the beam waist radius, and P0 is the transmission power. The measured water turbidity τ and light attenuation coefficient c are obtained through the sonar of the underwater robot, and the cross-medium channel measurement model C is constructed. Where P m is the transmission power of the mth mode, is the noise power; Then, the optimal power allocation strategy PM0 is solved according to the Lagrange multiplier method. The specific steps are as follows: Where δ is the multiplier that satisfies the total power constraint, (x) + Represents max(x, 0).
7. The method for optimizing underwater robot imaging algorithm based on deep learning according to claim 5, characterized in that: The specific working steps of step three also include the following: During the sonar signal transmission cycle, the transmission timing of the acoustic pulse and the optical pulse is staggered through time interleaving multiplexing technology to avoid mutual interference; Establish the acoustic and optical propagation delay difference model, according to the formula The time delay difference Δt is calculated, where C1 is the speed of sound in water and C2 is the speed of sound in light; Obtaining a sonar receiving signal Y1 and an optical image receiving signal Y2; According to the formula Get the parameters that maximize the joint probability of sonar and optical reception signals Where θ is the parameter to be estimated.
8. The method for optimizing underwater robot imaging algorithm based on deep learning according to claim 6, characterized in that: The specific working steps of step three also include the following: The specific method of obtaining the sonar receiving signal Y1 and the optical image receiving signal Y2 is as follows; For the sonar receiving signal Y1, the sonar transmitting signal is convolved with the channel impulse response through a convolution operation to obtain the receiving signal Y1. For the optical image receiving signal Y2, based on the acquisition of Y1, the real part (Re) and conjugate (*) operations are taken.
9. The method for optimizing underwater robot imaging algorithm based on deep learning according to claim 1, characterized in that: The specific steps of step 4 are: A metasurface phase control model is constructed, and the phase response of its unit structure is: Where μ1 is the equivalent dielectric constant, B is the working wavelength, which represents the wavelength of the electromagnetic wave, and μ2 is the unit height, which represents the physical height of the metasurface unit; Then the feed-forward neural network is trained to predict the wavefront distortion; Finally, a PID controller is designed to adjust the voltage of the metasurface unit, and the control law is: Where, eij is the phase error, K P is 0.8, K i is 0.2, K d is 0.
05.
10. The method for optimizing underwater robot imaging algorithm based on deep learning according to claim 9, characterized in that: The specific steps of training the feedforward neural network to predict wavefront distortion also include: The input layer receives the raw light field data and sonar point cloud data of the curved CMOS array. The raw light field data of the curved CMOS array contains the light intensity information of the underwater environment, while the sonar point cloud data provides the three-dimensional position information of underwater objects. These data are used as the input of the neural network to predict the wavefront phase error distribution. The hidden layer consists of 3 layers of fully connected networks, each with 512 neurons. A fully connected network means that each neuron is connected to each neuron in the previous layer, which can capture the complex nonlinear relationship of the input data. The activation function of each layer is LeakyReLU. The output layer predicts the wavefront phase error distribution.
Citation Information
Cited By
Intelligent benthonic animal identification method and system and storage medium
CN121582573A