Moving target real-time omnidirectional three-dimensional identification and distortion correction system based on deep knowledge prior diffraction neural network
Through a deep knowledge prior diffraction neural network system, combined with passive metasurfaces and deep learning, high-precision three-dimensional target recognition and distortion correction in dynamic environments are achieved, solving the problems of perspective dependence and insufficient real-time performance of traditional technologies. It is suitable for fields such as autonomous driving, industrial inspection and augmented reality.
Patent Information
- Application Number
- CN202510856843.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-09-26
AI Technical Summary
Existing three-dimensional object recognition technology has problems such as strong viewpoint dependence, insufficient real-time performance and weak distortion correction capability, making it difficult to achieve efficient three-dimensional target recognition and distortion correction in dynamic scenes.
A diffraction neural network system based on deep knowledge prior is adopted, combined with passive metasurface and deep learning, and the speed of light calculation and real-time distortion correction are realized through the diffraction effect of electromagnetic waves. The deep knowledge prior is used to generate network modules, diffraction neural network modules and dynamic monitoring modules for three-dimensional target recognition and distortion correction.
It achieves high-precision, low-power three-dimensional target recognition and distortion correction in dynamic environments. It is suitable for fields such as autonomous driving, industrial inspection and augmented reality. It overcomes the perspective, posture and distortion sensitivity of traditional technologies and provides reliable technical support.
Smart Images

Figure CN120707449A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the interdisciplinary field of computer vision, optical imaging and artificial intelligence, and more specifically to a real-time omnidirectional three-dimensional recognition and distortion correction system for moving targets based on a deep knowledge prior diffraction neural network. Background Art
[0002] Existing 3D object recognition technology has the following limitations: 1. Strong dependence on viewing angle: Traditional methods rely on fixed-view scanning and are difficult to cope with changes in object posture or occlusion.
[0003] 2. Insufficient real-time performance: Systems based on active devices require complex post-processing and cannot meet the needs of dynamic scenarios.
[0004] 3. Weak distortion correction capability: Existing diffraction neural networks are limited by the number of physical layers and are unable to handle 2D / 3D distortion simultaneously.
[0005] Therefore, to address the above problems, a system combining passive metasurfaces and deep knowledge prior DNN is proposed to achieve real-time omnidirectional recognition and distortion correction through light-speed calculations. Overcoming the efficiency and stability bottlenecks of traditional technologies is an urgent problem that technicians in this field need to solve. Summary of the Invention
[0006] In light of this, the present invention provides a real-time, omnidirectional, three-dimensional recognition and distortion correction system for moving targets based on a deep knowledge prior diffraction neural network. By leveraging the electromagnetic wave diffraction effect and deep learning optimization, this system achieves high-precision, low-power three-dimensional target recognition and distortion correction in dynamic environments. This system has broad applications in areas such as autonomous driving, industrial inspection, and augmented reality.
[0007] To achieve the above objectives, the present invention adopts the following technical solutions: a real-time omnidirectional three-dimensional recognition and distortion correction system for moving targets based on a deep knowledge prior diffraction neural network, comprising: a deep knowledge prior generation network module, a diffraction neural network module, and a dynamic monitoring module; The deep knowledge prior generation network module generates metasurface phase control parameters based on the input scattered field and Gaussian noise, and inputs the obtained phase control parameters into the diffraction neural network module; The diffraction neural network module processes the phase control parameters, implements light-speed signal processing based on a physically stacked metasurface array, obtains an electric field distribution, and inputs the electric field distribution into the dynamic monitoring module; The dynamic monitoring module analyzes the output electric field distribution in real time and reversely optimizes the parameters of the deep knowledge prior generation network module and the diffraction neural network module through a loss function during training.
[0008] Preferably, the deep knowledge prior generation network module inputs a random Gaussian noise sequence and generates optimized metasurface phase control parameters through a multi-layer convolutional neural network.
[0009] Preferably, the deep knowledge prior generation network module optimizes the phase distribution by using the phase distribution as prior knowledge to guide the training process to achieve the required transmission characteristics, and optimizes the transmission coefficient distribution as a two-dimensional image.
[0010] Preferably, the diffraction neural network module includes a multi-layer passive metasurface array, each layer of the diffraction neural network includes multiple metasurface units, the passive metasurface array realizes light speed calculation based on the Rayleigh-Sommerfeld diffraction model, and optimizes the parameters of the metasurface array through the MSE loss function.
[0011] Preferably, the dynamic monitoring module includes an electric field detection probe and a signal processing unit. Multiple electric field detection probes are provided to record the change of electric field intensity at each point in real time, and judge the target motion state through the change of field intensity distribution.
[0012] It can be seen from the above technical solution that compared with the existing technology, the present invention discloses a real-time omnidirectional three-dimensional recognition and distortion correction system for moving targets based on deep knowledge prior diffraction neural network. The system combines passive metasurface arrays with deep learning optimization to achieve efficient and low-power three-dimensional target recognition and dynamic distortion correction. It is suitable for real-time applications in complex environments, and can also achieve real-time and omnidirectional target recognition in complex dynamic environments, overcoming the sensitivity of traditional technologies to perspective, posture and distortion, and providing reliable technical support for autonomous driving, industrial inspection, augmented reality and other fields. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.
[0014] Figure 1 A module structure connection diagram provided by the present invention; FIG2( a ) is a diagram showing the working principle of the correction-perception diffraction neural network based on the deep knowledge prior generation network provided by the present invention; FIG2( b ) is a detailed network model diagram of the deep knowledge prior generation network provided by the present invention; FIG3( a ) is a schematic diagram of a polarization conversion unit in the transmission characteristics of a metasurface unit provided by the present invention; FIG3( b ) is a phase amplitude response diagram of the metasurface polarization conversion unit at 10.5 GHz in the transmission characteristics of the metasurface unit provided by the present invention; FIG4 (a) is a diagram showing the transmission coefficient training loss of each layer of the metasurface during the training process provided by the present invention; FIG4( b ) is a training loss diagram of the diffractive neural network provided by the present invention; FIG5( a ) is a diagram showing the planar distortion results in the multiple distortion correction experiment provided by the present invention; FIG5( b ) is a diagram showing the double random distortion results in the multiple distortion correction experiment provided by the present invention; Figure 6 (a) shows the left and right turning ( ) Electric field distribution diagram under changes; Figure 6 (b) shows the front and back tilt ( ) Electric field distribution diagram under changes; FIG6( c ) is a diagram showing the electric field distribution under omnidirectional movement in the dynamic three-dimensional perception experiment results provided by the present invention. DETAILED DESCRIPTION
[0015] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0016] like Figure 1 As shown, the embodiment of the present invention discloses a real-time omnidirectional three-dimensional recognition and distortion correction system for moving targets based on a deep knowledge prior diffraction neural network, including: a deep knowledge prior generation network module, a diffraction neural network module and a dynamic monitoring module; The deep knowledge prior generation network module generates metasurface phase control parameters based on the input scattered field and Gaussian noise, and inputs the obtained phase control parameters into the diffraction neural network module; The diffraction neural network module processes the phase control parameters, implements light-speed signal processing based on a physically stacked metasurface array, obtains an electric field distribution, and inputs the electric field distribution into the dynamic monitoring module; The dynamic monitoring module analyzes the output electric field distribution in real time and reversely optimizes the parameters of the deep knowledge prior generation network module and the diffraction neural network module through a loss function during training.
[0017] Furthermore, when a plane wave directly illuminates an object, it can be processed using a diffraction neural network module. There is no need to specifically collect the scattered field after the plane wave passes through the object. This scattered field propagates directly in the system.
[0018] Specifically, the deep knowledge prior generation network module inputs a random Gaussian noise sequence and generates optimized metasurface phase control parameters through a multi-layer convolutional neural network.
[0019] Specifically, the deep knowledge prior generation network module optimizes the phase distribution by using the phase distribution as prior knowledge to guide the training process to achieve the required transmission characteristics, and optimizes the transmission coefficient distribution as a two-dimensional image.
[0020] In a specific embodiment of the present invention, a 10-layer convolutional neural network structure is used, the input is a Gaussian random noise sequence (which conforms to the Gaussian distribution and is consistent with the number of metasurface units), and the output is the optimized metasurface phase distribution parameters; specifically, after being processed by the 10-layer convolutional neural network, the output is the optimized metasurface phase distribution. Furthermore, as shown in Figure 2(a) and Figure 2(b), the deep knowledge prior generation network module optimizes the phase distribution by using it as prior knowledge to guide the training process to achieve the desired transmission characteristics, and optimizes the transmission coefficient distribution as a two-dimensional image. Assume x is the transmission coefficient image generated during the optimization process, and the deep knowledge prior generation network is generated from the blurred image Start continuous optimization and generate the optimal transmission coefficient solution , the inverse problem of image restoration is transformed into an energy minimization problem, and the formula is as follows: ; The problem is divided into task-related data items and regular items. In the data items: ; Regularization term Corresponds to the prior in the generative network. This term makes the iterative image x Getting closer , so the objective function is x It has a certain smoothing effect and avoids overfitting. Convolutional networks are used to achieve image restoration (transfer coefficient generation) tasks, so the network is generated The random vector z Mapping to Image x . Replace the prior implicitly extracted by the convolutional neural network : ; ; in, zis a fixed random code initially input into the network, such as a random sequence of Gaussian distribution; are randomly initialized network parameters; is the optimal parameter solution obtained through training, corresponding to the optimal output of the network . In this method, the phase is regarded as an optimization target rather than a parameter to be trained directly. It is worth noting that the deep knowledge prior generation network does not require a large amount of training data, but instead iteratively updates the phase prior starting from an initial random sequence. Since the deep knowledge prior generation network is separated from the implementation of the physical system, its architecture can be flexibly customized, including the design of the number of layers, functions, and operations. This design increases the width and depth of the network and enhances the model's ability to represent and abstract complex information. Therefore, the network can learn more complex patterns and capture higher-order features. This flexible architecture provides a strong foundation for diffractive neural networks to utilize deep knowledge priors, enabling them to extract a large number of sample features and achieve excellent generalization performance.
[0021] As shown in Figure 3 (a)-Figure 3 (b), specifically, the diffraction neural network module includes a multi-layer passive metasurface array, each layer of the diffraction neural network includes multiple metasurface units, each layer contains multiple phase modulation units, and the passive metasurface array realizes light speed calculation based on the Rayleigh-Sommerfeld diffraction model, and optimizes the parameters of the metasurface array through the MSE loss function.
[0022] Furthermore, in Figure 3(a), α represents the arm length, and β represents the physical parameter that controls the phase. Changing these two parameters changes the phase. α and β are direct parameters optimized by the neural network, and these two parameters determine the transmission coefficient (which consists of amplitude and phase).
[0023] In a specific embodiment of the present invention, the metasurface array adopts a three-layer metal-two-layer dielectric structure; a plane wave irradiates the target object, undergoes primary modulation by the first layer of the metasurface, performs preliminary distortion compensation on the incident diffraction field, and obtains a preliminary modulated field distribution, and uses the second layer of the metasurface to perform secondary modulation on the preliminary modulated field distribution based on the Rayleigh-Sommerfeld diffraction model, and uses the deep knowledge prior to generate the phase control precise compensation parameters of the network module on the third layer of the metasurface to generate the final phase distribution for the metasurface, and outputs the electric field distribution at the output level according to the phase distribution.
[0024] Specifically, the metasurface unit adopts a three-layer metal-two-layer dielectric structure: Metal layer: 0.035mm thick copper layer with a specific grid pattern design Dielectric layer: F4B material, each layer thickness 2 mm Unit period: 6mm Overall size: 384mm×384mm (64×64 unit array) The design achieves a phase coverage of 2π and a transmittance above 0.8, with an operating frequency band of 8-12 GHz.
[0025] The Diffractive Neural Network workflow includes: (1) Incident light processing: A plane wave illuminates the target object to generate a diffraction field distribution, which is then modulated using the first metasurface layer. (2) Intermediate processing: After the initial modulation, the second metasurface is used for secondary modulation, and the final phase modulation is completed on the third metasurface. (3) Output imaging: A clear image is formed on the output plane, and the probe scans and records the electric field distribution, and the recognition results are displayed in real time.
[0026] As shown in Figure 4 (a) and Figure 4 (b), each layer of the diffractive neural network is composed of controllable metasurface units. For the propagation of time-harmonic electromagnetic waves in free space, its control effect can be equivalently expressed as a complex transmission function. Assume that l Layer ( m,n The complex transmission function of each unit is: ; in and They represent the unit's ability to control the amplitude and phase of the incident electromagnetic wave. After modulation in this layer, the wave propagates to the next layer or receiving surface in the Rayleigh–Sommerfeld form, and its spatial propagation kernel function is: ; in, , . Therefore the layer is at position ( x,y ) can be expressed as: ; By recursively carrying out the above process, we can obtain the final output electromagnetic field distribution at the output layer (perception layer): .
[0027] During the training phase, by constructing the loss function LOSS To measure the difference between the output field and the target response (such as the output plane distribution of multiple sources), and calculate the gradient of each unit in the network based on the complex back propagation algorithm. l Layer ( m, n ) units, the gradient calculation expression is as follows: ; By continuously iteratively updating the control parameters of each layer of metasurface units, the network's perception of electromagnetic information in complex multi-source scenarios is optimized.
[0028] Furthermore, the loss function uses the MSE function.
[0029] Specifically, the dynamic monitoring module includes an electric field detection probe and a signal processing unit. Multiple electric field detection probes are set to record the changes in electric field intensity at each point in real time, and judge the target movement state through the changes in field intensity distribution.
[0030] In a specific embodiment of the present invention, referring to Figures 5(a), 5(b), 6(a), 6(b), and 6(c), the system performance is verified by the following tests: (1) Static target recognition: aircraft models in different postures, with a recognition accuracy of 98%; (2) Dynamic target tracking: ±45° yaw angle change, tracking error <2°; (3) Distortion correction: Correct more than three types of distortion simultaneously, SSIM>0.9.
[0031] Furthermore, the distortion correction implementation process includes: (1) Input distorted image (e.g. 30° viewing angle tilt + rotation distortion); (2) Modulated by metasurface array; (3) Output the corrected image on the plane; (4) Calculate the structural similarity index SSIM (which can reach 0.95 using the calculation method of the present invention).
[0032] Furthermore, dynamic target tracking includes: Dynamic recognition process: (1) Set up 6 key monitoring points (such as nose, wings, etc.); (2) Real-time recording of the electric field intensity at each point; (3) When the target is moving (e.g. ±15° pitch change): When the field strength at each monitoring point remains stable (approximately -60dB), and the field strength in the non-target area is <-80dB (4) Determine the target's motion state through changes in field intensity distribution.
[0033] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Reference can be made to the common and similar parts between the various embodiments. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method description.
[0034] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not limited to the embodiments shown herein but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A real-time omnidirectional three-dimensional moving target recognition and distortion correction system based on a deep knowledge prior diffraction neural network, characterized by: include: Deep knowledge prior generation network module, diffraction neural network module and dynamic monitoring module; The deep knowledge prior generation network module generates metasurface phase control parameters based on the input scattered field and Gaussian noise, and inputs the obtained phase control parameters into the diffraction neural network module; The diffraction neural network module processes the phase control parameters, implements light-speed signal processing based on a physically stacked metasurface array, obtains an electric field distribution, and inputs the electric field distribution into the dynamic monitoring module; The dynamic monitoring module analyzes the output electric field distribution in real time and reversely optimizes the parameters of the deep knowledge prior generation network module and the diffraction neural network module through a loss function during training.
2. The system for real-time omnidirectional three-dimensional recognition and distortion correction of moving targets based on a deep knowledge prior diffraction neural network according to claim 1, characterized in that: The deep knowledge prior generation network module inputs a random Gaussian noise sequence and generates optimized metasurface phase control parameters through a multi-layer convolutional neural network.
3. The system for real-time omnidirectional three-dimensional recognition and distortion correction of moving targets based on a deep knowledge prior diffraction neural network according to claim 1, characterized in that: The deep knowledge prior generation network module optimizes the phase distribution by using the phase distribution as prior knowledge to guide the training process to achieve the required transmission characteristics, and optimizes the transmission coefficient distribution as a two-dimensional image.
4. The system for real-time omnidirectional three-dimensional recognition and distortion correction of moving targets based on a deep knowledge prior diffraction neural network according to claim 1, characterized in that: The diffraction neural network module includes a multi-layer passive metasurface array, each layer of the diffraction neural network includes multiple metasurface units, the passive metasurface array realizes light speed calculation based on the Rayleigh-Sommerfeld diffraction model, and optimizes the parameters of the metasurface array through the MSE loss function.
5. The system for real-time omnidirectional three-dimensional recognition and distortion correction of moving targets based on a deep knowledge prior diffraction neural network according to claim 1, characterized in that: The dynamic monitoring module includes an electric field detection probe and a signal processing unit. Multiple electric field detection probes are set to record the changes in electric field intensity at each point in real time and judge the target movement state through the changes in field intensity distribution.