A multi-modal phased array radar target recognition method and system

By aligning spatiotemporal baselines and caching synchronized multimodal data in the data pool, combined with the feature fusion model of phased array radar and complex functions, the problem of multimodal data fusion is solved, and high-precision and highly robust three-dimensional target detection is achieved, which is suitable for autonomous driving and automatic navigation.

CN120451635BActive Publication Date: 2025-10-21CHENGDU HETAICHENG TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510503523.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-10-21
Estimated Expiration
2045-04-22

AI Technical Summary

Technical Problem

Existing technologies are unable to effectively fuse data from multiple sensors, resulting in insufficient target detection accuracy and robustness, making it difficult to achieve high-precision and robust three-dimensional target detection in complex environments.

Method used

Multimodal data synchronization and calibration are achieved through spatiotemporal baseline alignment and data pool caching. Cross-modal geometric coarse registration guided by phased array radar and refined registration based on coupled partial differential equations are used, combined with the multimodal feature fusion mathematical model of complex functions to generate fused multimodal features, which are finally identified through a CNN target classifier.

Benefits of technology

It improves the accuracy and robustness of target detection, fully utilizes the advantages of multiple sensors, and significantly improves the overall performance of CNN neural network target detection. It is suitable for the fields of autonomous driving and automatic navigation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120451635B_ABST
    Figure CN120451635B_ABST
Patent Text Reader

Abstract

The application discloses a kind of multi-modal phased array radar target identification method and system, belong to radar signal processing field, the method includes: the data of different modalities is preprocessed, obtains the multi-modal data coupled in time;Coupled multi-modal data in time is preliminarily registered, and rough multi-modal data is obtained;Rough multi-modal data is finely registered, and higher accuracy spatial alignment is realized, obtains a multi-modal data cube;Through multi-modal feature fusion mathematical model, the generated multi-modal data cube is fused, and a final fused feature after fusion is calculated;Classification is carried out through combined CNN target classifier, and classification result is obtained, and the identification of target is realized.The application significantly improves the overall performance of phased array radar three-dimensional target detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of radar signal processing, and in particular to a multi-modal phased array radar target recognition method and system. Background Art

[0002] In existing technologies, although radar can provide high-precision three-dimensional spatial information, its resolution is low and it cannot obtain rich texture and color information. Optical sensors can capture image information rich in details, but they are greatly affected by lighting conditions and cannot directly obtain the depth information of the target. Therefore, a single sensor cannot meet the needs of target detection in complex environments. At the same time, existing technologies cannot fully utilize the advantages of multiple sensors, making it difficult to achieve high-precision and robust three-dimensional target detection. How to effectively combine the advantages of different modal sensors and fuse multi-source heterogeneous information such as radar, optical, and infrared to improve the accuracy and robustness of target detection has become a key issue that needs to be addressed. To this end, it is necessary to provide a three-dimensional target detection method based on multimodal sensor fusion. Through means such as spatiotemporal baseline alignment and data pool caching, it is necessary to achieve synchronization and calibration of multimodal data. The multimodal alignment data cube is obtained by using cross-modal geometric registration guided by phased array radar and refined registration based on coupled partial differential equations. Then, the multimodal data is fused through the multimodal feature fusion mathematical model of complex functions to generate fused multimodal features, and finally high-precision and high-robust three-dimensional target detection is achieved through the CNN target classifier. Summary of the Invention

[0003] One of the purposes of the present invention is to provide a multi-modal phased array radar target recognition method to solve the problem in the prior art that it is impossible to effectively fuse multiple sensor data and it is difficult to fully utilize the advantages of different sensors, resulting in insufficient target detection accuracy and robustness.

[0004] The present invention is implemented through the following technical solution, a multi-modal phased array radar target recognition method, comprising the following steps: S100, preprocessing data of different modes to obtain multi-modal data coupled in time, wherein the data of different modes include phased array radar data, optical sensor data and infrared sensor data; S200, performing preliminary alignment on the multi-modal data coupled in time, wherein the preliminary alignment implements cross-modal geometric coarse alignment guided by the phased array radar data through affine transformation to obtain coarse multi-modal data; S300, performing fine alignment on the coarse multi-modal data to implement A higher-precision spatial alignment is achieved, and recognizable feature points or areas are extracted from data of different modalities through feature combination, and the alignment results are fine-tuned through optimization means to obtain a multimodal data cube; S400, the generated multimodal data cube is fused through a multimodal feature fusion mathematical model, and a final fused feature is calculated; S500, the final fused feature is complex-decomposed and converted into a real feature suitable for processing by a CNN target classifier, and the classification is performed by the CNN target classifier to obtain a classification result, thereby realizing target recognition.

[0005] Furthermore, the preprocessing is a spatiotemporal baseline alignment operation to ensure that data of different modalities are aligned under the same time reference, including: time synchronization and spatial alignment. Among them, time synchronization is achieved by deploying a central clock or clock source containing a precise clock protocol at the data source end of different modalities to stamp the data of different modalities with the same timestamp, thereby achieving the unification of data timestamps of different modalities; spatial alignment is achieved by calculating the rigid body transformation matrix between the radar coordinate system and the optical camera coordinate system, and converting the phased array radar data coordinates to the optical sensor data coordinate system through the rigid body transformation matrix.

[0006] Furthermore, the spatiotemporal baseline alignment operation implements the timestamp alignment operation in the data pool cache by setting up a data pool cache to store and manage the raw data from different sensors, ensuring that the data from different sensors can be docked within the correct time window.

[0007] Furthermore, the data pool is cached as a ring buffer data pool. When the buffer of the ring buffer data pool is full, the new data will overwrite the oldest data to ensure that the latest time series data is always maintained, thereby ensuring that all data can be effectively managed and adjusting the buffer size and data processing method according to the sampling frequency of different sensors.

[0008] Furthermore, step S200 includes the following sub-steps: S210, projecting the radar point cloud of the phased array radar data along the depth axis to simulate the viewing angle of the optical sensor, forming a virtual projection view of the phased array radar data, and obtaining a two-dimensional radar virtual image; S220, performing an affine transformation on the optical image based on the rigid body transformation matrix and the radar virtual image, and roughly aligning the target in the optical / infrared image to the coordinate system of the radar point cloud.

[0009] Furthermore, step S200 may also include sub-step S230: considering that in actual application, the sensor platform serving as the source of different modal data is often dynamically moving, the sensor motion can be monitored and adjusted in real time through dynamic parameter estimation, and the Kalman filter can be used to fuse IMU data and other sensor information to estimate the instantaneous posture changes of the sensor. Through these estimated posture changes, the projection matrix of the radar virtual projection is updated in real time.

[0010] Furthermore, the refined registration in step S300 includes the following sub-steps: S310, extracting gradient information of different input modalities to reflect the local features of different modalities; S320, combining the gradient information of different modalities, iteratively updating the displacement field of the optical image and the displacement field of the infrared image, and finely aligning the images of different modalities; S330, using the updated displacement field of the optical image and the displacement field of the infrared image to align the images of each modality, aligning the optical image and the infrared image to the radar image, and generating a multimodal alignment data cube based on the alignment data: ,in, is a data cube, is a set of real numbers, indicating that the data cube is a four-dimensional tensor, and the value of each position is composed of real numbers; is the height of the image, is the width of the image, is the feature dimension, and 3 represents the number of multi-modal channels, including phased array radar R, optical O, and infrared I modes.

[0011] Furthermore, the update of the displacement field can be achieved by constructing a feature coupling-driven displacement field equation group, which includes: a displacement field update equation for the optical image and a displacement field update equation for the infrared image, wherein the displacement field update equation for the optical image is obtained by calculating the similarity gradient between the radar image and the optical image to adjust the displacement field of the optical image, and then combining the difference between the displacement fields of the optical image and the infrared image to coordinate the transformation of the two images, and finally using a smoothing term to ensure that the displacement field of the optical image changes smoothly, thereby obtaining the update; the displacement field update equation for the infrared image is obtained by calculating the similarity gradient between the radar image and the infrared image to adjust the displacement field of the infrared image, and then calculating the difference between the displacement fields of the infrared image and the optical image to coordinate the transformation of the two images, and using a smoothing term to ensure that the displacement field of the infrared image changes smoothly, thereby obtaining the update.

[0012] Furthermore, the displacement field equations driven by characteristic coupling can be expressed as follows:

[0013] ,

[0014] In the above formula, Update equations for the displacement field of the optical image; is the displacement field update equation of the infrared image; where, is the displacement field of the optical image, is the displacement field of the infrared image, is the time rate of change of the optical image displacement field, is the similarity gradient control parameter of the optical image, is the radar-optical signature similarity function; is the control parameter for the consistency of the displacement field between the infrared image and the optical image, is the control parameter of the smoothness of the displacement field of the optical image, is the time rate of change of the displacement field of the infrared image, is the similarity gradient control parameter of the infrared image, It is the control parameter of the smoothness of the displacement field of the infrared image; is the radar-infrared feature similarity function, is the gradient of the optical sensor data, is the gradient of the infrared sensor data, is the gradient of the projected characteristic field of the phased array radar.

[0015] Furthermore, the radar-optical signature similarity function can be expressed as follows:

[0016] ,in, is an exponential decay function, is the adjustment parameter of the radar-optical signature similarity function, is a transformation function that represents the gradient of the optical image through the displacement field The result after transformation.

[0017] Furthermore, the radar-infrared feature similarity function can be expressed as follows:

[0018] ,in, is the hyperbolic tangent function, is the adjustment parameter of the radar-infrared feature similarity function, is the mutual information between radar and infrared images, reflecting the statistical correlation between the two.

[0019] Furthermore, the multimodal feature fusion mathematical model in step S400 is constructed by the following sub-steps: S410, splitting the data cube into different modal data, and mapping each modal data to the complex plane to encode its physical properties; S420, constructing a holomorphic mapping fusion equation for realizing the holomorphic mapping, thereby fusing the complex forms of different modalities into a unified complex space to obtain the fused features; S430, projecting the fused features into a compact pseudo-Euclidean space in the complex hyperbolic space through the Klein model to obtain the final fused features, and the Klein model is expressed by the following formula:

[0020] ,in, is the final fusion feature, is the logarithmic sign, is the holomorphic mapping fusion equation, is a complex hyperbolic space, is the norm of the fusion feature in the complex hyperbolic space calculated by the holomorphic mapping fusion equation, which represents the size of the feature. It is the inner product of the fusion features calculated by the holomorphic mapping fusion equation in the complex hyperbolic space.

[0021] Furthermore, the data cube may include: radar mode, optical mode and infrared mode, among which, the radar mode can be represented by amplitude-phase, where the amplitude is the echo intensity and the phase is dynamically adjusted by beamforming to reflect the microscopic motion characteristics of the target, thereby constructing the representation of the radar mode in complex space; the optical mode, the real part is the texture mean and the imaginary part is the edge gradient, and the stability of the spatial structure information is reflected by complex numbers, thereby constructing the representation of the optical mode in complex space; the infrared mode, the thermal gradient amplitude is used as the mode and the dynamic heat diffusion phase is used as the argument, combined with the heat conduction model to describe the time domain characteristics of temperature change, thereby constructing the representation of the infrared mode in complex space.

[0022] Furthermore, the representation of radar mode in complex space can be expressed as follows:

[0023] ,in, is an imaginary unit; is a natural constant; is the complex value corresponding to the radar image at a certain spatial position (x, y) and time point t, is the amplitude of the radar image at a certain spatial position (x, y) and time point t, is the phase of the radar image at a certain spatial position (x, y) and time point t.

[0024] Furthermore, the representation of the optical mode in complex space can be characterized by the following formula:

[0025] ,in, is the complex value corresponding to the optical image at a certain spatial position (x, y) and time point t; is the texture information of the optical image at a certain spatial position (x, y) and time point t, is the edge gradient information of the optical image at a certain spatial position (x, y) and time point t.

[0026] Furthermore, the representation of the infrared modality in complex space can be characterized by the following formula:

[0027] ,in, is the complex value corresponding to the infrared image at a certain spatial position (x, y) and time point t; is the thermal gradient amplitude of the infrared image at a certain spatial position (x, y) and time point t, reflecting the changes between different temperature regions; is the dynamic temperature phase of the infrared image at a certain spatial position (x, y) and time point t.

[0028] Furthermore, the result of the holomorphic mapping can be obtained by the following holomorphic mapping fusion equation:

[0029] ,

[0030] in, is a holomorphic mapping fusion function, is the representation of radar mode in complex space, is the representation of the optical mode in complex space, is the representation of infrared mode in complex space, for The partial derivative of the complex conjugate of , for The complex conjugate of is the partial derivative condition, which requires that the optical mode satisfies holomorphism, meaning that the complex must vary along the real direction in the complex plane to preserve analytical properties; is the radar phase stability weight function, which is used to suppress the jump interference in the radar phase; is the norm of the radar mode, which is used to represent the size of the radar mode in the complex space. Represents a unit circle or a constrained space; is the optical edge continuity weight function, which is used to emphasize the edge information of the image; It is the infrared external phase modulation function, which is used to adjust the phase information of the infrared signal;

[0031] Furthermore, the radar phase stability weight function can be expressed as follows:

[0032] ,

[0033] in, It is an activation function used to process the second-order gradient of the radar signal, thereby reducing unnecessary changes in the radar image, ignoring negative gradient values, and retaining only the positive mutation part; A scale parameter that controls the smoothing sensitivity of radar signals. It is used to standardize the degree of mutation and prevent excessive Laplace values ​​from causing the exponential function to approach 0 (i.e., almost completely suppressing the relevant region). This parameter can be selected based on the actual phase changes of the radar image in the data, empirically set, or through methods such as cross-validation. is the radar modal complex amplitude The Laplace operator (second-order gradient) is used to measure the phase change rate or intensity change of a region in the image.

[0034] Furthermore, the optical edge continuity weight function can be expressed as follows:

[0035] , the weight function represents the gradient of the optical mode in the complex space , the value of the continuity weight will be larger in the edge part of the image and smaller in the smooth part, thus ensuring that the structure part in the image occupies a larger weight in the fusion process.

[0036] Furthermore, the infrared out-of-band phase modulation function can be expressed as follows:

[0037] ,in, is a complex exponential term used to modulate the phase of the infrared signal to ensure phase synchronization of the infrared signal, thus avoiding instability caused by phase difference during the fusion process; Indicates the change in the phase of the infrared signal; for about The derivative of , usually refers to the local gradient or rate of change of the infrared signal in image processing, which reveals the changing trend of the infrared signal in space. is the sign function, used to obtain the sign of the imaginary part.

[0038] Furthermore, step S500 includes the following sub-steps: S510, performing complex decomposition, splitting the final fusion feature into two real channels of amplitude and phase, the amplitude channel reflects the intensity or energy of the signal, and the phase channel describes the relative phase or time-space relationship of the signal, which is used to represent the directionality or periodic change of the signal; S520, splicing the real part, imaginary part, amplitude and phase features obtained by the complex decomposition into different input channels to form a multi-channel feature map; S530, inputting the multi-channel feature map into the CNN target classifier for target classification, and the CNN will extract high-level features through multiple convolutional layers, and finally perform classification through the fully connected layer.

[0039] Furthermore, the two real channels of amplitude and phase are expressed as:

[0040] ,in, is the final fusion feature, which is a complex signal; is the amplitude channel, is the amplitude; is the amplitude of the complex feature, indicating the size of the complex signal; is the phase channel, is the phase, is the phase of the complex characteristic.

[0041] Furthermore, the amplitude of the complex feature is calculated by the following formula:

[0042] ,in, is the real part of the complex characteristic, is the imaginary part of the complex characteristic.

[0043] Furthermore, the phase of the complex feature is calculated by the following formula:

[0044] ,in, is the inverse tangent function.

[0045] Furthermore, the CNN target classifier may include: convolution layer: the convolution layer is used to extract local patterns of input features. The convolution layer slides in space, extracts features by learning convolution kernels, and generates convolution feature maps; pooling layer: the pooling layer is used to downsample the convolution feature maps to reduce the amount of calculation while retaining important spatial information; fully connected layer: the features after convolution and pooling are sent to the fully connected layer for high-level abstraction and classification tasks. In the fully connected layer, the neural network calculates the probability distribution of the category based on the input features, and finally outputs the classification results to achieve target recognition.

[0046] On the other hand, the present invention provides a multi-modal phased array radar target recognition system, which includes a processor and a memory, wherein a computer program is stored in the memory. When the computer program is executed by the processor, the multi-modal phased array radar target recognition method as described above is implemented.

[0047] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0048] 1. Through spatiotemporal baseline alignment and data pool caching, the present invention achieves precise synchronization, calibration, and alignment of different modal data from phased array radars, optical radars, and infrared sensors, providing consistent, high-quality input for subsequent data processing and target recognition, and improving the accuracy and robustness of target detection.

[0049] 2. The present invention utilizes cross-modal geometric coarse registration guided by phased array radar and refined registration based on coupled partial differential equations to achieve high-precision spatial alignment of data of different modalities, obtain a multimodal alignment data cube, and lay the foundation for multimodal feature fusion.

[0050] 3. The present invention designs a multimodal feature fusion mathematical model based on complex functions, which effectively integrates three types of multi-source heterogeneous information, namely radar, optical and infrared, fully mines the feature information of different modal data, generates high-quality fusion features, and improves the accuracy and robustness of target detection.

[0051] 4. The present invention converts the fused complex features into real features suitable for CNN target classifier input, achieving high-precision and high-robust three-dimensional target detection, overcoming the shortcomings of insufficient detection accuracy and robustness of a single sensor, and making full use of the advantages of multiple sensors. Through effective multimodal data fusion, the overall performance of CNN neural network target detection is significantly improved, providing strong support for target recognition in complex environments, and can be widely used in the fields of autonomous driving and automatic navigation. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] The drawings described herein are used to provide a further understanding of the embodiments of the present invention, constitute a part of this application, and do not constitute a limitation of the embodiments of the present invention. In the drawings:

[0053] Figure 1 This is a flow chart of the method provided in Example 1 of the present invention. DETAILED DESCRIPTION

[0054] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions of the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings herein can be arranged and designed in various different configurations.

[0055] Example 1

[0056] In existing phased array radar target recognition, due to the lack of time synchronization, calibration and alignment between multiple modal data, it is difficult to geometrically align multimodal data. The only way is to effectively fuse multiple sensor data to achieve accurate target recognition. Therefore, in order to address the shortcomings of the existing technology, this embodiment discloses a multimodal phased array radar target recognition method. The multimodal phased array radar target recognition method disclosed in this embodiment achieves synchronization and calibration of multimodal data through means such as spatiotemporal baseline alignment and data pool caching. With phased array radar data as the core, the method guides cross-modal geometric precise alignment based on the constructed coupled partial differential equation to obtain a multimodal data cube. Then, the multimodal data is fused through a multimodal feature fusion mathematical model based on complex functions to generate fused multimodal features. Finally, a CNN target classifier is used to achieve high-precision and high-robust three-dimensional target detection based on phased array radar.

[0057] Figure 1 The flowchart of the overall method in this embodiment is shown. It can be seen from the figure that this embodiment includes the following steps:

[0058] Step 1: Synchronize, calibrate, and align the input data from the phased array radar, optical sensors (e.g., COMS or CCD sensors), and infrared sensors to provide consistent input for subsequent data processing and target recognition. In this embodiment, synchronization, calibration, and alignment can be achieved through spatiotemporal baseline alignment and data pool caching.

[0059] Specifically, the goal of spatiotemporal baseline alignment is to ensure that data from different sensors are aligned under the same time reference, and to ensure that data from radar, optical, and infrared sensors can be compared and fused in the same time coordinate system, including time synchronization and spatial alignment.

[0060] In this embodiment, precise time synchronization can be achieved through the Precision Time Protocol (IEEE.1588.PTP). In multi-modal sensor systems such as phased array radar, optical sensors, and infrared sensors, the sampling frequencies and clocks of the radar, optical, and infrared sensors often differ (for example, the radar may sample at 1ms / frame, while the optical sensor may sample at 30ms / frame). These differences in sampling intervals make it challenging to time-align the data from different sensors. To address this issue, the PTP protocol is used to unify sensor timestamps, ensuring that the clocks of the radar, optical, and infrared sensors remain synchronized. A central clock or clock source (such as a GPS clock) can be deployed at the backend of the phased array radar, optical, and infrared sensors, and this time source can be distributed to all sensors using the PTP protocol. This allows all sensors of all modalities to generate timestamps at the same time, ensuring that each sensor's data stream is accompanied by timestamp information. During subsequent processing, these timestamps can be used to precisely align measurement results from different sensors, thereby eliminating errors caused by asynchronous sensor clocks.

[0061] Furthermore, because radar, optical, and infrared sensor data typically exist in different coordinate systems (for example, radar data is typically expressed in a three-dimensional Cartesian coordinate system (XYZ), while optical sensor data is expressed in a two-dimensional pixel coordinate system (UV)), data fusion can be achieved by calculating the rigid body transformation matrix between the radar coordinate system and the optical / infrared coordinate system. This transformation matrix can be obtained through a calibration process. In this embodiment, a checkerboard calibration plate can be used for calibration. By capturing checkerboard images at different angles and aligning feature points in the radar and camera, the rigid body transformation matrix between the radar coordinate system and the optical / infrared coordinate system is calculated. This matrix can be used to convert the radar point cloud coordinates into the pixel coordinate system of the optical / infrared coordinate system, thereby providing accurate spatial alignment for subsequent data fusion.

[0062] To ensure smooth spatiotemporal baseline alignment, a data pool cache is set up to store and manage raw data from different sensors and align their timestamps, ensuring that data from different sensors can be connected within the correct time window. In this embodiment, a ring buffer (FIFO queue) is used to store real-time data streams. When the buffer is full, new data overwrites the oldest data, ensuring that the latest time series data is always available. For each sensor (such as radar, optical, and infrared), a separate ring buffer is established. Each buffer stores the corresponding data frame by timestamp. The capacity of the ring buffer can be set based on the sensor's sampling frequency and data latency.

[0063] In the ring buffer, data is stored using timestamps as key indexes. Each frame of the time-synchronized data stream (whether radar, optical, or infrared) is timestamped, and these timestamps are unified based on hardware synchronization. For each frame of radar data, the optical and infrared data with the corresponding timestamp are searched and read. Interpolation or delay compensation is performed on this data based on the timestamp, ensuring that all sensor data is processed and fused using the same time base. Ultimately, data from all sensors is aligned to a common timeline, ensuring spatial and temporal consistency, providing reliable input for subsequent target recognition, tracking, and fusion.

[0064] It's important to note that in this step, spatial alignment is achieved through hardware synchronization and coordinate calibration. The construction of a ring buffer (FIFO queue) ensures efficient data management, and the buffer size and data processing methods can be adjusted based on the sampling frequency of different sensors. This provides reliable data support for multimodal data fusion and target recognition, laying the foundation for accurate identification and efficient tracking.

[0065] Step 2: During the registration of multimodal sensor data, the phased array radar provides high-precision three-dimensional spatial information, while the optical and infrared sensors provide two-dimensional image information. Step 1, through coordinate calibration, has already obtained the rigid body transformation matrix between the radar coordinate system and the optical / infrared coordinate system. This rigid body transformation matrix describes the relative position and attitude between the radar coordinate system and the optical / infrared coordinate system. Based on this transformation matrix, an affine transformation can be performed on the optical or infrared image, achieving coarse cross-modal geometric registration guided by the phased array radar data. By geometrically aligning the three-dimensional point cloud data generated by the phased array radar with the optical / infrared image information, a coarse registration foundation is formed, resulting in coarsely registered multimodal data for subsequent fine registration and fusion.

[0066] The specific steps include: 1) generating a virtual projection view according to the phased array radar.

[0067] Phased array radars generate high-resolution 3D point cloud data, typically expressed in a radar coordinate system, containing the target's spatial coordinates (X, Y, Z). Optical and infrared cameras, on the other hand, typically work as 2D images, so it's necessary to convert the 3D point cloud data into a 2D view. Projecting a 3D point cloud onto a 2D plane can be accomplished by selecting a specific axis of the radar coordinate system (such as the depth axis). In this embodiment, the radar point cloud is projected along the depth axis (Z axis), simulating the perspective of an optical camera and forming a 2D virtual radar image.

[0068] 2) Based on the rigid body transformation matrix obtained in step 1 and the radar virtual image, an affine transformation is performed on the optical image. This roughly aligns the target in the optical / infrared image with the radar point cloud coordinate system. Affine transformations include rotation, translation, and scaling operations, and their goal is to roughly align the two (the optical / infrared image and the radar virtual image). In other words, an affine transformation can move, rotate, or scale the optical image in space so that it roughly matches the radar virtual image. This is only a rough alignment and may contain errors, but it lays the foundation for subsequent fine-grained alignment.

[0069] 3) Considering that sensor platforms often move dynamically in practical applications (for example, on airborne or vehicle-mounted platforms), radar, optical, and infrared sensors will experience position and attitude changes as the platform moves. Dynamic parameter estimation can also be used to monitor and adjust sensor motion in real time. In this embodiment, instantaneous sensor attitude changes are estimated by fusing IMU (Inertial Measurement Unit) data. IMU data provides platform acceleration and angular velocity information, thereby estimating the instantaneous attitude of the sensor platform (such as pitch, roll, and yaw). Using these estimated attitude changes, the projection matrix of the radar virtual projection is updated in real time. Even if the platform moves or rotates, the projection matrix still accurately reflects the current attitude and position of the sensors. Kalman filtering can also be considered to fuse IMU data with other sensor information to estimate attitude changes. Combining Kalman filtering with IMU data can adjust the projection matrix during the radar virtual imaging process in real time, ensuring that the virtual projection image is aligned with the optical / infrared image regardless of platform motion.

[0070] It's important to note that in this step, the virtual projection view provides a foundation for the transformation, coarse alignment is used to initially align the radar and optical images, and dynamic parameter estimation ensures consistency even when the sensors are moving. These three elements work together to effectively register the radar and optical images, facilitating more precise subsequent processing.

[0071] Step 3: Fine-tune the coarsely registered multimodal data (including radar, optical, and infrared images) to achieve higher-precision spatial alignment. In this embodiment, feature fusion is used to extract recognizable feature points or regions from the data of different modalities (radar, optical, and infrared images). These features can be used for cross-modal matching. Optimization is then used to fine-tune the registration results, resulting in a multimodal data cube for subsequent multimodal anti-interference feature fusion.

[0072] Specifically, in this embodiment, a refined registration mathematical model based on coupled partial differential equations is used to extract features from different modal data and fine-tune registration, thereby obtaining a multimodal data cube. The steps include:

[0073] 1) In the refined registration mathematical model based on coupled partial differential equations in this embodiment, the input data is the coarsely registered phased array radar projection feature field (i.e., radar image). , optical images and infrared images Then, the gradient of the phased array radar projection characteristic field is calculated respectively. , which represents the direction and degree of spatial change of the phased array radar projection feature field. For radar images, the gradient can help extract the area where the radar signal changes greatly; the gradient of the optical sensor data , represents the spatial variation of the optical image and is used to extract the edge or detail information of the optical image; the gradient of the infrared sensor data , represents the spatial variation of the infrared image.

[0074] In other words, we first need to extract gradient information from the input data. Gradients describe areas of strong change in the image, usually representing edges or details. Therefore, gradient information can reveal image details, especially edges, textures, and other information, which can serve as the driving force in the alignment process.

[0075] These gradient information reflects the local features of each image modality and will be used to calculate the similarity between different modalities in subsequent steps.

[0076] 2) Construct radar-optical feature similarity functions and radar-infrared feature similarity functions, and combine the calculated gradients of different modes to iteratively update the displacement fields of the optical image and the infrared image to finely align the images of different modes.

[0077] In this embodiment, the displacement field of the optical image is defined as: The displacement field of the infrared image is defined as: .

[0078] The displacement field reflects the deformation of the image during the registration process and can be optimized based on the similarity of the coupled features. The registration goal is to make the data of different modalities (radar, optical, infrared) as consistent as possible in space by adjusting these displacement fields.

[0079] In this embodiment, the displacement fields of the optical image and the infrared image are updated by a feature-coupling-driven displacement field equation set, which includes the displacement field update equations for the optical image and the displacement field update equations for the infrared image.

[0080] The displacement field update equation of the optical image can be constructed by adjusting the displacement field of the optical image based on the similarity gradient between the radar image and the optical image; coordinating the transformation of the two images by calculating the difference between the displacement fields of the optical image and the infrared image; and finally ensuring the smoothness of the displacement field of the optical image through the smoothing term.

[0081] The displacement field update equation of the infrared image can be constructed by calculating the similarity gradient between the radar image and the infrared image to adjust the displacement field of the infrared image, and coordinating the transformation of the two images by calculating the difference between the displacement fields of the infrared image and the optical image; and by using the smoothing term to ensure the smooth change of the displacement field of the infrared image.

[0082] Specifically, the displacement field equations driven by the characteristic coupling can be characterized by the following set of partial differential equations:

[0083] ,

[0084] In the above formula, Update the equation for the displacement field of the optical image.

[0085] in, is the time rate of change of the displacement field of the optical image, which represents the update amount of the displacement field of the optical image at each time step. By continuously updating the displacement field, the features of the optical image can be gradually aligned with the infrared image; is the similarity gradient control parameter of the optical image, which is used to control the influence of the similarity gradient on the displacement field update. If this control parameter is large, the gradient term will play a greater role in the displacement field update, and the displacement field of the optical image will adjust more quickly to the corresponding gradient direction. is the radar-optical signature similarity function; is the control parameter for the consistency of the displacement fields between the infrared image and the optical image. It is used to adjust the interaction between the displacement fields of the infrared image and the optical image. If this control parameter is large, the displacement fields of the two images will tend to be consistent. It is a parameter that controls the smoothness of the displacement field of the optical image. By applying regularization to the displacement field, the displacement field is made smoother in space and unnatural fluctuations are avoided.

[0086] It should be noted that This term is an update term that represents the gradient of the similarity between the optical image and the radar image. By taking the gradient of the radar-optical feature similarity function, the direction of change in the gradient similarity between the optical image and the radar image can be determined. Through this term, the update of the displacement field will guide the optical image to adjust toward the characteristic direction of the radar image. This term represents the difference between the displacement field of the optical image and the displacement field of the infrared image. By comparing the displacement fields of the infrared image and the optical image, the displacement field of the optical image can be adjusted to better align with the infrared image. This term helps the displacement fields between the two images become consistent. This term is a smoothing term that represents the second-order gradient of the displacement field of the optical image. This term is used to prevent excessive fluctuations in the displacement field, ensure a smooth transition of the displacement field in space, and play a regularization role, making the displacement field change smoother and avoiding unreasonable mutations.

[0087] In the above formula, Update the equations for the displacement field of the infrared image.

[0088] in, is the time rate of change of the displacement field of the infrared image, which indicates how the displacement field of the infrared image is updated over time. is the similarity gradient control parameter of the infrared image, which is used to control the influence of the similarity gradient on the displacement field update. If this control parameter is large, the gradient term will play a greater role in the displacement field update, and the displacement field of the infrared image will adjust more quickly to the corresponding gradient direction; It is the control parameter of the smoothness of the displacement field of the infrared image; is the radar-infrared signature similarity function.

[0089] It should be noted that This term, similar to the update term for optical images, represents the gradient of similarity between the infrared and radar images. Taking the gradient of the radar-infrared feature similarity function helps understand the changes in similarity between the infrared and radar images. Through this term, the update of the displacement field guides the adjustment of the infrared image features, gradually aligning them with the radar image features. This term represents the difference between the displacement field of the infrared image and the displacement field of the optical image. By comparing the displacement fields of the two, the displacement field of the infrared image can be adjusted to make it closer to the displacement field of the optical image, thereby strengthening the alignment of the two images. This term is similar to the smoothing term in the displacement field update equation of the optical image. It is used to smooth the displacement field of the infrared image. Through the second-order gradient term, it ensures the smooth transition of the displacement field and avoids unreasonable mutations in the displacement field.

[0090] It should also be noted that, as the reference modality, the radar image's displacement field generally does not need to be updated because it does not require alignment with other modalities. Optical and infrared images, on the other hand, are aligned to the radar image by updating their own displacement fields. Therefore, the feature-coupled-driven displacement field equations do not include displacement field update equations for the radar image. Instead, alignment with the radar image is achieved by updating the displacement fields of the optical and infrared images.

[0091] Specifically, in this embodiment, the radar-optical signature similarity function It can be expressed as:

[0092] ,

[0093] in, It is an exponential decay function used to convert gradient differences into similarity metrics. The smaller the gradient difference between the two images, the closer the function value is to 1 (i.e., the higher the similarity); the larger the difference, the closer the function value is to 0 (i.e., the lower the similarity). It is a tuning parameter of the radar-optical feature similarity function, which is used to control the "sensitivity" of the radar-optical feature similarity function. A larger tuning parameter will make the model more sensitive to gradient differences, thereby strongly penalizing smaller gradient differences; a smaller tuning parameter will reduce the intensity of this penalty. is a transformation function that represents the gradient of the optical image through the displacement field The result of the transformation is to transform the gradient of the optical image into a spatial displacement field so that it is aligned with the gradient of the radar image. It reflects the geometric transformation of the optical image relative to the radar image.

[0094] is the Euclidean distance between the radar and optical image gradients (i.e., the difference between the two). By calculating the difference in gradients, the structural similarity of the two images is measured.

[0095] It's important to note that the radar-optical signature similarity function measures the similarity between the radar and optical images based on the difference in gradients. By calculating the gradient difference and incorporating the effects of the displacement field, the function encourages spatial alignment of the radar and optical images, bringing them closer together at the gradient level. The adjustment parameter controls the strictness of this alignment.

[0096] Specifically, in this embodiment, the radar-infrared feature similarity function It can be expressed as:

[0097] ,

[0098] in, It is a hyperbolic tangent function, which is used to map the mutual information value to a fixed range (usually -1 to 1), thereby ensuring the smoothness of the mutual information effect and avoiding extreme cases where the value is too large or too small. is the adjustment parameter of the radar-infrared feature similarity function, which is used to control the contribution of mutual information to the final similarity function. A larger adjustment parameter will increase the impact of mutual information on the similarity function, and vice versa. The mutual information (Mutual Information) between radar and infrared images reflects the statistical correlation between the two. Mutual information is an important indicator for measuring the correlation between two images. Especially in image registration tasks, mutual information can be used to measure the degree of information sharing in the overlapping parts of two images. A higher mutual information value indicates that the two images are more consistent in space and therefore have stronger similarity.

[0099] The modulus of the cross product of the radar image and infrared image gradients (i.e., the length of the cross product) can measure the relative changes in the two gradient directions and reflect the structural differences between the radar and infrared images. If the gradient directions of the two images change more consistently, the modulus of the cross product will be smaller, indicating that their features are more consistent. Conversely, if the cross product modulus is larger, it means that the structures of the two images are significantly different.

[0100] It should be noted that in this embodiment, different functions are used to construct the radar-optical feature similarity function and the radar-infrared feature similarity function. The radar-optical feature similarity function is constructed using an exponential decay function because it measures the gradient difference between the radar image and the optical image and more strongly penalizes areas with large differences, allowing for more accurate image alignment. Exponential decay is well-suited for strongly penalizing large errors. When the gradient difference is large, the exponential decay function can quickly reduce the similarity, thereby guiding the model to reduce these large differences through iterative optimization, ensuring the accuracy of image alignment. In other words, for the radar-optical feature similarity function, exponential decay is necessary to process and penalize large gradient differences to ensure accuracy during image alignment.

[0101] The hyperbolic tangent function is used to construct the radar-infrared feature similarity function because, for radar and infrared images, it is necessary not only to measure the similarity between the two images but also to consider their structural alignment. The hyperbolic tangent function has a value range of [-1, 1], and its output values ​​are smooth and symmetrically distributed. This makes it particularly suitable for processing "correlation" metrics such as mutual information (MI). Using the hyperbolic tangent function ensures that the influence of mutual information is smooth, and large mutual information values ​​are neither overly amplified nor overly reduced. In other words, the hyperbolic tangent function effectively controls and adjusts the influence of similarity, giving higher weight to portions with larger mutual information values, while smaller portions are smoothed to avoid excessive influence.

[0102] 3) Use the updated displacement fields of the optical image and the infrared image to align the modal images, and align the optical image and infrared image to the radar image. Based on the aligned data of the three modalities (radar, optical, and infrared), a multimodal data cube is generated:

[0103] ,

[0104] in, is a data cube, is a set of real numbers, indicating that the data cube is a four-dimensional tensor, and the value of each position is composed of real numbers; is the height of the image, is the width of the image, is the feature dimension, 3 represents the number of multi-modal channels. In this embodiment, there are three modes: phased array radar R, optical O, and infrared I, so it is 3.

[0105] Step 4: Through the multimodal feature fusion mathematical model, the generated multimodal data cube is fused to calculate a fused multimodal feature.

[0106] In this step, the multimodal data cube is split and complex space expressions of different modalities are constructed. The complex signals of different modalities are fused into a unified complex space through holomorphic mapping, and then the fused features are converted into a format suitable for classification through projection into complex hyperbolic space.

[0107] Specifically, in this embodiment, a multimodal feature fusion mathematical model of a complex function can be used to achieve the fusion of multimodal data and the generation of anti-interference features. The following steps are included:

[0108] 1) First, the data cube is split into different modal data. Complex plane expressions of different modal data are constructed according to the physical characteristics of different modes, so as to map each modal data to the complex plane.

[0109] Among them, the radar mode can be represented by amplitude-phase, where the amplitude is the echo intensity and the phase is dynamically adjusted by beamforming to reflect the microscopic motion characteristics of the target, thereby obtaining the representation of the radar mode in complex space.

[0110] The optical mode, whose real part is the texture mean and imaginary part is the edge gradient, reflects the stability of spatial structure information through complex numbers, thereby obtaining the representation of the optical mode in complex space.

[0111] For infrared mode, the thermal gradient amplitude is used as the mode and the dynamic heat diffusion phase as the argument. The time domain characteristics of temperature change are described in combination with the heat conduction model, thereby obtaining the representation of infrared mode in complex space.

[0112] Specifically, the complex form of the radar mode can be expressed as follows:

[0113] ,

[0114] in, is an imaginary unit; is a natural constant; is the complex value corresponding to the radar image at a certain spatial position (x, y) and time point t. Each pixel point (or time point) has an amplitude and a phase, and these two pieces of information can be represented by complex numbers. It is the amplitude of the radar image at a certain spatial position (x, y) and time point t. This amplitude reflects the strength of the radar signal, usually indicating the strength of the radar return signal or the amplitude of the echo. is the phase of the radar image at a certain spatial position (x, y) and time point t. The phase reflects the propagation of the radar wave, such as the arrival time of the wave or the phase difference with other waves.

[0115] The complex form of the optical mode can be expressed as follows:

[0116] ,

[0117] in, is the complex value corresponding to the optical image at a certain spatial position (x, y) and time point t; is the texture information of the optical image at a certain spatial position (x, y) and time point t. Texture information refers to the color or intensity distribution of pixels in the image; It is the edge gradient information of the optical image at a certain spatial position (x, y) and time point t. The edge gradient can be calculated by the edge detection algorithm to reflect the area with strong changes in the image and is usually used to represent the boundary of the object in the image.

[0118] The plural form of the infrared mode can be expressed as follows:

[0119] ,in, is the complex value corresponding to the infrared image at a certain spatial position (x, y) and time point t; is the thermal gradient amplitude of the infrared image at a certain spatial position (x, y) and time point t, reflecting the changes between different temperature regions; It is the dynamic temperature phase of the infrared image at a certain spatial position (x, y) and time point t, reflecting the dynamic characteristics of temperature changes, such as the location and change trend of the heat source.

[0120] 2) After constructing the complex mode, the complex signals of the three complex modes are fused into a unified complex space using a holomorphic mapping to obtain the fused features.

[0121] Specifically, the result of the holomorphic mapping can be obtained by the following fusion equation:

[0122] ,

[0123] in, is a holomorphic mapping fusion function, is the representation of radar mode in complex space. The complex representation of radar signal is a combination of amplitude and phase; is the representation of the optical modality in complex space, which includes the complex representation of texture and edge gradient; The complex representation of infrared modalities in complex space usually consists of thermal gradient amplitude and dynamic temperature phase; these three complex representations provide a unified mathematical space for multimodal fusion, converting the features of each modality into points on the complex plane, so that further fusion can be performed in this space.

[0124] for The partial derivative of the complex conjugate of , for The complex conjugate of is the partial derivative condition, which requires that the optical mode satisfies the holomorphic condition, that is, in the complex function, the complex cannot have a change in the conjugate complex part, holomorphism means that the complex It must change along the real direction in the complex plane to maintain analytical properties; it is used to ensure that the optical mode changes smoothly along the real axis in the complex space, thus avoiding unnecessary perturbations or interference.

[0125] is the radar phase stability weight function, which is used to suppress the jump interference in the radar phase and ensure the phase stability of the radar. In this embodiment, it can be expressed by the following formula:

[0126] ,

[0127] in, It is an activation function used to process the second-order gradient of the radar signal, thereby reducing unnecessary changes in the radar image, ignoring negative gradient values, and retaining only the positive mutation part; A scale parameter that controls the smoothing sensitivity of radar signals. It is used to standardize the degree of mutation and prevent excessive Laplace values ​​from causing the exponential function to approach 0 (i.e., almost completely suppressing the relevant region). This parameter can be selected based on the actual phase changes of the radar image in the data, empirically set, or through methods such as cross-validation. is the radar modal complex amplitude The Laplace operator (second-order gradient) is used to measure the phase change rate or intensity change of a region in the image.

[0128] is the norm of the radar mode, and Represents a unit circle or a constrained space; the norm represents the size of the radar mode in complex space. Through normalization, the amplitude of the radar mode can be scaled to the range of the unit circle or unit sphere, thereby avoiding the impact of scale differences between different modes on the fusion result.

[0129] is the optical edge continuity weight function, which is used to emphasize the edge information of the image, reduce the weight of the blurred area, and enhance the contribution of the structured area. Optical images usually have many smooth areas and edge areas, and the edge areas usually contain more structural information. In this embodiment, it can be expressed by the following formula:

[0130] , which means that the gradient of the optical mode in complex space is , the value of the continuity weight will be larger in the edge part of the image and smaller in the smooth part, thus ensuring that the structure part in the image occupies a larger weight in the fusion process.

[0131] The infrared out-of-phase modulation function is used to adjust the infrared signal's phase information to ensure synchronization with the phases of other modalities. Since there may be a certain phase difference between the infrared signal and other modalities, the infrared out-of-phase modulation function can achieve phase synchronization between modalities by adjusting the infrared modal phase. It uses the imaginary part and gradient information of the infrared signal to generate a transformation factor, thereby modulating the infrared signal's phase to ensure consistency with the radar and optical modalities during fusion. In this embodiment, it can be expressed as follows: ,

[0132] in, is a complex exponential term used to modulate the phase of the infrared signal to ensure phase synchronization of the infrared signal, thus avoiding instability caused by phase difference during the fusion process; Indicates the change in the phase of the infrared signal; for about The derivative of , usually refers to the local gradient or rate of change of the infrared signal in image processing, and reveals the changing trend of the infrared signal in space.

[0133] is a sign function used to obtain the sign of the imaginary part, that is:

[0134] ,

[0135] plural The imaginary part represents the imaginary component of the infrared signal. The imaginary part usually carries the phase change information of the signal. Through this imaginary part, the phase change trend of the infrared image can be obtained.

[0136] It should be noted that the fusion equation disclosed in this embodiment is a mathematical equation for complex manifold fusion, which aims to fuse data from three different modalities: radar, optical, and infrared, through the mapping of complex functions (complex manifolds). By using complex analysis technology, multimodal information is mapped into a complex space, and the characteristics of this space are used to perform multimodal feature fusion. By weighting and modulating each term, the equation ultimately achieves a stable fusion of the three modalities in the complex feature space, ensuring the balance and stability of each modal contribution. This fusion method can improve the expressiveness of multimodal signals, so that the final complex features can better reflect the original modal information.

[0137] 3) Finally, the fused features are projected into a compact pseudo-Euclidean space in the complex hyperbolic space through the Klein model, thereby effectively mapping the multimodal fusion features into the final fusion features suitable for image classifier input.

[0138] In this embodiment, the Klein model can be expressed by the following formula:

[0139] ,

[0140] in, is the final fusion feature, is the logarithmic sign, is the holomorphic mapping fusion equation, is a complex hyperbolic space, The norm of the fusion feature in the complex hyperbolic space calculated by the holomorphic mapping fusion equation represents the size of the feature. The square of the norm is results. The inner product of the fused feature calculated by the holomorphic mapping fusion equation in the complex hyperbolic space, that is, the inner product of the fused feature and itself, is used to measure the "energy" or "size" of the feature in the complex hyperbolic space.

[0141] It should be noted that in the above formula, is a logarithmic term in which the numerator Indicates that the inner product of the feature is added by 1, which is intended to add a smooth offset to avoid the denominator being zero. The square of the feature norm minus 1 is used to adjust the result based on the feature size to ensure feature normalization. This logarithmic term allows for smooth adjustment of the final fused feature, avoiding numerical divergence or over-amplification, and ensuring the stability of the fusion process.

[0142] This term is a feature normalization term. Its function is to normalize features to unit values, i.e., to make the feature size (norm) equal to 1. This is used to eliminate the influence of different feature scales and prevent certain features with large values ​​from having a significant impact on the final result. Normalization ensures that the direction of the final fused features remains unchanged, allowing the fused features to be more balanced when combined with other modal features, improving the overall fusion effect.

[0143] Step 5: Perform complex decomposition on the final fusion features calculated in step 4, and convert the final fusion features into real features suitable for CNN target classification processing to ensure that these features are suitable for convolution calculations.

[0144] In this embodiment, the final fusion feature can be split into two real channels of amplitude and phase, and the two real channels obtained by splitting are input into a CNN target classifier for target recognition.

[0145] Specifically, it may include the following sub-steps:

[0146] 1) The two real channels of amplitude and phase of complex decomposition can be expressed as:

[0147] ,

[0148] in, The final fusion feature is a complex signal that contains real and imaginary parts, where the real part represents one direction of the signal and the imaginary part represents the information in the other direction. This complex feature often contains the amplitude and phase information of the signal.

[0149] is the amplitude channel, is the amplitude; The magnitude of the complex feature represents the size of the complex signal (i.e., the distance from the origin to the complex number on the complex plane). It is often used to indicate the importance or strength of the signal. The magnitude can be calculated using the following formula:

[0150] ,

[0151] in, is the real part of the complex characteristic, is the imaginary part of the complex characteristic.

[0152] is the phase channel, is the phase, The phase of the complex feature, that is, the angle between the complex number and the real axis on the complex plane, reflects the relative angle or phase information of the signal. In multimodal signal fusion, phase information can describe the spatiotemporal relationship or phase difference of the signal and is often used to capture information such as signal directionality and phase synchronization. The phase can be calculated using the following formula:

[0153] ,in, is the inverse tangent function.

[0154] It should be noted that by converting the complex signal Splitting the signal into two real-valued channels, amplitude and phase, allows for the extraction of the signal's energy and relative spatiotemporal structure, respectively. The amplitude channel reflects the signal's strength or energy, while the phase channel describes the signal's relative phase or spatiotemporal relationship, often used to indicate directional or periodic changes. This decomposition method can help better understand and process the diverse information contained in the signal. This is particularly true in multimodal signal fusion applications, where it effectively combines the amplitude and phase information of different modalities, improving the accuracy of subsequent analysis.

[0155] 2) After decomposition, channel concatenation is performed. The real, imaginary, amplitude, and phase features obtained from the complex decomposition are concatenated into different input channels to form a multi-channel feature map. Each channel can represent different information (for example, the real part can be one channel, the imaginary part can be another channel, and the amplitude and phase can also be different channels).

[0156] The features of each channel can also be normalized at the same time to ensure that they have similar scales in the CNN object classifier. In this embodiment, Z-score normalization or minimum-maximum normalization can be used to achieve normalization of each channel.

[0157] These features are formed into a multi-channel input dataset as the input layer of the CNN target classifier.

[0158] 3) The processed features are input into the CNN target classifier for target classification. CNN will extract high-level features through multiple convolutional layers and finally perform classification through the fully connected layer.

[0159] In this embodiment, the CNN object classifier may include:

[0160] Convolutional layer: The convolutional layer is used to extract local patterns of input features. The convolutional layer slides in space, extracts features by learning convolution kernels, and generates convolutional feature maps.

[0161] Pooling layer: The pooling layer is used to downsample the convolutional feature map, reducing the amount of computation while retaining important spatial information. In this embodiment, the pooling method that can be used is maximum pooling or average pooling.

[0162] Fully Connected Layer: After convolution and pooling, the features are fed into the fully connected layer for high-level abstraction and classification. In the fully connected layer, the neural network calculates the probability distribution of the category based on the input features and ultimately outputs the classification result to achieve target recognition.

[0163] The specific implementation methods described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific implementation method of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A multi-modal phased array radar target recognition method, characterized in that: The phased array radar target recognition method comprises: S100, pre-processing the data of different modes to obtain multi-modal data coupled in time, The data of different modalities include phased array radar data, optical sensor data and infrared sensor data; S200, performing preliminary registration on the temporally coupled multimodal data, The preliminary registration realizes cross-modal geometric coarse registration guided by phased array radar data through affine transformation to obtain coarse multi-modal data; S300, performs fine registration on the coarse multimodal data to achieve higher precision spatial alignment, By combining features, recognizable feature points or regions are extracted from data of different modalities, and the registration results are fine-tuned through optimization to obtain a multimodal data cube. S400, fusing the generated multimodal data cube using a multimodal feature fusion mathematical model to calculate a fused final fusion feature; S500, performing complex decomposition on the final fusion feature, converting the final fusion feature into a real number feature suitable for processing by the CNN target classifier, performing classification by the CNN target classifier, obtaining a classification result, and realizing target recognition.

2. The multi-modal phased array radar target recognition method according to claim 1, characterized in that: The preprocessing is a spatiotemporal baseline alignment operation to ensure that data of different modalities are aligned under the same time base, including: time synchronization and spatial alignment, wherein, Time synchronization achieves uniform timestamps for data of different modalities by deploying a central clock or clock source with a precise clock protocol at the data source of different modalities. Spatial alignment is achieved by calculating a rigid body transformation matrix between the radar coordinate system and the optical camera coordinate system, and converting the phased array radar data coordinates into the optical sensor data coordinate system through the rigid body transformation matrix.

3. The multi-modal phased array radar target recognition method according to claim 1, characterized in that: The step S200 includes the following sub-steps: S210, projecting the radar point cloud of the phased array radar data along the depth axis to simulate the viewing angle of the optical sensor to form a virtual projection view of the phased array radar data, thereby obtaining a two-dimensional radar virtual image; S220 , performing affine transformation on the optical image based on the rigid body transformation matrix and the radar virtual image, and roughly aligning the target in the optical / infrared image to the coordinate system of the radar point cloud.

4. The multi-modal phased array radar target recognition method according to claim 1, characterized in that: The refined registration in step S300 includes the following sub-steps: S310, extracting gradient information of different modes of input to reflect local features of different modes; S320, combining gradient information of different modalities, iteratively updating the displacement field of the optical image and the displacement field of the infrared image, and finely aligning the images of different modalities; S330: Use the updated displacement fields of the optical image and the infrared image to align the images of each modality, and align the optical image and the infrared image to the radar image. And based on the aligned data, a multimodal alignment data cube is generated: , in, is a data cube, is a set of real numbers, indicating that the data cube is a four-dimensional tensor, and the value of each position is composed of real numbers; is the height of the image, is the width of the image, is the feature dimension, and 3 represents the number of multi-modal channels, including phased array radar R, optical O, and infrared I modes.

5. The multi-modal phased array radar target recognition method according to claim 4, characterized in that: The displacement field update can be achieved by constructing a feature-coupling-driven displacement field equation group, which includes: an optical image displacement field update equation and an infrared image displacement field update equation, wherein: The displacement field update equation of the optical image is calculated by calculating the similarity gradient between the radar image and the optical image to adjust the displacement field of the optical image. The difference between the displacement fields of the optical image and the infrared image is then combined to coordinate the transformations of the two images. Finally, the smoothing term is used to ensure that the displacement field of the optical image changes smoothly, thus constructing it; The displacement field update equation of the infrared image is calculated by calculating the similarity gradient between the radar image and the infrared image to adjust the displacement field of the infrared image. The difference between the displacement fields of the infrared image and the optical image is then calculated to coordinate the transformations of the two images. The smoothing term is used to ensure that the displacement field of the infrared image changes smoothly, and thus it is constructed.

6. The multi-modal phased array radar target recognition method according to claim 1, characterized in that: The multimodal feature fusion mathematical model in step S400 splits the multimodal data cube, fuses the complex signals of different modes into a unified complex space using a holomorphic mapping, and obtains the final fusion feature through projection into the complex hyperbolic space. The following sub-steps are included: S410, splitting the data cube into different modal data, and mapping each modal data to a complex plane to encode its physical characteristics; S420. Construct a holomorphic mapping fusion equation to implement the holomorphic mapping, thereby fusing the complex forms of different modes into a unified complex space to obtain fused features; S430, projecting the fused features into a compact pseudo-Euclidean space in the complex hyperbolic space through the Klein model to obtain the final fused features. The Klein model is represented by the following formula: , in, is the final fusion feature, is the logarithmic sign, is the holomorphic mapping fusion equation, is a complex hyperbolic space, is the norm of the fusion feature in the complex hyperbolic space calculated by the holomorphic mapping fusion equation, which represents the size of the feature. It is the inner product of the fusion features calculated by the holomorphic mapping fusion equation in the complex hyperbolic space.

7. The multi-modal phased array radar target recognition method according to claim 6, characterized in that: The data cube includes: radar mode, optical mode and infrared mode, wherein, Radar modes are represented using amplitude-phase, where the amplitude represents the echo intensity and the phase is dynamically adjusted by beamforming to reflect the target's microscopic motion characteristics. This constructs a representation of the radar mode in complex space. The optical mode, whose real part is the texture mean and imaginary part is the edge gradient, reflects the stability of spatial structure information through complex numbers, thereby constructing the representation of the optical mode in complex space; For infrared mode, the thermal gradient amplitude is used as the mode and the dynamic heat diffusion phase as the argument. The time domain characteristics of temperature change are described in combination with the heat conduction model, thereby constructing the representation of infrared mode in complex space.

8. The multi-modal phased array radar target recognition method according to claim 1, characterized in that: The step S500 includes the following sub-steps: S510, perform complex number decomposition and split the final fusion feature into two real channels of amplitude and phase. The amplitude channel reflects the strength or energy of the signal. The phase channel describes the relative phase or spatiotemporal relationship of the signal and is used to indicate the directionality or periodic changes of the signal. S520, splicing the real part, imaginary part, amplitude, and phase features obtained by complex number decomposition into different input channels to form a multi-channel feature map; S530, input the multi-channel feature map into the CNN target classifier for target classification. CNN will extract high-level features through multiple convolutional layers and finally perform classification through the fully connected layer.

9. The multi-modal phased array radar target recognition method according to claim 8, characterized in that: The two real channels of amplitude and phase are expressed as: , in, is the final fusion feature, which is a complex signal; is the amplitude channel, is the amplitude; is the amplitude of the complex feature, indicating the size of the complex signal; is the phase channel, is the phase, is the phase of the complex characteristic.

10. A multi-modal phased array radar target recognition system, characterized in that: The target recognition system includes: processor; The memory stores a computer program, and when the computer program is executed by the processor, the multi-modal phased array radar target recognition method according to any one of claims 1 to 9 is implemented.

Citation Information

Patent Citations

  • Phased array radar target identification system based on positioning filtering algorithm and method

    CN106291539A

  • Spatial positioning method based on multi-modal visual fusion

    CN119251303A