Target identification method, system and device based on multi-modal data and storage medium
Through multimodal data fusion, data is collected using radar and infrared sensors, converted into a distance-Doppler map and extracted features, combined with environmental parameters to generate confidence, and finally fused features to improve the accuracy and reliability of high-speed aircraft target recognition.
Patent Information
- Application Number
- CN202510598666.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2045-05-09
AI Technical Summary
When high-speed aircraft use radar to identify targets in complex environments, they are affected by factors such as electromagnetic interference and clutter, resulting in low recognition accuracy.
The multimodal data recognition method is used to collect data using radar and infrared sensors, and convert it into a distance-Doppler diagram through the correlation between radar echo data and infrared images. The characteristics are extracted using convolutional neural networks, and feature confidence is generated by combining environmental parameters and infrared image quality evaluation coefficients. Finally, the radar and infrared features are fused to determine the target type.
It effectively overcomes the limitations of single radar identification due to complex environment interference, greatly improving the accuracy and reliability of target identification.
Smart Images

Figure CN120446940A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of target recognition, and in particular relates to a target recognition method, system, device and storage medium based on multimodal data. Background Art
[0002] Radar, a common means of target identification, plays an important role in many fields. However, using radar for target identification on high-speed aircraft presents numerous challenges. The complex high-speed flight environment and the presence of various interference factors, such as electromagnetic interference and clutter, can result in low radar recognition accuracy. Summary of the Invention
[0003] In view of the above-mentioned deficiencies in the prior art, the present invention provides a target recognition method, system, device and storage medium based on multimodal data to solve the above-mentioned technical problems.
[0004] In a first aspect, the present invention provides a method for target recognition based on multimodal data, comprising: The radar and infrared sensors at the same position are used to collect radar echo data and infrared images, and the radar echo data and infrared images of the same target are correlated to obtain the target radar echo data and target infrared image; Converting the target radar echo data into a range-Doppler map; extracting radar features from the range-Doppler map using a convolutional neural network, and extracting infrared features from the infrared image of the target; Acquiring environmental parameters and quality evaluation coefficients of the target infrared image, and generating feature confidence levels based on the environmental parameters and the quality evaluation coefficients; The radar features and infrared features are fused according to the feature confidence, and the target type is determined based on the fused features.
[0005] In an optional embodiment, radar echo data and infrared images are collected using radar and infrared sensors at the same location, and the radar echo data and infrared images of the same target are correlated to obtain target radar echo data and target infrared images, including: The radar echo data of multiple measuring points are acquired by the radar, and the infrared images of multiple measuring points are collected by the infrared sensor, wherein the radar and the infrared sensor are installed at the same position; Acquire radar position parameters and infrared position parameters of each measurement point, wherein the radar position parameters include azimuth and elevation angles relative to the radar, and the infrared position parameters include azimuth and elevation angles relative to the infrared sensor; The spherical distance between the radar position parameter of each measurement point in the radar echo data and the infrared position parameter of the measurement point collected by the infrared sensor is calculated, and the measurement point corresponding to the radar position parameter and infrared position parameter with the smallest spherical distance is taken as the matching measurement point. The radar echo data and infrared image corresponding to the matching measurement point are determined as the target radar echo data and target infrared image of the same target.
[0006] In an optional embodiment, converting the target radar echo data into a range-Doppler map includes: Perform pulse compression on the target radar echo data through convolution operation to obtain a compressed echo signal; Perform Fourier transform on the compressed echo signal of each range unit to obtain the frequency spectrum; The amplitude of the frequency spectrum is taken, and the zero-frequency component is moved to the center of the spectrum to obtain the range-Doppler map of the target.
[0007] In an optional embodiment, the convolutional neural network includes 4 convolutional layers, 8 residual modules and a global pooling layer.
[0008] In an optional embodiment, obtaining environmental parameters and a quality evaluation coefficient of a target infrared image, generating a confidence level of a radar feature based on the environmental parameters, and generating a confidence level of the infrared feature based on the quality evaluation coefficient includes: Acquiring environmental parameters and converting the environmental parameters into environmental parameter quantized values; Performing quality detection on the target infrared image to obtain a quality evaluation coefficient; Inputting the quantized value of the environmental parameter and the quality evaluation coefficient into a multi-layer perceptron to obtain the confidence of the radar feature and the confidence of the infrared feature; The multi-layer perceptron includes a first branch, a second branch and a confidence calculation module, and the first branch and the second branch are both composed of a fully connected layer and a ReLU layer; The first branch is used to perform a linear transformation on the radar feature to obtain a radar confidence vector; The second branch is used to perform linear transformation on the infrared feature to obtain an infrared confidence vector; The confidence calculation module is used to splice the radar confidence vector and the infrared confidence vector into a combination vector, and obtain the feature confidence by performing dimensionality reduction processing on the combination vector.
[0009] In an optional embodiment, radar features and infrared features are fused according to feature confidence, and the target type is determined based on the fused features, including: generating a radar feature weight and an infrared feature weight according to the feature confidence; The dot product of the radar feature and the radar feature weight is recorded as the radar feature vector, and the dot product of the infrared feature and the infrared feature weight is recorded as the infrared feature vector; Obtaining low-rank modal factors of radar features and low-rank modal factors of infrared features; Calculating a fusion feature vector according to the radar feature vector, the low-rank modal factor of the radar feature, the infrared feature vector, and the low-rank modal factor of the infrared feature; A classifier is used to determine the target type according to the fused feature vector.
[0010] In an optional embodiment, generating a radar feature weight and an infrared feature weight according to the feature confidence level includes: Perform norm normalization on the feature confidence to obtain the confidence coefficient; The confidence coefficient is used as the infrared feature weight; The difference between 1 and the confidence coefficient is used as the radar feature weight.
[0011] In a second aspect, the present invention provides a target recognition system based on multimodal data, comprising: The data matching module is used to collect radar echo data and infrared images using radar and infrared sensors at the same position, and to associate the radar echo data and infrared images of the same target to obtain the target radar echo data and target infrared image; A preprocessing module, configured to convert the target radar echo data into a range-Doppler map; a feature extraction module, configured to extract radar features from the range-Doppler map using a convolutional neural network, and to extract infrared features from the target infrared image; A confidence calculation module, configured to obtain environmental parameters and a quality evaluation coefficient of a target infrared image, and generate a feature confidence level based on the environmental parameters and the quality evaluation coefficient; The feature processing module is used to fuse radar features and infrared features according to feature confidence and determine the target type based on the fused features.
[0012] According to a third aspect, a device is provided, comprising: A memory for storing a target recognition program based on multimodal data; A processor is configured to implement the steps of the target recognition method based on multimodal data as provided in the first aspect when executing the target recognition program based on multimodal data.
[0013] In a fourth aspect, a computer-readable storage medium is provided, on which a target recognition program based on multimodal data is stored. When the target recognition program based on multimodal data is executed by a processor, the steps of the target recognition method based on multimodal data provided in the first aspect are implemented.
[0014] The beneficial effects of the present invention are that the target recognition method, system, device and storage medium based on multimodal data provided by the present invention realize the utilization of multimodal data by using radar and infrared sensors at the same position to collect radar echo data and infrared images, and determine the relevant data of the same target. After converting the radar echo data into a range-Doppler map, a convolutional neural network is used to extract radar features and infrared features respectively. At the same time, feature confidence is generated by combining environmental parameters and infrared image quality evaluation coefficients, thereby fusing the two features. This multimodal data fusion method effectively overcomes the limitations of single radar recognition being interfered with by complex environments, and greatly improves the accuracy and reliability of target recognition.
[0015] In addition, the present invention has a reliable design principle, a simple structure and a very broad application prospect. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0017] Figure 1 is a schematic flow chart of a method according to an embodiment of the present invention.
[0018] Figure 2 It is a schematic principle diagram of a method according to an embodiment of the present invention.
[0019] Figure 3 This is a ResNet18 network structure diagram of a method according to an embodiment of the present invention.
[0020] Figure 4 4 is a structural diagram of a residual module of a method according to an embodiment of the present invention.
[0021] Figure 5 The figure is a flow chart of infrared image self-checking of a method according to an embodiment of the present invention.
[0022] Figure 6 The figure is a flow chart of confidence calculation of a method according to an embodiment of the present invention.
[0023] Figure 7 It is a schematic diagram of feature fusion of a method according to an embodiment of the present invention.
[0024] Figure 8 FIG. 4 is a schematic block diagram of a system according to an embodiment of the present invention.
[0025] Figure 9 A schematic structural diagram of a device provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0026] In order to enable those skilled in the art to better understand the technical solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.
[0027] Unless otherwise defined, all technical and scientific terms used in this application have the same meaning as those commonly understood by those skilled in the art to which the present invention pertains. The terms used in this application and in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention.
[0028] The method for object recognition of multimodal data provided by the embodiment of the present invention is executed by a computer device. Accordingly, the system for object recognition of multimodal data runs in the computer device.
[0029] Figure 1 is a schematic flow chart of a method according to an embodiment of the present invention. Figure 1 The execution subject can be a multimodal data target recognition system. According to different requirements, the order of the steps in the flowchart can be changed, and some steps can be omitted.
[0030] like Figure 1 As shown, the method includes: S1. Collect radar echo data and infrared images using radar and infrared sensors at the same location, and associate the radar echo data and infrared images of the same target to obtain target radar echo data and target infrared images; S2. Converting the target radar echo data into a range-Doppler map; S3. Using a convolutional neural network to extract radar features from the range-Doppler map and extract infrared features from the target infrared image; S4. Obtaining environmental parameters and quality evaluation coefficients of the target infrared image, and generating feature confidence levels based on the environmental parameters and the quality evaluation coefficients; S5. Fuse the radar features and infrared features according to the feature confidence, and determine the target type based on the fused features.
[0031] Please refer to Figure 2The system as a whole can be divided into a self-detection module and a feature fusion module. In the self-detection module, the quality of the input infrared image and radar data is first evaluated. Secondly, the indicators quantified according to the evaluation system are encoded, and the data confidence is learned through the network. Finally, the learned confidence is sent to the feature fusion module, and the feature vector of the sensor is adjusted according to the confidence. In the feature extraction module, the CNN network is first used to extract features from the processed radar RD image and infrared image. Then, the obtained feature vector is adjusted by the confidence from the detection module, so that more reliable information can be used in subsequent fusion. Then, the LMF network is used to interact with the heterogeneous feature vectors, obtain the mutual dependence between the features, reduce redundancy, and adaptively fuse the feature vectors. Finally, the decoder is used to classify the features and identify the target type.
[0032] In an embodiment of the present invention, based on step S1, a possible embodiment will be given below to illustrate its specific implementation scheme in a non-limiting manner.
[0033] The radar echo data of multiple measuring points are acquired by a radar, and the infrared images of the multiple measuring points are acquired by an infrared sensor, wherein the radar and the infrared sensor are installed at the same position; the radar position parameters and the infrared position parameters of each measuring point are acquired, wherein the radar position parameters include the azimuth and the pitch angle relative to the radar, and the infrared position parameters include the azimuth and the pitch angle relative to the infrared sensor; the spherical distance between the radar position parameter of each measuring point in the radar echo data and the infrared position parameter of the measuring point acquired by the infrared sensor is calculated, and the measuring point corresponding to the radar position parameter and the infrared position parameter with the smallest spherical distance is used as a matching measuring point, and the radar echo data and the infrared image corresponding to the matching measuring point are determined as the target radar echo data and the target infrared image of the same target.
[0034] In a specific example, data registration is performed based on the angle measurement of the radar and infrared sensors. Assuming that the radar and infrared sensors are installed at the same position, Represents radar measurement data, Represents infrared measurement data, where each point and are the azimuth and elevation angles of radar and infrared respectively. For each radar measurement point , calculate it and all infrared measurement points The spherical distance between them is:
[0035] in Indicates the radar measurement Points and infrared measurement The spherical distance between points, for each radar measurement point , select the nearest infrared measurement point :
[0036] in Indicates the distance to the radar measurement point The nearest infrared measurement point. Based on this matching criterion, the best match is selected from the radar data and infrared data at the same time, thereby associating the radar data and infrared data to the same target.
[0037] In an embodiment of the present invention, based on step S2, a possible embodiment will be given below to illustrate its specific implementation scheme in a non-limiting manner.
[0038] The target radar echo data is pulse compressed through a convolution operation to obtain a compressed echo signal; the compressed echo signal of each range unit is Fourier transformed to obtain a frequency spectrum; the amplitude of the frequency spectrum is taken, and the zero-frequency component is moved to the center of the spectrum to obtain the range-Doppler map of the target.
[0039] Taking into account the heterogeneity of radar echo data and infrared image data, and the fact that the radar signal and infrared signal of the target are independent, directly fusing the raw data of the two sensors will result in weak feature correlation. Therefore, it is considered to perform a preliminary transformation on the radar echo data to obtain a representation that has a certain correlation with the infrared image before performing feature fusion.
[0040] A radar range-Doppler plot describes the reflection characteristics of a target object under radar wave illumination and is commonly used in radar signal processing and target detection. Targets vary significantly in speed, manifesting as significant differences in the Doppler frequency domain. Furthermore, targets vary significantly in physical size, resulting in significant differences in Doppler shift, Doppler spread, and the range units occupied by different targets.
[0041] For radar echo signals, pulse compression is first required to improve the range resolution. Assume that the transmitted signal is , the received signal is , the echo signal after pulse compression is obtained through convolution operation :
[0042] After pulse compression, the echo signal obtained has a shorter pulse width in the time domain, which can accurately reflect the distance of the target. Perform Fourier transform to obtain frequency spectrum :
[0043] in, is the Doppler shift, = represents the Fourier transform. Then, we take the amplitude, move the zero-frequency component to the center of the spectrum, and draw the RD graph of the target.
[0044] In an embodiment of the present invention, based on step S3, a possible embodiment will be given below to illustrate its specific implementation scheme in a non-limiting manner.
[0045] The ResNet18 network is selected as the feature extraction head to extract features of radar RD images and infrared images respectively. ResNet18 consists of 4 convolutional layers, 8 residual blocks and a global pooling layer. The overall network structure is as follows Figure 3 As shown in . The residual block structure is as follows Figure 4 shown.
[0046] The ResNet network can effectively solve the problem of gradient disappearance and degradation by introducing residual connections. The residual module mainly consists of two convolutional layers, batch normalization and activation functions. The specific implementation is as follows: the input feature map After the first layer of convolution, the output is obtained through batch normalization and nonlinear activation function. :
[0047] in represents a two-dimensional convolution, Represents the first layer convolution kernel weights. Then the output is obtained through the second layer of convolution and batch normalization :
[0048] in Represents the weight of the second layer convolution kernel. Finally, the residual connection is realized and the feature map of the input is With the second layer convolution output Connect and then perform nonlinear activation to get the final output of the residual block :
[0049] After 4 residual blocks, the final feature vector is obtained Therefore, the entire feature extraction module can be expressed as:
[0050] in, Indicates the residual blocks, Represents the original input image.
[0051] In an embodiment of the present invention, based on step S4, a possible embodiment will be given below to illustrate its specific implementation scheme in a non-limiting manner.
[0052] S401. Acquire environmental parameters and convert the environmental parameters into quantized values of the environmental parameters.
[0053] Quantified values of environmental parameters, such as sea state level, indicate that higher sea state levels lead to larger waves. Waves and uneven sea surfaces can alter the reflection path of radar signals, creating a multipath effect that reduces the clarity and accuracy of echo signals. This also creates more scattering, which interferes with radar.
[0054] Since the sea state level itself has been quantified using numerical values, when inputting into the network, the sea state level at the time of data collection is directly input as a separate one-dimensional information.
[0055] S402. Perform quality detection on the target infrared image to obtain a quality evaluation coefficient.
[0056] Please refer to Figure 5 , first calculate the brightness of the image. If it is not within the normal brightness range, it means that most of the information contained in the image is brightness information, and it is basically unable to provide information such as the position, texture and contour of the target itself. Therefore, the image is discarded and the image feature is no longer used; if it is within the normal brightness range, then calculate the contrast and input the calculated contrast as a one-dimensional parameter into the network.
[0057] In the brightness judgment module, the average brightness of the image is first calculated, and then the image is judged based on the brightness threshold. When the overall brightness of the image is too high or too low, it indicates that the image may be overexposed or dark. At this time, the effective information in the image is greatly reduced, affecting the subsequent network learning, so the image is directly discarded. In the contrast calculation module, the image is first Gaussian smoothed to reduce the impact of noise on subsequent calculations. Then, the Sobel operator is used to calculate the gradient of each pixel in the image:
[0058] in Indicates the horizontal gradient of the pixel. Indicates the vertical gradient of the pixel. Indicates the position in the image After obtaining the gradient, we continue to calculate the magnitude and direction of the gradient:
[0059]
[0060] in represents the magnitude of the gradient, Indicates the direction of the gradient. Then, non-maximum suppression is performed along the gradient direction to find the pixel with the local maximum gradient amplitude. At the same time, double thresholding is used to divide these pixels. If the gradient amplitude is greater than the high threshold, the pixel is divided into a strong edge; the pixel between the low threshold and the high threshold is divided into a weak edge; the pixel below the low threshold is divided into a non-edge. Finally, the weak edge points are screened. If the weak edge point is adjacent to the strong edge point, it is retained, otherwise it is removed. Finally, all reliable edge points are obtained. The gradient amplitude of these points is averaged to obtain the contrast between the target and the background of the image:
[0061] in represents the contrast of the infrared image, Represents the gradient amplitude value of the reliable edge points screened out by the above process. The calculated contrast is input into the network as one-dimensional information to better describe the reliability of the infrared image.
[0062] After obtaining the edges and contrast, the image must be further evaluated to determine whether it has been subjected to point or surface interference, such as infrared decoys. Based on the characteristics of surface source interference, the combustion unit forms an infrared radiation field with a certain intensity and area, generating multiple point or surface radiation sources. Based on this, the image edges are morphologically dilated, and then the dilated image is subjected to connected component analysis to calculate the total number of connected domains. A large number of connected domains indicates that the image has been subjected to point or surface source interference, and the confidence level should be lowered. Combining the above process, the image quality evaluation coefficient is:
[0063] in is the image quality evaluation coefficient, is the connected domain control coefficient, which controls the influence of the number of connected domains on the image quality evaluation coefficient. is the number of connected domains.
[0064] S403. Input the quantized value of the environmental parameter and the quality evaluation coefficient into a multi-layer perceptron to obtain the confidence level of the radar feature and the confidence level of the infrared feature.
[0065] The sea state level and infrared image quality evaluation coefficient will be converted into the confidence of adjusting the sensor features through this module. Considering that these two parameters are closely related to the sensor data, this application uses a multi-layer perceptron to deeply couple the relationship between parameters and features. The weights and biases of the multi-layer perceptron are updated and iterated in the network. This allows the network to learn the contribution of each feature in the two feature vectors to the recognition task under different circumstances. The input of the confidence calculation module is composed of the sea state level and infrared image quality evaluation coefficient It consists of two parts, and its structure is as follows Figure 6 First, the input is linearly encoded and nonlinearly activated through the information encoding module, so that the information encoding module consists of two identical branches, each of which consists of a fully connected layer and a Layer composition, output and Expressed as:
[0066] in , , and Represent the weight and bias of the linear transformation of sea state level and infrared image contrast respectively. Through this step, the dimension of the feature is changed from 1 to The two vectors after linear transformation are then concatenated to obtain a combined vector, which is then reduced to the dimension of the feature vector through the dimensionality reduction module. This confidence vector can interact with the feature vector to dynamically adjust the confidence of each feature in the feature vector. The output of the confidence calculation module is It can be expressed as:
[0067] in and are the weights and biases of the dimensionality reduction module.
[0068] In an embodiment of the present invention, based on step S5, a possible embodiment will be given below to illustrate its specific implementation scheme in a non-limiting manner.
[0069] generating a radar feature weight and an infrared feature weight according to the feature confidence; The dot product of the radar feature and the radar feature weight is recorded as a radar feature vector, and the dot product of the infrared feature and the infrared feature weight is recorded as an infrared feature vector; the low-rank modal factor of the radar feature and the low-rank modal factor of the infrared feature are obtained; a fused feature vector is calculated based on the radar feature vector, the low-rank modal factor of the radar feature, the infrared feature vector and the low-rank modal factor of the infrared feature; and a classifier is used to determine the target type based on the fused feature vector.
[0070] Existing methods commonly use feature concatenation as a basic fusion strategy, essentially linearly concatenating the feature vectors output by the encoders of each modality. However, such methods only achieve shallow feature combination and fail to effectively model nonlinear interactions between cross-modal features, resulting in a loss of potential correlation information. To overcome this limitation, this application introduces an Outer Product Interaction mechanism, which models high-order feature interactions by constructing a tensor product of bimodal features. Compared to traditional concatenation methods, the outer product operation explicitly captures second-order correlations between feature dimensions and establishes a mathematical representation of cross-modal feature associations. Compared to the currently mainstream cross-attention mechanism, this method has a unique theoretical advantage: when sensor data reliability degrades asymmetrically (such as radar signals in harsh sea conditions), the cross-attention mechanism may lead to erroneous association learning due to the noise sensitivity of its weight distribution mechanism. However, the outer product operation, by maintaining the symmetry of feature interactions, effectively mitigates the impact of unimodal data degradation on fusion performance.
[0071] The disadvantage of the outer product is its computational complexity and storage cost. Therefore, this paper adopts the Low-Rank Matrix Factorization (LMF) fusion network, which can effectively reduce the computational overhead and storage cost. Figure 7 shown.
[0072] The feature confidence is first norm-normalized and then adjusted for the infrared and radar feature vectors from the encoder: , , , in, is a vector of norm, is the normalized confidence vector. The adjusted feature vector is input to the LMF fusion network. Since the outer product needs to be calculated and converted into a one-dimensional feature vector for downstream tasks, for the traditional outer product fusion network, there are:
[0073]
[0074] And for the weight matrix , introduce the low-rank matrix, expressed as:
[0075] in and The infrared characteristics and radar characteristics are low-rank modal factors. Further derivation yields:
[0076] in" " represents the dot product between vectors. This transformation decouples the fusion of the two modalities, transforming the complex computation of the outer product of the eigenvectors and the conversion of the outer product feature fusion matrix into the feature fusion vector into a direct dot product between the eigenvectors and the modal factors. This reduces computational overhead and significantly increases efficiency. Furthermore, performing the outer product operation on the eigenvectors allows for simple and effective interaction between each feature in the vector, ensuring that the network can better capture the relationships between features.
[0077] In addition, low-rank modal factors are obtained by performing low-rank decomposition on the eigenvectors. Low-rank decomposition is the process of decomposing a high-dimensional data matrix into low-rank matrices and sparse matrices. This can be achieved through various methods, such as robust principal component analysis (RPCA) and augmented Lagrange multiplier method (ALM). In this application, low-rank decomposition is performed on radar and infrared eigenvectors to obtain their respective low-rank modal factors and sparse components.
[0078] In multimodal fusion, to further improve computational efficiency and fusion effectiveness, the weight tensor can be decomposed into modality-specific low-rank factors. This approach avoids the explicit creation of high-dimensional tensors, thereby reducing memory overhead and computational complexity. For the low-rank modal factors of radar and infrared signatures, modality-specific factor decomposition can be used to further extract low-rank factors specific to each modality.
[0079] The fusion feature f U Input the support vector machine and obtain the target type, such as ship.
[0080] In some embodiments, the target recognition system based on multimodal data may include multiple functional modules composed of computer program segments. The computer program of each program segment in the target recognition system based on multimodal data may be stored in a memory of a computer device and executed by at least one processor to perform (see Figure 1 Description) Functions for object recognition based on multimodal data.
[0081] In this embodiment, the target recognition system based on multimodal data can be divided into multiple functional modules according to the functions it performs, such as Figure 8As shown. The module referred to in the present invention refers to a series of computer program segments that can be executed by at least one processor and can perform fixed functions, which are stored in a memory. In this embodiment, the functions of each module will be described in detail in subsequent embodiments.
[0082] The data matching module is used to collect radar echo data and infrared images using radar and infrared sensors at the same position, and to associate the radar echo data and infrared images of the same target to obtain the target radar echo data and target infrared image; A preprocessing module, configured to convert the target radar echo data into a range-Doppler map; a feature extraction module, configured to extract radar features from the range-Doppler map using a convolutional neural network, and to extract infrared features from the target infrared image; A confidence calculation module, configured to obtain environmental parameters and a quality evaluation coefficient of a target infrared image, and generate a feature confidence level based on the environmental parameters and the quality evaluation coefficient; The feature processing module is used to fuse radar features and infrared features according to feature confidence and determine the target type based on the fused features.
[0083] Figure 9 The target recognition method based on multimodal data provided for the embodiment of the present application can be applied to a device. Those skilled in the art will understand that the device structure involved in the embodiment of the present invention does not constitute a limitation on the device, and the device may include more or fewer components than shown in the figure, or combine certain components, or arrange components differently. In an embodiment of the present invention, the device includes but is not limited to a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices and other similar computing devices. The components shown in this application, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the embodiments of the present application described and / or required in this application.
[0084] The device 900 may include a processor 910, a memory 920, and a communication unit 930. These components communicate via one or more buses. Those skilled in the art will appreciate that the server structure shown in the figure does not limit the present invention. The server structure may be a bus structure or a star structure, and may include more or fewer components than shown, or combine certain components, or arrange the components differently.
[0085] The memory 920 can be used to store execution instructions of the processor 910. The memory 920 can be implemented by any type of volatile or non-volatile memory device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk. When the execution instructions in the memory 920 are executed by the processor 910, the device 900 can perform some or all of the steps in the above-described method embodiments.
[0086] The processor 910 is the control center of the storage device, which uses various interfaces and lines to connect various parts of the entire electronic device. It executes various functions of the electronic device and / or processes data by running or executing software programs and / or modules stored in the memory 920, and calling data stored in the memory. The processor can be composed of an integrated circuit (IC), for example, it can be composed of a single packaged IC, or it can be composed of multiple packaged ICs with the same or different functions. For example, the processor 910 can only include a central processing unit (CPU). In an embodiment of the present invention, the CPU can be a single computing core or multiple computing cores.
[0087] The communication unit 930 is configured to establish a communication channel so that the storage device can communicate with other devices, receive user data sent by other devices, or send user data to other devices.
[0088] The present invention also provides a computer storage medium, wherein the computer storage medium may store a program that, when executed, may include some or all of the steps of each embodiment provided by the present invention. The storage medium may be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).
[0089] Those skilled in the art will clearly understand that the techniques in the embodiments of the present invention can be implemented using software and a necessary general-purpose hardware platform. Based on this understanding, the technical solutions in the embodiments of the present invention, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, among other media capable of storing program code, and includes instructions for causing a computer device (which can be a personal computer, a server, or a second device, a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention.
[0090] In this specification, the same or similar parts between the various embodiments can be referred to each other. In particular, for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method embodiment.
[0091] In the several embodiments provided by the present invention, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of the modules is merely a logical function division. In actual implementation, there may be other division methods, such as multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be through some interface, indirect coupling or communication connection of the system or module, which can be electrical, mechanical or other forms.
[0092] The modules described as separate components may or may not be physically separate, and the components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed across multiple network modules. Some or all of the modules may be selected to achieve the purpose of the present embodiment according to actual needs.
[0093] In addition, each functional module in each embodiment of the present invention may be integrated into one processing module, or each module may exist physically separately, or two or more modules may be integrated into one module.
[0094] Although the present invention has been described in detail with reference to the accompanying drawings and in conjunction with preferred embodiments, the present invention is not limited thereto. Without departing from the spirit and essence of the present invention, persons of ordinary skill in the art may make various equivalent modifications or substitutions to the embodiments of the present invention, and such modifications or substitutions shall be within the scope of the present invention. Any changes or substitutions that can be easily conceived by persons skilled in the art within the technical scope disclosed in the present invention shall be within the scope of protection of the present invention.
Claims
1. A target recognition method based on multimodal data, characterized in that: include: The radar and infrared sensors at the same position are used to collect radar echo data and infrared images, and the radar echo data and infrared images of the same target are correlated to obtain the target radar echo data and target infrared image; Converting the target radar echo data into a range-Doppler map; extracting radar features from the range-Doppler map using a convolutional neural network, and extracting infrared features from the infrared image of the target; Acquiring environmental parameters and quality evaluation coefficients of the target infrared image, and generating feature confidence levels based on the environmental parameters and the quality evaluation coefficients; The radar features and infrared features are fused according to the feature confidence, and the target type is determined based on the fused features.
2. The method according to claim 1, characterized in that The radar echo data and infrared image are collected by using the radar and infrared sensors at the same position, and the radar echo data and infrared image of the same target are correlated to obtain the target radar echo data and target infrared image, including: The radar echo data of multiple measuring points are acquired by the radar, and the infrared images of multiple measuring points are collected by the infrared sensor, wherein the radar and the infrared sensor are installed at the same position; Acquire radar position parameters and infrared position parameters of each measurement point, wherein the radar position parameters include azimuth and elevation angles relative to the radar, and the infrared position parameters include azimuth and elevation angles relative to the infrared sensor; The spherical distance between the radar position parameter of each measurement point in the radar echo data and the infrared position parameter of the measurement point collected by the infrared sensor is calculated, and the measurement point corresponding to the radar position parameter and infrared position parameter with the smallest spherical distance is taken as the matching measurement point. The radar echo data and infrared image corresponding to the matching measurement point are determined as the target radar echo data and target infrared image of the same target.
3. The method according to claim 1, characterized in that Converting the target radar echo data into a range-Doppler map, including: Perform pulse compression on the target radar echo data through convolution operation to obtain a compressed echo signal; Perform Fourier transform on the compressed echo signal of each range unit to obtain the frequency spectrum; The amplitude of the frequency spectrum is taken, and the zero-frequency component is moved to the center of the spectrum to obtain the range-Doppler map of the target.
4. The method according to claim 1, wherein The convolutional neural network includes 4 convolutional layers, 8 residual modules and a global pooling layer.
5. The method according to claim 1, wherein Acquiring environmental parameters and a quality evaluation coefficient of a target infrared image, generating a confidence level of a radar feature according to the environmental parameters, and generating a confidence level of the infrared feature according to the quality evaluation coefficient, including: Acquiring environmental parameters and converting the environmental parameters into environmental parameter quantized values; Performing quality detection on the target infrared image to obtain a quality evaluation coefficient; Inputting the quantized value of the environmental parameter and the quality evaluation coefficient into a multi-layer perceptron to obtain the confidence of the radar feature and the confidence of the infrared feature; The multi-layer perceptron includes a first branch, a second branch and a confidence calculation module, and the first branch and the second branch are both composed of a fully connected layer and a ReLU layer; The first branch is used to perform a linear transformation on the radar feature to obtain a radar confidence vector; The second branch is used to perform linear transformation on the infrared feature to obtain an infrared confidence vector; The confidence calculation module is used to splice the radar confidence vector and the infrared confidence vector into a combination vector, and obtain the feature confidence by performing dimensionality reduction processing on the combination vector.
6. The method according to claim 1, characterized in that The radar and infrared features are fused based on the feature confidence level, and the target type is determined based on the fused features, including: generating a radar feature weight and an infrared feature weight according to the feature confidence; The dot product of the radar feature and the radar feature weight is recorded as the radar feature vector, and the dot product of the infrared feature and the infrared feature weight is recorded as the infrared feature vector; Obtaining low-rank modal factors of radar features and low-rank modal factors of infrared features; Calculating a fusion feature vector according to the radar feature vector, the low-rank modal factor of the radar feature, the infrared feature vector, and the low-rank modal factor of the infrared feature; A classifier is used to determine the target type according to the fused feature vector.
7. The method according to claim 6, characterized in that Generating a radar feature weight and an infrared feature weight according to the feature confidence, including: Perform norm normalization on the feature confidence to obtain the confidence coefficient; The confidence coefficient is used as the infrared feature weight; The difference between 1 and the confidence coefficient is used as the radar feature weight.
8. A target recognition system based on multimodal data, characterized in that: include: The data matching module is used to collect radar echo data and infrared images using radar and infrared sensors at the same position, and to associate the radar echo data and infrared images of the same target to obtain the target radar echo data and target infrared image; A preprocessing module, configured to convert the target radar echo data into a range-Doppler map; a feature extraction module, configured to extract radar features from the range-Doppler map using a convolutional neural network, and to extract infrared features from the target infrared image; A confidence calculation module, configured to obtain environmental parameters and a quality evaluation coefficient of a target infrared image, and generate a feature confidence level based on the environmental parameters and the quality evaluation coefficient; The feature processing module is used to fuse radar features and infrared features according to feature confidence and determine the target type based on the fused features.
9. A device, characterized in that include: A memory for storing a target recognition program based on multimodal data; A processor, configured to implement the steps of the target recognition method based on multimodal data as described in any one of claims 1 to 7 when executing the target recognition program based on multimodal data.
10. A computer-readable storage medium storing a computer program, characterized in that: The readable storage medium stores a target recognition program based on multimodal data. When the target recognition program based on multimodal data is executed by the processor, the steps of the target recognition method based on multimodal data as described in any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Ship classification method based on HRRP and SAR data credible decision fusion
CN118587549A
Aviation video stream target identification processing method and system
CN119151984A
Radar infrared signal feature extraction and fusion method based on time sequence convolutional network
CN119399511A
Laser radar target detection method and system, terminal and storage medium
CN119418308A
Dynamic ship classification and identification method based on multi-modal mass change
CN119810564A
Cited By
Traffic accident video evidence analysis system and method based on multi-modal deep learning
CN121147821A
Distributed power distribution network fitting health condition assessment method, device and equipment under sky-ground monitoring system and storage medium
CN121684657A