Radar Target Multimodal Data Fusion Recognition Method Based on Knowledge Embedding Model
By introducing domain knowledge text data into radar target recognition and using knowledge embedding model for multi-mode data fusion recognition, the problem of insufficient utilization of external knowledge in the existing technology is solved, and the accuracy and accuracy of radar target recognition is improved.
Patent Information
- Application Number
- CN202510265440.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-03-07
AI Technical Summary
The existing radar target recognition methods are insufficiently utilized in complex application scenarios, which cannot meet the needs of robust target recognition, and the ability to extract and utilize multiple information features is insufficient.
The multi-mode data fusion recognition method of radar target based on knowledge embedding model is adopted. By obtaining radar echo data and domain knowledge data, the trained radar target recognition network extracts distance image feature tokens and domain knowledge text feature tokens, and performs fusion recognition to improve recognition performance.
Make full use of domain knowledge to improve the accuracy and performance of radar target recognition, enhance the interaction between HRRP echo and domain knowledge, and improve the recognition accuracy.
Smart Images

Figure CN119783042B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of radar, and particularly relates to a radar target multi-modal data fusion recognition method based on a knowledge embedding model. Background Art
[0002] The radar one-dimensional high-resolution range profile (HRRP) describes the distribution of target scattering centers along the radar line of sight. It includes relatively fine information such as the shape, structure, and radial size of the target, and can achieve high-resolution modeling of the target. Therefore, it is widely used in the field of radar target recognition. However, due to the increasingly complex actual application scenarios, there are problems such as insufficient utilization of external knowledge in target recognition based on a single HRRP, and it cannot meet the requirements of robust target recognition.
[0003] Patent CN202310427516.3 discloses a radar target recognition model training method, a radar target recognition method, and a device. By using a feature extraction network based on a series-parallel convolutional network and pooling attention, the efficiency of extracting effective information in the signal is improved. Patent CN202210201673.8 discloses a radar target fusion recognition method and system. By using a one-dimensional range image sub-network to extract one-dimensional range image features and a two-dimensional range image sub-network to extract two-dimensional range image features, and using a fusion network for fusion recognition, the accuracy of target recognition is improved. The above technologies only use radar echo data in applications, and the ability of the recognition method to extract and utilize features of various information still needs to be improved.
[0004] Therefore, there is an urgent need to provide a radar target recognition method to improve the above defects. Summary of the Invention
[0005] In order to solve the above problems existing in the prior art, the present invention provides a radar target multi-modal data fusion recognition method based on a knowledge embedding model. The technical problems to be solved by the present invention are realized through the following technical solutions:
[0006] In a first aspect, the present invention provides a radar target multi-modal data fusion recognition method based on a knowledge embedding model, including:
[0007] Obtain the radar echo data to be recognized and the corresponding domain knowledge data;
[0008] Input the radar echo data to be recognized and the corresponding domain knowledge data into a trained radar target recognition network for processing, extract the range image feature tokens in the radar echo data to be recognized, extract the domain knowledge text feature tokens of the domain knowledge data, and perform fusion recognition on the range image feature tokens and the domain knowledge text feature tokens to obtain the fusion recognition result of the target to be recognized;
[0009] Among them, the trained radar target recognition network is obtained by training the initial radar target recognition network with preset category data as the training data set.
[0010] Advantages of the present invention:
[0011] A radar target multi-modal data fusion recognition method based on a knowledge embedding model provided by the present invention introduces domain knowledge text data in the process of radar target recognition, enabling the domain knowledge text data to provide more target information complementary to the one-dimensional high-resolution range profile. The domain knowledge text data assists the trained radar target recognition network to better recognize the one-dimensional high-resolution range profile features, which helps to improve the radar target recognition performance.
[0012] The following will further elaborate on the present invention in conjunction with the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 is a flowchart of a radar target multi-modal data fusion recognition method based on a knowledge embedding model provided by an embodiment of the present invention;
[0014] Figure 2 is a schematic diagram of a trained radar target recognition network provided by an embodiment of the present invention;
[0015] Figure 3 is a schematic diagram of a trained first feature extraction module provided by an embodiment of the present invention;
[0016] Figure 4 is a schematic diagram of a trained second feature extraction module provided by an embodiment of the present invention;
[0017] Figure 5 is a schematic diagram of a trained mutual attention feature fusion recognition module provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0018] The following further describes the present invention in detail with specific embodiments, but the embodiments of the present invention are not limited thereto.
[0019] Please refer to Figure 1 , Figure 1 is a flowchart of a radar target multi-modal data fusion recognition method based on a knowledge embedding model provided by an embodiment of the present invention. A radar target multi-modal data fusion recognition method based on a knowledge embedding model provided by the present invention includes:
[0020] S101. Obtain the radar echo data to be recognized and the corresponding domain knowledge data.
[0021] Specifically, in this embodiment, the radar echo data to be recognized is one-dimensional high-resolution range profile, and the domain knowledge data is target measurement information, including the track information of the target and the Radar Cross Section (RCS). The radar echo data to be recognized and the corresponding domain knowledge data are three types of aircraft target data, and the three types of aircraft are composed of aircraft A, aircraft B, and aircraft C.
[0022] S102. Input the radar echo data to be recognized and the corresponding domain knowledge data into the trained radar target recognition network for processing, extract the range profile feature tokens in the radar echo data to be recognized, extract the domain knowledge text feature tokens of the domain knowledge data, and perform fusion recognition on the range profile feature tokens and the domain knowledge text feature tokens to obtain the fusion recognition result of the target to be recognized;
[0023] Among them, the trained radar target recognition network is obtained by training the initial radar target recognition network with the preset category data as the training data set.
[0024] Specifically, please refer to Figure 2 , Figure 2 which is a schematic diagram of the trained radar target recognition network provided by the embodiment of the present invention. In this embodiment, the trained radar target recognition network includes a trained first feature extraction module, a trained second feature extraction module, and a trained mutual attention feature fusion recognition module; input the radar echo data to be recognized and the corresponding domain knowledge data into the trained radar target recognition network for processing, extract the range profile feature tokens in the radar echo data to be recognized, extract the domain knowledge text feature tokens of the domain knowledge data, and perform fusion recognition on the range profile feature tokens and the domain knowledge text feature tokens to obtain the fusion recognition result of the target to be recognized, including:
[0025] Use the trained first feature extraction module to perform feature extraction on the radar echo data to be recognized to obtain range profile feature tokens;
[0026] Use the trained second feature extraction module to perform feature extraction on the domain knowledge data to obtain domain knowledge text feature tokens;
[0027] Use the trained mutual attention feature fusion recognition module to perform feature fusion and recognition on the range profile feature tokens and the domain knowledge text feature tokens to obtain the fusion recognition result of the target to be recognized.
[0028] It should be noted that the token is token.
[0029] Further, please refer to Figure 3 , Figure 3FIG. 0 is a schematic diagram of a trained first feature extraction module provided by an embodiment of the present invention. In this embodiment, the trained first feature extraction module includes a one-dimensional high-resolution range image preprocessing module and a range image feature extraction module; using the trained first feature extraction module to extract features from the radar echo data to be recognized, obtaining range image feature tokens, including:
[0030] Using the one-dimensional high-resolution range image preprocessing module to preprocess the radar echo data to be recognized, obtaining a processed one-dimensional range image;
[0031] Using the range image feature extraction module to extract features from the processed one-dimensional range image, obtaining range image feature tokens.
[0032] Optionally, the one-dimensional high-resolution range image preprocessing module performs normalization and centering alignment operations using the L2 norm. The range image feature extraction module adopts a Transformer structure, sequentially including 1 one-dimensional convolutional layer (convolution kernel size is 8, stride is 8), 1 layer normalization layer, 2 attention layers and 1 layer normalization layer; each attention layer sequentially includes 1 multi-head attention layer, 1 layer normalization layer, 1 feed-forward layer, 1 layer normalization layer; among them, the multi-head attention layer includes 3 linear layers (input nodes are 128, output nodes are 128), and the feed-forward layer sequentially includes 1 linear layer (input nodes are 128, output nodes are 512), 1 GELU activation layer and 1 linear layer (input nodes are 512, output nodes are 128).
[0033] Further, please refer to Figure 4 , Figure 4 FIG. 16 is a schematic diagram of a trained second feature extraction module provided by an embodiment of the present invention. In this embodiment, the trained second feature extraction module includes a domain knowledge text preprocessing module and a domain knowledge text feature extraction module; using the trained second feature extraction module to extract features from the domain knowledge data, obtaining domain knowledge text feature tokens, including:
[0034] Using the domain knowledge text preprocessing module to preprocess the domain knowledge data, obtaining processed domain knowledge text tokens;
[0035] Using the domain knowledge text feature extraction module to extract features from the processed domain knowledge text tokens, obtaining domain knowledge text feature tokens.
[0036] Optionally, the domain knowledge text preprocessing module is used to describe domain knowledge in text. For example, "The target attitude is 30 degrees" is described as "The target attitude is 30 degree", and the text descriptions of the remaining domain knowledge are similar. Then, a tokenizer is used to obtain the numerical vector corresponding to the text description. In this embodiment, the CLIP vocabulary and the Byte Pair Encoding tokenization method are used to obtain domain knowledge text tokens, with a dimension of 50×128, where 50 is the number of domain knowledge text tokens and 128 is the dimension of each token. The domain knowledge text feature extraction module adopts a Transformer structure, which sequentially includes 2 attention layers, 1 layer normalization layer, and 1 feed-forward layer. Each attention layer includes 1 multi-head attention layer. Among them, the multi-head attention layer includes 3 linear layers (with 128 input nodes and 128 output nodes); the feed-forward layer sequentially includes 1 linear layer (with 128 input nodes and 512 output nodes), 1 GELU activation layer, and 1 linear layer (with 512 input nodes and 128 output nodes).
[0037] Further, please refer to Figure 5 , Figure 5 FIG. is a schematic diagram of a trained mutual attention feature fusion recognition module provided by an embodiment of the present invention. In this embodiment, the trained mutual attention feature fusion recognition module includes a plurality of cascaded mutual attention fusion modules and a fusion classification layer. Among them, each mutual attention fusion module includes a first layer normalization layer, a second layer normalization layer, a mutual attention fusion layer, a third layer normalization layer, a fourth layer normalization layer, a first feed-forward layer, and a second feed-forward layer; the trained mutual attention feature fusion recognition module is used to perform feature fusion and recognition on the range image feature tokens and the domain knowledge text feature tokens to obtain the fusion recognition result of the target to be recognized, including:
[0038] For the first-level mutual attention fusion module, the first layer normalization layer is used to normalize the range image feature tokens to obtain a first feature;
[0039] The mutual attention fusion layer is used to process the first feature to obtain a second feature;
[0040] The second feature is added to the range image feature tokens to obtain a third feature;
[0041] The third layer normalization layer is used to normalize the third feature to obtain a fourth feature;
[0042] The first feed-forward layer is used to process the fourth feature to obtain a fifth feature;
[0043] The fifth feature is added to the third feature to obtain a one-dimensional high-resolution range image classification feature vector;
[0044] The second normalization layer is used to normalize the domain knowledge text feature tokens to obtain the sixth feature;
[0045] The cross-attention fusion layer is used to process the sixth feature to obtain the seventh feature;
[0046] The seventh feature is added to the domain knowledge text feature tokens to obtain the eighth feature;
[0047] The fourth normalization layer is used to normalize the eighth feature to obtain the ninth feature;
[0048] The second feed-forward layer is used to process the ninth feature to obtain the tenth feature;
[0049] The tenth feature is added to the eighth feature to obtain the domain knowledge text classification feature vector;
[0050] The one-dimensional high-resolution range image classification feature vector and the domain knowledge text classification feature vector are concatenated to obtain the fused feature;
[0051] The fused classification layer is used to classify the fused feature vector to obtain the fused recognition result of the target to be recognized.
[0052] In this embodiment, the training process of the trained radar target recognition network includes:
[0053] Obtain data of multiple preset categories to construct a training dataset; wherein, the training dataset includes multiple samples, and each sample includes radar echo data and corresponding domain knowledge data; optionally, each data sample includes a radar one-dimensional high-resolution range image and corresponding domain knowledge data, the dimension of the radar one-dimensional high-resolution range image is 1×256; the dimension of the domain knowledge data is 1×5, and the domain knowledge data includes target attitude, target pitch angle, target height, target speed, target RCS;
[0054] Input part of the samples in the training dataset into the th untrained radar target recognition network for training to obtain the prediction result classified and output during the th training process;
[0055] According to the prediction result classified and output during the th training process and the true label of the sample of the th untrained radar target recognition network during training, calculate the classification loss and use it as the classification loss of the th training process; meanwhile, according to the one-dimensional high-resolution range image classification feature vector and the domain knowledge text classification feature vector output during the th training process, calculate the matching loss and use it as the matching loss of the th training process;
[0056] According to the classification loss and matching loss of the th training process, perform backpropagation to update the th network parameters of the radar target recognition network to be trained, and obtain the th radar target recognition network to be trained; iterate in this way until the number of training times or the degree of convergence meets the preset conditions, and obtain the trained radar target recognition network.
[0057] It should be noted that use the sum of the classification loss and matching loss of the th training process to perform backpropagation to update the th network parameters of the radar target recognition network to be trained.
[0058] Optionally, the total number of samples in the training dataset is 3,000. Each sample includes 1 radar one-dimensional high-resolution range profile and 1 piece of domain knowledge data. The domain knowledge data includes target attitude, target pitch angle, target height, target speed, and target RCS. The training dataset includes 3 types of targets. Among them, there are 1,000 samples of aircraft A, 1,000 samples of aircraft B, and 1,000 samples of aircraft C. The total number of samples in the test dataset is 1,500, and the composition of each sample is the same as that of the training dataset. Among them, there are 500 samples of aircraft A, 500 samples of aircraft B, and 500 samples of aircraft C.
[0059] Furthermore, in this embodiment, the mutual attention feature fusion recognition module in the initial radar target recognition network includes an initial mutual attention fusion module, an initial one-dimensional high-resolution range profile classification layer, an initial domain knowledge text classification layer, and an initial fusion classification layer; the process of obtaining the classification loss of the th training process includes:
[0060] Use the initial mutual attention fusion module to generate the one-dimensional high-resolution range profile classification feature vector and the domain knowledge text classification feature vector in the th training process;
[0061] Use the initial one-dimensional high-resolution range profile classification layer to classify the one-dimensional high-resolution range profile classification feature vector generated in the th training process to obtain a first prediction result, and compare the first prediction result with the true label of the sample of the th radar target recognition network to be trained to obtain a first classification loss;
[0062] Use the initial fusion classification layer to classify the fusion feature vector generated in the th training process to obtain a second prediction result, and compare the second prediction result with the Compare with the true label of the sample of the radar target recognition network to be trained next time to obtain the second classification loss;
[0063] Use the initial domain knowledge text classification layer to classify the domain knowledge text classification feature vectors generated during the th training process to obtain the third prediction result, and compare the third prediction result with the true label of the sample of the radar target recognition network to be trained in the th time to obtain the third classification loss;
[0064] Add the first classification loss, the second classification loss, and the third classification loss to obtain the classification loss of the th training process.
[0065] Optionally, the initial mutual attention feature fusion recognition module includes 1 initial mutual attention fusion module, an initial one-dimensional high-resolution range profile classification layer, an initial domain knowledge text classification layer, and an initial fusion classification layer. Among them, the initial mutual attention fusion module sequentially includes 2 first-layer normalization layers, 1 mutual attention fusion layer, 2 second-layer normalization layers, and 2 feed-forward layers. The mutual attention fusion layer sequentially includes 2 linear layers (with 128 input nodes and 256 output nodes), 2 linear layers (with 128 input nodes and 512 output nodes), 2 linear layers (with 256 input nodes and 128 output nodes), 2 linear layers (with 128 input nodes and 512 output nodes), 2 GELU activation layers, and 2 linear layers (with 512 input nodes and 128 output nodes). The one-dimensional high-resolution range profile classification layer includes 1 linear layer (with 128 input nodes and 128 output nodes), 1 GELU activation layer, and 1 linear layer (with 128 input nodes and 3 output nodes). The domain knowledge text classification layer includes 1 linear layer (with 128 input nodes and 128 output nodes), 1 GELU activation layer, and 1 linear layer (with 128 input nodes and 3 output nodes). The fusion classification layer includes 1 linear layer (with 256 input nodes and 128 output nodes), 1 GELU activation layer, and 1 linear layer (with 128 input nodes and 3 output nodes).
[0066] Furthermore, in this embodiment, the expression of the first classification loss, the second classification loss, or the third classification loss is:
[0067] ;
[0068] Among them, represents the first classification loss, the second classification loss, or the third classification loss, represents the true label, represents the prediction result, represent different modalities, representing one-dimensional high-resolution range profile modality, domain knowledge text modality, and fusion modality respectively, represents the total number of categories, represents the category index.
[0069] Furthermore, in this embodiment, the expression of the matching loss is:
[0070] ;
[0071] wherein, represents the matching loss value, represents the one-dimensional high-resolution range profile classification feature vector, represents the domain knowledge text classification feature vector, represents the transpose, represents the vector L2 norm.
[0072] Optionally, during training, the batch size is 128, the initial learning rate is 1e-4, the optimizer is AdamW, and the total number of iterations is set to 100.
[0073] In summary, a radar target recognition method based on the knowledge embedding fusion of a large language model provided by the present invention has the following beneficial effects:
[0074] First, compared with the radar target recognition method that only uses HRRP echoes, the radar target multi-modal data fusion recognition method based on the knowledge embedding model constructed by the present invention makes full use of domain knowledge, including target motion state knowledge and target RCS knowledge. The domain knowledge and echoes describe the target characteristics from different dimensions, improving the accuracy of target modeling.
[0075] Second, the radar target multi-modal data fusion recognition method based on the knowledge embedding model constructed by the present invention fully considers the characteristics of domain knowledge and HRRP echoes, and designs an optimal way to comprehensively utilize information. For HRRP echoes, by using a feature extractor that perceives local and global structures, the scattering center distribution features are mined, and a HRRP echo feature token sequence is constructed. For domain knowledge, by constructing a text-based knowledge representation strategy, the physical meaning of each dimension of knowledge is clarified, and the deep knowledge meaning is mined by using the understanding and reasoning ability of the large language model to construct a domain knowledge token sequence. Through the targeted processing of HRRP echoes and domain knowledge, the ability to mine various information contained in HRRP echoes and domain knowledge is improved.
[0076] Third, the radar target multi-modal data fusion and recognition method based on the knowledge embedding model constructed by the present invention designs an HRRP-knowledge fusion module based on mutual attention. Compared with the conventional fusion strategy, it enhances the interaction between HRRP echoes and domain knowledge, improves the effectiveness of fusion, and enhances the radar target recognition performance.
[0077] In an optional embodiment of the present invention, simulation experiments are conducted to verify the effect of the radar target recognition method based on language large model knowledge embedding fusion provided in the above embodiment. Specifically:
[0078] I. Simulation Conditions
[0079] The hardware platform for the simulation experiment of the present invention is: the processor is an Intel(R) i7 CPU with a main frequency of 3.20 GHz, the memory capacity is 64 GB, the graphics card is an Nvidia RTX 3090, and the video memory capacity is 24 GB.
[0080] The software platform for the simulation experiment of the present invention is: Python 3.8, Pytorch 2.0.
[0081] II. Simulation Content and Result Analysis
[0082] In the simulation experiment of the present invention, the method proposed by the present invention and an existing technology are used to conduct target recognition experiments on three types of aircraft, and the recognition rates are compared. The existing technology refers to the paper "An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale" published by Alexey Dosovitskiy et al., which applies the Transformer network to target recognition. The original network is designed for two-dimensional image data. In this simulation experiment, its main structure is retained, and the two-dimensional operations are modified to one-dimensional operations to adapt to one-dimensional high-resolution range profile data. In the simulation experiment, the recognition rate is used to evaluate the performance of the recognition model, which is defined as the ratio of the number of correctly recognized samples to the total number of test samples. The higher the recognition rate, the more correctly recognized samples, and the better the recognition performance of the model, as shown in Table 1.
[0083] Table 1 Comparison of target recognition accuracy rates between the method proposed by the present invention and the existing technology
[0084]
[0085] By comparing the recognition rates of the present invention and the existing technology, it can be seen that the present invention effectively improves the radar target recognition performance.
[0086] It should be noted that in this text, relative terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprise", "include" or any other variant are intended to cover non-exclusive inclusion, so that an article or device comprising a series of elements includes not only those elements but also other elements not expressly listed. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the article or device comprising the element. Terms such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. Orientations or positional relationships indicated by "upper", "lower", "left", "right", etc. are based on the orientations or positional relationships shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation of the present invention.
[0087] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine the different embodiments or examples described in this specification.
[0088] The above content is a further detailed description of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can be made, and all should be regarded as belonging to the protection scope of the present invention.
Claims
1. A radar target multi-modal data fusion recognition method based on a knowledge embedding model, characterized in that Including: Obtain radar echo data to be recognized and corresponding domain knowledge data; Input the radar echo data to be recognized and the corresponding domain knowledge data into a trained radar target recognition network for processing, extract range image feature tokens from the radar echo data to be recognized, extract domain knowledge text feature tokens of the domain knowledge data, and fuse and recognize the range image feature tokens and the domain knowledge text feature tokens to obtain a fused recognition result of the target to be recognized; Among them, the trained radar target recognition network is obtained by training an initial radar target recognition network with preset category data as a training data set; among them, the trained radar target recognition network includes a trained mutual attention feature fusion recognition module, and the trained mutual attention feature fusion recognition module includes a plurality of cascaded mutual attention fusion modules and a fusion classification layer. Among them, each mutual attention fusion module includes a first normalization layer, a second normalization layer, a mutual attention fusion layer, a third normalization layer, a fourth normalization layer, a first feed-forward layer, and a second feed-forward layer; using the trained mutual attention feature fusion recognition module to perform feature fusion and recognition on the range image feature tokens and the domain knowledge text feature tokens to obtain a fused recognition result of the target to be recognized, including: Perform normalization processing on the range image feature tokens using the first normalization layer to obtain a first feature; Process the first feature using the mutual attention fusion layer to obtain a second feature; Add the second feature to the range image feature tokens to obtain a third feature; Perform normalization processing on the third feature using the third normalization layer to obtain a fourth feature; Process the fourth feature using the first feed-forward layer to obtain a fifth feature; Add the fifth feature to the third feature to obtain a one-dimensional high-resolution range image classification feature vector; Perform normalization processing on the domain knowledge text feature tokens using the second normalization layer to obtain a sixth feature; Process the sixth feature using the mutual attention fusion layer to obtain a seventh feature; Add the seventh feature to the domain knowledge text feature tokens to obtain an eighth feature; Perform normalization processing on the eighth feature using the fourth normalization layer to obtain a ninth feature; Process the ninth feature using the second feed-forward layer to obtain a tenth feature; Add the tenth feature to the eighth feature to obtain a domain knowledge text classification feature vector; Concatenate the one-dimensional high-resolution range image classification feature vector and the domain knowledge text classification feature vector to obtain a fused feature vector; Classify the fused feature vector using the fusion classification layer to obtain a fused recognition result of the target to be recognized.
2. The radar target multi-modal data fusion recognition method based on the knowledge embedding model according to claim 1, wherein, The trained radar target recognition network further includes a trained first feature extraction module and a trained second feature extraction module; Inputting the radar echo data to be recognized and the corresponding domain knowledge data into the trained radar target recognition network for processing, and extracting the range profile feature tokens from the radar echo data to be recognized and the domain knowledge text feature tokens from the domain knowledge data, includes: Using the trained first feature extraction module to extract features from the radar echo data to be recognized, obtaining range profile feature tokens; Using the trained second feature extraction module to extract features from the domain knowledge data, obtaining domain knowledge text feature tokens.
3. The radar target multi-modal data fusion recognition method based on the knowledge embedding model according to claim 2, wherein, The trained first feature extraction module includes a one-dimensional high-resolution range profile preprocessing module and a range profile feature extraction module; The step of using the trained first feature extraction module to extract features from the radar echo data to be recognized and obtaining range profile feature tokens includes: Using the one-dimensional high-resolution range profile preprocessing module to preprocess the radar echo data to be recognized, obtaining the processed one-dimensional range profile; Using the range profile feature extraction module to extract features from the processed one-dimensional range profile, obtaining range profile feature tokens.
4. The radar target multi-modal data fusion recognition method based on the knowledge embedding model according to claim 2, characterized in that The trained second feature extraction module includes a domain knowledge text preprocessing module and a domain knowledge text feature extraction module; The step of using the trained second feature extraction module to extract features from the domain knowledge data and obtaining domain knowledge text feature tokens includes: Using the domain knowledge text preprocessing module to preprocess the domain knowledge data, obtaining the processed domain knowledge text tokens; Using the domain knowledge text feature extraction module to extract features from the processed domain knowledge text tokens, obtaining domain knowledge text feature tokens.
5. The radar target multi-modal data fusion recognition method based on the knowledge embedding model according to claim 1, wherein, The training process of the trained radar target recognition network includes: Obtaining a plurality of data of the preset categories to construct a training data set; wherein, the training data set includes a plurality of samples, and each sample includes radar echo data and corresponding domain knowledge data; Input a part of the samples in the training dataset into the th radar target recognition network to be trained for training, and obtain the prediction results output by classification during the th training process; According to the predicted result of the classification output during the -th training process and the true label of the sample of the radar target recognition network to be trained in the -th training, calculate the classification loss and use it as the classification loss of the -th training process; meanwhile, according to the one-dimensional high-resolution range image classification feature vector and the domain knowledge text classification feature vector output during the -th training process, calculate the matching loss and use it as the matching loss of the -th training process; According to the classification loss and matching loss of the -th training process, perform backpropagation to update the -th network parameters of the radar target recognition network to be trained, and obtain the -th radar target recognition network to be trained; iterate in this way until the number of training times or the convergence degree meets the preset conditions, and obtain the trained radar target recognition network.
6. The method for radar target multi-modal data fusion recognition based on the knowledge embedding model according to claim 5, wherein The mutual attention feature fusion recognition module in the initial radar target recognition network includes an initial mutual attention fusion module, an initial one-dimensional high-resolution range profile classification layer, an initial domain knowledge text classification layer, and an initial fusion classification layer; the acquisition process of the classification loss in the $th$ training process includes: Generate the one-dimensional high-resolution range image classification feature vector and the domain knowledge text classification feature vector during the th training process by using the described initial mutual attention fusion module; Use the described initial one-dimensional high-resolution range image classification layer to classify the one-dimensional high-resolution range image classification feature vectors generated during the th training process to obtain a first prediction result, and compare the first prediction result with the true label of the sample of the radar target recognition network to be trained for the th time to obtain a first classification loss; Using the initial fusion classification layer to classify the fusion feature vectors generated during the th training process to obtain a second prediction result, and comparing the second prediction result with the true label of the sample for training the th radar target recognition network to be trained to obtain a second classification loss; Use the described initial domain knowledge text classification layer to classify the domain knowledge text classification feature vectors generated during the th training process to obtain a third prediction result, and compare the third prediction result with the true label of the sample for training the th radar target recognition network to be trained to obtain a third classification loss; Add the first classification loss, the second classification loss, and the third classification loss to obtain the classification loss of the th training process.
7. The method for radar target multi-modal data fusion recognition based on the knowledge embedding model according to claim 6, characterized in that The expressions of the first classification loss, the second classification loss or the third classification loss are all: ; Among them, represents the first classification loss, the second classification loss, or the third classification loss, represents the true label, represents the prediction result, represents different modalities, respectively represent the one-dimensional high-resolution range image modality, the domain knowledge text classification modality, and the fusion modality, represents the total number of categories, represents the category index.
8. The radar target multi-modal data fusion recognition method based on the knowledge embedding model according to claim 5, characterized in that The expression of the matching loss is: ; Among them, represents the matching loss value, represents the one-dimensional high-resolution range profile classification feature vector, represents the domain knowledge text classification feature vector, represents the transpose, represents the vector L2 norm.
Citation Information
Patent Citations
A radar target fusion recognition method and system
CN114578307B
Radar target identification model training method, radar target identification method and device
CN116593980A
Space micro-motion target identification method based on complex value dynamic fusion network
CN119293581A