Radar signal modulation mode open set identification method based on multi-modal fusion

By constructing a multimodal fusion radar signal recognition model, combining time-frequency images and text descriptions, and optimizing the embedding space structure and rejection mechanism, the problem of identifying unknown categories in radar signal recognition is solved. This achieves efficient rejection of unknown categories and recognition of known categories, thereby improving the recognition performance of radar signal modulation methods.

CN120804781APending Publication Date: 2025-10-17BEIJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510911180.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-02
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing radar signal modulation identification methods are limited in robustness and adaptability when faced with category uncertainty in open environments, making it difficult to effectively identify unknown categories. Furthermore, single-modal feature representations fail to fully exploit the complementarity of different information sources.

Method used

A dual-tower network model is constructed to integrate the two modal features of time-frequency distribution images and text descriptions. The embedding space structure is optimized using the triplet loss function. The extreme value theory is introduced to model the tail distribution of distances and a rejection mechanism is established to improve the recognition ability and robustness of unknown class signals.

Benefits of technology

It significantly improves the performance of open set recognition of radar signal modulation methods, while maintaining a high recognition rate of known categories and increasing the rejection rate of unknown categories, demonstrating high robustness and stability in low signal-to-noise ratio environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120804781A_ABST
    Figure CN120804781A_ABST
Patent Text Reader

Abstract

According to the radar signal modulation mode open set recognition method based on multi-modal fusion, time-frequency image and text features are fused, classification discrimination is performed through metric learning, an extreme value theory is introduced, a rejection mechanism is established, and the recognition capability of unknown class signals is improved. The method specifically comprises the following steps: generating a time-frequency graph according to different radar signal models, generating corresponding text description by using a large language model, and constructing an image-text pairing data set; building an image-text fusion model to realize modal feature interactive fusion, and optimizing an embedding space; calculating a category embedding center based on the fused joint embedding representation; for each known class, measuring the distance between the image-text joint embedding representation of the sample under the class and the embedding center of the class; selecting a tail sample based on the distance measurement result, and fitting a Weibull distribution model; and distance correction is carried out on the test samples by using the distribution model, difference values of distance measurement results before and after correction of each type are aggregated, unknown type distance measurement is constructed, and a softmax function is uniformly input to complete identification and rejection discrimination.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of electronic countermeasures, in particular to a radar signal modulation mode recognition method in electronic intelligence processing. BACKGROUND

[0002] Radar signal modulation mode recognition plays an increasingly important role in key tasks such as spectrum monitoring and target recognition. Accurate identification of the modulation type of radar signals not only helps to understand the working mechanism of the opponent radar system, but also is the basis for realizing intelligent electronic reconnaissance and jamming.

[0003] Most of the current mainstream radar modulation recognition methods are based on closed set assumption, that is, the class set of the training set and the test set is completely consistent. This kind of method usually characterizes the input data through deep neural network, and realizes high recognition accuracy within the known class range. However, due to the lack of modeling ability for unknown classes, when the modulation type not seen in the training stage appears in the test process, the model often has difficulty in making accurate judgment, and even misjudges as known class, resulting in significant performance degradation. Therefore, when facing the class uncertainty in the open environment, the robustness and adaptability of this kind of method are severely limited.

[0004] In recent years, some researches have begun to focus on the open set recognition task of radar signal modulation mode, trying to build a classification model with unknown class perception ability, mainly from the aspects of improving loss function, designing rejection mechanism and introducing uncertainty modeling. However, this kind of method generally relies on single modal feature representation, and cannot effectively mine the complementarity between different information sources. In contrast, multi-modal fusion has shown significant performance advantages in image recognition, video understanding and other fields, especially in handling high-dimensional heterogeneous data and improving model generalization ability. Therefore, in the open set scene, it is not a bad idea to build a recognition model with multi-modal fusion and uncertainty recognition ability, which is an effective way to improve the performance of radar signal modulation mode open set recognition. SUMMARY

[0005] A radar signal modulation mode open set recognition method based on multi-modal fusion is proposed in this paper, a double-tower network model is constructed to fuse time-frequency distribution image and text description two modal features, and triplet loss is combined to optimize the embedding space structure. Classification is performed through metric learning, and the tail distribution of distance is modeled by introducing extreme value theory to establish a rejection mechanism, thereby improving the recognition ability and robustness of unknown class signals.

[0006] The recognition method is as followsFigure 1 As shown, it mainly includes steps 200-250.

[0007] At step 200, data of time-frequency distribution modes are made according to different radar signal models, image data corresponding text data is completed by using a large language model, and a complete image-text mode data set is constructed.

[0008] According to the radar signal model, time series of different categories of radar signals are simulated, the Choi-Williams Distribution (CWD) result of the radar signal in the form of a two-dimensional matrix is calculated through formula (1), the matrix values are mapped to color intensity, and the time-frequency distribution of the radar signal is reflected in the form of an image.

[0009] CWD x (t, f) = ∫∫A x (θ, τ) · Φ CW (θ, τ) · e j2π(tθ+fτ) dθdτ (1)

[0010] Wherein, each parameter is defined as follows:

[0011] CWD x (t, f) is the CWD of the signal in time and frequency;

[0012] A x (θ, τ) is the ambiguity function of the signal, defined as:

[0013]

[0014] x(t) is a complex-valued signal;

[0015] x * (t) is the complex conjugate of the signal;

[0016] t is the time variable;

[0017] f is the frequency variable;

[0018] e -j2πθt is a complex exponential modulation term;

[0019] Φ CW (θ, τ) is the kernel function of CWD, defined as:

[0020]

[0021] Wherein σ is a smoothing control parameter;

[0022] e j2π(tθ+fτ) is a two-dimensional Fourier kernel function;

[0023] To realize the joint modeling of image modalities and text modalities, the application further constructs an image-text multi-modal data set. After obtaining the time-frequency image of each type of signal, a large language model with image understanding capability is used to generate the corresponding text description to complete the text modality feature. As shown in Figure 3 , the process is completed through interactive step-by-step questioning, mainly including the following three aspects:

[0024] 1) Describe the visual features of the images of each type of signal.

[0025] 2) Extract discriminative features of each type of signal relative to other types.

[0026] 3) Restate the discriminative features in diverse languages to enhance semantic robustness.

[0027] Step 210, construct a double-tower model architecture composed of an image encoder and a text encoder, and design a multi-modal interaction module to fuse image-text modality features, as shown in Figure 4 . The interaction module is based on cross-attention mechanism, which receives the input feature vectors of image modalities and text modalities, obtains weights through multi-head attention calculation, and the calculation process is shown in formula (4).

[0028]

[0029] Feature is the joint embedding representation of image and text;

[0030] is the text input feature vector;

[0031] is the image input feature vector;

[0032] B is the batch size;

[0033] d is the total dimension of the feature vector;

[0034] h is the number of attention heads;

[0035] d k = d / h is the dimension of each attention head;

[0036] is the linear projection matrix of the i-th head;

[0037] is the output linear mapping matrix;

[0038] || represents the concatenation of the outputs of each attention head in the feature dimension;

[0039] Further, according to formula (5), a triplet loss function is used to optimize the distribution structure of the training samples in the embedding space.

[0040]

[0041] where f a is the embedding representation of the anchor sample, f p is the embedding representation of the positive sample (same class as the anchor), f n is the embedding representation of the negative sample (different class from the anchor), d ap is the Euclidean distance calculation between the embedding representation of the anchor sample and the positive sample, and a is the margin.

[0042] Step 220, iterate the optimization process in step 210. After completing the optimization of the training sample joint embedding representation, the sample in the embedding space according to formula (6) is aggregated to obtain the class embedding center Feature center of the class, which is used as a high-dimensional representation for subsequent distance measurement.

[0043]

[0044] where, is the joint embedding representation of the i-th sample, is the number of samples in the c-th class, is the sample index set of all classes c, y i ∈{1,2,...,C} is the real class label of the sample, and there are C class labels in total.

[0045] Step 230, for each known class in the training set, measure the Euclidean distance d between the sample joint embedding representation and the class embedding center according to formula (7).

[0046] d = ||Feature-Feature center ||2 (7)

[0047] Step 240, select a part of the samples farthest from the class center as the tail data, and perform Weibull distribution fitting on the distance distribution results. The cumulative distribution function of the fitted Weibull distribution is shown in formula (10).

[0048]

[0049] where τ, λ, η are the position parameter, scale parameter and shape parameter obtained after distribution fitting, respectively.

[0050] Step 250, based on the fitted class-specific Weibull distribution, correct the distance measurement results of the test sample, such as Figure 5d is weighted and corrected according to formula (9) and (10), and the category correction distance d is obtained corrected .

[0051]

[0052] d corrected = alpha * d (10)

[0053] In the formula, alpha is a correction factor introduced for each category sample distance.

[0054] The difference between the distance measurement results before and after the aggregation of each category correction is constructed as the unknown class distance measurement d unknown , and the calculation process is shown in formula (11).

[0055]

[0056] The corrected distance of all known classes and the unknown class distance are input into the softmax function together, and are uniformly converted into a probability distribution, as shown in formula (12):

[0057]

[0058] Wherein, d c represents the distance between the input sample and the category c; when c = C + 1, it represents the probability distribution result of the "unknown class";

[0059] According to formula (13), if the unknown class probability is the smallest, the sample is determined as an unknown category; otherwise, it is classified into the known class with the smallest probability, realizing joint recognition and rejection.

[0060]

[0061] Beneficial effects

[0062] A radar signal modulation mode open set recognition method based on multi-modal fusion is provided, which fuses time-frequency images and text features, classifies and discriminates through metric learning, introduces an extreme value theory to establish a rejection mechanism, and improves the recognition ability of unknown class signals.

[0063] Specifically, the present application combines actual radar signal parameters to construct a simulation data set, and generates discriminative feature text description for the time-frequency distribution graph of each type of signal to form a picture-text paired sample. In the open set recognition scene, the influence of single modal input and picture-text multi-modal fusion input on recognition performance is compared and evaluated respectively, and the embedding feature discrimination ability of the method in the metric learning framework is further verified. The results show that the method has higher acceptance rate, and is superior to the traditional method in unknown class rejection rate, which significantly improves the open set recognition performance of radar signal modulation mode. BRIEF DESCRIPTION OF DRAWINGS

[0064] In order to clearly and unambiguously explain the technical steps of the present application, all the drawings used in the present application will be described simply below. It should be noted that the drawings described below are only some examples of the implementation of the present application, and other ordinary skilled persons in the art can still obtain other drawings in other different scenarios based on these drawings.

[0065] Figure 1 is an embodiment flow of the present application; Figure 1 is an embodiment flow of the present application;

[0066] Figure 2 is a radar signal sample parameter setting scheme provided by the present application; Figure 2 is a radar signal sample parameter setting scheme provided by the present application;

[0067] Figure 3 is a radar signal time-frequency distribution graph feature description text generation scheme provided by the present application; Figure 3 is a radar signal time-frequency distribution graph feature description text generation scheme provided by the present application;

[0068] Figure 4 is a radar signal multi-modal feature fusion process framework provided by the present application; Figure 4 is a radar signal multi-modal feature fusion process framework provided by the present application;

[0069] Figure 5 is a radar signal modulation mode open set identification strategy execution process framework provided by the present application; Figure 5 is a radar signal modulation mode open set identification strategy execution process framework provided by the present application;

[0070] Figure 6 is a confusion matrix of a multi-modal and single-modal model framework under different signal-to-noise ratios provided by the present application; Figure 6 is a confusion matrix of a multi-modal and single-modal model framework under different signal-to-noise ratios provided by the present application;

[0071] Figure 7 is a performance comparison graph of a multi-modal and single-modal model framework under an unknown class expansion scenario provided by the present application; Figure 7 is a performance comparison graph of a multi-modal and single-modal model framework under an unknown class expansion scenario provided by the present application; DETAILED DESCRIPTION

[0072] The steps and processes of the present application will be described completely and clearly below in conjunction with the drawings in the present application. It is obvious that the examples described in the present application are only one example application scenario of the present application, and other results based on the content of the present application without substantial changes are all within the protection scope of the present application.

[0073] Before implementing the method of the present application, it is usually necessary to configure the corresponding multi-modal data set, model structure and training strategy according to the specific application scenario, and to select reasonable evaluation indicators. Specifically, according to the characteristics of the radar signal modulation recognition task, paired samples containing image modalities and text modalities are constructed, and image encoders and text encoders suitable for the task are selected to extract and fuse features of the two modalities. In addition, the corresponding optimization algorithm, hyperparameter configuration and running platform need to be set to support the stable operation of the model in the training and inference process.

[0074] To further illustrate the technical solution of the present invention, the following, combined with specific examples, illustrates how to configure a multimodal dataset, model structure, training parameters, evaluation metrics, and the complete execution process of this method in the task of open-set recognition of radar signal modulation modes in practical applications.

[0075] Construct a picture-text pairing dataset for the modulated signal. Figure 2 and attached Figure 3 As shown in Figure 2, this paper designs a sample space containing 11 types of modulation methods, including BPSK, Costas, LFM, P1-P4, and T1-T4. During the data set construction process, different parameter combinations are set for each modulation method. The parameters include carrier frequency, symbol period, frequency hopping sequence length, bandwidth, number of step frequencies, etc. The specific parameters and value ranges are shown in the attached figure. Figure 2 Each signal sample is generated at a 150MHz sampling rate, and the signal-to-noise ratio range covers from -6dB to 10dB in 2dB intervals.

[0076] For each modulated signal sample, in addition to generating its corresponding time-frequency diagram, it is also necessary to construct a text description of the discriminative features between the sample and other known category samples, as shown in the attached Figure 3 As shown. This embodiment specifically uses the ChatGPT-4o large language model to conduct interactive question and answer to obtain the corresponding text description of the sample time-frequency distribution diagram. Taking into account that the time-frequency images of radar signals of the same category have similar visual features under the same signal-to-noise ratio environment, there is no need to generate text descriptions for all samples one by one. In order to improve the efficiency of the text generation process and save computing costs, in this embodiment, only 20% of the samples from each category of samples under each signal-to-noise ratio condition are randomly selected to perform the above three rounds of question-and-answer text description generation. Subsequently, through the different expressions given by ChatGPT-4o in question three, the generated text description is extended to the other 80% of samples of the same category under the signal-to-noise ratio environment, thereby efficiently completing the construction of the entire image and text dataset.

[0077] The training dataset consisted of 900 samples for each known modulation type; the test set contained 180 samples for both known and unknown types. During the testing phase, various scenarios with varying numbers of unknown types were tested. For a test with five unknown types, the known types were BPSK, Costas, and T1 to T4. For a test with four unknown types, LFM was added to the known types. For a test with three unknown types, P2 was added to the known types. For a test with two unknown types, P3 was added to the known types.

[0078] In this embodiment, the image modality and the text modality use independent coding structures for feature extraction, as shown in the attached figure. Figure 4The image modality input is processed by a visual encoder (ViT-B / 16) in the CLIP model; the text modality input is processed by a corresponding text encoder in the CLIP model. The above-mentioned image and text modality encoders are used for input feature extraction, and the extracted features are fused by a text-image interaction module to generate a text-image joint embedding representation for subsequent training.

[0079] During the training process, the Adam optimizer is used, the initial learning rate is 5e-6, the weight decay is set to 0.2, the β1 and β2 parameters are 0.9 and 0.98 respectively, and ε is 1e-5. In order to improve the stability and convergence effect of the model, the StepLR strategy is adopted, and the learning rate is multiplied by 0.1 every 10 epochs. The model is trained for a total of 50 epochs, and the training process is completed on an NVIDIA 3090 GPU platform.

[0080] To comprehensively evaluate the recognition effect of the method in the open set recognition task of radar modulation mode, two key performance indicators are set: (1) Acceptance Rate, which refers to the correct recognition rate of known class samples by the system, and is used to measure the recognition ability of the model for known signals; (2) Rejection Rate, which refers to the rejection ability of the system for unknown class samples, that is, the proportion of correctly identifying unknown samples as "unknown class". In order to ensure accurate recognition of known classes and effective rejection of unknown signals, the open set recognition task needs to achieve excellent performance in both indicators.

[0081] After completing the above data construction, model configuration, training parameter setting and evaluation index determination, the specific execution process of the method in the open set recognition task of radar signal modulation mode is described in steps as follows:

[0082] The specific steps of the radar signal modulation mode open set recognition method based on multi-modal fusion are as follows:

[0083] Step 300: Constructing a text-image multi-modal data set

[0084] First, according to the parameter settings in the accompanying drawings Figure 2 , generate radar signal samples including BPSK, Costas, LFM, P1-P4 and T1-T4, a total of 11 types of modulation modes, and calculate the time-frequency distribution of each signal matrix form by CWD. The distribution results are mapped into image form, and the resolution is unified to 224x224.

[0085] Interact with the ChatGPT-4o model through question and answer, and gradually extract the signal features and related domain knowledge contained in the image. The specific interaction question and answer process is shown in the accompanying drawings Figure 3 , including the following steps:

[0086] (1) Different categories of multiple time-frequency distribution images under the same signal-to-noise ratio condition are input to ChatGPT-4o model, and the first round of questions is proposed:

[0087] "Please describe the visual features of these time-frequency distribution images of different categories in combination with the knowledge of radar signals."

[0088] ChatGPT-4o combines the visual features of the input images with the built-in domain knowledge to automatically generate preliminary visual feature descriptions for this category of signals, such as spectral concentration, linear distribution, and energy distribution uniformity.

[0089] (2) Further, for the significant features of a specific category, the second round of questions is proposed to ChatGPT-4o:

[0090] "Please describe the discriminative features of this category compared to other categories."

[0091] ChatGPT-4o combines the features of the current category with those of other categories to automatically generate more targeted and in-depth discriminative feature descriptions. For example, it emphasizes the linear variation of signal trajectories, time-domain stability, or the difference in spectral energy concentration.

[0092] (3) Finally, to ensure that the text description has better robustness and diversity, the third round of questions is proposed to ChatGPT-4o:

[0093] "Please present the discriminative features of the above-mentioned category in different ways."

[0094] ChatGPT-4o again generates descriptions that are semantically consistent with the previous descriptions but have different expression forms, further refining and enhancing the category features from different angles to form a hierarchical text description result.

[0095] Step 310: Image-text encoding and feature fusion and embedding optimization

[0096] This step is based on the aforementioned image-text encoding structure and interaction module, and completes the feature fusion of image and text modalities to generate a joint image-text embedding representation, as shown in Fig. Figure 4 .

[0097] The specific parameter settings are as follows: batch size B = 64, feature dimension d = 512, attention head number h = 8, and each head dimension d k = 64. The expression form of the fused features is shown in equation (14):

[0098]

[0099] During training, the triplet loss function is used to optimize the distribution structure of the joint embedding representation feature, as shown in formula (15):

[0100]

[0101] Where α is the boundary interval, which is set to 0.7.

[0102] Step 320: Calculate category embedding center

[0103] After completing the optimization of the distribution structure of the joint embedding representation feature of the training samples, for the samples of known categories in the training set, the joint embedding representation feature of the image and text of this type of samples is aggregated to calculate the category embedding center feature center , as shown in formula (6). The obtained center will be used for subsequent distance measurement.

[0104] Step 330: Class distance metric

[0105] Calculate the joint embedding representation feature of each sample to the central feature of all categories center The Euclidean distance of is shown in formula (7).

[0106] Step 340: Tail data distribution fitting

[0107] The distance measurement results of the 20 samples farthest from the class center are selected from the training samples of each category, and the Weibull distribution is fitted to obtain the Weibull distribution parameter τ c ,λ c ,η c , construct a category-specific rejection CDF function:

[0108]

[0109] Step 350: Correction and Rejection Decision

[0110] Input the test set data into the model and execute step 330 again to obtain the distance measurement result of the test sample. Figure 5 As shown, the distance measurement results of each sample with the category embedding center of all known classes are sorted from large to small. Let the sorted index be j, then combined with the category-specific Weibull distribution fitting parameter τ c ,λ c ,η c , calculate the correction factor α for each category i :

[0111]

[0112] Using the correction factor α iThe original distance d of each category is corrected to obtain a corrected distance

[0113]

[0114] Further, the distance metric d of the unknown category is calculated using the original distance and the correction factor unknown :

[0115]

[0116] Next, the corrected known category distance and the unknown category distance d unknown are substituted into the softmax function to be uniformly converted into probabilities:

[0117] The known category probability is:

[0118]

[0119] The unknown category probability is:

[0120]

[0121] Finally, the following discrimination strategy is adopted:

[0122] If the unknown category probability P(y=unknown|Feature) is the smallest among all probabilities, it is determined that the sample belongs to the unknown category and is rejected; otherwise, the sample is classified into the known category with the smallest probability

[0123] Some result graphs obtained in the example scenario are explained below. For ease of explanation, the model framework proposed in the application based on multi-modal fusion is referred to as "multi-modal model", and the model framework using only the image modality is referred to as "single-modal model".

[0124] Attached Figure 6 shows the confusion matrix results of the multi-modal (image and text) model and the single-modal (image) model in the open set recognition task under different signal-to-noise ratios. The figure includes 8 subgraphs, and the signal-to-noise ratio conditions from left to right in the column direction are -6dB, -4dB, -2dB and 0dB, respectively. The upper row is the multi-modal (image and text) model result, and the lower row is the single-modal (image) model result. In each subgraph, the row represents the true category, and the column represents the model predicted category, including 6 known modulation modes (BPSK, Costas and T1-T4) and one "unknown class" (including LFM and P1-P4 signals).

[0125] Results show that at a -6dB signal-to-noise ratio (SNR) condition, the multimodal model achieved an 87.7% rejection rate for unknown classes, surpassing the 85.8% of the unimodal model. Although the unimodal model performed slightly better in recognizing some known classes, the multimodal model significantly reduced the number of unknown samples misclassified as T2 and T3, demonstrating stronger discriminative robustness. As the SNR increased to -4dB, the gap between the multimodal and unimodal models in recognizing known classes narrowed further, with both remaining above 98%. The multimodal model's 88.8% rejection rate for unknown classes still outperformed the unimodal model's 87.1%. At -2dB, both modal models achieved recognition rates exceeding 99.3% for known classes, with only one misclassified sample in individual classes, such as Costas and T4. The multimodal model achieved a 92.0% rejection rate for unknown classes, surpassing the 89.0% of the unimodal model, and continued to demonstrate its superiority in suppressing misclassifications of the T2 class. When the signal level drops to 0dB, except for two cases in the BPSK category that were misclassified as unknown, the remaining known categories all achieved 100% classification accuracy. The multimodal model's rejection rate for unknown categories was 90.8%, higher than the 88.8% of the single-modal model. Overall, from -6dB to 0dB, the multimodal model consistently outperformed the single-modal model in its ability to reject unknown categories, especially in low signal-to-noise ratio environments, demonstrating higher robustness and stability. By introducing text information, the multimodal model significantly reduced the tendency of T2 and T3 boundary fuzzy categories to attract unknown categories. While achieving a high recognition rate for known categories, it effectively improved the model's open-set discrimination ability, verifying the advanced nature of the open-set identification method for radar signal modulation based on multimodal fusion proposed in this invention.

[0126] Attachment Figure 7 The rejection rate and acceptance rate performance of the multimodal model and the single-modal model in the open set recognition of radar signal modulation modes are demonstrated under the condition of -6dB signal-to-noise ratio as the number of unknown modulation category signals gradually expands from 2 to 5.

[0127] In terms of rejecting signals of unknown modulation categories, the multimodal model combined with triplet loss achieved rejection rates of 0.8778, 0.9574, 0.9417, and 0.8767, respectively, maintaining an overall high level. Furthermore, it outperformed the single-modal model in all settings with the number of unknown categories (2 to 5), with corresponding rejection rates of 0.8361, 0.95, 0.9028, and 0.8578, respectively. This demonstrates that the multimodal model possesses stronger rejection capabilities and robust discrimination when faced with multi-category unknown interference, enabling it to more effectively identify and reject interference signals from non-target categories.

[0128] In terms of known class identification capability, the acceptance rates of the multi-modal model under four unknown class quantity settings are 0.9284, 0.9285, 0.9429 and 0.9556 respectively; the corresponding acceptance rates of the single-modal model are 0.9556, 0.9604, 0.9468 and 0.9676, which are slightly higher than those of the multi-modal model, indicating that the single-modal model has certain advantages in known class identification. However, the acceptance rates of the two models under each setting remain at a high level with a small fluctuation range, indicating that the multi-modal method proposed in the application can ensure strong unknown class rejection capability while still maintaining high and stable identification performance for known classes, and has good practicability.

[0129] Overall, the multi-modal model maintains a high acceptance rate while having excellent unknown class rejection capability, and achieves a good performance balance in the open set identification task of radar signal modulation mode.

Claims

1. A radar signal modulation mode open set identification method based on multimodal fusion, characterized in that: include: Image modal data of time-frequency distribution diagrams are generated according to different radar signal models, and corresponding text descriptions are generated using a large language model to construct an image-text modality pairing dataset. A graphic-text fusion model is constructed to achieve interactive fusion of image and text modal features and optimize the data embedding space. Based on the fused joint embedding representation, the category embedding center of the known class is calculated. For each known class, the distance between the graphic-text joint embedding representation of the sample under that class and its category embedding center is measured. Based on the above distance measurement results, tail samples are selected and a Weibull distribution model is fitted. The distance of the test samples is corrected according to the distribution model, and combined with the unknown class distance measurement results, the softmax function is uniformly input for classification and discrimination to achieve known class recognition and unknown class rejection.

2. The method for constructing an image-text modality pairing dataset according to claim 1, wherein: Based on different radar signal models, time series data of various radar signals are generated by simulation. The time series are analyzed using the Cui-Williams distribution method to obtain energy distribution results in the form of a two-dimensional matrix. The matrix is ​​then mapped into a color image to construct image modal data. The image modality data is fed into a large language model with image understanding capabilities, and text modality data is automatically generated through an interactive step-by-step question-answering process, which includes: (1) asking questions for each category of image data to obtain preliminary textual information describing the visual features of the image; (2) asking further questions to extract discriminative features between images of each category; (3) The above discriminative descriptions are restated in different ways to enhance the robustness and diversity of the language description. This allows us to obtain diverse semantic description texts corresponding to each type of image, forming a structured image-text modality pairing dataset.

3. The method for constructing a graphic-text fusion model and optimizing the embedding space according to claim 1, characterized in that: The dual-tower model architecture is composed of an image encoder and a text encoder. It extracts features from the image modality and text modality data respectively to obtain the corresponding modality feature vectors. The multimodal interaction module built based on the cross-attention mechanism receives the image modality feature vector and the text modality feature vector, calculates the attention weights between the modalities through the multi-head attention mechanism, and generates a joint image-text embedding representation after fusion, as shown in the following formula: Among them, Feature is the joint embedding representation of image and text; Input feature vector for text; is the image input feature vector; B is the batch size; d is the total dimension of the feature vector; h is the number of attention heads; d k = d / h is the dimension of each attention head; is the linear projection matrix of the i-th head; is the output linear mapping matrix; || represents the concatenation of the outputs of each attention head in the feature dimension; the triplet loss function is used to optimize the joint embedding representation of image and text (formula (5) in the manual), and the anchor sample, positive sample and negative sample triplet are constructed to minimize the Euclidean distance between the anchor sample and the positive sample, while maximizing the distance between it and the negative sample.

4. The method for computing the category embedding center of a known class according to claim 1, characterized in that After completing the optimization of the joint image-text embedding representation of the training samples, for each known category, the joint image-text embedding vectors of all training samples in that category are extracted; the embedding vectors of samples belonging to the same category are aggregated in the embedding space, and the category embedding center is obtained by using the average value calculation method.

5. The distance measurement method according to claim 1, characterized in that For each known category in the training set, the Euclidean distance between the joint image-text embedding representation of each sample in that category and the embedding center of its corresponding category is calculated.

6. The Weibull distribution fitting method according to claim 1, wherein: For each known category, based on the Euclidean distance between the joint embedding representation of the image and text of the samples in the training set and the center of its category embedding, a preset number of samples farthest from the class center are selected as tail samples; based on the tail sample set corresponding to each category, the Weibull distribution is fitted to its sample distance distribution to obtain the distribution parameters of each category, including the position parameter τ c , scale parameter λ c , and shape parameter η c ; Based on the fitted parameters, the cumulative distribution function (CDF) corresponding to each known category is constructed.

7. The correction and rejection decision method according to claim 1, characterized in that: For each test sample, calculate the original Euclidean distance between its joint image and text embedding representation and the embedding center of each known category, and calculate the distance correction factor α based on the Weibull distribution obtained by fitting the category i , the correction factor is obtained by the following formula: Among them, the distance measurement results of each sample and the category embedding center of all known categories are sorted from large to small, and the index after sorting is set to j; C is the total number of all known categories. According to the correction factor α i Correct the original distance to obtain the category-corrected distance The calculation method is: Furthermore, based on the difference between the original distance and the corrected distance, the distance index of the unknown category is calculated as follows: Compare all corrected distances to the unknown class distance d constructed by the original distance and the correction factor unknown Enter the softmax function at the same time: If the probability of the unknown class is the smallest, the sample is judged as an unknown class; otherwise, it is classified into the known class with the smallest probability, achieving joint recognition and rejection.