Radar communication radiation source identification method and system based on multi-modal alignment
By designing a domain-customized signal encoder and text encoder for collaborative training, and combining semantic and physical parameter constraints, high-precision identification of radar radiation source signals was achieved. This solved the problems of time-consuming rule base updates and low recognition accuracy in noisy environments in traditional methods, and improved the robustness and adaptability of the system.
Patent Information
- Application Number
- CN202511311184.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-15
- Publication Date
- 2025-11-14
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional radar radiation source signal identification methods are time-consuming to update the rule base when new radar signals frequently appear, have poor adaptability, cannot effectively fuse cross-modal information, and have a decreased identification accuracy in strong noise environments.
Customized signal encoders and text encoders are trained together in the design field. Through multimodal alignment technology, radar signal features and text descriptions are accurately mapped in a unified vector space. Combined with semantic information and physical parameter constraints, the robustness and accuracy of recognition are improved.
Achieving high-precision radiation source identification in low signal-to-noise ratio environments improves system development efficiency and scalability, and supports multi-task application scenarios.
Smart Images

Figure CN120949170A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of electronic reconnaissance and signal processing technology, specifically relating to a radar communication radiation source identification method and system based on multimodal alignment. Technical Background
[0002] In modern electronic warfare and spectrum management, rapid identification of radar radiation source signals is a core component of threat assessment and adversarial decision-making. Traditional methods often rely on manually constructed feature and rule libraries, which, while effective for identifying known signal types, have significant shortcomings in the following aspects: New radar signals frequently emerge, making manual rule library updates time-consuming and lacking adaptability; existing solutions often process radar signal features and textual descriptions separately, leading to ineffective fusion of cross-modal information and loss of important semantic relationships; general multimodal models (such as CLIP) perform well on natural images, but their accuracy drops significantly on specialized signal images such as radar time-frequency maps, mainly due to the significant differences between signal patterns and natural images; in noisy environments, traditional feature extraction methods struggle to reliably extract key features, resulting in a significant decrease in recognition accuracy.
[0003] Therefore, it is necessary to propose a multimodal alignment technique that combines signal features and semantic information, and to design a dedicated signal encoder for the characteristics of radar time-frequency maps, so as to achieve high-precision radiation source identification even in complex environments such as low signal-to-noise ratio. Summary of the Invention
[0004] This invention proposes a radiation source signal identification method based on multimodal alignment, which can be applied not only to radar signals but also to feature extraction from any radiation source signal. Its key feature is that, through the collaborative training of a domain-customized signal encoder and a text encoder, precise mapping between radar signal features and text descriptions is achieved in a unified vector space, thereby significantly improving robustness and accuracy in unknown radiation source identification tasks.
[0005] The technical solution of this invention is as follows:
[0006] S1. Preprocessing the Received Signal: This system first receives the raw radar signal input and performs time-domain segmentation through windowing and framing. A Hanning or Hamming window function is preferred, with a window length N configured to 256-1024 sampling points and a frame shift L set to the range of N / 4 to N / 2. A Short-Time Fourier Transform (STFT) is performed on the segmented signal to convert the one-dimensional time series into a two-dimensional time-frequency spectrum representation. Subsequently, amplitude normalization is performed on the time-frequency spectrum, optionally linearly normalized to the [0,1] or [-1,1] interval. To enhance data diversity, Gaussian noise injection, time-domain shift (±5%), and frequency-domain offset (±2%) data augmentation strategies are introduced. Finally, the processed time-frequency spectrum is uniformly scaled to a resolution of 128×128 pixels as the standardized input to the neural network in S2.
[0007] S2. Signal Encoder Construction: To address the characteristics of radar time-frequency maps (local time-frequency patterns, hardware nonlinear features), a hierarchical feature extraction network is designed as the backbone. The backbone consists of shallow and deep feature extraction layers. The shallow layer comprises two sets of 3×3 convolutional kernels (16 and 32 channels respectively) for extracting signal edge features and local time-frequency patterns. The deep layer contains three sets of 3×3 convolutional kernels (channel numbers increasing to 64, 128, and 256), responsible for capturing high-order spectral features. A channel attention mechanism (SE module) is embedded after the deep convolutions to enhance key frequency bands of the hardware fingerprint through feature reweighting. The network ends with a global average pooling layer and a fully connected layer, outputting a 256-dimensional signal embedding vector. The specific design process is as follows... Figure 2 As shown, during the training phase, the triplet loss and cross-entropy loss functions are jointly optimized. The triplet loss constrains the clustering of similar samples in the embedding space, while the cross-entropy loss improves the clarity of the classification boundary.
[0008] S3. Multimodal Vector Alignment: Based on the principle of multimodal vector alignment, the signal embedding vector and the text encoding vector are trained to align them in the same vector space. There are two encoders: a signal encoder and a text encoder. The samples of mode A in the signal image are encoded into vectors, and the radar radiation source signal extracted by the CNN in S2 is embedded into the vectors; text encoder. The samples of text modality B are encoded as vectors, and the text is encoded using a Transformer. Each training batch contains N sets of strictly paired signal-text samples (such as time-frequency graphs and their technical document fragments), while generating N×(N-1) negative sample pairs. Based on the trained signal encoder... and text encoder To build a recognition knowledge base, follow these steps: collect technical documents and corresponding signal samples of known radar signals, and use... The text description is encoded into a 256-dimensional vector; a standardized text description containing key parameters is extracted from the document, using... Encode the signal samples into 256-dimensional vectors; store the vectors along with the corresponding signal types and parameters in the database; when adding a new signal type, simply encode its text description and add it to the database.
[0009] Physically constrained contrastive loss is used: The cosine similarity of positive sample pairs is maximized, while the similarity of negative sample pairs is minimized. Predefined parameter parsing rules (such as regular expressions for pulse width and repetition frequency extraction) are used to force the mean square error between the deconvolution output of the signal vector and the physical parameters of the text description (such as a 1μs pulse width) to be minimized. Cosine alignment loss (semantic level) and spectral constraint loss (physical level) are simultaneously optimized, and gradient descent is used to make both types of losses converge. After training, the time-frequency maps of similar radars and their text descriptions have high similarity in 256-dimensional space, and the vector space distance reflects the differences in actual physical parameters. The pre-trained dual encoder is frozen, supporting tasks such as zero-shot signal classification and cross-modal retrieval (signal-text bidirectional search), achieving unified mapping of semantic-physical features between modalities without additional fine-tuning.
[0010] S4. Unknown Radar Source Signal Identification: For the unknown radar signal to be identified, perform the following operations in sequence: 1) Generate and normalize the time-frequency diagram according to the S1 module process; 2) Extract a 256-dimensional signal embedding vector through the encoder network of the S2 module; 3) Calculate the cosine similarity between this vector and all text description vectors in the knowledge base; 4) Select the text description with the highest similarity as the identification result output, and add a confidence score. The system supports dynamic expansion of the knowledge base; when adding a new signal type, only the corresponding text description needs to be added and the encoder parameters updated.
[0011] Compared to traditional multimodal reasoning tasks for radar signal radiation sources, this signal identification reasoning method adds the following elements:
[0012] First, addressing the issue of poor adaptability of visual encoders (ViT) in general multimodal models (such as CLIP) to specialized signal image domains like radar time-frequency maps, this invention abandons the traditional approach of directly using pre-trained models from natural images. Instead, it innovatively designs a dedicated hierarchical convolutional encoder tailored to the time-frequency characteristics of radar. This encoder enhances the capture of local time-frequency patterns and signal edge features through shallow convolutional modules and embeds a channel attention mechanism (SE module) in the deep network to achieve adaptive enhancement of key frequency bands of hardware fingerprints, significantly improving the recognition and representation capabilities of signal features.
[0013] Secondly, this invention overcomes the limitation of traditional multimodal alignment relying solely on semantic similarity by innovatively introducing a physical parameter constraint mechanism. By constructing a joint loss function, the cosine alignment loss between the signal vector and the text vector at the semantic level is applied. Compared with spectral constraint loss at the physical parameter level This combination forces the model to maintain both semantic relevance and physical consistency in the latent space, thereby significantly improving the reliability and robustness of recognition results in low signal-to-noise ratio environments.
[0014] Third, this invention employs an end-to-end joint training framework, achieving fully automated modeling from raw signals to recognition results. This avoids the cumbersome process of manually designing features and rule bases in traditional methods, significantly improving system development efficiency and scalability. The framework adopts a modular design, allowing for adaptation to radiation source identification tasks in different fields such as communication and electronic countermeasures by changing the encoder structure, demonstrating strong versatility.
[0015] Finally, the 256-dimensional unified feature representation generated by this invention possesses both good interpretability and compatibility with downstream tasks, directly supporting various application scenarios such as threat level assessment, fine-grained signal classification, and cross-modal retrieval. Through a feature sharing mechanism, it effectively improves computational efficiency and resource utilization in multi-task systems. Attached Figure Description
[0016] Figure 1 This is a flowchart of acquiring radar signals and performing preprocessing;
[0017] Figure 2 It is a signal encoder Build the method architecture diagram;
[0018] Figure 3 This is a diagram of the signal-text multimodal alignment training architecture;
[0019] Figure 4 This is a schematic diagram of the unknown signal identification process. Detailed Implementation
[0020] The present invention will now be described in detail with reference to the accompanying drawings:
[0021] This invention is divided into three parts: data preprocessing, signal feature recognition processing, multimodal knowledge graph construction, and intelligent multimodal question answering system construction. The specific steps are S1-4:
[0022] S1. Preprocessing and feature extraction of the received signal: In radar source identification tasks, the received signal is often affected by additive noise, multipath effects, and other interferences, especially under low signal-to-noise ratio (SNR) conditions. Effective preprocessing is required to improve the robustness of feature extraction. First, the original signal... Perform windowing and frame splitting processing: .in Here, L is the frame shift and N is the window length. Radar signals are typically one-dimensional time series, but they can be converted into a two-dimensional image matrix to better utilize the spatial feature extraction capabilities of CNNs. A time-spectrum matrix is generated through Short Time Fourier Transform (STFT). The signal is normalized, with the amplitude normalized to [0,1] or [-1,1]. Due to the limited radar signal data, data augmentation techniques can be used to add noise, time shift, and frequency shift. Finally, the time-frequency map is normalized to 128×128 resolution using bilinear interpolation, forming the input matrix for the CNN. The detailed steps are shown in the flowchart below. Figure 1 As shown, the processed time-frequency image is ready for CNN feature extraction.
[0023] S2. Signal Encoder Construction: To address the characteristics of radar time-frequency maps (local time-frequency patterns, hardware nonlinear features), a hierarchical feature extraction network is designed as the backbone. The backbone network consists of shallow and deep feature extraction layers. The shallow layers (Conv1-2) use 3×3 convolutions (16, 32 channels), followed by ReLU, BatchNorm, and pooling to capture low-level features such as edges. The deep layers (Conv3-5) use 3×3 convolutions (64, 128, 256 channels); important frequency bands are weighted after Conv3-5 using a channel attention mechanism. First, global average pooling is applied to the feature map to obtain channel descriptions, followed by fully connected layers and Sigmoid activation. Generate weights and multiply them channel by channel with the original feature map. , This enhances the sensitivity to key areas of hardware fingerprints. High-order features are extracted using global average pooling, and the output layer is mapped to a 256-dimensional feature vector. Triple loss is used for joint training to improve robustness. The specific design process is as follows: Figure 2 As shown.
[0024] S3. Multimodal Vector Alignment: Based on the principle of multimodal vector alignment, a loss function incorporating physical constraints is used to train the signal embedding vector and the text encoding vector, aligning them in the same vector space to ensure that semantically related signals and text vectors are close to each other. This invention has two encoders:
[0025] signal encoder The samples of signal image mode A are encoded into vectors. Using the radar signal embedding vector extracted by the CNN in S2, a 256-dimensional embedding vector is generated, denoted as... ,in Let i be the i-th radar signal sample.
[0026] Text encoder Encode the samples of text modality B into vectors, and encode the text using an existing Transformer, denoted as . ,in To and The corresponding text description.
[0027] Collect technical documents for known radar models, extract standardized descriptive text, and ensure each description includes key parameters such as signal type, operating frequency band, pulse width, and repetition frequency; then use a trained text encoder. All standardized text descriptions are encoded into 256-dimensional text embedding vectors; real signal samples from each known model are collected and a signal encoder is used. Convert the signal into a signal embedding vector and store it along with the corresponding text vector; create a structured database table with fields including: signal model ID, original text description, text embedding vector (stored in the vector database), physical parameters (JSON), and signal embedding vector; design a knowledge base update interface so that when a new radar model is added, its text description only needs to be updated... The data is encoded into vectors and inserted into the database to enable dynamic expansion of the knowledge base.
[0028] Subsequently, contrastive loss was used to train the two encoders. During training, the encoder parameters were adjusted based on the contrastive loss by calculating the similarity between sample pairs. End-to-end training was employed, and the signal encoder was updated simultaneously. and text encoder Parameters, such as Figure 3 As shown. In each iteration, all values within the batch are calculated. and The cosine similarity matrix is used as input for loss calculation. The joint loss function is then used to calculate the loss:
[0029] Cosine similarity contrast loss (InfoNCE form):
[0030]
[0031] in: The dot product similarity represents the similarity between positive sample pairs (matched signal-text pairs). This represents the dot product similarity of negative sample pairs (mismatched pairs). The temperature hyperparameter controls the sharpness of the distribution; N is the batch size.
[0032] Physical constraint loss:
[0033]
[0034] in This represents a function that decodes the k-th physical parameter from a signal vector (e.g., reconstructing the pulse width using a deconvolutional network). A function that parses the k-th physical parameter from a text vector (e.g., extracting "1μs" using a regular expression); K represents the total number of physical parameters (e.g., pulse width, repetition frequency, main lobe width, etc.).
[0035] Joint loss function;
[0036]
[0037] in , To balance the hyperparameters (suggested initial value of 0.5), control the relative importance of the two losses, and continuously optimize the vector distribution, the semantically related radar signal vectors and text vectors are finally precisely aligned in the same space and stored in the training set database, providing a reliable vector matching foundation for the subsequent identification of unknown signals.
[0038] S4. Identification of Unknown Radar Source Signals: For unknown radar signals... Through the constructed signal encoder Extract its embedding vector The similarity between this vector and the aligned text encoding vector is calculated, and the text with the highest similarity is selected as the recognition result. Figure 4 The flowchart illustrates the process of identifying the primary signal. This system is applied to signal identification tasks in the field of electronic reconnaissance. When an unknown radiation source radar signal is input, the system provides a description of the most similar signal features and gives the probability of similarity.
Claims
1. A radar communication radiation source identification method based on multimodal alignment, characterized in that, Includes the following steps: S1. Signal preprocessing: Receive the raw radar signal, perform windowing, framing, and short-time Fourier transform on it to generate a time-frequency spectrum, and perform normalization and resolution standardization on the time-frequency spectrum to generate a standardized time-frequency image. S2, Signal Feature Extraction: The standardized time-frequency image is input into the signal encoder. signal encoder A hierarchical convolutional neural network with embedded channel attention mechanism is used to extract and output a fixed-dimensional signal embedding vector; S3. Multimodal Alignment Training: Building a Text Encoder This is used to encode text descriptions into text embedding vectors; the signal encoder is trained synchronously using a joint loss function. and text encoder This aligns the paired signal embedding vectors with the text embedding vectors in a unified vector space. The joint loss function includes a contrastive loss based on cosine similarity. Physical constraint loss based on physical parameter analysis This aligns the radar signal embedding vector with its corresponding text description vector in the same vector space, and builds a recognition knowledge base based on the trained encoder. S4. For the unknown radar signal to be identified, its signal embedding vector is obtained after processing by S1 and S2. The similarity between the obtained signal embedding vector and the text embedding vector in the knowledge base is calculated, and the text description with the highest similarity is output as the recognition result.
2. The radar communication radiation source identification method based on multimodal alignment according to claim 1, characterized in that, The windowing and framing process in S1 uses a Hanning window or a Hamming window as the window function, with a window length N ranging from 256 to 1024 sampling points and a frame shift L ranging from 1 / 4 to 1 / 2 of the window length N.
3. The radar communication radiation source identification method based on multimodal alignment according to claim 1, characterized in that, S1 also includes data augmentation, which can be achieved by injecting Gaussian noise into the time-frequency spectrum, performing time-domain shifting, and frequency-domain offsetting, or by one or more combinations thereof.
4. The radar communication radiation source identification method based on multimodal alignment according to claim 1, characterized in that, Constructing a signal encoder The hierarchical convolutional neural network includes: a shallow feature extraction module, consisting of at least two sets of convolutional layers with a kernel size of 3×3, used to extract edge features and local time-frequency patterns of the signal; a deep feature extraction module, consisting of at least three sets of convolutional layers with a kernel size of 3×3, and the number of channels increases layer by layer, used to capture high-order spectral features; and a channel attention module, embedded in the deep feature extraction module, used to reweight the channels of the feature map to enhance feature extraction of key frequency bands of the hardware fingerprint.
5. The radar communication radiation source identification method based on multimodal alignment according to claim 1, characterized in that, The physical constraint loss The physical parameter values are calculated as follows: the physical parameter values are decoded from the signal embedding vector and parsed from the corresponding text description, and the mean square error between the two is calculated; the physical parameters include one or more of pulse width, repetition frequency, and main lobe width.
6. The radar communication radiation source identification method based on multimodal alignment according to claim 1, characterized in that, The joint loss function , represented as: ,in and This is a hyperparameter used to balance the weights of the two loss parameters.
7. A radar communication radiation source identification system based on multimodal alignment, used to execute a radar communication radiation source identification method based on multimodal alignment as described in any one of claims 1 to 6, characterized in that, include: The preprocessing module is configured to perform step S1; The signal encoding module has the signal encoder built-in. It is configured to execute step S2; The text encoding and alignment module has the text encoder built-in. It also stores a pre-trained knowledge base, is configured to execute step S3 and provide text embedding vectors during inference; The recognition reasoning module is configured to execute step S4 and output the recognition result and confidence level.
8. A radar communication radiation source identification system based on multimodal alignment according to claim 7, characterized in that, The knowledge base supports dynamic expansion. When a new radar signal type is added, only the corresponding text description needs to be added and the text encoder parameters need to be updated. There is no need to retrain the entire system.
9. A computer-readable storage medium having a computer program stored thereon, the program being executed by a processor to implement the method as claimed in any one of claims 1 to 7.
Citation Information
Cited By
HRRP large model multi-scene identification method based on double-domain expert knowledge
CN121388821A
Open set specific radiation source identification method and system based on multi-granularity embedding guidance
CN121412777A
A Method and System for Identifying Specific Radiation Sources Based on Multi-Granularity Embedding
CN121412777B