Similarity self-adaption-based semi-supervised open world radiation source individual identification method

Through a semi-supervised method based on similarity adaptation, using feature extractors and similarity loss transformation clustering problems, combined with pseudo-labels and entropy regularization, the problem of identifying new categories of radiation sources in open world scenarios is solved, and high accuracy and stable radiation source individual recognition is achieved.

CN120561673APending Publication Date: 2025-08-29ARMY ENG UNIV OF PLA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510594608.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2025-08-29

AI Technical Summary

Technical Problem

The existing individual radiation source identification methods cannot effectively identify and utilize new categories of radiation sources in open world scenarios, and rely on manual intervention, lack of automation and flexibility, resulting in poor accuracy and insufficient stability.

Method used

A semi-supervised method based on similarity adaptation is adopted, and the clustering problem of feature extractors and similarity loss transformation is combined with pseudo-labels and entropy regularization to achieve accurate identification of known and new radiation sources, improving the generalization and stability of the model in open world scenarios.

Benefits of technology

Effectively utilize limited labeled data and a large amount of unlabeled data to accurately identify known and new radiation sources, improve the recognition accuracy and adaptability of the model in complex electromagnetic environments, and reduce dependence on labeled data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120561673A_ABST
    Figure CN120561673A_ABST
Patent Text Reader

Abstract

The invention discloses a semi-supervised open world radiation source individual identification method based on similarity self-adaption, and aims to solve the problems that unmarked data cannot be effectively utilized and unknown classes cannot be subdivided in an open world scene in existing open set and semi-supervised radiation source individual identification. The method comprises the following steps: acquiring a data set containing marked samples and unmarked samples, and extracting sample features by using a feature extractor; clustering is converted into sample pair similarity prediction through pairwise similarity loss, and a new category is identified; identifying a known class by using self-adaptive cross entropy loss, and balancing the learning speeds of the new class and the known class; introducing an entropy regularization item to prevent model overfitting; and through minimizing an overall objective function training model formed by pairwise similarity loss, adaptive cross entropy loss and entropy regularization, identification of known and new radiation sources is realized. According to the method, finite marked data and a large amount of unmarked data are effectively utilized, known classes and new classes can be accurately identified, and the adaptability to new class radiation sources is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of radiation source identification, and in particular to a semi-supervised open-world radiation source individual identification method based on similarity adaptation. Background Art

[0002] Emitter identification is a physical layer authentication technology that identifies wireless devices by extracting the radio frequency fingerprint of received signals. Open-set radiator identification involves identifying and rejecting unknown class samples while classifying known class samples. This approach typically relies on sufficient training samples of known classes. However, in open-world scenarios, labeled samples are limited, and unlabeled samples may contain unknown classes. Furthermore, open-world identification requires not only detecting unknown class samples but also identifying specific new classes from these unknown class samples and accumulating them into the recognition model. Existing open-set radiator identification methods typically only classify all unknown class samples into one category and are unable to further subdivide the unknown class.

[0003] To address this challenge, the technical solutions used in existing technologies are mainly a two-step approach of "open set identification followed by clustering", that is, the model first distinguishes known classes from unknown classes through open set identification, and then uses a clustering algorithm on the unknown class data to identify specific new classes. However, the "open set identification followed by clustering" method has problems such as poor stability and insufficient adaptability when processing known and new class data. First, open set identification and clustering are independent steps, which may lead to information fragmentation and error propagation, affecting the clustering results. Second, the effectiveness of clustering depends on the accuracy of open set identification, and traditional clustering algorithms are sensitive to initial conditions, which may lead to unstable results. In addition, this method also relies on manual intervention, such as threshold setting and parameter adjustment, and lacks automation and flexibility.

[0004] This shows that when faced with new types of radiation sources and insufficient labeled samples, individual radiation source identification faces challenges such as poor accuracy and an inability to effectively identify and utilize new types. Therefore, it is necessary to explore more flexible and effective identification methods to improve the model's ability to adapt to new types of radiation sources. Summary of the Invention

[0005] This application provides a semi-supervised open-world radiation source individual identification method based on Similarity-Adaptive (SAA), which can be used to solve the technical problem of overcoming the shortcomings of existing open-set and semi-supervised radiation source individual identification methods. This application effectively utilizes limited labeled data and a large amount of unlabeled data to accurately identify known and new types of radiation sources, improve the generalization and stability of the model in open-world scenarios, and enhance its adaptability to new categories of radiation sources.

[0006] This application provides a semi-supervised open-world radiation source individual identification method based on similarity adaptation, the method comprising the following steps:

[0007] Step 1: Get the labeled sample set D l and the unlabeled sample set D u Data set D, where D l Contains N l Labeled samples x i , corresponding to the known class y i ∈C l , D u Contains N u unlabeled samples, true labels y u From the unknown class label set Y u , and D u Contains D l New classes that do not exist in ;

[0008] Step 2: Input the labeled samples and unlabeled samples into the feature extractor to obtain the features of the input samples; the feature extractor includes a cascade of convolutional layers, pooling layers, and residual units;

[0009] Step 3: Use pairwise similarity loss to transform the clustering problem into a sample pair similarity prediction problem. Use labeled samples and unlabeled samples to train the feature extractor and similarity predictor to determine the sample pair similarity. Calculate the pairwise similarity loss through binary cross entropy loss. optimization Identify new categories;

[0010] Step 4: Use adaptive cross entropy loss to identify known classes and balance the learning speed between new classes and known classes. Generate pseudo labels based on sample similarity to calculate cross entropy loss. Introduce uncertainty adaptive boundary mechanism to adjust classification boundaries by estimating the uncertainty of the model, determine the uncertainty of unlabeled dataset samples and obtain overall uncertainty estimation. Then calculate the adaptive cross entropy loss L ace ;

[0011] Step 5: Introduce entropy regularization term Apply entropy regularization to the mean of the output probability of the entire batch of samples to avoid overfitting of the model;

[0012] Step 6: By minimizing the overall objective function The model is trained and used to identify radiation source samples to achieve accurate classification of known and new types of radiation sources.

[0013] Furthermore, the operation process of the feature extractor is as follows:

[0014] The network first receives the original I and Q signals through the input layer, and initially extracts features through a one-dimensional convolution layer using 64 convolution kernels of size 7. The feature map size is then reduced by a maximum pooling layer. The feature extraction capability is then further enhanced by two residual units that each contain two one-dimensional convolution layers using 64 convolution kernels of size 7. The residual unit alleviates the gradient vanishing problem in deep network training through jump connections. Finally, the average pooling layer performs feature integration to generate signal features. Each convolution layer is followed by a batch normalization layer, and the rectified linear unit (ReLU) is used as the activation function.

[0015] Furthermore, when calculating the pairwise similarity loss, for the labeled dataset, similar instance pairs are identified based on the true labels. For the unlabeled dataset, the cosine distance of the sample pairs in the feature space is calculated to evaluate the similarity. For each unlabeled dataset, the corresponding sample is assigned a pseudo label with the most similar neighbor to itself. For each small batch of feature representations, the labeled sample Z is determined. l and unlabeled samples Z u The set of closest neighbors Z l ′ and Z u ′, calculated using binary cross entropy loss For any pair of samples (z i ,z j ), Determined by:

[0016]

[0017] in, Represents pairwise similarity loss, which is used to measure the accuracy of the model's similarity prediction for paired samples; m represents D l The number of labeled samples in . n represents D u The number of unlabeled samples in Φ. T represents the transpose of the similarity predictor weight parameters; z i and z i ′ respectively represent the feature representation of the sample after passing through the feature extractor, where z i From the collection z i 'From the set ζ represents the softmax function, which converts the result into a probability distribution through normalization; <·,·> represents the inner product operation, which measures the similarity between two samples.

[0018] Furthermore, when calculating the adaptive cross entropy loss, pseudo labels are generated based on the similarity of sample pairs to calculate the cross entropy loss. The uncertainty adaptive boundary mechanism is introduced to set a larger boundary at the beginning of training to reduce the misclassification of new categories. As the training progresses, the boundary value is gradually reduced. By calculating the unlabeled dataset D u The uncertainty of the sample in the quantile is used to estimate the overall uncertainty Then calculate L ace , so that the model can accurately classify known categories while avoiding excessive bias towards known categories; L ace The method for determining is as follows:

[0019]

[0020] Among them, m represents the number of samples with labeled data; z i Represents the feature representation of labeled data; Z l is the set of all labeled data feature representations; Φ T represents the transpose of the predictor weight parameters; y i It is z i Corresponding sample category; y j Represents and z i Different labels for other sample features; uncertainty The strength of is controlled by the regularization parameter λ, which is set to 1; the parameter s is used to adjust the temperature of the cross entropy loss; e represents a natural constant; and log represents the natural logarithm function.

[0021] Furthermore, the entropy regularization term L reg It is to apply entropy regularization to the mean of the output probability of the entire batch of samples as follows:

[0022]

[0023] Among them, c a Represents the total number of all known classes and possible new classes; is the average output probability of each batch of samples, b represents the number of samples in each batch, Represents the output probability of the i-th sample in each batch; represents the average output probability vector The kth component of .

[0024] The beneficial effects of the present invention are:

[0025] 1. Effective use of data: This paper combines semi-supervised learning and open-world methods to effectively utilize limited labeled data and a large amount of unlabeled data containing known and new classes, reducing dependence on labeled data and improving data utilization.

[0026] 2. Accurate category identification: Through the synergistic effect of pairwise similarity loss, adaptive cross-entropy loss, and entropy regularization, it can not only accurately identify known types of radiation sources, but also effectively distinguish and identify new types of radiation sources, overcoming the defect of existing methods that cannot subdivide unknown categories.

[0027] 3. Improved generalization and stability: In open-world scenarios, the method of the present invention demonstrates good generalization and stability, has strong adaptability to different label ratios and new class ratios, and can maintain a high recognition accuracy rate in complex and changing electromagnetic environments. It enhances the model's adaptability to new categories of radiation sources and helps promote the application of individual radiation source identification technology in open environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 This is a framework diagram of the similarity-adaptive semi-supervised open-world radiator individual identification method proposed in this embodiment;

[0029] Figure 2 This is a structural diagram of the feature extractor proposed in this embodiment. DETAILED DESCRIPTION

[0030] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.

[0031] The following first introduces the embodiments of the present application with reference to the accompanying drawings.

[0032] Taking the experiment on the ZigBee dataset as an example, the specific implementation process of the method of the present invention is described:

[0033] 1. Data Preparation: We used a ZigBee dataset containing 56,000 signal samples from 20 devices. We randomly divided the data into training and test sets in a 7:3 ratio. We assumed that the first 50% of the classes in the training set were known, and the remaining 50% were unknown. We used 50% of the known classes as labeled samples, and the remaining known and unknown class samples as unlabeled data. The test set had the same ratio of known to unknown classes as the training set, and both were labeled.

[0034] 2. Model Training Parameter Settings: In the similarity-based adaptive recognition framework, 100 epochs of training were performed using the SGD optimizer with a learning rate of 0.1 and a batch size of 128. The cosine annealing algorithm was used to adjust the learning rate. Experiments were conducted in PyTorch 1.12.1 and Python 3.8.16 using a single NVIDIA RTX 6000 Ada Generation GPU. The experimental results are the average of 10 runs.

[0035] 3. Implementation of the method steps of the present invention

[0036] Step 1: Feature extraction: Input the divided training set and test set samples into the feature extractor in sequence. According to the structure and operation flow of the feature extractor, the original I / Q signal is processed to obtain the feature representation of the sample.

[0037] Step 2: New category identification: Based on the calculation method of pairwise similarity loss, during the training process, the similarity predictor is trained using labeled sample sets and unlabeled sample sets to calculate the similarity of sample pairs. By minimizing the pairwise similarity loss, the model can recognize new categories.

[0038] Step 3: Known class recognition and learning balance: Generate pseudo labels for unlabeled samples based on the similarity of sample pairs, combine the uncertainty adaptive boundary mechanism, calculate the adaptive cross entropy loss, adjust the model's learning of known classes and new classes, so that the model can accurately recognize known classes while balancing the learning speed of new classes and known classes.

[0039] Step 4: Prevent overfitting: During the training process, entropy regularization is applied to the mean output probability of batch samples according to the calculation method of the entropy regularization term to avoid overfitting of the model.

[0040] Step 5: Model training and identification: The model is trained by minimizing the overall objective function. After training, the test set samples are fed into the trained model. The radiant samples are classified based on the model output. The known class accuracy (ACC), clustering accuracy (CA), and harmonic overall accuracy (HOA) are calculated to evaluate model performance.

[0041] The relevant experimental results are shown in Tables 1, 2, and 3. It can be found that the method of the present invention can effectively utilize limited labeled data and a large amount of unlabeled data containing known classes and new classes, reducing dependence on labeled data. At the same time, it can not only accurately identify radiation sources of known classes, but also effectively distinguish and identify radiation sources of new classes, and has strong adaptability to different label ratios and new class ratios.

[0042] To evaluate the recognition performance of the SAA method under different labeled sample ratios, we fixed the known class ratio at 0.5 and conducted extensive experiments on the data from Step 1 under conditions where the labeled sample ratio ranged from 0.1 to 0.9. The results are shown in Table 1. As can be seen, the proposed SAA, as an end-to-end framework, performs better than the traditional two-step method of opening the set first and then clustering when recognizing known and new classes. It demonstrates excellent recognition performance and stability under different labeled sample ratios. More labeled data helps improve the generalization ability of SAA.

[0043] Table 1 The impact of the proportion of labeled samples on recognition performance (%) (the proportion of known categories is 0.5)

[0044]

[0045] To evaluate the recognition performance of the SAA method under novel class ratios, we fixed the known sample labeled ratio at 0.5 and conducted experiments with novel class ratios ranging from 0.1 to 0.9. We conducted extensive experiments based on the data from Step 1, and the results are shown in Table 2. This validates the effectiveness of the SAA method under varying novel class ratios. The SAA method demonstrates stable and good performance for both known and unknown class recognition under varying novel class ratios, and demonstrates generalization capabilities despite changes in novel class ratios.

[0046] Table 2 Impact of new class ratio on recognition performance (%) (labeling ratio 0.5)

[0047]

[0048] To evaluate the impact of each component of the SAA optimization objective on the model's recognition performance, ablation experiments were conducted on the data from Step 1, using adaptive cross-entropy loss, pairwise similarity loss, and entropy regularization, both individually and in combination, with a known class ratio of 50% and a labeled sample ratio of 50%. The experimental results are detailed in Table 3. The ablation results validate the effectiveness of each component of the loss function. Pairwise similarity loss helps improve the model's recognition of novel classes, adaptive cross-entropy loss enhances recognition of known classes, and regularization improves the model's generalization ability to both known and novel classes. The synergistic effect of these components enables the model to accurately recognize both known and novel classes, with each contributing to the model's ultimate performance.

[0049] Table 3 Ablation analysis of SAA method for 50% known categories and 50% new categories (%)

[0050]

[0051] The above-described embodiments of the present application do not constitute a limitation on the scope of protection of the present application.

Claims

1. A similarity-adaptive semi-supervised open-world radiator individual identification method, characterized by: The method comprises the following steps: Step 1: Get the labeled sample set D u and the unlabeled sample set D u Data set D, where D l Contains N l Labeled samples x i , corresponding to the known class y i ∈C l , D u Contains N u unlabeled samples, true labels y u From the unknown class label set Y u , and D u Contains D l New classes that do not exist in ; Step 2: Input the labeled samples and unlabeled samples into the feature extractor to obtain the features of the input samples; the feature extractor includes a cascade of convolutional layers, pooling layers, and residual units; Step 3: Use pairwise similarity loss to transform the clustering problem into a sample pair similarity prediction problem. Use labeled samples and unlabeled samples to train the feature extractor and similarity predictor to determine the sample pair similarity. Calculate the pairwise similarity loss through binary cross entropy loss. optimization Identify new categories; Step 4: Use adaptive cross entropy loss to identify known classes and balance the learning speed between new classes and known classes. Generate pseudo labels based on sample similarity to calculate cross entropy loss. Introduce uncertainty adaptive boundary mechanism to adjust classification boundaries by estimating the uncertainty of the model, determine the uncertainty of unlabeled dataset samples to obtain the overall uncertainty estimate u, and then calculate the adaptive cross entropy loss L. ace ; Step 5: Introduce entropy regularization term Apply entropy regularization to the mean of the output probability of the entire batch of samples to avoid overfitting of the model; Step 6: By minimizing the overall objective function Train the model and use the trained model to identify radiation source samples.

2. The similarity-adaptive semi-supervised open-world radiation source individual identification method according to claim 1 is characterized in that: The operation process of the feature extractor is as follows: The network first receives the original I and Q signals through the input layer, and initially extracts features through a one-dimensional convolution layer using 64 convolution kernels of size 7. The feature map size is then reduced by a maximum pooling layer. The feature extraction capability is then further enhanced by two residual units that each contain two one-dimensional convolution layers using 64 convolution kernels of size 7. The residual unit alleviates the gradient vanishing problem in deep network training through jump connections. Finally, the average pooling layer performs feature integration to generate signal features. Each convolution layer is followed by a batch normalization layer, and the rectified linear unit (ReLU) is used as the activation function.

3. The similarity-adaptive semi-supervised open-world radiation source individual identification method according to claim 1 is characterized in that: When calculating the pairwise similarity loss, for the labeled dataset, similar instance pairs are identified based on the true labels. For the unlabeled dataset, the cosine distance of the sample pairs in the feature space is calculated to evaluate the similarity. For each unlabeled dataset, the corresponding sample is assigned a pseudo label with the most similar neighbor to itself. For each small batch of feature representations, the labeled sample Z is determined. l and unlabeled samples Z u The set of closest neighbors Z l ' and Z u ', calculated using binary cross entropy loss For any pair of samples (z i ,z j ), Determined by: in, Represents pairwise similarity loss, which is used to measure the accuracy of the model's similarity prediction for paired samples; m represents D l The number of labeled samples in u The number of unlabeled samples in Φ; T represents the transpose of the similarity predictor weight parameters; z i and z i ′ respectively represent the feature representation of the sample after passing through the feature extractor, where z i From the collection z i 'From the set ζ represents the softmax function, which converts the result into a probability distribution through normalization; <·,·> represents the inner product operation, which measures the similarity between two samples.

4. The similarity-adaptive semi-supervised open-world radiation source individual identification method according to claim 1, characterized in that: When calculating the adaptive cross entropy loss, pseudo labels are generated based on the similarity of sample pairs to calculate the cross entropy loss. The uncertainty adaptive boundary mechanism is introduced to set a larger boundary at the beginning of training to reduce the misclassification of new categories. As the training progresses, the boundary value is gradually reduced. By calculating the unlabeled dataset D u The uncertainty of the sample in the quantile is used to estimate the overall uncertainty Then calculate L ace : Among them, m represents the number of samples with labeled data; z i Represents the feature representation of labeled data; Z l is the set of all labeled data feature representations; Φ T represents the transpose of the predictor weight parameters; y i It is z i Corresponding sample category; y j Represents and z i Different labels for other sample features; uncertainty The strength of is controlled by the regularization parameter λ, which is set to 1; the parameter s is used to adjust the temperature of the cross entropy loss; e represents a natural constant; and log represents the natural logarithm function.

5. The similarity-adaptive semi-supervised open-world radiation source individual identification method according to claim 1, characterized in that: The entropy regularization term It is to apply entropy regularization to the mean of the output probability of the entire batch of samples as follows: Among them, c a Represents the total number of all known classes and possible new classes; is the average output probability of each batch of samples, b represents the number of samples in each batch, Represents the output probability of the i-th sample in each batch; represents the average output probability vector The kth component of .