Time-frequency positioning model self-supervised learning method based on mask auto-encoder
Through the self-supervised learning method of the time-frequency positioning model based on the masked self-encoding machine, the influence of signal time-frequency domain overlap on the time-frequency positioning performance is solved, and the robustness and performance of the detection model are improved.
Patent Information
- Application Number
- CN202510139688.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-08
- Publication Date
- 2025-05-30
AI Technical Summary
In concurrent cognitive scenarios, the overlap of the time frequency domain of the signal seriously reduces the performance of the target detector, and the prior art fails to systematically analyze and evaluate the impact of signal overlap on time frequency positioning.
The self-supervised learning method of the time-frequency positioning model based on masked self-encoding machine is adopted to reduce the impact of signal time-frequency domain overlap on time-frequency positioning through time-frequency map reconstruction pre-training and downstream fine-tuning of time-frequency positioning.
The performance of the object detection model is improved, the robustness of the time-frequency positioning model is enhanced, and the problem of signal overlap and labeling is more efficiently handled.
Smart Images

Figure CN120068944A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of wireless communication technologies, and particularly to a self-supervised learning method for a time-frequency localization model based on a masked autoencoder. Background Art
[0002] In a wireless communication system, spectrum sensing technology can significantly improve spectrum utilization efficiency and plays a key role. With the development of broadband technology and the enhancement of bandwidth usage flexibility, spectrum sensing technology needs to detect the usage of spectrum resources within a wider frequency range. Therefore, broadband spectrum sensing technology has become an extremely important development direction. In particular, time-frequency localization technology can simultaneously detect the time range and frequency range in the time-frequency representation, which is beneficial to achieving precise dynamic spectrum management in broadband cognitive scenarios.
[0003] Currently, a target detector trained on a large-scale labeled spectrogram dataset can achieve excellent time-frequency localization performance in a simple electromagnetic scenario. However, in an actual concurrent cognitive scenario, users are allowed to reuse the same frequency band under specific interference constraints, and the overlapping phenomenon of signals in the time-frequency domain will seriously degrade the performance of the target detector. So far, there is no systematic analysis and evaluation of the impact of signal overlap on time-frequency localization.
[0004] It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of the present invention, and thus may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0005] The present invention provides a self-supervised learning method for a time-frequency localization model based on a masked autoencoder, which is used to reduce the impact of signal overlap in the time-frequency domain on time-frequency localization and improve the performance of the target detection model.
[0006] Other features and advantages of the present invention will become apparent through the following detailed description, or will be learned in part through the practice of the present invention.
[0007] According to a first aspect of the present invention, there is provided a self-supervised learning method for a time-frequency localization model based on a masked autoencoder, the method including pre-training of time-frequency spectrogram reconstruction and downstream fine-tuning of time-frequency localization;
[0008] The pre-training of time-frequency spectrogram reconstruction: After masking the time-frequency spectrogram, it is input into a Transformer-based autoencoder network, and the Transformer-based autoencoder network includes an encoder and a decoder. Among them, the encoder extracts features from the masked spectrogram to obtain a feature map, and the decoder decodes the feature map to output the reconstructed time-frequency spectrogram;
[0009] The downstream fine-tuning of time-frequency localization: Use the pre-trained encoder as the backbone network of the time-frequency localization model and share the network parameters; extract the features of the labeled time-frequency spectrogram through the backbone network, perform feature fusion through the feature pyramid, and obtain the time-frequency localization result by decoding the fused features through the head network; based on the labeled information and the time-frequency localization result, train a robust time-frequency localization model through the loss function.
[0010] In some exemplary embodiments, the method further includes generating a time-frequency spectrogram, specifically: obtaining the time-frequency spectrogram by applying STFT to the time series signal sequence.
[0011] In some exemplary embodiments, the masking process includes:
[0012] Decompose the time-frequency spectrogram into several grids, and specify each grid as the basic unit for generating the mask;
[0013] In each training cycle, generate a random mask for each spectrogram, randomly specify each grid as masked or retained, thus forming a mask matrix.
[0014] In some exemplary embodiments, when pre-training the time-frequency spectrogram reconstruction, select the mean square error MSE between the reconstructed time-frequency spectrogram and the input time-frequency spectrogram Φ as the loss function, and use the loss function to optimize the network parameters of the autoencoder. The expression of the loss function is:
[0015]
[0016] where |·| 2 represents the l 2 norm.
[0017] In some exemplary embodiments, the time-frequency localization model includes: YOLOv3, YOLOv4, FasterRCNN, and SSD models.
[0018] In some exemplary embodiments, when performing downstream fine-tuning of time-frequency localization, optimize for the presence detection, classification, and time-frequency localization performance of the signal. The designed expression of the loss function is:
[0019] L TFL = λ d L det + λ c L cls + λ T L CIoU
[0020] where the loss terms L det 、L cls and L CIoUfor optimizing detection performance, classification performance, and time-frequency localization performance respectively; λ d , λ c and λ T are hyperparameters used to adjust the weights of each loss term.
[0021] According to a second aspect of the present invention, there is provided a storage medium having stored thereon a computer program which, when executed by a processor, implements the self-supervised learning method of the time-frequency localization model based on a masked autoencoder described in the first aspect above.
[0022] According to a third aspect of the present invention, there is provided a computer program product having stored thereon a computer program which, when executed by a processor, implements the self-supervised learning method of the time-frequency localization model based on a masked autoencoder described in the first aspect above.
[0023] According to a fourth aspect of the present invention, there is provided an electronic device, comprising:
[0024] a processor; and
[0025] a memory for storing executable instructions of the processor;
[0026] wherein the processor is configured to implement the self-supervised learning method of the time-frequency localization model based on a masked autoencoder described in the first aspect above when executing the executable instructions.
[0027] The self-supervised learning method of the time-frequency localization model based on a masked autoencoder provided by the embodiments of the present invention, based on the analysis of the influence of signal overlap, establishes a self-supervised learning (Self-Supervised Learning, SSL) framework based on a masked autoencoder (Masked Autoencoder, MAE), and uses an unlabeled dataset to pre-train the backbone network to have excellent feature extraction capabilities, so as to address the problems of overlapping diversity, annotation difficulties, and feature destruction in the time-frequency spectrogram, and improve the robustness of the time-frequency localization model.
[0028] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] The drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present invention, and are used together with the specification to explain the principles of the present invention. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.
[0030] Figure 1 It is a schematic diagram of time-frequency overlap;
[0031] Figure 2 It is a schematic diagram of an overlapping dataset generation system;
[0032] Figure 3 It is an architecture diagram of a self-supervised learning system based on MAE;
[0033] Figure 4 It is a structural diagram of an autoencoder network based on Transformer;
[0034] Figure 5 It is a numerical curve of the loss function of TRTFL under whether to use self-supervised learning pre-training;
[0035] Figure 6 It is a comparison of the recognition performance of different time-frequency localization models and TRTFL-SSL;
[0036] Figure 7 It is a comparison of the performance of each time-frequency localization algorithm under different overlap rates;
[0037] Figure 8 It is the performance impact of SSL on TRTFL under different amounts of training data;
[0038] Figure 9 It is the detection performance impact of SSL on each time-frequency localization algorithm. Detailed implementation manners
[0039] Now, the example embodiments will be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this invention will be more complete and comprehensive, and will fully convey the concept of the example embodiments to those skilled in the art. The features, structures, or characteristics described may be combined in any suitable manner in one or more embodiments.
[0040] In addition, the accompanying drawings are only schematic illustrations of the present invention and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and thus their repeated description will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0041] In the related art, given the similarity between time-frequency localization in spectrograms and object detection in images, a series of object detection algorithms have been applied to time-frequency localization tasks, such as YOLOv2, YOLOv3, YOLOv5, YOLOR, SSD, and CenterNet. Although the existing technologies have effectively achieved wideband spectrum sensing in a simple electromagnetic environment, they have not considered the signal overlap phenomenon. The concurrent cognitive mechanism allows different existing users to transmit signals simultaneously on the same frequency band as long as the interference is below the interference temperature threshold. Therefore, the perceived signals will overlap in the time-frequency domain. In addition, although cognitive radio networks have access mechanisms for selecting spectrum holes to avoid interference, when the distance between existing users exceeds their sensing range, they may still choose the same frequency band for signal transmission. In this case, the signals perceived by secondary users located between existing users will overlap in the time-frequency domain.
[0042] Multiple studies have pointed out the problem of signal overlap. Traditional methods are often difficult to effectively handle the complex situations brought about by signal overlap and interference due to strict assumptions and limited adaptability and learning ability. In contrast, deep learning methods are excellent in processing complex non-linear feature extraction tasks, making them particularly suitable for processing high-dimensional time-frequency data. Similarly, in the field of computer vision, object detectors based on deep learning also face challenges related to object overlap, and usually use techniques such as multi-scale feature extraction, non-maximum suppression, and attention mechanisms to solve them.
[0043] Although the above technologies consider the situation of signal overlap, the impact of signal overlap in the time-frequency domain on time-frequency localization has not been explicitly considered. The new challenges brought about by the signal overlap phenomenon to the learning framework and detector structure are as follows: In different signal overlap scenarios, the current time-frequency localization algorithms based on supervised learning require a larger-scale labeled dataset to ensure model performance and generalization ability. In addition, since the overlapping signals are intertwined in the time-frequency domain, it is difficult to determine the time and frequency ranges of the overlapping parts, which brings great difficulties to the manual annotation work. These two factors lead to insufficient annotation quantities for model training. Therefore, it has become a major challenge to train a detector with strong feature extraction ability using limited labeled samples. Moreover, the overlapping parts of the overlapping signals are intertwined in the time-frequency domain, resulting in feature degradation and incompleteness. Therefore, the robustness of the feature extraction backbone network based on convolutional neural networks will be weakened to a certain extent. Therefore, designing a detector with strong feature extraction ability to cope with the problem of damaged features has become another major challenge.
[0044] The present invention comprehensively analyzes the impact of signal overlap in time-frequency localization, and highlights the challenges brought about by the diverse forms of signal overlap, difficult annotation, and impaired features. To alleviate these problems, the present invention proposes a feasible solution: a self-supervised learning framework. To overcome the limitation of relying on large-scale annotated datasets, a new framework based on MAE is adopted. This framework introduces self-supervised learning in the pre-training stage, aiming to improve the model's ability to learn effective feature representations in the absence of annotated samples; and utilizes knowledge transfer in the fine-tuning stage to enhance the robustness of the detector.
[0045] First, analyze the characteristics of time-frequency overlap: As current wireless systems need to operate flexibly within the shared wideband radio frequency spectrum, a comprehensive perception of the signal presence in the broadband electromagnetic spectrum becomes particularly crucial. Time-frequency localization methods based on object detection can estimate the presence, category, start and end times, and frequencies of signals, which is of great significance for achieving precise dynamic spectrum management. Different from traditional object detection tasks of detecting objects in images, the detection objects in time-frequency localization are all signals in the spectrogram, and the distributions of different types of signals in the time-frequency domain constitute the detection targets. In addition, due to the characteristics of the Concurrent Cognitive MechAnistom (CCM) and wide-range perception, signals are prone to overlap in the time-frequency domain, which has a certain impact on the performance of the detector. For example, inaccurate time-frequency localization may lead to problems of missed detection and false alarms. Missed detection means that the target signal exists but is not detected, which may cause co-channel interference, thereby reducing the signal quality and increasing the bit error rate. False alarm means that the target signal does not exist but is wrongly detected as existing, which will hinder the optimal utilization of the idle spectrum and lead to a decrease in spectrum efficiency.
[0046] As Figure 1 shown in the time-frequency overlap schematic diagram, by analyzing the characteristics of time-frequency overlap, some key points of the impact of signal overlap on time-frequency localization performance can be obtained:
[0047] 1) Diversity of overlap: Signal overlap can occur at multiple locations, different scales, and various ratios, which poses significant challenges for effective processing. Object detectors rely on a large number of diverse training samples and their annotation information, covering different forms of overlap, to train a network with good generalization ability. However, due to the inherent complexity of overlap, the detector may not be able to learn all types of overlap instances, resulting in a decline in the performance of the convolutional neural network.
[0048] 2) Difficulty in annotation: Object detection algorithms require a large number of annotated samples to effectively train neural networks. However, manually annotating the time-frequency localization information of time-frequency blocks of signals in spectrograms is both time-consuming and laborious, often resulting in the scarcity of annotations in practical scenarios. In addition, due to the overlap between signals, this annotation difficulty is further exacerbated, further expanding the problem of annotation scarcity. As Figure 1 shown in (a), when the time-frequency blocks of two signals highly overlap, it is difficult to accurately determine the boundaries of the targets.
[0049] 3) Feature destruction: Overlap will destroy the features of target signals, resulting in blurred features between similar target signals and the incompleteness of the features of individual target signals. As Figure 1 shown in (b), when two similar target signals overlap, their features are almost indistinguishable, and the object detector may incorrectly select a larger prior box, thus producing incorrect detection results. In addition, as Figure 1 shown in (c), overlap will lead to the incompleteness of the features of target signals, affecting the accuracy of the lengths and widths of the detection boxes predicted by the object detector.
[0050] Secondly, generate a time-frequency overlap simulation dataset: To facilitate the study of time-frequency localization problems in cognitive networks, we propose a concurrent cognitive dataset generation system, as Figure 2 shown. In our cognitive environment, accessing users follow CCM, allowing different radio frequency devices to transmit simultaneously on the same frequency band, provided that the transmission power of the transmitted signals remains below the interference temperature threshold. In this case, the transmitted signals will overlap in the time-frequency domain. In addition, considering practicality and generality, we select LTE and NR signals as target signals and use the MATLAB toolbox to generate their waveforms.
[0051] (1) Cognitive environment configuration
[0052] The cognitive network system consists of N RF radio frequency devices randomly distributed within a circular area with a diameter of 1000 meters. N RFA radio frequency device follows the CCM access mechanism, which allows different communication systems to transmit simultaneously on the same frequency band. Each generated radio frequency signal is concentrated within the 2.4 GHz frequency band, has a broadband bandwidth of 61.44 MHz, and a signal duration of 90 ms. Different from existing cognitive datasets, we assume that the accessing users have certain cognitive abilities and follow specific access rules. In the cognitive environment we propose, the accessing users can dynamically determine the frequency band for their signal transmission according to the spectrum usage of other accessing users. The access order of the accessing users is randomly determined, and the bandwidth and signal type are randomly selected from the candidate set. In addition, we assume that external accessing users can actively judge whether they can reuse the frequency resources assigned to other cognitive targets based on the received signal sensitivity, which may lead to signal overlap. Specifically, the spectrum access process is summarized as follows:
[0053] Step 1: Generate the frequency order of the accessing users and initialize the available frequency range for each accessing user.
[0054] Step 2: According to the frequency order, each accessing user then determines the parameters of the transmitted signal, including the signal type, bandwidth, and transmission power level.
[0055] Step 3: Each accessing user independently determines the center frequency of its signal transmission through a random selection process, and the frequency is within its available frequency band. If the available frequency band is insufficient, the accessing user will avoid access.
[0056] Step 4: Update the available frequency band range for each accessing user according to the reuse relationship.
[0057] Step 5: Repeat the process described in Steps 2 - 4.
[0058] (2) Signal waveform
[0059] For practicality and generality, we selected LTE and NR signals and used the MATLAB~ toolbox to generate their waveforms. For LTE signals, we generated LTE waveforms by using the MATLAB LTE toolbox. For LTE data generation and channel modeling, we adopted the lteRMCDLTool function and lteFadingChannel function that conform to the 3GPP TS36.101 standard.
[0060] For NR signals, we generated NR waveforms by using the MATLAB NR toolbox. To generate NR data and perform channel modeling, we adopted the nrWaveformGenerator function and nrCDLChannel function that conform to the 3GPP TR38.901 standard.
[0061] (3) Data generation
[0062] This dataset includes spectrograms obtained by applying STFT to a time-series signal sequence. Specifically, the signal received at time i can be expressed as:
[0063]
[0064] where s j (i) is the wireless signal of the j-th access user, h j follows the Rayleigh fading channel distribution, and z(i) represents additive white Gaussian noise. To comprehensively evaluate the performance of the detector, it is usually necessary to evaluate its detection ability at different SNRs. To obtain quantized SNR samples from -20 dB to 20 dB with an interval of 2 dB, we choose to adjust the noise power z. The SNR can be expressed as:
[0065]
[0066] The time-series signal with added Gaussian white noise is processed by STFT with a Hanning window of size 512 and a 512-point FFT to generate a time-frequency spectrogram and then annotated. Specifically, the generated dataset contains 126,000 spectrograms, covering SNR levels from -20 dB to 20 dB with a step of 2 dB. The dataset for each SNR level contains 6,000 spectrograms, which are specifically classified as follows: time-frequency spectrograms of 500 individual LTE signals, time-frequency spectrograms of 500 individual NR signals, and 5,000 mixed time-frequency spectrograms. Five different overlap ratios are designed for the mixed spectrograms: 0%, 25%, 50%, 75%, and 100%, and 1,000 time-frequency spectrograms are assigned to each overlap ratio. The overlap ratio is defined as the ratio of the length of the overlapping bandwidth between two signals to the total length of the union of their bandwidths.
[0067] Based on the above analysis of time-frequency overlap characteristics and the generation of time-frequency overlap simulation datasets, in this exemplary embodiment, a self-supervised learning method for a time-frequency localization model based on a masked autoencoder is provided. Similar to traditional supervised learning methods, current data-driven intelligent time-frequency localization detectors rely heavily on a large number of annotated samples during the process of effectively optimizing their network parameters and achieving excellent performance. Although radio frequency sampling devices can quickly collect a large number of signal samples, manually annotating these samples accurately is a time-consuming and labor-intensive task. This results in a shortage of annotated samples for model training. However, detectors trained with insufficient annotated samples exhibit limited generalization ability. To address this issue and enhance the feature extraction ability of the detector, we propose a self-supervised learning framework based on MAE to effectively utilize the unannotated dataset. As Figure 3 shown, the proposed system architecture includes two stages: spectrogram reconstruction pre-training and downstream time-frequency localization fine-tuning.
[0068] In the pre-training stage, we apply the random masking technique to the spectrogram and input it into a Transformer-based autoencoder network. During this process, the encoder extracts features from the masked spectrogram to obtain the feature map. Subsequently, this feature map is input into the decoder to complete the reconstruction task. Through pre-training, we can obtain an encoder network with excellent feature representation capabilities.
[0069] In the fine-tuning stage, we use the pre-trained encoder as the backbone network of the time-frequency localization model and share the network parameters. Then, the features of the labeled samples are extracted from the backbone network and fused through the feature pyramid. Subsequently, the processed features are decoded through the head network to obtain accurate time-frequency localization results. Finally, by using the labeled information and the time-frequency localization results, a robust time-frequency localization model is trained through the loss function.
[0070] Specifically, the pre-training stage of time-frequency spectrogram reconstruction is described as follows:
[0071] To address the problem of scarce labeled samples, we propose to use MAE to construct a self-supervised learning task, enabling the model to obtain feature extraction capabilities without relying on the amount of labeled data. MAE uses a simple and efficient method by partially masking the input spectrogram and then reconstructing its content.
[0072] (1) Mask processing
[0073] To enable random masking operations on the training spectrogram, we decompose each input spectrogram into grids, where H and W represent the height and width of the spectrogram respectively. The parameter k is used to control the granularity of these grids. In addition, we specify each grid as the basic unit for generating the mask. To expand the exploration scope of the mask, a random mask is dynamically generated for each spectrogram in each training epoch. Subsequently, each grid is randomly designated as masked or retained, thus forming a mask matrix named Ω. Through this random masking operation, each spectrogram is enhanced into a set of diverse training triples where Φ is the input spectrogram, Ω is the generated spatial mask, is the spectrogram after masking processing, represents the Hadamard product in the spatial domain. Through this random masking, each spectrogram can learn to be reconstructed in different sizes and shapes, thus enriching the diversity of the training data.
[0074] (2) Autoencoder structure
[0075] Initially, inspired by the MAE architecture, we established a time-frequency spectrogram reconstruction framework for the autoencoder. Subsequently, through extensive experiments, we used a trial-and-error method to find the best balance between performance and computational efficiency, thereby determining specific network architecture parameters, such as the number of dual-attention Transformer blocks. Finally, we designed a Transformer-based autoencoder network structure, as Figure 4 shown, which mainly consists of two basic components: an encoder and a decoder. The main role of the encoder is to extract significant features from the masked spectrogram, thus providing support for the decoder to reconstruct the spectrogram. First, shallow information is extracted through the chunking module. Subsequently, this shallow information is passed into the dual-attention Transformer block for correlation-based feature extraction. Then, these correlation features are input into the chunk merging module for downsampling. After iterative operations of the dual-attention Transformer block and the chunk merging module, a feature map is finally obtained. The role of the decoder is to reconstruct the spectrogram from the feature map, and it uses the upsampling module to complete this task. In this study, we set the number of dual-attention Transformer blocks in the encoder to 1, 1, 2, and 1, where each block contains two multi-head sub-attention layers. We adopted a chunk partitioning module with a chunk size of 4×4 and a dimension of 96. In addition, we used chunk merging modules with dimensions of 192, 384, 768, and 1536 respectively.
[0076] (3)MAE Loss Function
[0077] To effectively perform the time-frequency spectrogram reconstruction task, a loss function must be adopted to optimize the network parameters of the autoencoder. Specifically, we selected the mean squared error (MSE) between the reconstructed time-frequency spectrogram and the input time-frequency spectrogram Φ as the loss function, and its expression is as follows:
[0078]
[0079] where, |·| 2 represents the l 2 norm.
[0080] Specifically, the time-frequency localization downstream fine-tuning stage is described as follows:
[0081] Through pre-training on reconstructing time-frequency spectrograms from a large amount of unlabeled data, the encoder, as the backbone network of the time-frequency localization detector, demonstrates excellent feature extraction capabilities. During the subsequent fine-tuning stage for the time-frequency localization downstream task, the main objective is to fine-tune and train the entire time-frequency localization detector using a small amount of labeled data. Specifically, to avoid the disruption of the pre-trained backbone network caused by the instability of the previous training, we freeze the parameters of the backbone network in the initial stage. After the training reaches a certain level of stability, we gradually unfreeze the parameters of the backbone network, ultimately obtaining a robust time-frequency localization detection network. The time-frequency localization model used in this invention is the TRTFL model. Note that the proposed self-supervised learning architecture is general and can be applied to all time-frequency localization models. In subsequent simulations, common models such as YOLOv3, YOLOv4, FasterRCNN, and SSD were used.
[0082] Since the signal recognition and time-frequency localization tasks under study require optimizing the detection of signal presence, classification, and time-frequency localization performance during the training phase. For this purpose, we designed a loss function, and its expression is as follows:
[0083] L TFL = λ d L det + λ c L cls + λ T L CIoU
[0084] Among them, the loss terms L det 、L cls and L CIoU are used to optimize the detection performance, classification performance, and time-frequency localization performance respectively. λ d = 1, λ c = 3, and λ T = 7 are hyperparameters used to adjust the weights of each loss term.
[0085] To optimize the detection performance, we use the cross-entropy loss function as the optimization objective for L det . In addition, due to the imbalance problem between positive and negative samples in the time-frequency map, focal loss is introduced as an improvement scheme to solve the sample imbalance problem in object detection, and its expression is:
[0086]
[0087] In the network, each spectrogram is divided into a grid of N S ×N S , and each grid predicts N B bounding boxes. Indicates that the object falls into the j-th bounding box of grid i, otherwise it is 0. Indicates the confidence of the network in predicting the presence of an object. γ = 2 and β = 3 are hyperparameters of the focal loss function, used to balance positive and negative samples.
[0088] The classification loss term is defined as:
[0089]
[0090] where p i (c) represents the true probability that the object in grid i belongs to class c. If grid i belongs to class c, then p i (c) = 1, otherwise 0. Represents the probability that the model predicts grid i belongs to class c. If grid i contains an object, then otherwise 0.
[0091] To optimize the time-frequency localization performance of the signal, we design the loss function as:
[0092]
[0093] where Γ CIoU is the Complete Intersection over Union (CIoU) between the bounding box and the ground truth.
[0094]
[0095] Self-supervised learning plays a role in the pre-training stage. The model learns from unlabeled data by reconstructing the masked part of the time-frequency spectrogram. This process is crucial for extracting useful features, which provide a basis for the subsequent fine-tuning stage and can work effectively even when the labeled data is limited. In the pre-training stage, the unlabeled spectrograms are randomly masked and reconstructed. The loss function L MSE , represents the Mean Squared Error (MSE) between the reconstructed spectrogram and the original spectrogram. By participating in the reconstruction task, the encoder can obtain a more efficient latent representation, thereby enhancing its feature extraction ability. In the fine-tuning stage, labeled samples are used to adjust the fine-tuning model to adapt to the time-frequency localization task. First, the corresponding parameters of the pre-trained model are transferred to the fine-tuning model. Subsequently, these parameters are updated by minimizing the loss function defined in Equation L TFL . We summarize the training process of the proposed framework in the above algorithm.
[0096] The following describes the effect of this application in detail in combination with simulation experiments
[0097] Simulation experiment parameter settings
[0098] To implement the proposed self-supervised learning framework and the time-frequency localization detector, we used a workstation configured with an Intel(R) i7-8700K CPU and an NVIDIA GTX-1080Ti GPU (CUDA 10.0 and cuDNN 7.4.1.5). In addition, we adopted the machine learning framework Pytorch to build and train the proposed detector. In the pre-training stage, we used the entire unlabeled proposed dataset to train the spectrogram reconstruction pre-training module. In the fine-tuning stage, we used 10% of the labeled spectrograms to train the downstream time-frequency localization fine-tuning module. The datasets for both were divided in the ratio of training set: test set: validation set as 8:1:1. In addition, the training of both modules was carried out for a total of 150 epochs, with a learning rate of 0.001 used in the first 50 epochs and the learning rate reduced to 0.0001 in the following 100 epochs. Adam optimizer was used for optimization in this paper.
[0099] 4.4.2 Performance Comparison
[0100] Figure 5 Shows the convergence of the proposed TRTFL detector with and without pre-training in the fine-tuning stage. It can be clearly seen from the loss curve that our TRTFL detector can converge stably within 150 epochs whether pre-trained or not. In addition, the pre-trained detector converges faster and reaches a lower loss. This indicates that the proposed SSL pre-training architecture can accelerate model convergence and improve model performance to a certain extent.
[0101] Table 1 Evaluation Metrics of the Model on the Benchmark Database
[0102]
[0103] To verify the generalization ability of the proposed method, we conducted detailed performance simulation experiments on the current state-of-the-art SPREAD-small dataset. This dataset mainly contains 99,330 time-frequency spectrograms and their corresponding five signal type labels: Wi-Fi, Bluetooth, ZigBee, Lightbridge, and XPD. In the pre-training stage, we used the entire SPREAD-small dataset (without labels) to pre-train the time-frequency spectrogram reconstruction module. In the fine-tuning stage, we used 10% of the labeled time-frequency spectrograms to train the downstream time-frequency localization fine-tuning module. The simulation results are shown in Table 1, indicating that the performance of both the convolutional neural network-based model and the Transformer-based model has been improved after using self-supervised learning. In addition, the Transformer-based model is slightly better than the convolutional neural network-based model, further verifying the effectiveness of the proposed model.
[0104] To evaluate the performance, we use the mAP with an IoU threshold of 0.5 as the evaluation metric. As Figure 6 shown, the mAP performance of the proposed TRTFL-SSL model and other existing convolutional neural network-based object detectors under different SNR conditions is plotted. We train all detectors using 10% of the labeled time-frequency spectrograms in the proposed CR dataset. During testing, we construct 16 time-frequency spectrogram test sets, and the SNR level of each test set is selected from the interval [-10:2:20] dB. The results show that as the SNR increases, the performance gap between the detectors gradually decreases. In addition, the proposed TRTFL-SSL model achieves the best performance under both high SNR and low SNR conditions.
[0105] To statistically evaluate the recognition performance at different overlap rates, we plot the mAP vs. SNR curves of each detector under different SNR conditions, as Figure 7 shown. Among them, FasterRCNN, SSD, YOLOv3, and YOLOv4 are convolutional neural network-based models, while TRTFL-SSL is a Transformer-based model. It can be clearly seen from the figure that these four detectors all show excellent performance under high SNR conditions, while under low SNR conditions, the performance of the proposed TRTFL-SSL is better than that of the other three detectors. This verifies the robustness of the proposed model in challenging scenarios such as low SNR. In addition, it can be observed that the mAP of TRTFL-SSL is always higher than that of other CNN-based object detectors. This further proves that the proposed TRTFL model can effectively improve the recognition performance.
[0106] To verify the effectiveness of the proposed SSL framework, we conduct a self-comparison study based on the TRTFL detector, comparing the performance of two models with and without using the SSL framework during the pre-training stage of the Transformer backbone network. As Figure 8 shown, the recognition performance of different detectors changes with the increase in the number of training samples. It can be observed that as the number of training samples increases, the recognition performance of both detectors improves, but the recognition performance of TRTFL-SSL using the SSL framework is always better than that of TRTFL without using the SSL framework. In addition, we find that when the number of training samples reaches 1000, TRTFL-SSL has converged to a relatively high recognition performance range, while TRTFL only converges when the number of samples reaches 5000. These results indicate that the proposed SSL framework can not only alleviate the urgent need for the number of labeled samples but also further improve the recognition performance of the detector.
[0107] To verify the universality of the proposed SSL framework in improving the performance of detectors, we also conducted SSL-based pre-training on existing methods. As Figure 9 shown, the recognition performance of each detector is compared under the conditions of whether SSL pre-training is performed. It can be seen from the figure that the performance of all detectors pre-trained with SSL is better than that of detectors without SSL pre-training. In addition, regardless of whether SSL pre-training is adopted, the performance of the proposed TRTFL detector is better than that of other object detectors based on convolutional neural networks. These simulation results show that the proposed SSL framework can effectively improve the recognition performance of time-frequency localization detectors and has strong generality and practical value.
[0108] To break through the limitation of over-reliance on large-scale labeled datasets, the present invention adopts an innovative framework based on MAE. In the pre-training stage, this framework introduces a self-supervised learning mechanism, aiming to strengthen the model's ability to learn effective feature representations in the absence of labeled samples. By cleverly designing self-supervised tasks, such as masked image reconstruction, autoregressive prediction, etc., the model can mine valuable information from a large amount of unlabeled data, capture the internal structure and patterns of the data, and then construct a feature space with high representational ability for various signals (especially complex signals in signal overlap scenarios).
[0109] It should be noted that, on the other hand, the present application also provides a storage medium, which may be included in an electronic device; or may exist separately without being assembled into the electronic device. The above storage medium carries one or more programs, and when the above one or more programs are executed by an electronic device, the electronic device is enabled to implement the methods described in the following embodiments.
[0110] In one embodiment, the present application provides a computer program product, including a computer program, which when executed by a processor implements the steps in the above method embodiments.
[0111] In addition, the above drawings are only schematic illustrations of the processes included in the method according to the exemplary embodiments of the present invention, rather than for limiting purposes. It is easy to understand that the processes shown in the above drawings do not indicate or limit the time sequence of these processes. Additionally, it is also easy to understand that these processes can be executed, for example, synchronously or asynchronously in multiple modules.
[0112] Other embodiments of the present invention will be readily apparent to those skilled in the art upon consideration of the specification and practice of the invention herein. This application is intended to cover any variations, uses, or adaptations of the invention following the general principles of the invention and including known or customary technical means in the technical field not disclosed herein. The specification and examples are only to be considered as exemplary, and the true scope and spirit of the invention are pointed out by the claims.
[0113] It should be understood that the present invention is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present invention is only defined by the appended claims.
Claims
1. A self-supervised learning method for a time-frequency localization model based on a masked autoencoder, characterized in that: The method includes pre-training of time-frequency spectrogram reconstruction and downstream fine-tuning of time-frequency localization; The time-frequency spectrum reconstruction pre-training: the time-frequency spectrum is masked and then input into a Transformer-based autoencoder network, wherein the Transformer-based autoencoder network includes an encoder and a decoder, wherein the encoder extracts features from the masked spectrum to obtain a feature map, and the decoder decodes the feature map to output a reconstructed time-frequency spectrum; The time-frequency localization downstream fine-tuning is as follows: the pre-trained encoder is used as the backbone network of the time-frequency localization model, and the network parameters are shared; the features of the annotated time-frequency spectrum are extracted through the backbone network, and the features are fused through the feature pyramid. The fused features are decoded through the head network to obtain the time-frequency localization results; based on the annotation information and the time-frequency localization results, a robust time-frequency localization model is trained through the loss function.
2. The self-supervised learning method of the time-frequency localization model based on the masked autoencoder according to claim 1, characterized in that: The method also includes generating a time-frequency spectrum, specifically: obtaining the time-frequency spectrum by applying STFT to the time series signal sequence.
3. The self-supervised learning method of the time-frequency positioning model based on the masked autoencoder according to claim 1 or 2, characterized in that: The mask processing includes: Decompose the time-frequency spectrum into several grids, and designate each grid as a basic unit for generating a mask; In each training cycle, a random mask is dynamically generated for each spectrogram, randomly assigning each grid as masked or retained, thus forming a mask matrix.
4. The self-supervised learning method of the time-frequency positioning model based on the masked autoencoder according to claim 1, characterized in that: When the time-frequency spectrum is reconstructed and pre-trained, the time-frequency spectrum is selected to be reconstructed. The mean square error MSE between the input time-frequency spectrum Φ is used as the loss function, and the loss function is used to optimize the network parameters of the autoencoder. The loss function expression is: Here, |·|2 represents the l2 norm.
5. The self-supervised learning method of the time-frequency positioning model based on the masked autoencoder according to claim 1, characterized in that: The time-frequency positioning models include: YOLOv3, YOLOv4, FasterRCNN and SSD models.
6. The self-supervised learning method of the time-frequency positioning model based on the masked autoencoder according to claim 1, characterized in that: When fine-tuning the time-frequency positioning downstream, the signal presence detection, classification and time-frequency positioning performance are optimized, and the loss function expression designed is: L TFL =λ d L det +λ c L cls +λ T L CIoU Among them, the loss term L det , L cls and L CIoU They are used to optimize the detection performance, classification performance and time-frequency positioning performance respectively; d , c and λ T is a hyperparameter used to adjust the weight of each loss term.
7. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the self-supervised learning method of the time-frequency localization model based on a masked autoencoder as described in any one of claims 1 to 6 is implemented.
8. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the self-supervised learning method of the time-frequency positioning model based on a masked autoencoder according to any one of claims 1 to 6 is implemented.
9. An electronic device, characterized in that: include: processor; as well as A memory, configured to store executable instructions of the processor; The processor is configured to execute the self-supervised learning method of the time-frequency positioning model based on a masked autoencoder according to any one of claims 1 to 6 by executing the executable instructions.
Citation Information
Cited By
Multi-task electromagnetic model based on hybrid expert network
CN121981190A