A Wireless Signal Representation Learning Method Based on Masked Modeling
The mask modeling method for wireless signal representation learning addresses the limitations of supervised training by leveraging unlabeled data to uncover intrinsic features, facilitating efficient and versatile signal processing across diverse applications.
Patent Information
- Application Number
- CN202311784956.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-24
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2043-12-24
AI Technical Summary
The existing wireless signal analysis methods rely on labeled samples for supervised training, and the intrinsic characteristics of wireless signals are not fully mined. The trained network model performs well on specific signals but is difficult to migrate to other application scenarios, and has poor generalization.
The wireless signal representation learning method based on mask modeling is adopted, and some fragments of the wireless signal are randomly masked through the masking mechanism. Signal features are extracted using the Transformer model, and signal reconstruction and prediction tasks are constructed to perform network model pre-training under label-free information.
Implement universal representation learning of wireless signals under label-free information, fully tap the intrinsic features of the signal, and the pre-trained model is suitable for downstream tasks such as signal prediction, classification and denoising, improving the universality and efficiency of the model.
Smart Images

Figure CN117750411B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of signal processing, and particularly relates to a wireless signal representation learning method based on masked modeling. Background Art
[0002] In a wireless communication system, in addition to necessary message transmission and processing, it is also necessary to monitor and manage the spectrum resources occupied by users, extract information such as signal modulation patterns, idle spectrum, and unknown interference sources from the detected electromagnetic signals, so as to dynamically adjust the working parameters of wireless devices to efficiently utilize limited spectrum resources. In the DZZ field, it is necessary to search, measure, analyze, and identify the electromagnetic signals of the electronic information systems and electronic devices of non-cooperating parties, and obtain information such as technical parameters, functions, types, locations, and platform categories therein, in order to further implement effective measures and actions. It can be seen that in both civilian and JY fields, it is necessary to capture key information from the received electromagnetic signals. With the rapid development and progress of the electronic information field, various types of radio devices have emerged continuously, and the electromagnetic environment has become increasingly complex, bringing greater challenges to the analysis and identification of wireless signals.
[0003] Traditional signal analysis and identification methods mainly rely on signal processing tools such as cyclic stationary feature detection and high-order moment feature extraction, combined with machine learning techniques such as support vector machines, decision trees, and k-nearest neighbors. However, these techniques usually rely on the extraction of expert features, require a large amount of domain knowledge and engineering experience, and the analysis and processing process is relatively complex and time-consuming. The rapid development of the field of artificial intelligence has provided new solutions for signal processing and key information extraction. Xu Yuqing et al. proposed a low-latency automatic modulation recognition method based on a temporal convolutional network (TCN), and further reduced the computational complexity by applying the principal component analysis method and the uniform downsampling method. Change shuo et al. represented communication signals as in-phase quadrature and amplitude-phase vectors, respectively extracted the feature information of the two representations through a convolutional neural network and a recurrent neural network, and designed a signal-to-noise ratio discrimination module for feature fusion, improving the accuracy of automatic modulation recognition. Gao Yong et al. proposed a hierarchical fusion network for narrowband radar target recognition. The intra-domain network uses an autoencoder to learn low-dimensional features and reduce intra-domain feature redundancy, and the inter-domain network fuses different domain features and obtains a classification result through a neural network. Jiang Wangkui et al. proposed a framework for low probability of intercept radar signal waveform recognition. Specifically, the radar signal is transformed into a time-frequency image, the signal features are extracted by using a locally densely connected U-net, and finally the recognition task is achieved through a deep convolutional neural network. Zhang Weilong et al. designed an underwater acoustic signal recognition model based on a recurrent neural network and a convolutional neural network, improving the recognition accuracy and efficiency.
[0004] Existing research has designed effective network models for various wireless signals to extract intrinsic features, improving the efficiency and accuracy of signal analysis and recognition. However, existing work usually only relies on labeled samples for supervised training and does not fully exploit the intrinsic features of wireless signals. In addition, the trained network model has good performance on specific signals but is difficult to transfer to other related application scenarios and has poor generalization ability. Summary of the Invention
[0005] (1) Technical Problems to be Solved
[0006] The technical problem to be solved by the present invention is how to provide a method for learning the representation of wireless signals based on masked modeling to solve the problem that existing work usually only relies on labeled samples for supervised training and does not fully exploit the intrinsic features of wireless signals. In addition, the trained network model has good performance on specific signals but is difficult to transfer to other related application scenarios and has poor generalization ability.
[0007] (2) Technical Solutions
[0008] To solve the above technical problems, the present invention proposes a method for learning the representation of wireless signals based on masked modeling, which includes the following steps:
[0009] S1. Preprocess the collected wireless signals, convert them into a representation sequence input to the network model, and divide the signal sequence into two subsequences: past and future;
[0010] S2. Randomly mask some segments of the past sequence through a masking mechanism, input it into the embedding module and the Transformer encoder to extract the intrinsic features of the signal, and obtain an encoded sequence;
[0011] S3. Obtain a reconstructed sequence through linear transformation of the encoded sequence, and calculate the masked reconstruction loss;
[0012] S4. Map the future sequence to a high-dimensional space through embedding, and jointly input it into the Transformer decoder containing masked attention to model the temporal relationship of the signal, and output a decoded sequence;
[0013] S5. Obtain a prediction matrix through linear transformation of the decoded sequence, and calculate the signal prediction loss;
[0014] S6. Construct a joint loss containing the masked reconstruction loss and the signal prediction loss, and update the model parameters in reverse;
[0015] S7. Save the trained network model, and fine-tune the network model using the sample data of the downstream task for signal reconstruction, generation, prediction, and classification tasks.
[0016] (3) Advantageous Effects
[0017] The present invention proposes a wireless signal representation learning method based on masked modeling. The advantages and positive effects of the present invention are as follows: Compared with the prior art, the present invention can perform general representation learning of wireless signals without label information, and fully exploit the intrinsic features of the signals; Compared with the prior art, the present invention uses a masking mechanism to construct signal reconstruction and prediction tasks to pre-train a network model, which is applicable to downstream tasks such as signal prediction, classification, and denoising. The present invention can utilize a large amount of unlabeled signal data collected to pre-train a network model with a large scale and general feature extraction ability, and apply the pre-trained model to downstream tasks such as signal analysis, recognition, and prediction through techniques such as fine-tuning, pruning, and quantization, so as to improve the efficiency and performance of downstream tasks. Brief Description of the Drawings
[0018] Figure 1 It is a flowchart of a wireless signal representation learning method based on a masking mechanism according to the present invention;
[0019] Figure 2 It is a schematic diagram of a Transformer encoder provided by an embodiment of the present invention;
[0020] Figure 3 It is a schematic diagram of a self-attention mechanism provided by an embodiment of the present invention;
[0021] Figure 4 It is a schematic diagram of a Transformer decoder provided by an embodiment of the present invention;
[0022] Figure 5 It is a schematic diagram of masked multi-head attention provided by an embodiment of the present invention. Detailed Embodiments
[0023] To make the objectives, content, and advantages of the present invention clearer, the following further describes in detail the specific embodiments of the present invention with reference to the accompanying drawings and embodiments.
[0024] To address the above problems, the present invention provides a wireless signal representation learning method based on masked modeling. Using various wireless signals as the input of the network model, a part of the wireless signal is randomly masked through a masking mechanism, and the Transformer model is used to extract the feature information of the unmasked signal and reconstruct and predict the masked part, so as to achieve the representation learning of wireless signals without label information.
[0025] Furthermore, the wireless signal representation learning method based on masked modeling includes the following steps:
[0026] S1. Preprocess the collected wireless signals, convert them into a representation sequence for input to the network model, and divide the signal sequence into two subsequences: past and future;
[0027] S2. Randomly mask some segments of the past sequence through the masking mechanism, input it into the embedding module and the Transformer encoder to extract the intrinsic features of the signal, and obtain the encoded sequence;
[0028] S3. The encoded sequence undergoes a linear transformation to obtain the reconstructed sequence, and calculate the masked reconstruction loss;
[0029] S4. The future sequence is mapped to a high-dimensional space through embedding, and jointly input into the Transformer decoder with masked attention to model the temporal relationship of the signal, and output the decoded sequence;
[0030] S5. The decoded sequence undergoes a linear transformation to obtain the prediction matrix, and calculate the signal prediction loss;
[0031] S6. Construct a joint loss containing the masked reconstruction loss and the signal prediction loss, and update the model parameters backward;
[0032] S7. Save the trained network model, and fine-tune the network model using the sample data of the downstream task for tasks such as signal reconstruction, generation, prediction, and classification.
[0033] Furthermore, the specific steps of S1 are as follows:
[0034] S1.1. The received signal is first amplified, mixed, and low-pass filtered, and then sent to an analog-to-digital converter to sample the continuous-time signal and generate a discrete signal. Assuming that the total sampling time of the continuous signal is T and the discrete time length is L, the corresponding discrete signal is expressed as:
[0035] s[n] = s I [n] + js Q [n], n = 0, 1, …, L - 1,
[0036] where s I [n], s Q [n] are the in-phase and quadrature components of the discrete signal respectively.
[0037] S1.2. The discrete signal can be expressed as sequences of types such as IQ, amplitude and phase (AP), and Fourier transform (FT). The real-valued IQ sequence is:
[0038]
[0039] where, Correspondingly, the AP sequence of the signal is:
[0040]
[0041] where, The amplitude and phase of the i-th element are as follows:
[0042]
[0043] The FT sequence representation of the signal is:
[0044]
[0045] Among them, is the spectral amplitude vector of the signal, is the bispectrum of the signal, is the quad spectrum of the signal. The value corresponding to the k-th element in the FT vector is:
[0046]
[0047]
[0048]
[0049] It should be noted that the discrete signal can also be preprocessed into other representation forms, such as time-frequency diagrams, Cauchy value constellation diagrams, etc. Different masks and embedding methods are set to process the token sequences of the same dimension and input them into the network model.
[0050] S1.3. The signal representation sequence is divided into a past sequence and a future sequence in a certain proportion. w is the feature dimension of the representation sequence. The two subsequences are used as the inputs of the encoder and decoder respectively to perform the reconstruction and prediction tasks.
[0051] Furthermore, the specific steps of step S2 are as follows:
[0052] S2.1. Randomly replace some segments of the past sequences of types such as IQ, AP, and FT with zero values or noise. The mask sequence is expressed as:
[0053]
[0054] Among them, is a binary mask matrix, w is the feature dimension of the input series, and ⊙ is the Hadamard product of multiplying the elements at the corresponding positions. Considering the correlation of each feature dimension of the input sequence and the randomness of signal changes, randomly and uniformly draw L p *r index values without replacement from the discrete time set {0, 1..., L p -1}, and set the corresponding feature values to 0 or random noise. Here, r is the random mask ratio, and r ∈ [0, 1].
[0055] S2.2. The low-dimensional mask sequence is passed through a convolutional layer to obtain a high-dimensional embedding sequence One-dimensional convolution is used to model the local dependencies of the input vector, with a kernel size of 3, and patches are used to keep the length of the output after convolution unchanged.
[0056] S2.3. The masked embedding sequence is superimposed with absolute sinusoidal encoding to add positional information to the sequence, where L max is the maximum length of all signal samples. The Transformer structure supports variable-length sequence inputs.
[0057] S2.4. The masked embedding sequence with added positional information is modeled for global dependencies through the G-layer Transformer encoder, and the encoded sequence is output Each layer of the Transformer encoder mainly consists of two sub-blocks: the multi-head attention MSA(·) and the feed-forward network FFN(·), which are used to extract the temporal and spatial correlations of the signal. The formula for the Transformer encoder is:
[0058] e' = LN(Drop(MSA(e)) + e)
[0059] e” = LN(Drop(FFN(e')) + e')
[0060] where Drop and LN are the dropout and layer normalization operations respectively, which together with the residual connection operation improve the stability of the network. The formula for the MSA block is expressed as:
[0061] MSA(e) = Concat(head1, …, head H )W O
[0062] where Concat represents the concatenation of the feature dimensions of H sub-heads head, and W O is the linear transformation weight. The h-th sub-head head h is expressed as:
[0063]
[0064] where is the query matrix, is the key matrix, is the value matrix, and are the linear transformation matrices. is the learnable attention bias matrix, which is used to add relative position information to the higher-layer encoders. Usually
[0065] The forward feedback network is constructed by two fully connected layers and the ReLu activation function, and is used to extract the correlation between different channels. The dimension of the hidden layer is d mlp = 2d m 。
[0066] Further, the specific steps of step S3 are as follows:
[0067] S3.1. The high-dimensional encoded sequence Obtains a reconstructed sequence with the same dimension as the input through linear transformation
[0068] S3.2. Calculate the reconstructed mean square error of the masked part, that is:
[0069]
[0070] where n is the signal sample index for model training, B is the number of batch samples; m is the index of the mask value, M = L p *r. In the formula, the 2 below is the second norm, and the 2 above is the square
[0071] Further, the specific steps of step S4 are as follows:
[0072] S4.1. The future sequence Passes through the convolutional layer to model the local dependencies of the signal, and superimposes the absolute sinusoidal encoding to add position information to the sequence, obtaining a high-dimensional unmasked embedding sequence The convolutional kernel size is set to 1
[0073] S4.2. The unmasked embedding sequence passes through the G-layer Transformer decoder to obtain the decoded sequence. Each layer of the decoder contains masked multi-head attention, multi-head attention, and forward feedback network; the masked multi-head attention MMSA(·) models the temporal relationship of the sequence, specifically:
[0074] ue' = LN(Drop(MMSA(ue)) + ue)
[0075] MMSA(ue) = Concat(mhead1,…,mhead H ) W O
[0076] where Concat represents the concatenation of the feature dimensions of H sub-heads mhead, W O is the linear transformation weight; the h-th sub-head mhead h is expressed as:
[0077]
[0078] where, is the query matrix, is the key matrix, is the value matrix, and is the linear transformation matrix; is a learnable attention bias matrix for adding relative position information; Mask(·) is a weight matrix of size with the upper triangular elements set to zero to mask future information in the sequence;
[0079] The output ue' of the masked multi - head attention and the encoded sequence are input into the multi - head attention sub - block to further model the forward and backward dependencies of the signal, and its formula is:
[0080]
[0081] where Drop and LN are dropout and layer normalization operations respectively, and they, together with the residual connection operation, improve the stability of the network; The MSA block formula is expressed as:
[0082]
[0083] where Concat represents the concatenation of the feature dimensions of H sub - heads chead, W O is the linear transformation weight; The h - th sub - head chead h is expressed as:
[0084]
[0085] where, is the query matrix, is the key matrix, is the value matrix, and is the linear transformation matrix;
[0086] The output matrix of the multi - head attention passes through a forward feedback layer composed of two fully - connected layers and the ReLu activation function to select and process the features, and the dimension of the hidden layer is d mlp = 2d m .
[0087] Furthermore, the specific steps of step S5 are as follows:
[0088] S5.1. The decoded sequence undergoes a linear transformation to obtain a predicted sequence with the same dimension as the input
[0089] S5.2. Calculate the prediction error of the unmasked future sequence x f , that is:
[0090]
[0091] Among them, n is the signal sample index in the training set, and B is the number of batch samples during training; s(·) represents the right shift of the original sequence, and the input at each moment corresponds to the output at the next moment.
[0092] Further, the specific steps of step S6 are as follows:
[0093] Construct a joint loss based on the calculated reconstruction loss and prediction loss, that is:
[0094]
[0095] Among them, α is the loss weight. Set hyperparameters such as the optimizer, learning rate, and number of epochs of the model, and use the existing samples to train the network model.
[0096] Example 1:
[0097] The wireless signal representation learning method based on masked modeling provided by the present invention divides various wireless signals represented in forms such as IQ, AP, and FT into past sequences and future sequences. The past sequence is randomly masked and input into the Transformer encoder to obtain an encoded sequence, and the reconstructed sequence is output after linear transformation; the future sequence is input into the Transformer decoder, and the predicted sequence is output after linear transformation; construct a joint loss constructed by weighted summation of the reconstruction loss and the prediction loss, and use a large number of unlabeled signal samples to update the network model to fully extract the internal features of the signal for various downstream tasks.
[0098] The representation learning method of the present invention will be described in detail below with reference to the accompanying drawings. As Figure 1 shown, taking the wireless signal represented as an IQ sequence as an example, a masked modeling wireless signal representation learning method provided by an embodiment of the present invention includes the following steps:
[0099] S1: Preprocess the collected wireless signal, convert it into an IQ sequence input to the network model, and divide the IQ sequence into two subsequences: past and future;
[0100] S2: Randomly mask some segments of the past sequence through a masking mechanism, input it into the embedding module and the Transformer encoder to extract the internal features of the signal, and obtain an encoded sequence;
[0101] S3: The encoded sequence undergoes linear transformation to obtain a reconstructed sequence, and calculate the masked reconstruction loss;
[0102] S4: The future sequence is embedded and mapped into a high-dimensional space, and input into the Transformer decoder containing masked attention together with the encoded sequence to model the temporal relationship of the signal, and output a decoded sequence;
[0103] S5: The decoded sequence is linearly transformed to obtain a prediction matrix, and the signal prediction loss is calculated;
[0104] S6: Construct a joint loss that includes the masked reconstruction loss and the signal prediction loss, and update the model parameters in reverse;
[0105] S7: Save the trained network model, and fine-tune the network model using the sample data of the downstream task for tasks such as signal reconstruction, generation, prediction, and classification.
[0106] Furthermore, the specific steps of step S1 are as follows:
[0107] S1.1: The received signal is first amplified, mixed, and low-pass filtered, and then sent to an analog-to-digital converter to sample the continuous-time signal and generate a discrete signal. Assuming that the total sampling time of the continuous signal is T and the discrete time length is L, the corresponding discrete signal is expressed as:
[0108] s[n] = s I [n] + js Q [n], n = 0, 1, …, L - 1,
[0109] where s I [n], s Q [n] are the in-phase and quadrature IQ components of the discrete signal, respectively.
[0110] S1.2: The real-valued IQ sequence corresponding to the discrete signal is:
[0111]
[0112] where,
[0113] It should be noted that the discrete signal can also be expressed in other forms, such as AP sequence, FT sequence, time-frequency diagram, Cauchy value constellation diagram, etc. Different masks and embedding methods are set to process it into a token sequence of the same dimension and input it into the network model. The AP sequence of the signal is:
[0114]
[0115] where, The amplitude and phase of the i-th element are:
[0116]
[0117] The FT sequence of the signal is expressed as:
[0118]
[0119] where, is the spectral amplitude vector of the signal, is the bispectrum of the signal, is the quad spectrum of the signal. The value corresponding to the k-th element in the FT vector is:
[0120]
[0121]
[0122]
[0123] S1.2. Divide the IQ sequence into a past sequence and a future sequence in a certain proportion over time. w represents the characteristic dimension of the sequence. The two subsequences are used as the inputs of the encoder and decoder respectively to perform reconstruction and prediction tasks.
[0124] Furthermore, the specific steps of step S2 are as follows:
[0125] S2.1. Randomly replace some segments of the IQ past sequence with zero values or noise. The mask sequence is expressed as:
[0126]
[0127] where is a binary mask matrix, w is the characteristic dimension of the input series, and ⊙ is the Hadamard product of multiplying the elements at the corresponding positions. Considering the correlation of each characteristic dimension of the input sequence and the randomness of signal changes, randomly and uniformly draw L p *r index values without replacement from the discrete time set {0, 1…, L p -1}, and set the corresponding characteristic values to 0 or random noise. Here, r is the random mask ratio, and r ∈ [0, 1].
[0128] S2.2. The low-dimensional mask sequence is passed through a convolutional layer to obtain a high-dimensional embedded sequence One-dimensional convolution is used to model the local dependencies of the input vector. The kernel size is 3, and patches are used to keep the length of the output after convolution unchanged.
[0129] S2.3. The masked embedded sequence is superimposed with absolute sinusoidal encoding to add position information to the sequence. L max is the maximum length of all signal samples. The Transformer structure supports variable-length sequence inputs.
[0130] The masked embedded sequence after adding position information is modeled by the Transformer encoder of G layers to model the global dependencies, and the encoded sequence is output Each layer of the Transformer encoder mainly consists of two sub-blocks, namely the multi-head attention MSA(·) and the feed-forward network FFN(·), which are used to extract the temporal and spatial correlations of the signal. As Figure 2 shown, the formula of the Transformer encoder is:
[0131] e' = LN(Drop(MSA(e)) + e)
[0132] e” = LN(Drop(FFN(e')) + e')
[0133] where Drop and LN are the dropout and layer normalization operations respectively, and they, together with the residual connection operation, improve the stability of the network. The formula of the MSA block is expressed as:
[0134] MSA(e) = Concat(head1, …, head H )W O
[0135] where Concat represents the concatenation of the feature dimensions of H sub-heads head, and W O is the linear transformation weight. As Figure 2 shown, the h-th sub-head head h is expressed as:
[0136]
[0137] where, is the query matrix, is the key matrix, is the value matrix, and are the linear transformation matrices. is the learnable attention bias matrix, which is used to add relative position information to the high-level encoder.
[0138] The feed-forward network is constructed by two fully-connected layers and the ReLu activation function, which is used to extract the correlations between different channels, and the dimension of the hidden layer is d mlp = 2d m .
[0139] Furthermore, the specific step S3 is as follows:
[0140] S3.1. The high-dimensional encoded sequence is linearly transformed to obtain a reconstructed sequence
[0141] with the same dimension as the input.
[0142]
[0143] Among them, m is the index of the mask value, and B is the batch size during the training process.
[0144] Furthermore, the specific step S4 is as follows:
[0145] S4.1. Future sequence After embedding and position encoding operations similar to (S2.2) and (S2.3), the convolutional kernel size is set to 1.
[0146] S4.2. Unmasked embedded sequence After passing through the G-layer Transformer decoder as shown in Figure 4 , a decoded sequence is obtained. Each layer of the decoder contains masked multi-head attention, multi-head attention, and a feed-forward network. As shown in Figure 5 , the masked multi-head attention MMSA(·) is used to model the temporal relationship of the sequence. Specifically:
[0147] ue' = LN(Drop(MMSA(ue)) + ue)
[0148] MMSA(ue) = Concat(mhead1, …, mhead H )W O
[0149] Among them, Concat represents the concatenation of the feature dimensions of H sub-heads mhead, and W O is the linear transformation weight. The h-th sub-head mhead h is:
[0150]
[0151] Among them, is the query matrix, is the key matrix, is the value matrix, and are the linear transformation matrices. is a learnable attention bias matrix for adding relative position information. Mask(·) is to set the upper triangular elements in the weight matrix of size to zero to mask the future information in the sequence.
[0152] The output ue' of the masked multi-head attention is input into the multi-head attention sub-block described in S2.4 to further model the forward and backward dependencies of the signal. The difference from S2.4 lies in the input of the attention. The formula here is:
[0153]
[0154] Among them, is a query matrix, is a key matrix, is a value matrix, and are linear transformation matrices.
[0155] The output matrix of the multi-head attention passes through the forward feedback network described in S2.4 to select and process the features.
[0156] Furthermore, the specific steps of S5 are as follows:
[0157] S5.1. The decoded sequence undergoes a linear transformation to obtain a predicted sequence with the same dimension as the input
[0158] S5.2. Calculate the prediction error of the unmasked future sequence x f That is:
[0159]
[0160] where s(·) represents the right shift of the original sequence, and the input at each moment corresponds to the output at the next moment.
[0161] Furthermore, the specific steps of S6 are as follows:
[0162] According to the calculated reconstruction loss and prediction loss, construct a joint loss reality, that is:
[0163]
[0164] where α is the loss weight. Set hyperparameters such as the optimizer, learning rate, and number of epochs of the model, and use the existing samples to train the network model.
[0165] In summary, the advantages and positive effects of the present invention are as follows: Compared with the prior art, the present invention can perform general representation learning of wireless signals without label information and fully exploit the intrinsic features of the signals; compared with the prior art, the present invention uses a masking mechanism to construct signal reconstruction and prediction tasks to pre-train the network model, which is applicable to downstream tasks such as signal prediction, classification, and denoising. The present invention can use a large amount of unlabeled signal data collected to pre-train a network model with a large scale and general feature extraction ability, and apply the pre-trained model to downstream tasks such as signal analysis, recognition, and prediction through techniques such as fine-tuning, pruning, and quantization, improving the efficiency and performance of downstream tasks.
[0166] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and deformations can be made, and these improvements and deformations should also be regarded as the protection scope of the present invention.
Claims
1. A method for learning the representation of wireless signals based on masked modeling, characterized in that, The method includes the following steps: S1. Preprocess the collected wireless signals, convert them into a representation sequence for input to the network model, and divide the signal sequence into two subsequences: past and future; S2. Randomly mask some segments of the past sequence through a masking mechanism, input it into the embedding module and the Transformer encoder to extract the intrinsic features of the signal, and obtain an encoded sequence; S3. The encoded sequence undergoes a linear transformation to obtain a reconstructed sequence, and calculate the masked reconstruction loss; S4. The future sequence is mapped to a high-dimensional space through embedding, and jointly input into the Transformer decoder containing masked attention to model the temporal relationship of the signal, and output a decoded sequence; S5. The decoded sequence undergoes a linear transformation to obtain a prediction matrix, and calculate the signal prediction loss; S6. Construct a joint loss containing the masked reconstruction loss and the signal prediction loss, and update the model parameters in reverse; S7. Save the trained network model, and fine-tune the network model using the sample data of the downstream task for signal reconstruction, generation, prediction, and classification tasks.
2. The method for learning wireless signal representation based on mask modeling according to claim 1, wherein The specific steps of step S1 include: S1.
1. First, the received signal is amplified, mixed, and low-pass filtered, and then sent to an analog-to-digital converter to sample the continuous-time signal and generate a discrete signal. Assuming that the total sampling time of the continuous signal is T and the discrete time length is L, the corresponding discrete signal is expressed as: s[n] = s I [n] + js Q [n], n = 0, 1, …, L - 1; where s I [n], s Q [n] are the in-phase and quadrature components of the discrete signal, respectively; S1.
2. Represent the discrete signal as sequences of IQ, amplitude and phase AP, and spectral transform FT types; The real-valued IQ sequence is: Among them, Correspondingly, the AP sequence of the signal is: Among them, The amplitude and phase of the i-th element are as follows: The FT sequence of the signal is expressed as: Among them, is the spectral amplitude vector of the signal, is the bispectrum of the signal, is the quad spectrum of the signal; The value corresponding to the k-th element in the FT vector is: S1.
3. The signal representation sequence is divided into a past sequence and a future sequence where w is the feature dimension of the representation sequence. The two subsequences are used as the inputs to the encoder and decoder respectively to perform reconstruction and prediction tasks.
3. The method for learning wireless signal representation based on masked modeling according to claim 2, wherein, The specific steps of step S2 include the following steps: S2.
1. Randomly replace some segments of the past sequences of IQ, AP, and FT types with zero values or noise; the masked sequence is expressed as: Among them, is a binary mask matrix, w is the feature dimension of the input sequence, and ⊙ is the Hadamard product of multiplying the elements at the corresponding positions; considering the correlation of each feature dimension of the input sequence and the randomness of signal changes, from the discrete time set {0, 1..., L p - 1}, L p *r index values are randomly and uniformly drawn without replacement, and the corresponding feature values are set to 0 or random noise; here r is the random mask ratio, r ∈ [0, 1]; S2.
2. The low-dimensional masked sequence is passed through a convolutional layer to obtain a high-dimensional embedded sequence One-dimensional convolution is used to model the local dependencies of the input vectors, and patches are used to keep the length of the output after convolution unchanged; d m Is the feature dimension of the embedded sequence e; S2.
3. Mask Embedding Sequence Superimposed with Absolute Sinusoidal Encoding to add position information to the sequence, L max is the maximum length of all signal samples; S2.
4. After the masked embedding sequence with the added location information is modeled for global dependencies by the G-layer Transformer encoder, an encoded sequence is output. Each layer of the Transformer encoder includes two sub-blocks: the multi-head self-attention (MSA(·)) and the feed-forward network (FFN(·)), which are used to extract the temporal and spatial correlations of the signals. The formula for the Transformer encoder is: e' = LN(Drop(MSA(e)) + e) e” = LN(Drop(FFN(e')) + e') Where Drop and LN are the dropout and layer normalization operations respectively, which together with the residual connection operation improve the stability of the network; the MSA block formula is expressed as: MSA(e) = Concat(head1, …, head H )W O Among them, Concat represents the concatenation of the feature dimensions of H sub - heads head, and W O is the weight of the linear transformation; the h - th sub - head head h is expressed as: Among them, is the query matrix, is the key matrix, is the value matrix, and are linear transformation matrices; is a learnable attention bias matrix for adding relative position information to the high-level encoder; The forward feedback network is constructed by two fully connected layers and the ReLu activation function, and is used to extract the correlation between different channels. The dimension of the hidden layer is d mlp = 2d m .
4. The method for learning wireless signal representation based on mask modeling according to claim 3, wherein, In S2.2, the kernel size of the one-dimensional convolution is 3.
5. The method for learning wireless signal representation based on masked modeling according to claim 3, characterized in that In the S2.4, 6. The method for learning the representation of wireless signals based on masked modeling according to any one of claims 3-5, characterized in that The specific steps of step S3 include the following steps: S3.
1. High-dimensional coding sequence Obtain a reconstructed sequence with the same dimension as the input through linear transformation S3.
2. Calculate the reconstructed mean square error of the masked part, that is: where n is the index of the signal sample for model training, B is the number of batch samples; m is the index of the mask value, and M = L p *r.
7. The method for learning wireless signal representation based on mask modeling according to claim 6, wherein The specific steps of step S4 include the following steps: S4.1, Future Sequence Model the local dependencies of the signal through the convolutional layer, and superimpose the absolute sinusoidal encoding to add position information to the sequence, obtaining a high-dimensional unmasked embedded sequence S4.
2. The unmasked embedded sequence passes through the G-layer Transformer decoder to obtain a decoded sequence. Each layer of the decoder contains masked multi-head attention, multi-head attention, and a feed-forward network; the masked multi-head attention MMSA(·) models the temporal relationship of the sequence, specifically: ue' = LN(Drop(MMSA(ue)) + ue) MMSA(ue) = Concat(mhead1, …, mhead H )W O Among them, Concat represents the feature dimension concatenation of H sub - heads mhead, and W O is the weight of the linear transformation; the h - th sub - head mhead h is expressed as: Among them, is the query matrix, is the key matrix, is the value matrix, and are linear transformation matrices; is a learnable attention bias matrix that adds relative position information; Mask(·) is a weight matrix of size with the upper triangular elements set to zero to mask future information in the sequence; The output ue' of the masked multi-head attention and the encoded sequence are input into the multi-head attention sub-block to further model the forward and backward dependencies of the signals, and its formula is: Where Drop and LN are the dropout and layer normalization operations respectively, which together with the residual connection operation improve the stability of the network; the MSA block formula is expressed as: Among them, Concat represents the feature dimension concatenation of H sub - heads chead, and W O is the weight of the linear transformation; the h - th sub - head chead h is expressed as: Among them, is the query matrix, is the key matrix, is the value matrix, and is the linear transformation matrix; The output matrix of the multi-head attention passes through a forward feedback layer composed of two fully connected layers and the ReLu activation function to select and process features, and the dimension of the hidden layer is d mlp = 2d m .
8. The method for learning wireless signal representation based on masked modeling according to claim 7, characterized in that, In S4.1, the convolution kernel size is set to 1.
9. The method for learning wireless signal representation based on mask modeling according to claim 7, characterized in that, The specific steps of step S5 include: S5.
1. The decoded sequence is linearly transformed to obtain a predicted sequence with the same dimension as the input S5.
2. Calculate the prediction error of the unmasked future sequence x f That is: Where n is the signal sample index in the training set, and B is the number of batch samples during the training process; s(·) represents the right shift of the original sequence, and the input at each moment corresponds to the output at the next moment.
10. The method for learning wireless signal representation based on masked modeling according to claim 9, characterized in that, The specific steps of step S6 include: Construct a joint loss according to the calculated reconstruction loss and prediction loss, that is: Where α is the loss weight, set the optimizer, learning rate, and number of epochs hyperparameters of the model, and train the network model using the existing samples.
Citation Information
Patent Citations
Multi-element time sequence anomaly detection method for intelligent Internet of Things system
CN116663613A
Power grid time sequence data decoupling self-supervision pre-training method and system
CN116776228A