Radar signal feature representation method based on self-supervised contrastive learning

By using self-supervised contrast learning and dual-view data enhancement technology in radar signal processing, the problem of difficulty in feature extraction in complex signal environments is solved, efficient and accurate feature representation is achieved, and the understanding ability and application flexibility of radar signals are enhanced.

CN119416150BActive Publication Date: 2025-05-02NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411497212.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-25
Publication Date
2025-05-02
Estimated Expiration
2044-10-25

AI Technical Summary

Technical Problem

The prior art is difficult to accurately extract the effective characteristics of radar signals in complex signal environments, especially in the case of dynamic changes and severe noise interference, and lack of understanding of the global characteristics of the signal, resulting in insufficient generalization capabilities.

Method used

Using a radar signal feature representation method based on self-supervised contrast learning, an efficient and accurate feature extraction pre-training model is constructed through dual-view data augmentation technology and pre-training scheme to generate a feature representation that integrates global context information.

Benefits of technology

It effectively and accurately captures the timing dependence, contextual relationship and cross-view consistency of radar signals in complex environments, significantly enhancing the ability to understand radar signals, and improving the flexibility and practicality of signal processing and analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119416150B_ABST
    Figure CN119416150B_ABST
Patent Text Reader

Abstract

The present invention discloses a radar signal feature representation method based on self-supervised contrastive learning, comprising: obtaining radar signal time series data and performing preprocessing; enhancing the data to obtain strong enhanced views and weak enhanced views, inputting the views into an encoder to obtain a potential embedding representation, inputting the views into a time contrast module to perform a task of predicting time steps across views, outputting a corresponding context vector, and calculating a time contrast loss; inputting the context vectors corresponding to the strong enhanced views and the weak enhanced views into a context contrast module to perform self-supervised contrastive learning and calculate a context contrast loss; integrating to obtain a total loss function and training a model composed of the time contrast module and the context contrast module to obtain a radar signal time series feature representation model, which can efficiently and accurately capture the temporal dependency, contextual relationship and consistency across views in the radar signal time series, thereby enhancing the ability to understand radar signals and more flexibly coping with complex radar application scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of radar signal feature representation, and in particular relates to a radar signal feature representation method based on self-supervised contrast learning. Background Art

[0002] With the continuous improvement of radar system performance and the rapid enhancement of data processing capabilities, the complexity and diversity of radar signals are also increasing, resulting in many challenges such as diverse radar signal data sources and complex environments, scarce label data, severe noise interference, difficulty in meeting real-time processing requirements, insufficient feature generalization capabilities, and high model complexity.

[0003] The traditional manual feature extraction process relies on expert experience for feature design, but in complex signal environments, it is easy to ignore the dynamic changes and nonlinear characteristics of the signal. Although deep learning technology can automatically learn features, it often requires a large amount of labeled data and is sensitive to noise and interference. In addition, many existing methods lack an understanding of the global characteristics of the signal, resulting in insufficient generalization capabilities in diverse application scenarios. These limitations make it difficult to accurately extract effective features in complex environments.

[0004] Unsupervised (or self-supervised) contrastive learning has become an important and effective technique for learning valuable representations from time series data without relying on labels. This approach learns more abstract and hierarchical feature representations in radar signals by prompting the model to generate similar representations for different views or enhancements of the same data instance while pushing the representations of different instances further apart.

[0005] Modern radar systems often have many sensing and control components, and the processing and representation of radar signals are crucial for target recognition and detection. If the system misjudges and false alarms due to unreasonable or inefficient feature representation, it will cause a huge waste of resources. Therefore, unlike feature extraction technologies for other types of data, the feature processing of radar systems must be able to meet the requirements of real-time and accuracy, have good adaptability to target changes in dynamic environments, and avoid misjudgments caused by noise interference. This ensures that the radar system can work effectively in various complex situations and ensure the successful implementation of the mission. Summary of the invention

[0006] The technical problem to be solved by the present invention is to provide a radar signal feature representation method based on self-supervised contrastive learning in view of the deficiencies of the above-mentioned prior art. Through effective dual-view data enhancement technology and pre-training scheme, an efficient and accurate feature extraction pre-training model based on context consistency is realized, ensuring that the feature representation generated by the model can be used for subsequent downstream tasks such as radar signal recognition and classification.

[0007] In order to achieve the above technical objectives, the technical solution adopted by the present invention is:

[0008] The radar signal feature representation method based on self-supervised contrastive learning includes:

[0009] Step S1, obtaining radar signal time series data and preprocessing it to construct an unlabeled data set;

[0010] Step S2, using strong enhancement and weak enhancement to perform data enhancement on radar time series samples in the unlabeled data set to obtain a strong enhancement view and a weak enhancement view;

[0011] Step S3: input the strongly enhanced view and the weakly enhanced view into the encoder respectively to obtain corresponding potential embedding representations;

[0012] Step S4: input the latent embedding representations corresponding to the strong enhanced view and the weak enhanced view into the temporal contrast module to perform the task of cross-view prediction time step, output the context vectors corresponding to the strong enhanced view and the weak enhanced view, respectively, and calculate the temporal contrast loss corresponding to the strong enhanced view and the weak enhanced view, respectively;

[0013] Step S5: input the context vectors corresponding to the strong enhanced view and the weak enhanced view into a context contrast module to perform self-supervised contrast learning and calculate the context contrast loss;

[0014] Step S6, integrating the time contrast loss and the context contrast loss to obtain a total loss function, and training the model composed of the time contrast module and the context contrast module based on the total loss function to obtain a radar signal time series feature representation model, so as to generate a more comprehensive and robust feature representation for the radar signal time series that integrates the global context information.

[0015] To optimize the above technical solutions, the specific measures taken also include:

[0016] The above step S1 includes:

[0017] S101, obtaining radar signal time series data from the radar system, including timestamps collected by different sensors and object distance, speed, and intensity characteristics;

[0018] S102, denoising and filtering the acquired radar signal time series data;

[0019] S103, standardizing and normalizing the radar signal time series data;

[0020] S104, setting the window size W and the step size S, dividing the radar signal time series data into multiple segments of fixed length, each segment being used as an independent radar time series sample, to form an unlabeled data set.

[0021] The above step S2 includes: for the radar time series sample x, using the dithering method and the arrangement method as a strong enhancement strategy to generate a strong enhanced view x s ; For radar time series sample x, use scaling method and jittering method as weak enhancement strategy to generate weak enhancement view x w ;

[0022] The jitter adding method adds random noise to the original time series data to simulate the subtle changes that may occur in the actual acquisition process, and then adds it to the data value at each time point; the arrangement method first divides the time series into multiple subsequence fragments, fixes the size of these subsequence fragments, shuffles these subsequence fragments in a random order, keeps the order of data in the fragments unchanged, and finally recombine the shuffled fragments to form a new time series; the scaling method multiplies the entire time series by a random scaling factor to achieve amplification or reduction of the amplitude of the time series data.

[0023] The above step S3 uses the encoder f enc Map the input augmented view x′ to a d-dimensional latent embedding representation z=f enc (x′); where z=[z1,z2,…z T ], T is the total time step, z i It represents the d-dimensional embedding vector obtained after encoding the i-th time step, where d is the feature length.

[0024] The temporal contrast module in step S4 above adopts an autoregressive model f for the potential embedding representation z ar The embedding representation z of the time steps less than time t ≤t Summarized as context vector c t =f ar (z ≤t ),c t ∈R h , where h is f ar The hidden dimension of the context vector c t To complete the task of cross-view prediction time step to learn the intrinsic laws and characteristics of radar signal time series, the process is as follows:

[0025] First, add an h-dimensional token c, whose state in the output acts as a context vector that can represent the entire time series information. ≤t Input to a linear projection layer W Tran , output feature vector z′=W Tran (z ≤t );

[0026] The context vector c is then appended to the feature vector z′, so that the input feature becomes ψ0=[c;z′], where the subscript 0 indicates that ψ0 is the input feature of the first layer of Transformer;

[0027] Then ψ0 is passed to the Transformer layer. Each layer of Transformer contains the calculation of the multi-head attention machine layer and the multi-layer perceptron. After each layer of Transformer, the input is updated to a new feature vector. The formula is as follows:

[0028]

[0029] Among them, MHA is the multi-head attention layer, Norm is the normalization layer, MLP is the multi-layer perceptron, l is the number of layers where the input features are calculated in the Transformer, L is the maximum number of Transformer layers that can be stacked, and ψ l Represents the feature vector input to the lth layer transformer;

[0030] After ψ0 passes through L layers of Transformer, the network output is a combined vector ψ containing the context vector and other input features L , extract the first part of the combined vector, which is the updated version of the context vector c that was originally input after passing through the Transformer network And use it as the updated context vector c t , to complete the task of cross-view prediction time steps in order to learn the intrinsic laws and characteristics of radar signal time series.

[0031] In the above step S4, the enhanced view The corresponding temporal contrast loss is:

[0032]

[0033] Weakly enhanced view The corresponding temporal contrast loss is:

[0034]

[0035] Among them, K is the maximum step size for predicting future time steps; log is the logarithmic function; W k is a linear mapping function; N t,k Represents a set of negative samples in a batch, including samples other than positive samples; is the potential representation of the weakly enhanced view of the current sample x to be learned at the future time step t+k, used as a positive sample, It is the potential representation of the strongly enhanced view of the current sample x to be learned at the future time step t+k, which is used as a positive sample; A strongly enhanced view latent representation of all other samples in the same batch except the positive samples, used as negative samples; A weakly enhanced view latent representation representing all other samples in the same batch except the positive samples is used as negative samples.

[0036] In the above-mentioned step S5, the context contrast module first uses a nonlinear projection head to apply nonlinear transformation to the context vectors corresponding to the strong enhanced view and the weak enhanced view respectively, maps the context vectors corresponding to the strong enhanced view and the weak enhanced view output by the temporal contrast module in step S4 to the context contrast space, and performs self-supervised contrast learning based on the contrast space, as follows:

[0037] For a given batch of N input samples, after step S4, each sample generates two context vectors from different enhanced views, so there are 2N context vectors. For the context vector of one of the views Will Represented as the context vector of another augmented view from the same input sample The positive sample of As positive pairs, the remaining (2N-2) context vectors of other input samples in the same batch are considered Negative samples of It can form (2N-2) negative pairs with its negative samples.

[0038] The context contrast loss in step S5 above is:

[0039]

[0040] in is the context contrast loss;

[0041] is the contrast loss function between the positive sample pairs of sample x, used to measure and similarities between;

[0042] They represent the loss between the positive sample pairs of the weakly enhanced view and the strongly enhanced view of sample x, and the loss between the positive sample pairs of the strongly enhanced view and the weakly enhanced view, respectively;

[0043] It is the total contextual contrast loss function for all samples in a batch;

[0044] Represents normalization and The dot product of Represents normalization and The dot product of

[0045] 1 m≠i ∈{0,1} is the indicator function; τ is the temperature parameter.

[0046] The total loss function in step S6 above is:

[0047]

[0048] in is the temporal contrast loss corresponding to the strongly enhanced view; is the temporal contrast loss corresponding to the weakly enhanced view; is the context contrastive loss; λ1 and λ2 are fixed scalar hyperparameters.

[0049] In the above step S6, batch processing is performed on the data when training the model. The data is input into the model in batches, and the model parameters are continuously updated through back propagation and Adam optimizer to minimize the value of the total loss function. After the training process is completed, the weight of the model is saved for subsequent supervised learning or downstream tasks.

[0050] The present invention has the following beneficial effects:

[0051] The present invention utilizes the pre-training task of self-supervised contrastive learning to learn highly abstract and robust feature representations of radar signal time series, which can efficiently and accurately capture the timing dependencies, contextual relationships, and cross-view consistency in radar signal time series, thereby enhancing the ability to understand radar signals and more flexibly responding to complex radar application scenarios.

[0052] The present invention is based on the feature extraction of echo signals of ground radars. By using the pre-training task of self-supervised contrastive learning, the present invention can learn highly abstract and robust feature representations of radar signal time series. This method efficiently and accurately captures the timing dependencies, contextual relationships, and cross-view consistency in radar signals, thereby significantly enhancing the ability to understand radar signals. This enables more convenient and flexible signal processing and analysis in complex radar application scenarios, improving the practicality and efficiency of the application. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 is a flow chart of the steps of the present invention;

[0054] Figure 2 A schematic diagram of an encoder architecture provided by an embodiment of the present invention;

[0055] Figure 3 A schematic diagram of a time comparison module provided by an embodiment of the present invention;

[0056] Figure 4 A schematic diagram of a context comparison module provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0057] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0058] Although the steps in the present invention are arranged with numbers, they are not used to limit the order of the steps. Unless the order of the steps is clearly stated or the execution of a certain step requires other steps as a basis, the relative order of the steps can be adjusted. It is understood that the term "and / or" used in this article involves and covers any and all possible combinations of one or more of the associated listed items.

[0059] like Figure 1-4 As shown, the radar signal feature representation method based on self-supervised contrast learning of the present invention includes:

[0060] Step S1, obtaining radar signal time series data and preprocessing it to construct a labeled or unlabeled data set;

[0061] Step S2: generating two different but related views of the input radar signal time series data based on the strong enhancement and the weak enhancement;

[0062] Step S3: Input the two enhanced views into the encoder respectively to obtain high-dimensional (d-dimensional) potential embedding representations;

[0063] Step S4: Input the embedding representations of the two views obtained in step S3 into the time contrast module to perform the future time step prediction task and calculate the time contrast loss. The module outputs the context vectors of the two views.

[0064] Step S5, input the two context vectors obtained in step S4 into the context comparison module for comparison learning and calculate the context comparison loss;

[0065] Step S6, integrating the three loss functions to obtain a total loss function, and training the model in stages based on the loss function, and the model after training is a radar signal time series feature representation model;

[0066] In the embodiment, the step S1 acquires the original radar signal time series from the radar system, performs preprocessing such as denoising and standardization on the acquired original data, and constructs an unlabeled data set for the subsequent feature extraction model pre-training process, including the following sub-steps:

[0067] S101. Collect raw radar signal time series data from the radar system, including timestamps and relevant features such as object distance and speed. Convert the data collected by different sensors into a two-dimensional matrix or tensor form. Each row represents data at a different time step, and each column represents a different feature (distance, speed, intensity, etc.).

[0068] S102. The acquired original radar signal usually carries noise, and denoising and filtering operations are used to eliminate noise and interference. Kalman filtering is used to smooth the time series signal and eliminate random noise and instantaneous interference. Low-pass filtering is applied to remove high-frequency noise and retain the low-frequency signal part to extract the motion information of the target object. If it is necessary to remove the low-frequency trend, high-pass filtering is used.

[0069] S103, standardize and normalize the radar signal time series. Subtract the mean value of each feature and divide it by the standard deviation so that all features have the same distribution, as shown in formula (1). Then, scale the data to the range of [0,1] or [-1,1], as shown in formula (2), to ensure numerical stability and prevent scale differences between features from affecting model performance.

[0070]

[0071] S104, select a suitable window size W and step size S, divide the time series into multiple fixed-length segments, and each segment is used as an independent training sample. The windows can overlap partially, and the step size is set to 50% of the window size. After the above processing, the data will be input into the model as a fixed-size window sample, and each sample is a time series segment after denoising, filtering, standardization and normalization.

[0072] In the embodiment, the step S2 uses two methods to perform data enhancement on the radar signal time series samples to obtain a strong enhancement view and a weak enhancement view, including the following sub-steps:

[0073] S201. For a given input radar signal time series sample x, a strong enhancement view x is generated using a jittering and permutation method as a strong enhancement strategy. s , where x s ~T s , T s It is a strong data enhancement method.

[0074] For the jitter method, tiny random noise is added to the original time series data to simulate the subtle changes that may occur in the actual acquisition process. These noises are random values ​​sampled from the normal distribution and then added to the data value at each time point. For the permutation method, the time series is first divided into multiple subsequence fragments, and the sizes of these fragments are fixed. These subsequence fragments are shuffled in a certain random order, keeping the order of the data in the fragments unchanged, and finally the shuffled fragments are recombined to form a new time series.

[0075] S202: for a given input radar signal time series sample x, use a scaling and jittering method as a weak enhancement strategy to generate a weakly enhanced view x w , where x w ~T w , T w It is a weak data enhancement method.

[0076] For the scaling method, the entire time series is multiplied by a random scaling factor (a value between 0.5 and 2) to achieve the amplification or reduction of the amplitude of the time series data. The jittering method is the same as the strong enhancement method.

[0077] In the embodiment, the step S3 inputs the strong enhanced view and the weak enhanced view of the original time series sample into the encoder to obtain a potential embedded representation as the input of the subsequent time contrast module, including the following sub-steps:

[0078] S301, passing these enhanced views to the encoder f enc To extract their high-dimensional potential representations. The encoder f enc It has a 3-block convolutional architecture. Each convolutional block contains a one-dimensional convolutional layer to extract local features.

[0079] The convolution kernel operates on the input in a sliding window manner to capture local spatial information. The convolution kernel size is {8, 5, 3}. The batch normalization layer is connected after the convolution block to accelerate convergence and improve generalization ability. Finally, the ReLU activation function is used to introduce nonlinearity.

[0080] S302: For the input enhanced time series sample x, the encoder maps x' to a high-dimensional potential representation z = f enc (x′). Where z=[z1,z2,…z T ], T is the total time step, z i ∈R d , z i represents the embedded representation obtained after encoding the i-th time step, where d is the feature length. s represents a strongly enhanced view, z wrepresents the weakly enhanced view, and then these two views are input into the temporal contrast module.

[0081] In the embodiment, step S4 inputs the embedding obtained in the previous step into the time comparison module, which learns more robust time features by constructing a cross-view time step prediction task, including the following sub-steps:

[0082] S401, the time comparison module, aims to learn the intrinsic laws and highly abstract features of radar signal time series.

[0083] For the latent representation z obtained in step S3, the autoregressive model f ar All z ≤t (i.e., the embedding representation of the time step less than time t) is summarized as the context vector c t =f ar (z ≤t ),c t ∈R h , c t is an h-dimensional vector, where h is f ar The hidden dimension of .

[0084] Specifically, Transformer is used as the autoregressive model because of its efficiency and speed. It mainly consists of Multi-Head Attention (MHA) and Multi-Layer Perceptron (MLP) blocks;

[0085] The MLP block consists of two fully connected layers with a nonlinear ReLU function and a dropout in the middle.

[0086] In addition, this Transformer model uses a pre-normalized residual connection that produces more stable gradients, and stacks L identical layers to generate the final features.

[0087] Similar to the BERT model, we first add a token to the input Its state acts as a context vector in the output that can represent the entire time series information.

[0088] First, the feature z ≤t Input to a linear projection W Tran layer, which maps the feature z to the hidden dimension, i.e. The output of this linear projection layer is represented as

[0089] The context vector c is then appended to the feature vector z′, so that the input feature becomes ψ0 = [c; z′], where the subscript 0 indicates that it is the feature input to the first layer of Transformer.

[0090] Then, ψ0 is passed to the Transformer layer, as shown in the following formulas (3) and (4):

[0091]

[0092] Among them, MHA is the multi-head attention layer, Norm is the normalization layer, MLP is the multi-layer perceptron, l is the number of layers of the input feature in the Transformer operation, L is the maximum number of stacked Transformers, and ψ l represents the feature vector input to the l-th layer of the transformer.

[0093] Each layer of the Transformer contains the calculations of the multi-head attention mechanism and the multi-layer perceptron. After each layer, the input vector is updated into a new state. After ψ0 passes through L layers of the Transformer, the network output is a combined vector ψ containing the context vector and other input features L , extract the first part of this combined vector, that is, the updated version of the initially input context vector c after passing through the Transformer network and use it as the updated context vector c t to complete the task of the cross-view prediction time step to learn the internal laws and features of the radar signal time series.

[0094] S402. Use the context vector c t to complete the task of the cross-view prediction time step to learn the internal laws and features of the radar signal time series.

[0095] First, use a log-bilinear model, which will retain the mutual information between the input x t+k and c t , that is where W k is a linear function that maps c t back to the same dimension as z, that is, W k : R h→d .

[0096] By using the context of the strongly augmented view obtained in the previous step to predict the subsequent time steps of the weakly augmented , that is, the time steps from z t+1 to z t+k (1 < k ≤ K);

[0097] Similarly, use the context of the weakly augmented to predict the subsequent time steps of the strongly augmented .

[0098] By maximizing the dot product of the predicted time step and the true time step, and minimizing the dot product of the predicted time step and the other samples N in the batch t,k The dot product between them is used to minimize the temporal contrast loss.

[0099] Two kinds of losses and The calculation of is shown in formula (5) (6):

[0100]

[0101] Where K is the maximum step size for predicting future time steps, and the sum in the formula is performed for each future time step k, from t+1 to t+K; log is a logarithmic function used to calculate part of the loss, minimizing the contrast loss by maximizing the similarity of the same samples. k is a linear mapping function that transforms the context vector c t Projected back to the same dimension as the latent representation z, that is: W k :R h →R d , N t,k

[0102] Represents a set of negative samples in a small batch, which contains samples other than positive samples; is the potential representation of the weakly enhanced view of the current sample x to be learned at the future time step t+k, used as a positive sample, Similarly; A strongly enhanced view latent representation of all other samples in the same batch except the positive samples, used as negative samples; Same reason.

[0103] In the embodiment, the step S5 uses the context vector for a comparative learning process after nonlinear projection, and the model learns more discriminative features through this process, including the following sub-steps:

[0104] S501, context contrast module, aims to learn more discriminative representations. First, a nonlinear projection head is used to apply a nonlinear transformation to the context vectors of the two enhanced views output in step S4, and the context vectors are mapped to a context contrast space. Self-supervised contrast learning is performed based on this contrast space.

[0105] Given N input samples in a batch, after step S4, each sample generates two context vectors from different enhanced views, so there are 2N context vectors.

[0106] For the context of one of the views Will Represented as the context of another augmented view from the same input sample The positive sample of As a positive opposite;

[0107] The remaining (2N-2) contexts of other input samples in the same batch are considered Negative samples of It can form (2N-2) negative pairs with its negative samples.

[0108] S502, define context contrast loss To maximize the similarity between a sample and its positive counterpart, while minimizing its similarity with negative samples in the mini-batch.

[0109] Specifically, by comparing it with its positive sample The loss is normalized by dividing the similarity with all other (2N-1) samples (including 1 positive pair and (2N-2) negative pairs). The calculation of is shown in formula (7)(8).

[0110]

[0111] in is the context contrast loss;

[0112] is the contrast loss function between the positive sample pairs of sample x, used to measure and similarities between;

[0113] They represent the loss between the positive sample pairs of the weakly enhanced view and the strongly enhanced view of sample x and the loss between the positive sample pairs of the strongly enhanced view and the weakly enhanced view, respectively. For each sample in the batch, the loss function considers the directions of the two positive sample pairs to ensure the similarity between the two different enhanced views.

[0114] is the loss function between the context vectors of the strong and weak enhanced views of sample x, It is a function obtained by summing and averaging the losses calculated based on this function, representing the total contextual contrast loss of all samples in a batch, that is, the loss function of the contextual contrast module.

[0115] sim(u,v)=u T v / ∥u∥∥v∥ represents the normalized dot product of u and v (i.e. cosine similarity), 1 m≠i ∈{0,1} is an indicator function that takes the value 0 when m=i and takes the value 1 when m≠i, and τ is a temperature parameter. Minimizing this loss encourages the contextual embeddings of positive pairs to be closer and separates the embeddings of negative pairs from each other.

[0116] In the embodiment, step S6 obtains a total loss by integrating three losses and minimizes the total loss during the pre-training process, and finally obtains a radar signal time series feature representation model with the best performance, including the following sub-steps:

[0117] S601. Integrate three loss functions, namely two time contrast losses and one context contrast loss, to obtain a total self-supervised loss function, and use the total self-supervised loss function to train a feature representation model consisting of a time contrast module and a context contrast module.

[0118] The total loss function is composed of the three loss functions defined in step S5. The definition of the total loss function is shown in formula (9):

[0119]

[0120] Among them, λ1 and λ2 are fixed scalar hyperparameters, which represent the relative weight of each loss and are used to control the relative importance of temporal contrast loss and contextual contrast loss. During the training process, the model continuously optimizes the parameters to minimize the value of the total loss function.

[0121] By combining all the triple loss terms, the pre-trained model is encouraged to learn features that better reflect the internal laws of the radar signal time series during model optimization and capture more distinctive feature representations.

[0122] S602. In order to improve the training efficiency, the data is batch processed and input into the model in batches, each batch size is 64, and the model parameters are continuously updated through back propagation and Adam optimizer to improve the similarity of positive sample pairs and reduce the similarity of negative sample pairs. After the pre-training process is completed, the weights of the model are saved, and these weights will be used in subsequent supervised learning or downstream tasks. In the subsequent radar signal classification task, the use of this pre-trained model will be able to generate a more comprehensive and robust feature representation for the radar signal time series samples that integrates global context information (information in the time and frequency domains), thereby improving the pain point of the current feature representation difficulty.

[0123] It will be apparent to those skilled in the art that the invention is not limited to the details of the exemplary embodiments described above and that the invention can be implemented in other specific forms without departing from the spirit or essential features of the invention. Therefore, the embodiments should be considered exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description, and it is intended that all variations falling within the meaning and scope of the equivalent elements of the claims be included in the invention. Any reference numeral in a claim should not be considered as limiting the claim to which it relates.

[0124] In addition, it should be understood that although the present specification is described according to implementation modes, not every implementation mode contains only one independent technical solution. This description of the specification is only for the sake of clarity. Those skilled in the art should regard the specification as a whole. The technical solutions in each embodiment may also be appropriately combined to form other implementation modes that can be understood by those skilled in the art.

Claims

1. A radar signal feature representation method based on self-supervised contrastive learning, characterized in that: include: Step S1, obtaining radar signal time series data and preprocessing it to construct an unlabeled data set; Step S2, using strong enhancement and weak enhancement to perform data enhancement on radar time series samples in the unlabeled data set to obtain a strong enhancement view and a weak enhancement view; Step S3: input the strongly enhanced view and the weakly enhanced view into the encoder respectively to obtain corresponding potential embedding representations; Step S4: input the latent embedding representations corresponding to the strong enhanced view and the weak enhanced view into the temporal contrast module to perform the task of cross-view prediction time step, output the context vectors corresponding to the strong enhanced view and the weak enhanced view, respectively, and calculate the temporal contrast loss corresponding to the strong enhanced view and the weak enhanced view, respectively; The temporal contrast module uses an autoregressive model f for the potential embedding representation z ar The embedding representation z of the time steps less than time t ≤t Summarized as context vector c t =f ar (z ≤t ),c t ∈R h , where h is f ar The hidden dimension of the context vector c t To complete the task of cross-view prediction time step to learn the intrinsic laws and characteristics of radar signal time series, the process is as follows: First, add an h-dimensional token c, whose state in the output acts as a context vector that can represent the entire time series information. ≤t Input to a linear projection layer W Tran , output feature vector z′=W Tran (z ≤t ); The context vector c is then appended to the feature vector z′, so that the input feature becomes ψ0=[c;z′], where the subscript 0 indicates that ψ0 is the input feature of the first layer of Transformer; Then ψ0 is passed to the Transformer layer. Each layer of Transformer contains the calculation of the multi-head attention machine layer and the multi-layer perceptron. After each layer of Transformer, the input is updated to a new feature vector. The formula is as follows: Among them, MHA is the multi-head attention layer, Norm is the normalization layer, MLP is the multi-layer perceptron, l is the number of layers where the input features are calculated in the Transformer, L is the maximum number of Transformer layers that can be stacked, and ψ l Represents the feature vector input to the lth layer transformer; After ψ0 passes through L layers of Transformer, the network output is a combined vector ψ containing the context vector and other input features L , extract the first part of the combined vector, which is the updated version of the context vector c that was originally input after passing through the Transformer network And use it as the updated context vector c t , to complete the task of predicting time steps across views in order to learn the intrinsic laws and characteristics of radar signal time series; Step S5: input the context vectors corresponding to the strong enhanced view and the weak enhanced view into the context contrast module to perform self-supervised contrast learning and calculate the context contrast loss; Step S6, integrating the time contrast loss and the context contrast loss to obtain a total loss function, and training the model composed of the time contrast module and the context contrast module based on the total loss function to obtain a radar signal time series feature representation model, so as to generate a more comprehensive and robust feature representation for the radar signal time series that integrates the global context information.

2. The radar signal feature representation method based on self-supervised contrastive learning according to claim 1 is characterized in that: The step S1 comprises: S101, obtaining radar signal time series data from the radar system, including timestamps collected by different sensors and object distance, speed, and intensity characteristics; S102, denoising and filtering the acquired radar signal time series data; S103, standardizing and normalizing the radar signal time series data; S104, setting the window size W and the step size S, dividing the radar signal time series data into multiple segments of fixed length, each segment being used as an independent radar time series sample, to form an unlabeled data set.

3. The radar signal feature representation method based on self-supervised contrastive learning according to claim 1 is characterized in that: The step S2 includes: for the radar time series sample x, using the dithering method and the arrangement method as a strong enhancement strategy to generate a strong enhancement view x s ; For radar time series sample x, use scaling method and jittering method as weak enhancement strategy to generate weak enhancement view x w ; The jitter adding method adds random noise to the original time series data to simulate the subtle changes that may occur in the actual acquisition process, and then adds it to the data value at each time point; the arrangement method first divides the time series into multiple subsequence fragments, fixes the size of these subsequence fragments, shuffles these subsequence fragments in a random order, keeps the order of data in the fragments unchanged, and finally recombine the shuffled fragments to form a new time series; the scaling method multiplies the entire time series by a random scaling factor to achieve amplification or reduction of the amplitude of the time series data.

4. The radar signal feature representation method based on self-supervised contrastive learning according to claim 1 is characterized in that: The step S3 uses the encoder f enc Map the input augmented view x′ to a d-dimensional latent embedding representation z=f enc (x′); where z=[z1,z2,…z T ], T is the total time step, z i It represents the d-dimensional embedding vector obtained after encoding at the i-th time step, where d is the feature length.

5. The radar signal feature representation method based on self-supervised contrastive learning according to claim 1, characterized in that: In step S4, the enhanced view The corresponding temporal contrast loss is: Weakly enhanced view The corresponding temporal contrast loss is: Among them, K is the maximum step size for predicting future time steps; log is the logarithmic function; W k is a linear mapping function; N t,k Represents a set of negative samples in a batch, including samples other than positive samples; is the potential representation of the weakly enhanced view of the current sample x to be learned at the future time step t+k, used as a positive sample, It is the potential representation of the strongly enhanced view of the current sample x to be learned at the future time step t+k, which is used as a positive sample; A strongly enhanced view latent representation representing all other samples in the same batch except the positive samples, used as negative samples; A weakly enhanced view latent representation representing all other samples in the same batch except the positive samples is used as negative samples.

6. The radar signal feature representation method based on self-supervised contrastive learning according to claim 1, characterized in that: The context contrast module in step S5 first uses a nonlinear projection head to apply nonlinear transformation to the context vectors corresponding to the strong enhanced view and the weak enhanced view respectively, maps the context vectors corresponding to the strong enhanced view and the weak enhanced view output by the temporal contrast module in step S4 to the context contrast space, and performs self-supervised contrast learning based on the contrast space, as follows: For a given batch of N input samples, after step S4, each sample generates two context vectors from different enhanced views, so there are 2N context vectors. For the context vector of one of the views Will Represented as the context vector of another augmented view from the same input sample The positive sample of As positive pairs, the remaining (2N-2) context vectors of other input samples in the same batch are considered Negative samples of It can form (2N-2) negative pairs with its negative samples.

7. The radar signal feature representation method based on self-supervised contrastive learning according to claim 1, characterized in that: The context contrast loss in step S5 is: in is the context contrast loss; is the contrast loss function between the positive sample pairs of sample x, used to measure and similarities between; They represent the loss between the positive sample pairs of the weakly enhanced view and the strongly enhanced view of sample x, and the loss between the positive sample pairs of the strongly enhanced view and the weakly enhanced view, respectively; Represents normalization and The dot product of Represents normalization and The dot product of 1 m≠i ∈{0,1} is the indicator function; τ is the temperature parameter.

8. The radar signal feature representation method based on self-supervised contrastive learning according to claim 1, characterized in that: The total loss function in step S6 is: in is the temporal contrast loss corresponding to the strongly enhanced view; is the temporal contrast loss corresponding to the weakly enhanced view; is the context contrastive loss; λ1 and λ2 are fixed scalar hyperparameters.

9. The radar signal feature representation method based on self-supervised contrastive learning according to claim 1, characterized in that: In step S6, batch processing is performed on the data during model training. The data is input into the model in batches, and the model parameters are continuously updated through back propagation and Adam optimizer to minimize the value of the total loss function. After the training process is completed, the weight of the model is saved for subsequent supervised learning or downstream tasks.

Citation Information

Patent Citations

  • Scene text recognition system based on consistency regular training

    CN114529904A

  • NDVI time series data reconstruction method and system

    CN116434050A