A Self-Sustaining Incremental Recognition Method for Radar One-Dimensional Range Profiles Based on Guided Representation

By building dynamic query navigation and structural expansion modules in radar one-dimensional distance image target recognition, the problem of forgetting when new targets appear is solved, the recognition accuracy is improved, and the old targets are not forgotten.

CN116068518BActive Publication Date: 2025-07-04UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310057069.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-16
Publication Date
2025-07-04
Estimated Expiration
2043-01-16

AI Technical Summary

Technical Problem

The existing radar one-dimensional distance-image target recognition method is prone to catastrophic forgetting when facing the emergence of new targets and cannot effectively identify old targets.

Method used

Using a self-sustaining incremental recognition method based on vision transformer, we use dynamic query navigation module and structural expansion module to expand the input dimensions and model structure, retain old category information and learn under the guidance of new and old category information.

Benefits of technology

It improves the accuracy of radar one-dimensional recognition like new targets, prevents the model from forgetting old targets, and realizes the recognition of not forgetting old targets when new targets appear.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116068518B_ABST
    Figure CN116068518B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of target recognition, and particularly relates to a self-sustaining incremental recognition method for radar one-dimensional range profiles based on guided representation. Aiming at the catastrophic forgetting problem in the recognition of new targets in radar one-dimensional images, the present invention proposes a self-sustaining guided representation method according to the structural characteristics of vision transformers. On the one hand, learnable navigation is used to retain class-specific information in the form of one-dimensional blocks, and the input of the extended encoding layer is obtained by querying the navigation with the highest similarity during the training process, so that the model can obtain more additional information and improve the recognition accuracy. On the other hand, a structure extension module is constructed to retain all old class information, and the structure of the model is expanded by element-wise summation during each training, so that the model can learn under the joint guidance of new and old class information. At the same time, in order to prevent the model from losing the learning ability of new classes during memorization, an unfrozen model is used to enable the model to fully learn new classes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of target recognition, and particularly relates to a self-sustaining incremental recognition method for radar one-dimensional range profiles based on guiding representation. Background Art

[0002] The radar one-dimensional range profile (HRRP) is the scattering feature distribution of the target echo signal occupying multiple range cells along the range direction when the radar emits narrow pulses or broadband signals (such as linear frequency modulation signals and stepped frequency signals, etc.). Due to its advantages of strong real-time performance, small computational complexity, easy acquisition and storage, etc., it has been widely studied and is a main method for target recognition in the modern war environment. At present, deep learning methods have achieved good results in HRRP classification and recognition, but they do not consider the growth learning of data in the future stage. When new targets appear, the parameters for the model to recognize old targets are masked and the old targets cannot be recognized, that is, catastrophic forgetting. Common methods train with a small number of old target samples replayed together with new targets, but only a small amount of information can be obtained from a small number of old target samples, and forgetting still exists. Under this background, the present invention conducts research on the incremental recognition of new targets for the problem of radar one-dimensional image target recognition, aiming to not forget old targets while recognizing new targets and realizing automatic radar target recognition. Summary of the Invention

[0003] Based on the background of radar one-dimensional range profile target recognition and aiming at the problem that the existing replay-based incremental recognition method cannot solve the forgetting problem, the present invention proposes a self-sustaining guiding representation method. This method prevents forgetting by expanding the input dimension and the model structure at different learning stages. Specifically, the target category-specific information is retained in the form of tokens in vision transformer, and then the retained task information is used to expand the input to assist the model in recognition. In addition, the old class information is used to expand the model structure, enabling the model to learn under the joint guidance of new and old information and retain the memory of old class targets.

[0004] The model of the present invention uses vision transformer (ViT) as the basic framework. The new category samples Dt of the current task and a small number of sampled samples of all old categories constitute the training data set of the current task. After the training data set of the current task is subjected to feature extraction by vision transformer, it is input into a classifier to obtain the recognition result. The core idea of the network is as follows: on the one hand, a dynamic query guidance is constructed before the transformer encoder to obtain additional information of the current sample, and together with the patch token, it is input into the transformer encoder to extract features. On the other hand, in the transformer encoder part, through structural expansion and integration with the structure carrying old category information, the data of the current task is jointly learned to achieve the retention of old category information. The data of new and old categories are used for testing, and indicators such as average accuracy rate and average forgetting rate are statistically calculated.

[0005] The technical solution of the present invention is as follows:

[0006] A radar one-dimensional range profile self-sustaining incremental recognition method based on guided representation, comprising the following steps:

[0007] S1. Construct a data set:

[0008] The learning process is divided into two stages: a basic task and an incremental task. There is only one task T1 in the basic task, and the incremental task consists of multiple non-overlapping tasks {T2, T3,..., T T}; Take the amplitude image data of the radar one-dimensional range profile data as the input, allocate different numbers of categories to each task, and the number of samples of each type of target is the same; Define the number of target categories of the basic task as N1, the number of incremental tasks as T - 1, the number of target categories of each incremental task as N2, and the total number of categories N = N1 + N2×(T - 1);

[0009] S2. Data preprocessing:

[0010] Normalize all samples in S1, map the signal intensity to (-1, 1), and obtain the training set of task t:

[0011]

[0012] where t represents the t-th task, T represents the number of tasks; N represents the total number of categories; B represents the number of sampling points; K i represents the number of training samples of the i-th type of target; the sample label set of the training set is expressed as:

[0013]

[0014] In the formula, y trainij Denote the sample as x train ij and its class label;

[0015] S3. Construct a self-sustaining guided representation network: The basic framework of the network is the vision transformer. On this basis, a dynamic query navigation module is added to expand the input dimension of the vision transformer encoding layer; a structure expansion module is added to expand the network structure of the last few layers in the encoding layer. Specifically:

[0016] S31. Construct a pre-trained basic vision transformer network as where φ prem (·) represents the linear projection module, and φ trans (·) represents the transformer encoding layer. Specifically:

[0017] The linear projection module splits the radar one-dimensional range image into multiple one-dimensional blocks. The input x ∈ R 1×B is segmented into N e vectors, denoted as and then converted to the dimension D suitable for ViT through a linear transformation:

[0018]

[0019] Add a learnable token before x p and add the position information of each embedding block:

[0020]

[0021] The encoding layer has L layers, and each layer contains: two normalizations LN, multi-head attention MSA, two dropouts, and an MLP block. The overall output of the model is the output corresponding to the last layer x class :

[0022] z l ' = dropout(MSA(LN(z l-1 )))+z l-1 , l = 1,..., L

[0023] z l = dropout(MLP(LN(z l ')))+z l , l = 1,..., L

[0024]

[0025] S32. Construct a dynamic query navigation module. The input of the dynamic query navigation module is the preprocessed result X in S2 t , and by finding the navigation most similar to the input sample, output M navigations A with the same dimension as the token, specifically including:

[0026] Construct a navigation pool, which contains two parts: key value K and navigation A. The key value and the navigation are in one-to-one correspondence, expressed as {(k1, A1), (k2, A2),..., (k N , A N ), with a total of N; design the key value to have the same dimension as the output corresponding to the x of the vision transformer class Design the navigation to have the same dimension as a single input of the encoding layer A ∈R i ; D

[0027] Query to obtain the navigation, input X t , and after passing through a pre-trained vision transformer, obtain the output at the corresponding position of x class Calculate the similarity with all key values in the navigation pool, and find the indices of the M key values with the highest similarity:

[0028]

[0029] Expand the input. Since the key value and the navigation are in one-to-one correspondence, find M navigations according to the index of the key value, expressed as A q ; use A q to replace x class to expand x e :

[0030]

[0031] S33. Construct a structure expansion module, divide the encoding layer of the vision transformer into layers of fixed layers and layers of expansion layers The fixed layers are frozen after pre-training, and the expansion layers are expanded by the structure expansion module in an element-wise summation manner, specifically including:

[0032] Construct an expansion structure, design the expansion structure g to be the same as the expansion layer: MSA layers, each layer contains two normalization LNs, multi-head attention MSA, two dropouts, and an MLP block; after each task training ends, the expansion structure is updated using the old class data in the memory database;​​

[0033] The extended model structure, the inputs of the extended layer and the extended structure are the outputs of the fixed layer For the extended structure, the input of each layer is the output of the previous layer:

[0034] p i = g i (p i-1 )

[0035] For the extended layer, the input of each layer is the element-wise sum of the output of the previous layer of the model and the output of the extended structure:

[0036]

[0037]

[0038] where q i represents the output of the i-th layer of the model, represents the element-wise sum of two vectors;

[0039] S4. Train the network constructed in S3 with the training samples of the basic dataset, and the loss function of the model is:

[0040] L = L class + λL query

[0041] where λ is a hyperparameter; specifically including:

[0042] S41. The specific form of L class :

[0043]

[0044] In the above formula, L ce represents the cross-entropy loss, y represents the label of the sample, is the prediction result of the model on the current sample:

[0045]

[0046] Its meaning is: the output result of the average value of the outputs at the corresponding positions of the navigation taken out in S32 on the classifier;

[0047] S42. The specific form of L query :

[0048]

[0049] In the above formula, q(x) represents the output at the corresponding position of x class calculated in S32, k qDenote the M key values with the highest similarity to q(x). This loss function is used to update the key values so that the selected M key values are less different from q(x).

[0050] S5. After each task learning is completed, select H samples for each category in the task and retain them in the sample database.

[0051] S6. Input the one-dimensional image data of the radar target to be recognized and all the previously retained samples into the trained model for classification and recognition.

[0052] The beneficial effects of the present invention are as follows: Aiming at the catastrophic forgetting problem existing in the recognition of new radar one-dimensional images, a self-sustaining guided representation method is proposed according to the structural characteristics of the vision transformer. On the one hand, learnable navigation is used to retain category-specific information in the form of one-dimensional blocks. By querying during the training process to obtain the navigation with the highest similarity to expand the input of the encoding layer, the model can obtain more additional information and improve the recognition accuracy. On the other hand, a structure expansion module is constructed to retain all old class information. During each training, the structure of the model is expanded through element-wise summation, enabling the model to learn under the joint guidance of new and old category information. At the same time, in order to prevent the model from losing the learning ability of new categories during memory, an unfrozen model is used, allowing the model to fully learn new categories. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 Schematic diagram of the overall network structure.

[0054] Figure 2 Schematic diagram of the design of the dynamic query navigation module.

[0055] Figure 3 Schematic diagram of the structure expansion design. DETAILED DESCRIPTION OF THE INVENTION

[0056] The technical solutions of the present invention will be described in detail below in conjunction with the accompanying drawings:

[0057] The present invention specifically includes the following steps:

[0058] S1. Construct a data set:

[0059] Divide the entire learning process into two stages: a basic task and an incremental task. The basic task has only one task T1, and the incremental task consists of multiple non-overlapping tasks {T2, T3,..., T6}. The data of the radar one-dimensional range image contains 10 types of aircraft targets, with 1801 samples for each type, and the size of each sample is 1×320. The number of target categories in the basic task is 5, and the number of target categories in each incremental task is 1.

[0060] S2. Data preprocessing:

[0061] Normalize all samples in S1, map the signal strength to (-1, 1), and obtain the training set for task t:

[0062]

[0063] where t represents the t-th task. The sample label set of the training set is expressed as:

[0064] Y t train ={y train ij | i = 1, 2,..., 8 i ; j = 1, 2,..., 1801 i} ∈ R 1801×1 , t = 1, 2..., T

[0065] In the formula, y train ij represents the class label of sample x train ij .

[0066] Similarly, obtain the test set for task t and the test label set Y t test .

[0067] S3. Construct a pre-trained basic vision transformer network as where φ prem (·) represents a linear projection module, and φ trans (·) represents a transformer encoding layer.

[0068] The linear projection module splits the radar one-dimensional range image into multiple one-dimensional blocks. The input x ∈ R 1×320 is split into 20 vectors, denoted as x e ∈ R 20×16 . Then, through a linear transformation, it is converted into a dimension D suitable for ViT use:

[0069] x p = x e × W, k ∈ {1,..., N e}, W ∈ R 16×224

[0070] Before x p , add a learnable token and add the position information of each embedding block:

[0071] z0 = [x class , x e + E pos, E pos ∈R 20×224 , N p = 20

[0072] The encoding layer has 12 layers, and each layer contains: two normalization LNs, multi-head attention MSA (12 heads), two dropouts, and an MLP block. The specific structure is shown in the example Figure 1 . The overall output of the model is the last layer x class The corresponding output:

[0073] z l ' = dropout(MSA(LN(z l-1 )))+z l-1 , l = 1,..., 12

[0074] z l = dropout(MLP(LN(z l ')))+z l ', l = 1,..., 12

[0075]

[0076] Design a dynamic query navigation module, see the example Figure 2 . Its input is the preprocessed result X in S2 t , and by finding the navigation most similar to the input sample, it outputs 5 navigations A with the same dimension as the token

[0077] Construct a navigation pool. The pool contains two parts: key-value K and navigation A. The key-value and navigation are in one-to-one correspondence and are represented as {(k1, A1), (k2, A2),..., (k 10 , A 10 )}, and there are 10 in total. Design the key-value to have the same dimension k as the output corresponding to x of the vision transformer class ; Design the navigation to have the same dimension A as the single input of the encoding layer i ∈R 224 ; Design the navigation to have the same dimension A as the single input of the encoding layer i ∈R 224 .

[0078] Query to get the navigation. Input X t , and through a pre-trained vision transformer, obtain the output at the corresponding position of x class Calculate the similarity with all key-values in the navigation pool and find the indices of the 5 key-values with the highest similarity degree

[0079] ​

[0080] Expand the input. Since the key-value pairs and the navigation are in one-to-one correspondence, 5 navigations are found according to the index of the key-value pair, denoted as A q . Use A q to replace x class Expand x e :

[0081] p0 = [A q , x e + E pos , E pos ∈R 25×224 , N A = 25

[0082] After expanding the input using the navigation, the dimension change table for each layer is as follows:

[0083]

[0084] Design a structure expansion module, see the example Figure 3 . Divide the encoding layer of the vision transformer into 6 fixed layers and 6 expansion layers The fixed layers are frozen after pre-training, and the expansion layers are expanded by the structure expansion module in an element-wise summation manner.

[0085] Construct an expansion structure. Design the expansion structure g to be the same as the expansion layer: 6 MSA layers, each layer containing two normalizations LN, multi-head attention MSA, two dropouts, and an MLP block. After each task training is completed, the expansion structure is updated using the old class data in the memory database.

[0086] Expand the model structure. The inputs of the expansion layer and the expansion structure are both the output z6 of the fixed layer. For the expansion structure, the input of each layer is the output of the previous layer:

[0087] p i = g i (p i-1 )

[0088] For the expansion layer, the input of each layer is the element-wise summation of the output of the previous layer of the model and the output of the expansion structure:

[0089]

[0090]

[0091] where, q i represents the output of the i-th layer of the model, represents the element-wise summation of two vectors.

[0092] S4. Use the training samples of the basic data set to train the network constructed in S3. The loss function of the model is:

[0093] L = L class + 0.1×L query

[0094] L class The specific form of:

[0095]

[0096] In the above formula, L ce represents the cross-entropy loss, y represents the label of the sample, is the prediction result of the model on the current sample:

[0097]

[0098] Its meaning is: the output result of the average value of the outputs at the corresponding positions of the navigation taken out in S3 on the classifier.

[0099] L query The specific form of:

[0100]

[0101] In the above formula, q(x) represents the output at the corresponding position of x calculated in S3 class corresponding position, k q represents the 5 key values with the highest similarity to q(x). This loss function is used to update the key values to make the difference between the selected 5 key values and q(x) smaller.

[0102] S5. After each task learning is completed, select 20 samples for each category in the task and retain them in the sample database.

[0103] S6. Input the one-dimensional image data of the radar target to be recognized and all the previously retained samples into the trained model for classification and recognition.

Claims

1. A self-sustaining incremental recognition method for radar one-dimensional range profiles based on guiding representations, comprising the following steps: S1. Construct a data set: The learning process is divided into two stages: a basic task and an incremental task. There is only one task T1 in the basic task, and the incremental task consists of multiple disjoint tasks {T2, T3,..., T T}. The amplitude image data of the radar one-dimensional range image data is taken as the input, and different numbers of categories are assigned to each task, with the same number of samples for each type of target. The number of target categories in the basic task is defined as N1, the number of incremental tasks is T - 1, the number of target categories in each incremental task is N2, and the total number of categories N = N1 + N2 × (T - 1); S2. Data preprocessing: Perform normalization processing on all samples in S1, map the signal intensity to (-1, 1), and obtain the training set for task t: where t represents the t-th task, T represents the number of tasks; N represents the total number of categories; B represents the number of sampling points; K i represents the number of training samples of the i-th type of target; the sample label set of the training set is expressed as: where y train ij represents the class label of the sample x train ij ; S3. Construct a self-sustaining guiding representation network: The basic framework of the network is vision transformer. On this basis, add a dynamic query navigation module to expand the input dimension of the vision transformer encoding layer; add a structure expansion module to expand the network structure of the last few layers in the encoding layer; specifically: S31. Construct the pre-trained basic vision transformer network as where φ prem (·) represents a linear projection module, and φ trans (·) represents a transformer encoding layer; specifically: The linear projection module splits the radar one-dimensional range image into multiple one-dimensional blocks, with the input \(x\in\mathbb{R}\) 1×B which is segmented into \(N\) e vectors, denoted as and then through a linear transformation, it is converted into the dimension \(D\) suitable for ViT: Add a learnable token before x p and add the position information of each embedding block: The encoding layer has a total of L layers, and each layer contains: two layer norms LN, multi-head self-attention MSA, two dropouts, and an MLP block. The overall output of the model is the output x of the last layer. class The corresponding output: z l ' = dropout(MSA(LN(z l-1 )))+z l-1 , l = 1, ..., L z l = dropout(MLP(LN(z l '))) + z l ', l = 1, ..., L S32. Build a dynamic query navigation module. The input of the dynamic query navigation module is the preprocessed result X in S2 t , and by finding the navigation most similar to the input sample, output M navigations A with the same dimension as the token, specifically including: Construct a navigation pool, which contains two parts: key value K and navigation A. The key values and navigations are in one-to-one correspondence, expressed as {(k1, A1), (k2, A2),..., (k N , A N )}, with a total of N; Design that the output corresponding to the key value and the x of the vision transformer class has the same dimension Design that the single input of the navigation and the encoding layer has the same dimension A i ∈R D ; Retrieve the navigation and input X t , pass through a pre-trained vision transformer to obtain x class The output at the corresponding position Calculate The similarity with all key-values in the navigation pool, and find the indices of the M key-values with the highest similarity degree: Expand the input. Since the key values and the navigation are in one-to-one correspondence, find M navigations according to the index of the key values, denoted as A q ; Use A q to replace x class Expand x e : S33. Construct a structure expansion module and divide the encoding layer of the vision transformer into layers of fixed layers and layers of expansion layers The fixed layers are frozen after pre-training, and the expansion layers are expanded by the structure expansion module in an element-wise summation manner, specifically including: Construct an extended structure, and design the extended structure g to be the same as the extended layer: There are MSA layers, and each layer contains two normalization LNs, multi-head attention MSA, two dropouts, and an MLP block; after the training of each task is completed, the extended structure is updated using the old class data in the memory database; The extended model structure, and the inputs of the extended layer and the extended structure are the outputs of the fixed layer For the extended structure, the input of each layer is the output of the previous layer: p i = g i (p i-1 ) For the expansion layer, the input of each layer is the element-wise sum of the output of the previous layer model and the output of the expansion structure: where q i represents the output of the i-th layer of the model, represents the summation of the corresponding elements of two vectors; S4. Use the training samples of the basic data set to train the network constructed in S3. The loss function of the model is: L = L class + λL query where λ is a hyperparameter; specifically including: S41, L class The specific form of In the above formula, L ce represents the cross-entropy loss, y represents the label of the sample, is the prediction result of the model on the current sample: Its meaning is: the output result on the classifier of the average value of the outputs at the corresponding positions of the navigation taken out in S32; S42, L query The specific form is as follows: In the above formula, q(x) represents x calculated in S32 class The output at the corresponding position, k q represents the M key values with the highest similarity to q(x). This loss function is used to update the key values so that the difference between the selected M key values and q(x) is smaller; S5. After each task learning is completed, select H samples for each category in the task and retain them in the sample database; S6. Co-input the radar target one-dimensional image data to be recognized and all the previously retained samples into the trained model for classification and recognition.

Citation Information

Patent Citations

  • Radar target identification method based on Transform and time convolution network

    CN115079116A

  • Non-coherent signal radar target micro-motion feature extraction and imaging method

    CN115184933A