Quasi-delta modulation recognition method based on adaptive feature distillation prototype playback

By adopting the class-incremental modulation recognition method of adaptive feature distillation prototype playback in the field of modulation recognition, the catastrophic forgetting problem of deep learning models in this field is solved, and the effects of high recognition accuracy and low memory consumption are achieved.

CN118694641BActive Publication Date: 2025-05-16NAT UNIV OF DEFENSE TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410870835.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-01
Publication Date
2025-05-16
Estimated Expiration
2044-07-01

AI Technical Summary

Technical Problem

The prior art is difficult to effectively solve the catastrophic forgetting problem of deep learning models in the field of modulation recognition, especially in class incremental learning. The model's feature extractor will cause the old class features to fail to fully match the latest feature extractor, resulting in feature space offset.

Method used

A class-incremental modulation recognition method based on adaptive feature distillation prototype playback is adopted to design a deep neural network with mixed structures for feature extraction, combining feature superposition and adaptive distillation technology to limit the drift of parameters and alleviate the problem of forgetting.

Benefits of technology

In the case of saving a small number of samples, it effectively alleviates the catastrophic forgetting problem of deep learning models in the field of modulation recognition, improves the recognition accuracy of modulated signals, and reduces memory consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118694641B_ABST
    Figure CN118694641B_ABST
Patent Text Reader

Abstract

The present invention discloses a quasi-incremental modulation recognition method based on adaptive feature distillation prototype playback, and relates to the technical field of communication signal processing methods. The method comprises the following steps: designing a deep neural network with a hybrid structure to extract features of modulated signals; quasi-incremental learning tasks and training and test data set division; training and hyperparameter update of quasi-incremental learning models: setting the training parameters and initialization example set of quasi-incremental learning models, and after each task training is completed, determining whether there are tasks that have not been completed, and if so, training the next task, otherwise outputting the model and related parameters; classifier design and modulation classification. The method described in the present application can alleviate the problem of catastrophic forgetting of deep learning models in the field of modulation recognition by limiting prototype parameters while saving a small number of samples, and has high recognition performance for modulated signals.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of communication signal processing methods, and in particular to a quasi-incremental modulation recognition method based on adaptive feature distillation prototype playback. Background Art

[0002] In recent years, deep learning has achieved great success in image classification, object detection, speech recognition, signal recognition, natural language processing and other fields. However, with the continuous expansion of application scenarios and the increase in task complexity, the traditional static training method can no longer meet the needs of continuous model learning and adaptation to new tasks. Incremental learning has gradually become one of the research hotspots. Its core goal is to enable the model to continuously learn the features of new tasks while retaining the knowledge it has learned in the new task sequences it is constantly exposed to. This means that the model needs to continue to learn and adjust when exposed to new data, rather than simply reusing the previously trained model. The model obtained by this learning method can adapt to the characteristics and requirements of new tasks in a timely manner and is still effective for the original tasks. However, incremental learning also faces some challenges, the most important of which is catastrophic forgetting. Catastrophic forgetting refers to the phenomenon that when a model learns a new task, the recognition or prediction performance on the old tasks that have been learned decreases. This is because when learning a new task, the parameters of the model change, changing the features learned on the old tasks. Therefore, when designing an incremental learning algorithm, it is necessary to take into account the balance between new and old tasks, minimize the impact of catastrophic forgetting, and achieve continuous learning and evolution of the model.

[0003] The field of incremental learning has been looking for various strategies to solve the problem of catastrophic forgetting. Early incremental learning methods mainly focused on improvements in regularization, parameter isolation, and sample replay. However, these methods all have some obvious limitations. Regularization methods focus too much on restricting model parameters, which may cause the model to become less adaptable to new tasks because they may hinder the learning of new knowledge while emphasizing the retention of old knowledge. Parameter isolation methods face the problem of a sharp increase in the number of parameters as new tasks continue to increase, which may lead to insufficient computing resources and reduced training efficiency. Sample replay methods cannot be applied in some scenarios, such as when data privacy is emphasized and memory space is insufficient, because they need to store a large amount of historical data.

[0004] From the perspective of quasi-incremental modulation recognition, there are still two problems in the current work: 1) Quasi-incremental learning methods are mainly used in the field of computer vision, and there is still little practice in the field of modulation recognition, and the effect is not good; 2) Although the current quasi-incremental learning methods have solved the problem of catastrophic forgetting to a certain extent, during the incremental learning process, the feature extractor of the model will continue to learn, resulting in the original retained old class features cannot fully match the latest feature extractor, resulting in an offset in the feature space. Summary of the invention

[0005] The technical problem to be solved by the present invention is how to provide a quasi-incremental modulation recognition method with high modulation signal recognition accuracy, which can alleviate the catastrophic forgetting problem of deep learning models in the field of modulation recognition by limiting prototype parameters while saving a small number of samples.

[0006] In order to solve the above technical problems, the technical solution adopted by the present invention is: a method for identifying quasi-incremental modulation based on adaptive feature distillation prototype playback, comprising the following steps:

[0007] S101, deep learning feature extraction network model design: for the modulation signal with single-channel input and the time series characteristics of the signal data being correlated, multi-modal calculation is performed on each signal data, and a hybrid structure deep neural network is designed to extract the features of the modulation signal;

[0008] S102, quasi-incremental learning tasks and training and test data set division: the entire data set is divided into n tasks, one task is selected as the initial task, that is, the task of learning the original model, and the incremental learning is implemented by completing one task and then continuing to the next task; for each type of sample signal in each task, after multi-modal calculation of the signal, it is combined with the original sample into a new signal set, and divided into a training set and a test set according to a certain ratio;

[0009] S103, training of the quasi-incremental learning model and updating of hyperparameters: setting the training parameters and initialization example set of the quasi-incremental learning model, cyclically obtaining the training data set of each task, arbitrarily selecting a task as the task for training the original model, and using the remaining tasks for incremental learning. Before learning each task, the example set needs to be compressed, the output layer of the feature extraction model is incrementally updated, and the example set is saved after learning; after each task training is completed, it is determined whether there are tasks that have not been completed. If so, the next task is trained, otherwise the model and related parameters are output;

[0010] S104, classifier design and modulation classification: After each task is completed, the trained model parameters are used to calculate the feature vector of each sample in the test set and the average feature vector of each type of signal in the example set, and the Gaussian kernel between the two is compared to obtain the modulation category of each sample in the test set.

[0011] The beneficial effects of adopting the above technical solution are: the method described in the present application saves the example set by the feature superposition method, and concentrated features can be obtained from a small number of old classes. The adaptive distillation method is used to limit the drastic drift of parameters in each layer to alleviate the forgetting of old class features, which not only reduces memory consumption but also alleviates the problem of catastrophic forgetting. An efficient feature extraction network is also designed to achieve high-precision recognition of modulation categories. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments.

[0013] Figure 1 is a main flow chart of the method described in the embodiment of the present invention;

[0014] Figure 2 is a schematic diagram of the structure of the deep learning model designed in the method described in the embodiment of the present invention;

[0015] Figure 3 It is a flowchart of the incremental learning task and training test data set division method in the method described in the embodiment of the present invention;

[0016] Figure 4 is a schematic diagram of signal data set division in the method according to an embodiment of the present invention;

[0017] Figure 5 It is a schematic diagram of a method for training a quasi-incremental learning model and updating hyperparameters in the method according to an embodiment of the present invention;

[0018] Figure 6 is a schematic diagram of updating of example sets and parameters in the method according to an embodiment of the present invention;

[0019] Figure 7 It is a schematic diagram of classifier design and modulation classification in the method described in an embodiment of the present invention. DETAILED DESCRIPTION

[0020] The following is a clear and complete description of the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0021] In the following description, many specific details are set forth to facilitate a full understanding of the present invention, but the present invention may also be implemented in other ways different from those described herein, and those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0022] like Figure 1 As shown, the present invention discloses a method for identifying quasi-incremental modulation based on adaptive feature distillation prototype playback, the method comprising the following steps:

[0023] S101, deep learning feature extraction model design: for the modulation signal with single-channel input and the time series characteristics of the signal data being correlated, multi-modal calculation is performed on each signal data, and a hybrid structure deep neural network is designed to extract the features of the modulation signal;

[0024] The hybrid structure refers to a neural network structure composed of two or more networks for extracting features from multiple aspects. For example, a hybrid structure can be composed of a convolutional structure, a residual structure, and a long short-term memory network structure (LSTM), which can extract features of input targets from multiple aspects such as space and time sequence.

[0025] S102, quasi-incremental learning tasks and division of training and test data sets: divide the entire data set into n tasks, select one task as the initial task, that is, the learning task of the original model, and the incremental learning is implemented by completing one task and then continuing to the next task; for each type of sample signal in each task, perform multi-modal calculations on the signal, combine it with the original sample into a new signal set, and divide it into training set and test set according to a certain ratio.

[0026] The original model refers to a deep learning model used to learn the characteristics of signal data in the initial task; the original model training refers to the process of training the original model using the signal data in the initial task; the incremental learning refers to the process of learning a new modulation signal that is different from the previous modulation signal after the quasi-incremental deep learning model has learned a variety of modulation signals; the multimodal calculation refers to the calculation of different modes such as amplitude phase, time-frequency transformation and constellation diagram transformation of the IQ data of given signal data.

[0027] S103, training of the quasi-incremental learning model and update of hyperparameters: setting the training parameters and initialization example set of the quasi-incremental learning model, after looping to obtain the training data set for each task, arbitrarily select a task as the task for training the original model, and use the remaining tasks for incremental learning. Before learning each task, the example set needs to be compressed, the output layer of the feature extraction model is incrementally updated, and the example set is saved after learning; after each task training is completed, determine whether there are tasks that have not been completed. If so, train the next task, otherwise output the model and related parameters.

[0028] The example set refers to a collection of a small number of typical old class signals saved before incremental learning of new class signals; the new class signal refers to the modulation category that has not been learned by the incremental deep learning model, and each new class signal is composed of several new class signal samples; the old class signal refers to the modulation category that has been learned by the incremental deep learning model. As before, each old class signal is composed of several old class signal samples. An old class sample refers to any signal in a certain old class signal set, and the same applies to the new class sample.

[0029] S104, classifier design and modulation classification: After each task is completed, the trained model parameters are used to calculate the feature vector of each sample in the test set and the average feature vector of each type of signal in the example set, and the Gaussian kernel between the two is compared to obtain the modulation category of each sample in the test set.

[0030] Furthermore, the method of designing a deep learning feature extraction model in step S101 includes the following steps:

[0031] The convolution module combined with the residual module is used as part of the neural network. In addition, the signal has temporal features that the image does not have, so the long short-term memory network (LSTM) is needed to extract the temporal features of the signal data;

[0032] Exemplary: An improved deep convolutional neural network (AlexNet) can be used as a feature extractor, and the original feature extraction module composed of 5 convolutional layers, a Dropout layer and a ReLU activation function is replaced with a new feature extraction module composed of a convolutional layer and three residual modules, a maximum pooling layer and an LSTM layer, and the ReLU activation function in the residual module is replaced with a SeLU function;

[0033] A possible structure of the feature extraction network is as follows Figure 2 As shown, it mainly includes 1 convolution module, 3 residual modules connected in sequence, a maximum pooling layer module, an LSTM layer module and a fully connected layer module, wherein the convolution module includes a first two-dimensional convolution layer, a first Dropout layer and a first SeLU activation function layer connected in sequence; the residual module includes a second two-dimensional convolution layer, a second Dropout layer, a second SeLU activation function layer, a third two-dimensional convolution layer, a third Dropout layer and a third SeLU activation function layer connected in sequence; the output of the convolution module is the input of the first residual module, and the sum of the output of the first residual module after passing through the residual block network and the input of the first residual block is the output of the first residual block and the input of the second residual block; the feature extraction network includes 3 residual modules, and the output of the third residual module passes through the maximum pooling layer module, the LSTM layer module and the fully connected layer module in sequence to obtain the feature vector of the current data sample.

[0034] Further, such as Figure 3 As shown, the step S102 incremental learning task and training test data set division method includes the following steps:

[0035] S1021, data set reading and multi-modal conversion: reading the signal data set, which contains all modulation categories to be identified, calculating common signal modes such as time-frequency transform, amplitude and phase, and constellation diagram according to the data in the task, and adding the calculated data and the copied labels to the original data set according to the principle of combining the same category nearby;

[0036] Taking amplitude (Amplitude) and phase (Phase) information as an example, the amplitude (Amplitude) and phase (Phase) information of the original IQ two-way signal are calculated and used together with the original IQ two-way signal as the model input, and the related elements of the amplitude vector and the phase vector can be calculated by the following formula.

[0037]

[0038] Among them, i represents each data vector, X I , X Q Respectively represent all I-channel signals and Q-channel sequence vectors in the signal, i represents a single data vector, n represents the sequence length of the sample data, and the data are combined to obtain a new data set X∈{X I ,X Q ,X A ,X P}; The label is in one-hot encoding form. The IQ signal usually refers to the in-phase and quadrature signal. The phase difference between the in-phase and quadrature signals is 90°.

[0039] S1022, task definition and data set segmentation: define several training tasks, determine the number of categories and specific categories of each task, and thereby determine the task data set corresponding to each task;

[0040] Further, if Figure 4 As shown in FIG. 1 , X represents the entire signal data set in step S1021, which includes a total of D signal categories, whose labels are C0, C1 to C D , β Cj Indicates the signal category C in the signal data set j The corresponding number of signal samples; the model training process corresponding to all modulation categories of the signal data set is divided into n tasks to form a task library {T0, T1, T2…T n}, the number of categories in each task can be set as needed and the number of categories in different tasks can be different. Represents task T iThe number of categories in ;

[0041] Figure 4 An example division method is given, in which each of the first n-1 tasks contains 4 or 5 types of modulation signals, and the nth task contains 3 types of modulation signals;

[0042] For each task, its task data set consists of samples corresponding to all modulation categories in the task. For example, all signal samples with label categories of C0, C1, C2 and C4 in the signal data set constitute the task data set of the T0 task.

[0043] S1023: Division of training and test data sets: dividing the task data set of each training task into a task training set and a task test set according to a specific ratio;

[0044] For example, a given task dataset can be divided in a ratio of 7:3;

[0045] Further, such as Figure 5 As shown, the method for training the incremental learning model and updating the hyperparameters in step S103 specifically includes the following steps:

[0046] S1031, hyperparameter initialization: setting parameters such as loss function, optimizer and adjustment factor of the training task, using num to represent the number of signal categories that have participated in the training process, initially set to 0, and setting the current task index δ to 0;

[0047] For example, the loss function can be set to the cross entropy loss function to calculate the classification loss during training; the optimizer can be set to the Adam optimizer with a weight decay factor of 0.95 and a learning rate of 0.0001;

[0048] S1032, example set compression: Take out a task with index δ from the task library, called the current task, and determine the value of δ; if δ is 0, the task taken out is T0, which is in the training stage of the original model, and initializes an example set P as an empty set; when δ is not 0, the current task taken out is represented by T δ , in the incremental learning process, the example set P is compressed according to the value of num;

[0049] Specifically, when δ is 0, the task taken out is T0, which is used as the training task of the original model. Let num be equal to Update the output vector dimension of the model output layer to num; when δ is not 0, the current task is T δ , update num to be equal to Update the example set E. To prevent the example set from exceeding the set memory constraint, compress it.

[0050] in, Indicates that the task is T δ , the number of all categories included in the task, Represents the i-th sample of signal category C0 in the signal data set, and its corresponding label is

[0051] The specific process of example set compression is as follows: according to the set memory constraint and the size of each sample, the total number of samples K of all categories in the example set and the sum of the number of all categories in the current task stage, i.e., the value of the parameter num, are determined, and the maximum number of samples M of each category after the old class samples participating in the training are compressed is calculated. The specific calculation formula is as follows;

[0052] M=K / num

[0053] This ensures that the device's available memory budget is always fully utilized but never exceeded. The compression process removes elements in a fixed order starting from the end, and it is a requirement of the example set construction routine to ensure that the example set meets the required approximation properties even after the removal process is completed;

[0054] S1033, training data set update: according to the task division in step S1032, the corresponding task data set divided in step S102 is obtained, and it is determined whether the δ value is 0. If it is 0, it is in the original model training stage and there is no need to update the data set, and the obtained task data set is directly used as the training data set of the current task; if δ is not 0, the compressed example set is combined with the obtained task training data set to update the training data set of the current task;

[0055] Specifically, the updating process of the task training data set is based on the example set of the current task, which is composed of P = (P1, P2, ..., P s-1 ) and receive the task training data set X=(X s ,…,X t ), we can combine the example set with the received new class sample data as follows to obtain the task training data set for the current task

[0056]

[0057] S1034, neural network model training: determine whether the parameter δ is 0. If it is 0, read the training data set in step S1033 to train the original model. The training of the original model only uses the classification loss calculated by the cross entropy loss function to update the parameters of the original model; if the parameter δ is not 0, on the basis of the original model, call the task data set in the task call sequence described in step S1033 to perform class increment training. In addition to the cross entropy loss between the prediction vector of the model itself and the true label during training, the present application also proposes a method for adaptive feature distillation. In addition to the output layer, the present application also uses the teacher model to distill the current model for the output of the hidden layer, thereby combining the adaptive feature distillation loss and the cross entropy loss to perform back propagation optimization of the model parameters;

[0058] Calculation of adaptive feature distillation loss: When the model is in the tth task stage, the backbone network part calculates the current training data set based on the current stage model and the t-1 stage model For feature extraction, the features extracted by the backbone network of the two-stage model are r t-1 With r t In order to minimize the feature difference between the previous and next stages and prevent significant changes in the model, the L2 norm is used to construct the feature difference loss:

[0059]

[0060] Where N represents the number of samples in the model dataset at stage t; for the classification output layer, the modified cross entropy loss is used, which is defined as follows:

[0061]

[0062] Among them, l is the number of labels, that is, the number of categories, i represents the number of samples calculated, and y0 is Corresponding to the predicted output of the teacher model and the student model for the current input data, y′0 and They represent the predicted output after distillation coefficient and normalization, respectively, and are calculated as follows:

[0063]

[0064] Where T represents the distillation coefficient and j represents the value of each element in the output vector.

[0065] The total loss function of the adaptive feature distillation method is obtained as:

[0066]

[0067] Where λ0 and μ0 are the adjustment factor hyperparameters between the losses;

[0068] The overall loss of model training is the sum of the loss of the predicted results and the actual label and the loss of adaptive feature distillation. The learning loss of new class samples adopts cross entropy loss, and the total loss is:

[0069]

[0070] in, represents the cross entropy loss between the true label and the predicted label, Represents the true label.

[0071] S1035, updating the example set: receiving the model parameters trained in step S1033, extracting features from the sample data in the task data set prepared in step S1032 according to categories, that is, extracting features from each category separately using the model, obtaining the average feature vector of all samples in each category and saving it, obtaining the Gaussian kernel between the feature vector of each sample and the average feature vector, obtaining M old example samples close to the average feature vector of each category, updating the example set and adding the obtained M samples to the example set;

[0072] Example set update steps combined Figure 6 As shown, first, obtain the input new training data set X = X s ,…,X t , the example set P saved after the last task = {P1, P2, ... P s}; Secondly, the feature extractor of the trained neural network Extract features and obtain the average feature vector μ for each newly added category i , the calculation formula is as follows:

[0073]

[0074] In order to meet the requirements of the example set construction in step S1032, a feature superposition method is adopted, that is, the previously saved sample features are added to the feature vector extracted from the sample to be saved in proportion, so that samples similar to the previously saved sample features can be found and saved by this method. When the sample is compressed, the previous sample can contain all the saved sample information; in the process of obtaining the update of each type of example set sample, the calculation formula of the superposition feature of the current sample feature and the previously saved example set sample is as follows:

[0075]

[0076] Among them, g represents the number of samples that have been saved for each type of signal;

[0077] Finally, the feature vector of each sample of the new class is calculated The feature vector of each sample and the average eigenvector μi The value of the Gaussian kernel function is obtained by comparing it with the average eigenvector μ i The M samples with the smallest Gaussian kernel are used as all the example set samples of each new class, and finally the new class example set P is made. e = {P s ,P s+1 ,…P t}, the calculation method is as follows. Adding the obtained new class example set to the example set of the previous task can obtain the final example set of the current task P = {P1, P2, ... P t}.

[0078]

[0079] S1036, determine whether to continue incremental training: after each task is completed, determine whether the current task is the last task based on whether the parameter δ reaches the set threshold, i.e., the number of tasks n. If it is the last task, output the final model. If not, repeat steps S1032 to S1035 until the incremental model training is completed, and add 1 to the parameter δ;

[0080] S1037, model parameter output: returns the final model parameters of class incremental learning.

[0081] Furthermore, the step S104 of classifier design and modulation classification specifically includes the following steps:

[0082] like Figure 7 As shown, obtain the example set P of the current task = {P1, P2, ... P t} and the trained feature extractor Traverse all samples of each type of modulation mode in the example set, calculate the average feature vector of each type and normalize it, and get all the average feature vectors μ in the example set = {μ1,μ2,…μ t}, the calculation formula is:

[0083]

[0084] Traversing each sample in the test set, the feature extractor extracts the feature vector of each sample The feature vector Sequentially compare all the average eigenvectors μ = {μ1,μ2,…μ t}Use the Gaussian kernel function to calculate the minimum value to obtain the category of each test set sample. The calculation formula is:

[0085]

[0086] Among them, σ represents the smoothness parameter, y * Indicates the total category of samples obtained.

[0087] The method described in the present application can alleviate the problem of catastrophic forgetting of deep learning models in the field of modulation recognition by limiting prototype parameters while saving a small number of samples, and has high recognition performance for modulated signals.

Claims

1. A method for quasi-incremental modulation recognition based on adaptive feature distillation prototype playback, characterized in that The steps include: S101, deep learning feature extraction network model design: for the modulation signal with single-channel input and the time series characteristics of the signal data being correlated, multi-modal calculation is performed on each signal data, and a hybrid structure deep neural network is designed to extract the features of the modulation signal; S102, quasi-incremental learning tasks and training and test data set division: the entire data set is divided into n tasks, and one task is selected as the initial task, that is, the task of learning the original model. The incremental learning is implemented by completing one task and then continuing to the next task; for each type of sample signal in each task, the signal is multi-modally calculated, combined with the original sample into a new signal set, and divided into a training set and a test set according to the ratio; S103, training of the quasi-incremental learning model and updating of hyperparameters: setting the training parameters and initialization example set of the quasi-incremental learning model, cyclically obtaining the training data set of each task, arbitrarily selecting a task as the task for training the original model, and using the remaining tasks for incremental learning. Before learning each task, the example set needs to be compressed, the output layer of the feature extraction model is incrementally updated, and the example set is saved after learning; after each task training is completed, it is determined whether there are tasks that have not been completed. If so, the next task is trained, otherwise the model and related parameters are output; S104, classifier design and modulation classification: After each task is completed, the trained model parameters are used to calculate the feature vector of each sample in the test set and the average feature vector of each type of signal in the example set, and the Gaussian kernel between the two is compared to obtain the modulation category of each sample in the test set; The specific steps of step S103 are as follows: S1031, hyperparameter initialization: set the loss function, optimizer and adjustment factor of the training task, use num to represent the number of signal categories that have participated in the training process, initially set to 0, and set the current task index δ to 0; S1032, example set compression: take out a task with index δ from the task library, called the current task, and determine the value of δ; If δ is 0, the task taken out is T0, which is in the training stage of the original model, and a sample set P is initialized as an empty set; When δ is not 0, the currently retrieved task is represented by T δ , in the incremental learning process, the example set P is compressed according to the value of num; S1033, training data set update: according to the result of task division, the corresponding divided task data set is obtained, and it is determined whether the δ value is 0. If it is 0, it is in the original model training stage and there is no need to update the data set. The obtained task data set is directly used as the training data set of the current task; If δ is not 0, the compressed example set is combined with the acquired task training data set to update the training data set of the current task; S1034, neural network model training: determine whether the parameter δ is 0. If it is 0, read the training data set to train the original model. The training of the original model uses the classification loss calculated by the cross entropy loss function to update the parameters of the original model. If the parameter δ is not 0, based on the original model, call the task data set in the task call order to perform class increment training. The training includes the cross entropy loss between the prediction vector of the model itself and the true label, and use the adaptive feature distillation method to distill the current model using the teacher model for the output of the hidden layer, so as to combine the adaptive feature distillation loss and the cross entropy loss to perform back propagation to optimize the model parameters. S1035, updating the example set: receiving the trained model parameters, extracting features of the sample data in the prepared task data set according to categories, obtaining the average feature vector of all samples in each category and saving it, obtaining the Gaussian kernel between the feature vector of each sample and the average feature vector, obtaining M old example samples of the average feature vector of each category, updating the example set and adding the obtained M samples to the example set; S1036, determine whether to continue incremental training: after each task is completed, determine whether the current task is the last task based on whether the parameter δ reaches the number of tasks n. If it is the last task, output the final model; if not, repeat steps S1032 to S1035 until the incremental model training is completed, and add 1 to the parameter δ; S1037, model parameter output: return the final model parameters of class incremental learning; In step S1034: Calculation of adaptive feature distillation loss: When the model is in the tth task stage, the backbone network part calculates the current training data set based on the current stage model and the t-1 stage model For feature extraction, the features extracted by the backbone network of the two-stage model are r t-1 With r t , use the L2 norm to construct feature difference loss: Where N represents the number of samples in the model dataset at stage t; for the classification output layer, the modified cross entropy loss is used, which is defined as follows: Among them, l is the number of labels, that is, the number of categories, i represents the number of samples calculated, and y0 is Corresponding to the predicted output of the teacher model and the student model for the current input data, y′0 and They represent the predicted output after distillation coefficient and normalization, respectively, and are calculated as follows: Where T represents the distillation coefficient, and j represents the value of each element in the output vector; The total loss function of the adaptive feature distillation method is obtained as: Where λ0 and μ0 are the adjustment factor hyperparameters between the losses; The overall loss of model training is the sum of the loss of the predicted results and the actual label and the loss of adaptive feature distillation. The learning loss of new class samples adopts cross entropy loss, and the total loss is: in, represents the cross entropy loss between the true label and the predicted label, Represents the true label.

2. The method for identifying quasi-incremental modulation based on adaptive feature distillation prototype playback according to claim 1, characterized in that: The deep learning feature extraction network model includes: 1 convolution module, 3 residual modules connected in sequence, a maximum pooling layer module, an LSTM layer module and a fully connected layer module, wherein the convolution module includes a first two-dimensional convolution layer, a first Dropout layer and a first SeLU activation function layer connected in sequence; the residual module includes a second two-dimensional convolution layer, a second Dropout layer, a second SeLU activation function layer, a third two-dimensional convolution layer, a third Dropout layer and a third SeLU activation function layer connected in sequence; the output of the convolution module is the input of the first residual module, and the sum of the output of the first residual module after passing through the residual block network and the input of the first residual block is the output of the first residual block and the input of the second residual block; the feature extraction network includes 3 residual modules, and the output of the third residual module passes through the maximum pooling layer module, the LSTM layer module and the fully connected layer module in sequence to obtain the feature vector of the current data sample.

3. The method for identifying quasi-incremental modulation based on adaptive feature distillation prototype playback according to claim 1, characterized in that: The specific steps of step S102 are as follows: S1021, data set reading and multimodal conversion: reading a signal data set, which contains all modulation categories to be identified, performing time-frequency transformation, amplitude and phase calculation, and constellation diagram calculation on the data in the task according to the category, and adding the calculated data and the copied labels to the original data set according to the principle of combining the same category nearby; S1022, task definition and data set segmentation: define several training tasks, determine the number of categories and specific categories of each task, and thereby determine the task data set corresponding to each task; S1023, division of training and test data sets: dividing the task data set of each training task into a task training set and a task test set according to a certain ratio.

4. The method for identifying quasi-incremental modulation based on adaptive feature distillation prototype playback as claimed in claim 3, characterized in that: In step S1021: The amplitude and phase information of the original IQ two-way signal are calculated and used together with the original IQ two-way signal as the model input. The related elements of the amplitude vector and the phase vector are calculated by the following formula: Among them, X I , X Q Respectively represent all I-channel signals and Q-channel sequence vectors in the signal, i represents a single data vector, n represents the sequence length of the sample data, and the data are combined to obtain a new data set X∈{X I ,X Q ,X A ,X P }; The label is in one-hot encoding form. The IQ signal refers to the in-phase and quadrature signal. The phase difference between the in-phase and quadrature signals is 90°.

5. The method for quasi-incremental modulation recognition based on adaptive feature distillation prototype playback as claimed in claim 3, characterized in that: In step S1022: X represents the entire signal dataset, which includes a total of D signal categories, whose labels are C0, C1 to C D , β Cj Indicates the signal category C in the signal data set j The corresponding number of signal samples; the model training process corresponding to all modulation categories of the signal data set is divided into n tasks to form a task library {T0, T1, T2…T n }, the number of categories in each task is set as needed and the number of categories in different tasks is different, let Represents task T i The number of categories in the task; the first n-1 tasks, each task contains 4 or 5 types of modulation signals, and the nth task contains 3 types of modulation signals; for each task, its task data set consists of samples corresponding to all modulation categories in the task.

6. The incremental modulation recognition method based on adaptive feature distillation prototype playback according to claim 1, characterized in that: In step S1032: When δ is 0, the task taken out is T0, which is used as the training task of the original model. Let num be equal to Update the output vector dimension of the model output layer to num; when δ is not 0, the current task is T δ , update num to be equal to Update the example set E. To prevent the example set from exceeding the set memory constraint, compress it. in, Indicates that the task is T δ , the number of all categories included in the task, Represents the i-th sample of signal category C0 in the signal data set, and its corresponding label is According to the set memory constraints and the size of each sample, determine the total number of samples K of all categories in the example set, as well as the sum of the number of categories in the current task stage, that is, the value of the parameter num, and calculate the maximum number of samples M of each category after the old class samples participating in the training are compressed. The specific calculation formula is as follows; M=K / num.

7. The method for identifying quasi-incremental modulation based on adaptive feature distillation prototype playback according to claim 1, characterized in that: The example set in step S1035 further includes the following steps: Get input new training data set X = X s ,…,X t , the example set P saved after the previous task = {P1, P2, ... P s }; Feature extractor of the trained neural network Extract features and obtain the average feature vector μ for each newly added category i , the calculation formula is as follows: The feature superposition method is adopted. The previously saved sample features are added to the feature vector extracted from the sample to be saved in proportion. Samples with similar features to the previously saved samples are found and saved. During the sample update process of each type of example set, the calculation formula for the superposition features of the current sample features and the previously saved example set samples is as follows: Among them, g represents the number of samples that have been saved for each type of signal; Calculate the feature vector for each sample of the new input class The feature vector of each sample and the average eigenvector μ i The value of the Gaussian kernel function is obtained by comparing it with the average eigenvector μ i The M samples with the smallest Gaussian kernel are used as all the example set samples of each new class, and finally the new class example set P is made. e = {P s ,P s+1 ,…P t }, the calculation method is as follows, Add the obtained new class example set to the example set of the previous task to obtain the final example set of the current task P = {P1, P2, ... P t }.

8. The method for quasi-incremental modulation recognition based on adaptive feature distillation prototype playback as claimed in claim 1, characterized in that: The classifier design and modulation classification in step S104 specifically include the following steps: Get the example set P of the current task = {P1, P2, ... P t } and the trained feature extractor Traverse all samples of each type of modulation mode in the example set, calculate the average feature vector of each type and normalize it, and get all the average feature vectors μ in the example set = {μ1,μ2,…μ t }, the calculation formula is: Traversing each sample in the test set, the feature extractor extracts the feature vector of each sample The feature vector Sequentially compare all the average eigenvectors μ = {μ1,μ2,…μ t }Use the Gaussian kernel function to calculate the minimum value to obtain the category of each test set sample. The calculation formula is: Among them, σ represents the smoothness parameter, y * Indicates the total category of samples obtained.

Citation Information

Patent Citations

  • Communication signal incremental modulation open set identification method based on deep learning

    CN117459356A

  • New and old feature compatible learning method for structure expansion and distillation

    CN117934923A