A Knowledge Distillation and Communication Modulation Recognition Method Based on a Mixture of Experts Model
Through the knowledge distillation method based on the mixed expert model, the existing communication modulation recognition model is solved, and the recognition rate of low signal-to-noise ratio is achieved is achieved, and the recognition rate and calculation efficiency are improved.
Patent Information
- Application Number
- CN202411808888.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-09
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2044-12-09
AI Technical Summary
There are two main problems in the existing communication modulation identification model: one is that the model is too large and cannot meet the needs of lightweight deployment; the other is that under the low signal-to-noise ratio, the recognition rate is low, making it difficult to effectively extract signal characteristics.
The knowledge distillation method based on the mixed expert model is adopted, and the data set is divided by a method without repeated random sampling, and the initial expert model, selection model and student model are constructed. Training is performed through the Adam optimizer and cross entropy loss function, combining batch normalization and fully connected layers, gradually reducing the number of model layers and dimensions. Use the T-softmax function to optimize the selection model to achieve lightweight and efficient feature extraction of the model.
The model is lightweighted, and the student model parameters are reduced to less than 30% of the original model, solving the problem of increasing model, and improving the recognition rate under low signal-to-noise ratio, ensuring the accurate extraction of high signal-to-noise ratio signal characteristics.
Smart Images

Figure CN119652715B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of communication technologies, and particularly relates to a knowledge distillation and communication modulation recognition method based on a mixture of experts model. Background Art
[0002] In a communication system, in order to achieve efficient long-distance transmission, the transmitting end needs to modulate the signal, and this process can convert the original information into an analog or digital signal suitable for propagation in the transmission medium. The receiving end then needs to use demodulation technology to restore the original information.
[0003] In cooperative communication, the transmitting end can inform the receiving end of the required modulation format for demodulating the data. However, in non-cooperative communication, the receiving end cannot determine the modulation method of the signal, so it must analyze the characteristics of the signal to determine its modulation method for subsequent signal demodulation operations and applications, and at the same time perform spectrum monitoring to ensure the orderly use of the wireless spectrum.
[0004] As Figure 1 is a schematic diagram of an existing communication modulation recognition process system. The signal is collected by an antenna and sent to a computer at the receiving end. After signal detection and preprocessing, the preprocessed signal is obtained. It is input into a modulation recognition model, and by extracting the characteristics of the signal, recognition is performed, so as to achieve the purpose of recognizing the modulation method of the signal.
[0005] Since O'shea published Convolutional radio modulation recognition networks in 2016, the networks for communication modulation recognition have gradually become more diverse and the accuracy has been continuously improved. Various deep learning structures have gradually been introduced to improve the accuracy of signal recognition.
[0006] However, with the continuous emergence of the above-mentioned models and the continuous increase in their technical recognition rates, there are the following defects: First, generally, the models for communication modulation recognition are relatively large and require a large amount of computing power, which cannot meet the increasingly required lightweight deployment; Second, in the common lightweight communication modulation recognition in the prior art, the recognition rate is still relatively low under low signal-to-noise ratio conditions. Summary of the Invention
[0007] The present invention provides a knowledge distillation and communication modulation recognition method based on a mixture of experts model. The technical solution disclosed by the present invention achieves lightweighting. Compared with common convolutional neural network models, the finally trained student model can achieve less than 30% of its parameters, solving the problem of the increasing size of communication modulation recognition models. At the same time, the present invention sets up a mixture of experts model divided according to the signal-to-noise ratio of different communication signals and integrates it with the selection model, solving the problem that the signal feature extraction of high signal-to-noise ratio signals is interfered by low signal-to-noise ratio signals.
[0008] To achieve the above object, the present invention proposes a knowledge distillation and communication modulation recognition method based on a mixture of experts model, including:
[0009] Step 1: Use the method of non-repetitive random sampling to divide and extract the signals in the wireless communication field dataset according to a preset ratio, and cut them into a training set, a validation set, and a test set;
[0010] Step 2: Construct a first training module through the Adam optimizer and the cross-entropy loss function, and construct a second training module through the Adam optimizer and the knowledge distillation loss function; respectively construct an initial expert model, an initial selection model, and an initial student model by adding batch normalization and fully connected layers in the convolutional layer; the number of layers and dimensions corresponding to the convolutional layer and the fully connected layer in the initial expert model, the initial student model, and the initial selection model gradually decrease; set association points associated with the intermediate layer of the initial expert model between the layers of the initial student model through the hook function in pytorch;
[0011] Step 3: Cut the training set according to the signal-to-noise ratio in the training set to obtain multiple training subsets; use the first training module to train the initial expert model through multiple training subsets respectively to obtain multiple expert models capable of communication modulation recognition, and the multiple expert models form a mixture of experts model;
[0012] Step 4: Use the first training module to train the initial selection model through the training set and the mixture of experts model. If the selection model does not assign weights to the mixture of experts model, use the T-softmax function to identify the optimal expert model matching the training set to obtain the trained selection model; integrate the mixture of experts model with the trained selection model to obtain the teacher model; if the selection model assigns weights to the mixture of experts model, then through the aggregation model, weight-sum the output weights of the selection model and the output weights of the expert model, and output the recognition result with the highest probability, and the recognition result is the output result of the teacher model;
[0013] Step 5: Use the second training module to train the initial student model through the correlation points of the optimal expert model and the training set to obtain the intermediate layer of the trained student model capable of communication modulation recognition; then use the second training module to train the output result of the initial student model through the output result of the teacher model and the training set, and combine it with the trained intermediate layer to form the trained student model capable of communication modulation recognition;
[0014] Step 6: Iterate Step 5. After reaching the preset number of times, the internal data of the student model is locked, and the student model is verified using the validation set. If the knowledge distillation loss function in the second training module converges and the mean squared error value shows overfitting, stop the verification to obtain a qualified student model; input the test set into the qualified student model to output the communication modulation recognition result.
[0015] In Step 2, the construction method of the first training module includes:
[0016] Create an Adam optimizer using the torch.optim.adam function in Python;
[0017] Input the cross-entropy loss function value of the cross-entropy loss function into the Adam optimizer. Through the Adam optimizer, make the cross-entropy loss function value gradient descend to train the model to be trained, dynamically adjust the learning rate of each parameter in the model to be trained, optimize the convergence of the cross-entropy loss function, and obtain the constructed first training module;
[0018] Among them, the cross-entropy loss function is:
[0019] CE = -Σ - tilog(pi);
[0020] Among them, CE is the cross-entropy loss function value; ti represents the correct answer, and pi represents the predicted value.
[0021] In Step 2, the construction method of the second training module includes:
[0022] Create an Adam optimizer using the torch.optim.adam function in Python;
[0023] Input the knowledge distillation loss function value of the knowledge distillation loss function into the Adam optimizer. Through the Adam optimizer, make the knowledge distillation loss function value gradient descend to train the model to be trained, dynamically adjust the learning rate of each parameter in the model to be trained, optimize the convergence of the knowledge distillation loss function, and obtain the constructed second training module;
[0024] Among them, the knowledge distillation loss function is:
[0025] FKD_LOSS = a * MSE_Loss_result + b * MSE_Loss_hook1 + c * MSE_Loss_hook2;
[0026] Where FKD_LOSS is the value of the knowledge distillation loss function, a represents the weight of the difference value between the self-learned result features of the student model, b represents the weight of the difference value between the layer features of the convolutional layers, and c represents the weight of the difference value between the layer features of the fully connected layers, which plays the role of reconciling the order of magnitude and weighting; MSE_Loss_result represents the MSE of the student model's self-learning; MSE_Loss_hook1 represents the MSE of the inter-layer features corresponding to the same dimension between the convolutional layers, and MSE_Loss_hook2 represents the MSE of the inter-layer features corresponding to the same dimension extracted between the fully connected layers. Among them, MSE is the mean squared error value, and the mean squared error function is expressed as follows:
[0027]
[0028] Where Za represents the true value, za represents the predicted value, and n represents the total number of predictions performed.
[0029] In step two, the construction methods of the initial expert model, the initial selection model, and the initial student model include:
[0030] Set the nn.nn.covd function in pytorch to determine the layer dimension of the convolutional layer;
[0031] Flatten and reduce the dimension of the output convolution through the torch.reshape function;
[0032] Then construct the intermediate layer through the nn.Linear function, nn.RELU function, and nn.BatchNorm2d function. After performing full connection, taking the positive value, and batch normalization on the intermediate layer respectively, construct the initial expert model, the initial selection model, and the initial student model; the layer dimensions corresponding to the convolutional layer and the fully connected layer in the initial expert model, the initial student model, and the initial selection model gradually decrease; and set the association points associated with the intermediate layer of the initial expert model between the layers of the initial student model through the hook function in pytorch.
[0033] In step three, the method of cutting the training set according to the signal-to-noise ratio in the training set to obtain multiple training subsets includes:
[0034] Cut the training sets corresponding to the lowest and highest segments of the signal-to-noise ratio in the training set, and evenly cut the remaining training sets corresponding to the intermediate segments to obtain multiple training subsets.
[0035] In step 3, the method of using the first training module to train the initial expert model through multiple training subsets respectively to obtain multiple trained expert models capable of communication modulation recognition, and the multiple expert models forming a mixture of experts model includes:
[0036] Input the multiple training subsets into the initial expert model respectively, and output the predicted probability values of the batch signals;
[0037] Take the class label corresponding to the maximum predicted probability value of the signal as the predicted class label, input the true class label and the predicted class label of the signal into the cross-entropy loss function in the first training module, and calculate the cross-entropy loss function value;
[0038] Input the cross-entropy loss function value into the Adam optimizer, and use the Adam optimizer to perform gradient descent on the cross-entropy loss function value to train the initial expert model, dynamically adjust the learning rate of each parameter in the initial expert model, optimize the convergence of the cross-entropy loss function, and obtain multiple trained expert models capable of communication modulation recognition corresponding to the multiple training subsets. The multiple expert models form a mixture of experts model.
[0039] In step 4, the method of using the first training module to train the initial selection model through the training set and the mixture of experts model includes:
[0040] Extract the features of each expert model in the mixture of experts model through the convolutional layer of the initial selection model, and at the same time extract the features of the input training set;
[0041] Use the first training module to calculate the cross-entropy loss function value between the features of each expert model in the mixture of experts model and the features of the training set, and use the Adam optimizer to perform gradient descent on the cross-entropy loss function value to train the initial selection model, dynamically adjust the learning rate of each parameter in the initial selection model, and optimize the convergence of the cross-entropy loss function.
[0042] In step 4, the T-softmax function is:
[0043]
[0044] where pT-i is the recognition rate distribution of the knowledge distillation result; i is a single signal target; T is the temperature parameter, and the value range is 0 < T < 1; pi is the recognition rate of the original teacher's single output signal target; pj is the recognition rate of all signal targets output by the original teacher; n is the total number of predictions;
[0045] The aggregation model is:
[0046]
[0047] Among them, A is the output result of the teacher model; k is the number of expert models in the teacher model; N is the total number of expert models in the teacher model; Bk is the output weight of the k-th expert model; PBk is the output weight of the k-th expert model corresponding to the selection model.
[0048] In step five, the method of using the second training module to train the initial student model through the correlation points of the optimal expert model and the training set to obtain the intermediate layer of the trained student model capable of communication modulation recognition includes:
[0049] Extract the features of the training set through the convolutional layer of the initial student model. The initial student model extracts the output result of the intermediate layer of the optimal expert model in the teacher model through the correlation points. Use the second training module to calculate the knowledge distillation loss function value between the output result of the intermediate layer of the optimal expert model and the features of the training set. Use the Adam optimizer to make the knowledge distillation loss function value decline in gradient to train the initial student model, dynamically adjust the learning rate of each parameter in the initial student model, optimize the convergence of the knowledge distillation loss function, and obtain the intermediate layer of the trained student model capable of communication modulation recognition.
[0050] In step five, then use the second training module to train the output result of the initial student model through the output result of the teacher model and the training set, and the method of forming the trained student model capable of communication modulation recognition with the trained intermediate layer includes:
[0051] Extract the features of the training set through the convolutional layer of the initial student model, and extract the output result of the teacher model through the convolutional layer of the initial student model. Use the second training module to calculate the knowledge distillation loss function value between the output result of the teacher model and the features of the training set. Use the Adam optimizer to make the knowledge distillation loss function value decline in gradient to train the initial student model, dynamically adjust the learning rate of each parameter in the initial student model, optimize the convergence of the knowledge distillation loss function value, and form the trained student model capable of communication modulation recognition with the trained intermediate layer of the student model capable of communication modulation recognition.
[0052] Advantages of the present invention:
[0053] 1. A method for knowledge distillation and communication modulation recognition based on a mixture of expert models disclosed in the present invention adopts a mixture of expert models to identify the modulation mode of communication signals, further refines the dataset in the wireless communication field, enables each expert model in the ensemble learning to more fully identify signal features, and the extraction of high signal-to-noise ratio signal features is not interfered by low signal-to-noise ratio signals.
[0054] 2. The present invention adopts feature-based knowledge distillation to lightweight the communication signal modulation recognition. By learning the inter-layer relationship of the optimal expert model in the teacher model, the inter-layer information of the communication signal is further utilized. The core of this patent adopts a hybrid expert model integrated selection model as the teacher model for lightweighting the communication modulation recognition network. The selection model in it is used to evaluate the recognition effect of the hybrid expert model, so as to integrate the hybrid expert model, and then the hybrid expert model is used to guide the training of a single student model in turn, greatly simplifying the computational complexity of the model. To a certain extent, it solves the problem that the signal feature extraction of high signal-to-noise ratio will be interfered by the signal of low signal-to-noise ratio, making the whole invention achieve the lightweight of the model. Compared with the common convolutional neural network model, the finally trained student model can reach less than 30% of its parameters, solving the problem of the increasing size of the model.
[0055] 4. By introducing the T-softmax function, this patent optimizes the selection model in the expert model, converting the serial operation of sorting on the CPU into the parallel operation on the GPU, increasing its inference speed at one step, optimizing the network performance, and reducing the network training time. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] Figure 1 It is a schematic diagram of the system for the communication modulation recognition process of the signal.
[0057] Figure 2 It is a schematic diagram of the structure of the knowledge distillation learning model based on response.
[0058] Figure 3 Schematic diagram of the structure of the feature-based knowledge distillation learning model.
[0059] Figure 4 It is a schematic diagram of the structure of a single expert model.
[0060] Figure 5 It is a schematic diagram of the structure of the hybrid expert model.
[0061] Figure 6 It is a schematic diagram of the structure of the selection model.
[0062] Figure 7 It is a flow chart of a knowledge distillation and communication modulation recognition method based on a hybrid expert model.
[0063] Figure 8 It is a schematic diagram of the modulation recognition accuracy of each network model under different signal-to-noise ratios. DETAILED DESCRIPTION OF THE INVENTION
[0064] To enable those skilled in the art to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application.
[0065] In the process of research, the inventors found that in the communication modulation recognition of signals, the key technology for lightweight construction of the model is knowledge distillation. Knowledge distillation algorithms are as Figure 2 , Figure 3 shown, and are divided into response-based and feature-based knowledge distillation. Both are to construct a "teacher model" with a large scale and high recognition rate and a "student model" with a small scale and low recognition rate, calculate the mean square error value MSE of the two, and then optimize the algorithm model through gradient descent iteration, so that the "student model" can learn from the "teacher model" to achieve the effect of model lightweight. At the same time, in the previous research, in order to solve the problem of how to select the "educational qualification" of the expert model in the case of a multi-expert model and realize the lightweight of the modulation recognition model, the present invention chooses to introduce a selection model in the hybrid expert model to better integrate knowledge distillation of the network, so as to achieve the purpose of lightweight. Below, the technical solutions of the present invention will be described in detail through specific embodiments.
[0066] The embodiment of the present invention provides a knowledge distillation and communication modulation recognition method based on a hybrid expert model, including:
[0067] Step 1: Use the method of non-repetitive random sampling to divide and extract the signals in the wireless communication field dataset according to a preset ratio, and cut them into a training set, a validation set, and a test set;
[0068] The preset ratio is preferably 7:2:1. Specifically, it is loaded through the cPickle.load function in the cPickle library in the computer programming language python, as well as np.zeros in the numpy library and the torch.tensor function in pytorch to set the input data array, and convert the format of the (modulation method, signal-to-noise ratio) key in the dataset into a larger data of (number of data, (modulation method, signal-to-noise ratio), data sampling points), and then convert the 10 modulation classifications into a 10-bit binary Onehot vector form.
[0069] Step 2: Construct a first training module using the Adam optimizer and the cross-entropy loss function, and construct a second training module using the Adam optimizer and the knowledge distillation loss function; respectively construct an initial expert model, an initial selection model, and an initial student model by adding batch normalization to the convolutional layer and a fully connected layer; the number of layers and dimensions corresponding to the convolutional layer and the fully connected layer in the initial expert model, the initial student model, and the initial selection model gradually decrease; set association points associated with the intermediate layers of the initial expert model between the layers of the initial student model using the hook function in pytorch; extract the output between the layers of the expert model through the hook function.
[0070] As Figure 4 shown, the single expert model constructed by the present invention can be composed of 4 convolutional layers and 7 fully connected layers. At the same time, batch normalization is added to the convolutional layer to enhance the feature extraction ability of the convolutional layer and enhance its convergence. Each individual expert model is mainly responsible for learning the signal data set with a fixed signal-to-noise ratio. As Figure 6 shown, the selection model, for example, has 2 convolutional layers and 3 fully connected layers, and the dimensions of all five layers are relatively small.
[0071] Step 3: Cut the training set according to the signal-to-noise ratio in the training set to obtain multiple training subsets; use the first training module to train the initial expert model through multiple training subsets respectively to obtain multiple expert models that are trained and capable of communication modulation recognition, and the multiple expert models form a mixture of expert models;
[0072] In the learning of the data set by the traditional signal modulation recognition technology, the data set is not divided. Since signal modulation recognition depends on the extraction and recognition of signal features, in the case of a large signal-to-noise ratio span, the learning of low signal-to-noise ratio may interfere with the learning under high signal-to-noise ratio conditions, resulting in a decrease in the signal recognition rate; in order to improve the signal recognition rate, this step divides the data set based on the signal-to-noise ratio.
[0073] Step 4: Use the first training module to train the initial selection model through the training set and the mixture of expert models. If the selection model does not weight the mixture of expert models, use the T-softmax function to identify the optimal expert model that matches the training set to obtain a trained selection model; integrate the mixture of expert models and the trained selection model to obtain a teacher model; if the selection model weights the mixture of expert models, then through the aggregation model, weight and sum the output weights of the selection model and the output weights of the expert models, and output the recognition result with the highest probability, and the recognition result is the output result of the teacher model;
[0074] In order to enable a single expert model to better learn datasets under similar signal-to-noise ratio conditions, the present invention proposes to use a mixture of experts model, that is, in the case of parallel multi-expert models, a selection model is further introduced to integrate multiple experts, so as to realize the recognition of signal modulation modes.
[0075] However, the sorting of the original mixture of experts model is based on the pre-weighting of the selection network, which belongs to CPU serial operation in the sorting problem and consumes a large amount of operation time. Therefore, as Figure 5 shown, the present invention combines the temperature parameter T-softmax function to further improve the selection of the mixture of experts model by adjusting the probability gap between different categories output by the mixture of experts model. In the knowledge distillation method based on response, the temperature parameter T is greater than 0, which serves to reduce the probability gap between different categories output by the mixture of experts model, enabling the student model to learn more knowledge from the teacher model. On the contrary, in the current selection of the mixture of experts model, it is necessary to select 0 < T < 1, which can further exacerbate the probability gap between different categories output by the mixture of experts model, making the weight assigned to the discarded mixture of experts model approach 0 and the weight of the adopted mixture of experts model approach 1, thereby achieving the effect of both being able to select the "optimal" mixture of experts model and performing parallel operations, further accelerating the inference speed. Then, the initial student model is trained through the optimal expert model, enabling the student model to learn the convolutional layer, fully connected layer, and output of the optimal expert model respectively, and iteratively adjusting the relevant coefficients through the Adam optimizer by itself to achieve the purpose of learning.
[0076] Step Five: Use the second training module to train the initial student model through the correlation points of the optimal expert model and the training set to obtain the intermediate layer of the trained student model capable of communication modulation recognition; then use the second training module to train the output result of the initial student model through the output result of the teacher model and the training set, and combine it with the trained intermediate layer to form the trained student model capable of communication modulation recognition, as Figure 7 shown.
[0077] Step Six: Iterate Step Five. After reaching the preset number of times, the internal data of the student model is locked, and the student model is verified using the validation set. If the knowledge distillation loss function in the second training module converges and the mean square error value shows overfitting, stop the verification to obtain a qualified student model; input the test set into the qualified student model to output the communication modulation recognition result.
[0078] After this step is iterated multiple times, setting the learning parameter of the student model to the teacher model to gradually decrease and the learning parameter of the student model to the dataset to gradually increase can enable the student model to independently learn the modulation mode.
[0079] In step two, the construction method of the first training module includes:
[0080] Create an Adam optimizer using the torch.optim.adam function in Python;
[0081] Input the cross-entropy loss function value of the cross-entropy loss function into the Adam optimizer. Through the Adam optimizer, make the cross-entropy loss function value gradient descent to train the model to be trained, dynamically adjust the learning rate of each parameter in the model to be trained, optimize the convergence of the cross-entropy loss function, and obtain the constructed first training module;
[0082] Among them, the cross-entropy loss function is:
[0083] CE = -Σ - tilog(pi);
[0084] Among them, CE is the cross-entropy loss function value; ti represents the correct answer, and pi represents the predicted value.
[0085] In step two, the construction method of the second training module includes:
[0086] Create an Adam optimizer using the torch.optim.adam function in Python;
[0087] Input the knowledge distillation loss function value of the knowledge distillation loss function into the Adam optimizer. Through the Adam optimizer, make the knowledge distillation loss function value gradient descent to train the model to be trained, dynamically adjust the learning rate of each parameter in the model to be trained, optimize the convergence of the knowledge distillation loss function, and obtain the constructed second training module;
[0088] Among them, the knowledge distillation loss function is:
[0089] FKD_LOSS = a * MSE_Loss_result + b * MSE_Loss_hook1 + c * MSE_Loss_hook2;
[0090] Among them, FKD_LOSS is the knowledge distillation loss function value, a represents the weight of the difference value between the self-learned result features of the student model, b represents the weight of the difference value between the convolutional layer features, c represents the weight of the difference value between the fully connected layer features, which plays a role in reconciling the order of magnitude and weighting; MSE_Loss_result represents the MSE of the student model's self-learning; MSE_Loss_hook1 represents the MSE of the inter-layer features corresponding to the same dimension between convolutional layers, MSE_Loss_hook2 represents the MSE of the inter-layer features corresponding to the same dimension extracted between fully connected layers. Among them, MSE is the mean square error value, and the mean square error function is expressed as follows:
[0091]
[0092] Among them, Za represents the true value, za represents the predicted value, and n represents the total number of predictions performed.
[0093] In step two, the construction methods of the initial expert model, the initial selection model, and the initial student model include:
[0094] Set the nn.nn.covd function in pytorch to determine the layer dimension of the convolutional layer;
[0095] Flatten and reduce the dimension of the output convolution through the torch.reshape function;
[0096] Then, construct the intermediate layer through the nn.Linear function, nn.RELU function, and nn.BatchNorm2d function. After performing full connection, taking the positive value, and batch normalization on the intermediate layer respectively, construct the initial expert model, the initial selection model, and the initial student model; the layer dimensions corresponding to the convolutional layer and the full connection layer in the initial expert model, the initial student model, and the initial selection model gradually decrease; and set the association points associated with the intermediate layer of the initial expert model between the layers of the initial student model through the hook function in pytorch.
[0097] In step three, the method of cutting the training set according to the signal-to-noise ratio in the training set to obtain multiple training subsets includes:
[0098] Cut the training sets corresponding to the lowest and highest segments of the signal-to-noise ratio in the training set, and evenly cut the remaining training sets corresponding to the intermediate segments to obtain multiple training subsets.
[0099] When setting the signal-to-noise ratio grouping, it should also be noted that when the grouping signal-to-noise ratio span in the signal-to-noise ratio grouping is too large, the recognition rate of this network will be affected due to increased interference, resulting in the selection model not being inclined to select, and the network recognition effect will be poor; similarly, if the setting is too detailed, it will lead to an increase in the calculation amount of the selection model, resulting in a waste of model computing resources. Therefore, according to the solution disclosed in the present invention, the RML2016.10b dataset can be divided into 6 expert models according to the signal-to-noise ratio. For the lowest 4 segments (-20dB, -18dB, -16dB, -14dB) and the highest 4 segments (12dB, 14dB, 16dB, 18dB) of the signal-to-noise ratio, 1 expert model is set for learning and recognition respectively, and the rest are evenly distributed, with each 3 segments corresponding to 1 expert model, so that it can perform learning modulation classification recognition within a relatively reasonable range.
[0100] Since the overall recognition rate of the 4 groups with the lowest signal-to-noise ratio is relatively low, such as Figure 8As shown, it contains less information and is divided into a group mainly for identification in low signal-to-noise ratio situations. The overall characteristics of the top 4 groups are obvious. Unifying them for an expert to learn is beneficial to increasing the accuracy of feature extraction, improving the role of the expert in lower signal-to-noise ratios, and making up for the problem of unclear identification at some intersection signal-to-noise ratios. One expert model is set for each group to learn and identify, and the rest are evenly divided, each corresponding to one expert model, so that they can learn, modulate, and classify and identify within a relatively reasonable range.
[0101] In step three, using the first training module, training the initial expert model through multiple training subsets respectively to obtain multiple trained expert models capable of communication modulation identification. The method of forming a mixture of expert models by multiple expert models includes:
[0102] Input multiple training subsets into the initial expert model respectively, and output the predicted probability values of the batch signals;
[0103] Take the class label corresponding to the maximum predicted probability value of the signal as the predicted class label, and input the true class label and predicted class label of the signal into the cross-entropy loss function in the first training module to calculate the cross-entropy loss function value;
[0104] Input the cross-entropy loss function value into the Adam optimizer, and use the Adam optimizer to make the cross-entropy loss function value gradient descent to train the initial expert model, dynamically adjust the learning rate of each parameter in the initial expert model, optimize the convergence of the cross-entropy loss function, and obtain multiple trained expert models capable of communication modulation identification corresponding to multiple training subsets. Multiple expert models form a mixture of expert models.
[0105] In step four, using the first training module, the method of training the initial selection model through the training set and the mixture of expert models includes:
[0106] Extract the features of each expert model in the mixture of expert models respectively through the convolutional layer of the initial selection model, and at the same time extract the features of the input training set;
[0107] Use the first training module to calculate the cross-entropy loss function value of the features of each expert model in the mixture of expert models and the features of the training set, and use the Adam optimizer to make the cross-entropy loss function value gradient descent to train the initial selection model, dynamically adjust the learning rate of each parameter in the initial selection model, and optimize the convergence of the cross-entropy loss function.
[0108] In step four, the T-softmax function is:
[0109]
[0110] Among them, pT-i is the recognition rate distribution of knowledge distillation results; i is a single signal target; T is the temperature parameter, and the value range is 0 < T < 1; pi is the recognition rate of the original teacher's single output signal target; pj is the recognition rate of all signal targets output by the original teacher; n is the total number of predictions performed.
[0111] The selection model learns from the input data to select which expert the data matches for the best learning effect. That is, the smaller the value of the calculated cross-entropy loss function, the closer the distribution of the two probabilities. Select the expert model corresponding to the minimum value to integrate multiple experts; at the same time, classify the integration situation: if it is weighted, directly output the weights and perform weighted averaging with the experts to output the recognition result. If not selected, use the above T-softmax function. By presetting the value of T in the range (0,1), make its output performance as a selection.
[0112] Among them, the aggregation model is:
[0113]
[0114] Among them, A is the output result of the teacher model; k is the number of expert models in the teacher model; N is the total number of expert models in the teacher model; Bk is the output weight of the k-th expert model; PBk is the output weight of the k-th expert model corresponding to the selection model.
[0115] In step five, the method of using the second training module to train the initial student model through the correlation points of the optimal expert model and the training set to obtain the intermediate layer of the trained student model capable of communication modulation recognition includes:
[0116] Extract the features of the training set through the convolutional layer of the initial student model. The initial student model extracts the output result of the intermediate layer of the optimal expert model in the teacher model through the correlation points. Use the second training module to calculate the knowledge distillation loss function value between the output result of the intermediate layer of the optimal expert model and the features of the training set. Use the Adam optimizer to make the gradient of the knowledge distillation loss function decrease to train the initial student model, dynamically adjust the learning rate of each parameter in the initial student model, and optimize the convergence of the knowledge distillation loss function to obtain the intermediate layer of the trained student model capable of communication modulation recognition.
[0117] In step five, then use the second training module to train the output result of the initial student model through the output result of the teacher model and the training set, and the method of forming the trained student model capable of communication modulation recognition by combining with the trained intermediate layer includes:
[0118] Extract the features of the training set through the convolutional layer of the initial student model, and extract the output results of the teacher model through the convolutional layer of the initial student model. Use the second training module to calculate the value of the knowledge distillation loss function between the output results of the teacher model and the features of the training set. Train the initial student model by making the gradient of the knowledge distillation loss function value decrease through the Adam optimizer, dynamically adjust the learning rate of each parameter in the initial student model, optimize the convergence of the knowledge distillation loss function value, and form a trained student model capable of communication modulation recognition with the intermediate layer of the trained student model capable of communication modulation recognition.
[0119] To better illustrate the effect of the present invention, the experimental verification results of a knowledge distillation and communication modulation recognition method based on a mixture of experts model proposed by the present invention are now disclosed, as Figure 8 shown.
[0120] This experiment lists the mixture of experts model, a single teacher model, a student model without knowledge distillation, a student model with knowledge distillation for a single teacher model, a student model after FKD with serial operation, and a student model after FKD with parallel operation. It can be Figure 8 found that by using the technical solution disclosed in the present invention through the weight-selection mechanism of the selection network in the mixture of experts model, the accuracy of the model has been significantly improved. At the same time, compared with the serial network, the parallel network introduces the T-softmax function, making the weight assigned to the discarded network approach 0, the weight of the adopted network approach 1, and some critical weight points will change a single selection into multiple selections, which can not only improve the training speed but also slightly improve the accuracy of the model, solving the problem of slow serial training speed.
[0121] The beneficial effects of the embodiments of the present invention:
[0122] 1. A knowledge distillation and communication modulation recognition method based on a mixture of experts model disclosed in the present invention adopts a mixture of experts model to identify the modulation mode of communication signals, further refines the dataset in the field of wireless communication, enables each expert model in the ensemble learning to more fully identify the signal features, and the extraction of signal features with high signal-to-noise ratio is not interfered by signals with low signal-to-noise ratio.
[0123] 2. The present invention adopts feature-based knowledge distillation to lightweight the communication signal modulation recognition. By learning the inter-layer relationship of the optimal expert model in the teacher model, the inter-layer information of the communication signal is further utilized. The core of this patent adopts the hybrid expert model integration selection model as the teacher model for lightweighting the communication modulation recognition network. The selection model in it is used to evaluate the recognition effect of the hybrid expert model, so as to integrate the hybrid expert model, and then the hybrid expert model is used to alternately guide the training of a single student model, greatly simplifying the computational complexity of the model. To a certain extent, it solves the problem that the signal feature extraction of high signal-to-noise ratio is interfered by the signal of low signal-to-noise ratio, making the whole of the present invention achieve the lightweight of the model. Compared with the common convolutional neural network model, the finally trained student model can reach less than 30% of its parameters, solving the problem of the increasing size of the model.
[0124] 4. By introducing the T-softmax function, this patent optimizes the selection model in the expert model, converting the serial operation of sorting on the CPU into a parallel operation on the GPU, increasing its inference speed at one step, optimizing the network performance, and reducing the network training time.
[0125] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. A knowledge distillation and communication modulation recognition method based on a hybrid expert model, characterized in that: Including: Step 1: Use the method of non-repetitive random sampling to divide and extract the signals in the wireless communication field dataset according to a preset ratio, and cut them into a training set, a validation set, and a test set; Step 2: Construct a first training module through the Adam optimizer and the cross-entropy loss function, and construct a second training module through the Adam optimizer and the knowledge distillation loss function; respectively construct an initial expert model, an initial selection model, and an initial student model by adding batch normalization in the convolutional layer and a fully connected layer; the number of layers and dimensions corresponding to the convolutional layer and the fully connected layer in the initial expert model, the initial student model, and the initial selection model gradually decrease; set association points associated with the intermediate layer of the initial expert model between the layers of the initial student model through the hook function in pytorch; Step 3: Cut the training set according to the signal-to-noise ratio in the training set to obtain multiple training subsets; Use the first training module to train the initial expert model through multiple training subsets respectively to obtain multiple trained expert models capable of communication modulation recognition, and the multiple expert models form a mixture of expert models; Step 4: Use the first training module to train the initial selection model through the training set and the mixture of expert models. If the selection model does not weight the mixture of expert models, use the T-softmax function to identify the optimal expert model matching the training set to obtain a trained selection model; integrate the mixture of expert models and the trained selection model to obtain a teacher model; if the selection model weights the mixture of expert models, then through the aggregation model, weight and sum the output weights of the selection model and the output weights of the expert models to output the recognition result with the highest probability, and the recognition result is the output result of the teacher model; where the T-softmax function is: Where pT-i is the recognition rate distribution of the knowledge distillation result; i is a single signal target; T is the temperature parameter, and the value range is 0 < T < 1; pi is the recognition rate of the original teacher's single output signal target; pj is the recognition rate of all signal targets output by the original teacher; n is the total number of predictions made; Step 5: Use the second training module to train the initial student model through the association points of the optimal expert model and the training set to obtain the intermediate layer of the trained student model capable of communication modulation recognition; then use the second training module to train the output result of the initial student model through the output result of the teacher model and the training set, and form a trained student model capable of communication modulation recognition with the trained intermediate layer; Step 6: Iterate Step 5. After reaching the preset number of times, lock the internal data of the student model, use the validation set to verify the student model. If the knowledge distillation loss function in the second training module converges and the mean square error value shows overfitting, stop the verification to obtain a qualified student model; input the test set into the qualified student model to output the communication modulation recognition result.
2. The method for knowledge distillation and communication modulation recognition based on a hybrid expert model as claimed in claim 1, characterized in that: In Step 2, the construction method of the first training module includes: Create an Adam optimizer using the torch.optim.adam function in python; The cross entropy loss function value of the cross entropy loss function is input into the Adam optimizer, and the model to be trained is trained by making the cross entropy loss function value gradient decrease through the Adam optimizer, and the learning rate of each parameter in the model to be trained is dynamically adjusted to optimize the convergence of the cross entropy loss function, thereby obtaining the first training module constructed; Among them, the cross entropy loss function is: CE = -Σ-tilog(pi); Among them, CE is the cross entropy loss function value; ti represents the correct answer, and pi represents the predicted value.
3. The method for knowledge distillation and communication modulation recognition based on a hybrid expert model as claimed in claim 1, characterized in that: In step 2, the method for constructing the second training module includes: Create the Adam optimizer in Python using the torch.optim.adam function; The knowledge distillation loss function value of the knowledge distillation loss function is input into the Adam optimizer, and the knowledge distillation loss function value is gradient-decreased to train the to-be-trained model through the Adam optimizer, and the learning rate of each parameter in the to-be-trained model is dynamically adjusted to optimize the convergence of the knowledge distillation loss function, thereby obtaining the constructed second training module; Among them, the knowledge distillation loss function is: FKD_LOSS=a*MSE_Loss_result+b*MSE_Loss_hook1+c*MSE_Loss_hook2; Among them, FKD_LOSS is the value of the knowledge distillation loss function, a represents the weight of the difference between the features of the student model's self-learning results, b represents the weight of the difference between the features of the convolutional layers, and c represents the weight of the difference between the features of the fully connected layers, which plays a role in reconciling the order of magnitude and empowering; MSE_Loss_result represents the MSE of the student model's self-learning; MSE_Loss_hook1 represents the MSE of the inter-layer features of the same dimension between convolutional layers, and MSE_Loss_hook2 represents the MSE of the inter-layer features of the same dimension extracted between fully connected layers. MSE is the mean square error value, and the mean square error function is expressed as follows: Among them, Za represents the true value, za represents the predicted value, and n represents the total number of executed predictions.
4. The method for knowledge distillation and communication modulation recognition based on a hybrid expert model according to claim 1, characterized in that: In step 2, the construction methods of the initial expert model, the initial selection model and the initial student model include: Set the nn.nn.covd function in pytorch to determine the number of dimensions of the convolutional layer; The output convolution is flattened and reduced in dimension through the torch.reshape function; Then, the intermediate layer is constructed by nn.Linear function, nn.RELU function and nn.BatchNorm2d function, and the intermediate layer is fully connected, positively correlated and batch normalized respectively, and then the initial expert model, the initial selection model and the initial student model are constructed respectively; the number of layer dimensions corresponding to the convolutional layer and the fully connected layer in the initial expert model, the initial student model and the initial selection model gradually decreases; and the hook function in pytorch is used to set the association points associated with the intermediate layer of the initial expert model between the layers of the initial student model.
5. The method for knowledge distillation and communication modulation recognition based on a hybrid expert model as claimed in claim 1, characterized in that: In step 3, the training set is cut according to the signal-to-noise ratio in the training set to obtain multiple training subsets, including: The training sets corresponding to the lowest and highest signal-to-noise ratio segments in the training set are cut, and the training sets corresponding to the remaining middle segments are evenly cut to obtain multiple training subsets.
6. The method for knowledge distillation and communication modulation recognition based on a hybrid expert model as claimed in claim 5, characterized in that: In step 3, the first training module is used to train the initial expert model through multiple training subsets to obtain multiple trained expert models capable of communication modulation recognition. The method of forming a hybrid expert model with multiple expert models includes: Input multiple training subsets into the initial expert model respectively, and output the predicted probability values of batch signals; The category label corresponding to the maximum predicted probability value of the signal is used as the predicted category label, the real category label and the predicted category label of the signal are input into the cross entropy loss function in the first training module, and the cross entropy loss function value is calculated; The cross entropy loss function value is input into the Adam optimizer, and the initial expert model is trained by making the cross entropy loss function value gradient descend through the Adam optimizer. The learning rate of each parameter in the initial expert model is dynamically adjusted to optimize the convergence of the cross entropy loss function, and multiple trained expert models that can recognize communication modulation corresponding to multiple training subsets are obtained. Multiple expert models constitute a hybrid expert model.
7. The method for knowledge distillation and communication modulation recognition based on a hybrid expert model as claimed in claim 1, characterized in that: In step 4, the method of using the first training module to train the initial selection model through the training set and the hybrid expert model includes: The features of each expert model in the mixed expert model are extracted through the convolution layer of the initial selection model, and the features of the input training set are extracted at the same time; Using the first training module, the cross entropy loss function value of the features of each expert model in the hybrid expert model and the features of the training set is calculated. The initial selection model is trained by gradient descent of the cross entropy loss function value through the Adam optimizer. The learning rate of each parameter in the initial selection model is dynamically adjusted to optimize the convergence of the cross entropy loss function.
8. The method for knowledge distillation and communication modulation recognition based on a hybrid expert model as claimed in claim 1, characterized in that: In step 4, the aggregation model is: Among them, A is the output result of the teacher model; k is the number of expert models in the teacher model; N is the total number of expert models in the teacher model; Bk is the output weight of the kth expert model; PBk is the output weight of the kth expert model corresponding to the selection model.
9. The method for knowledge distillation and communication modulation recognition based on a hybrid expert model as claimed in claim 1, characterized in that: In step 5, the method of using the second training module to train the initial student model through the associated points of the optimal expert model and the training set to obtain the intermediate layer of the trained student model capable of communication modulation recognition includes: The features of the training set are extracted through the convolutional layer of the initial student model. The initial student model extracts the output results of the intermediate layer of the optimal expert model in the teacher model through the association points. The second training module is used to calculate the knowledge distillation loss function value of the output results of the intermediate layer of the optimal expert model and the features of the training set. The initial student model is trained by gradient descent of the knowledge distillation loss function value through the Adam optimizer. The learning rate of each parameter in the initial student model is dynamically adjusted to optimize the convergence of the knowledge distillation loss function, and the trained intermediate layer of the student model capable of communication modulation recognition is obtained.
10. The method for knowledge distillation and communication modulation recognition based on a hybrid expert model according to claim 9, characterized in that: In step 5, the method of using the second training module to train the output result of the initial student model through the output result of the teacher model and the training set, and forming a trained student model capable of communication modulation recognition with the trained intermediate layer includes: The features of the training set are extracted through the convolution layer of the initial student model, and the output results of the teacher model are extracted through the convolution layer of the initial student model. The second training module is used to calculate the knowledge distillation loss function value of the output results of the teacher model and the features of the training set. The initial student model is trained by gradient descent of the knowledge distillation loss function value through the Adam optimizer. The learning rate of each parameter in the initial student model is dynamically adjusted to optimize the convergence of the knowledge distillation loss function value. The trained student model capable of communication modulation recognition is composed of the intermediate layer of the trained student model capable of communication modulation recognition.
Citation Information
Patent Citations
Radar signal modulation mode identification method based on knowledge distillation
CN113343796A
Method and device for automatically modulating and identifying fine-grained signal based on distillation learning
CN117993477A