Low-cost signal detection model deployment method based on learning genes
By extracting the learning gene module and assisting in the design of lightweight signal detectors, the problem of high cost and low efficiency of deep learning-based signal detection models in CSI acquisition and training time is solved, and the rapid training and high generalization deployment of signal detection models are realized, reducing the training costs of manufacturers and protecting intellectual property rights.
Patent Information
- Application Number
- CN202510251783.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-06-06
AI Technical Summary
Deep learning-based signal detection models have high cost and low efficiency in acquiring a large number of CSIs and training time, especially in the field of multi-input and multi-output (MIMO) signal detection. The model has insufficient generalization ability in different scenarios, resulting in frequent retraining, which increases training overhead.
The low-cost signal detection model deployment method based on learning genes is adopted. By designing a large-scale population model and extracting plug-and-play learning gene modules from it, it assists in designing a lightweight intelligent signal detector, supporting meta-knowledge transfer between heterogeneous networks, and achieving rapid training and deployment of the model.
It effectively reduces the training costs of equipment manufacturers, improves the generalization ability and adaptability of the model, realizes the rapid convergence and high-performance deployment of the signal detection model, and protects the intellectual property rights of equipment manufacturers.
Smart Images

Figure CN120106174A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of communication technology, and in particular to a low-cost signal detection model deployment method based on learning genes. Background Art
[0002] With the increasing communication demands and environmental complexity, traditional solutions for communication tasks in the wireless physical layer have encountered bottlenecks, while deep learning (DL) has shown great potential in tasks such as channel estimation, modulation, and signal detection. In the field of multiple-input multiple-output (MIMO) signal detection, a large number of studies have shown that the performance of DL-based signal detection models is better than traditional methods in a variety of scenarios. However, the premise for ensuring the performance of DL-based signal detection models is to have a rich training data set and a properly configured neural network architecture. In addition, for DL-based signal detection models, channel state information (CSI) is also indispensable, and a large number of studies have used CSI as an explicit input for the detection model.
[0003] In view of the above characteristics, the design of DL-based signal detection models faces three major challenges. First, it is very expensive to obtain a large amount of CSI in practice, and there is often a problem of insufficient training samples. Second, when the system is configured with a large number of antennas, the NN architecture tends to become quite large, resulting in a long training time. Third, the statistical characteristics of CSI vary greatly in different scenarios, such as line-of-sight and non-line-of-sight conditions, path loss and delay, etc., resulting in a sharp drop in the performance of the trained NN when applied to new scenarios. At this time, it is necessary to use the new scenario data set for retraining, and training from scratch will incur a large training overhead, which will impose a burden on equipment manufacturers. Therefore, DL-based signal detection models often have obvious generalization errors or incur high training overhead.
[0004] In order to solve the above problems, knowledge transfer technology can be combined when designing a DL-based signal detection model, including pre-training fine-tuning, meta-learning, etc. The core idea of knowledge transfer technology is to transfer knowledge from previous fields or tasks to strengthen the learning of the target task, thereby reducing the need for large-scale training data sets, accelerating the training process, improving generalization capabilities, and reducing training costs. However, existing DL detectors based on knowledge transfer technology lack flexibility and are difficult to flexibly adapt to devices of different sizes and computing powers (such as mobile phones, smart watches, etc.), and require a lot of artificial intelligence expertise. In summary, how to reduce the training costs and technical barriers of manufacturers to develop strong generalization models, build and train highly generalized detection models for various types of devices, and achieve rapid convergence of models has become a key issue. Summary of the invention
[0005] The present invention provides a low-cost signal detection model deployment method based on learning genes, trains a highly generalized population model, extracts plug-and-play learning gene modules from it, and uses the modules to assist in the design and deployment of lightweight intelligent signal detectors, thereby supporting meta-knowledge migration between heterogeneous networks and realizing rapid training of detection models, effectively reducing manufacturers' training costs while protecting their intellectual property rights.
[0006] The embodiment of the present invention provides a low-cost signal detection model deployment method based on learning genes, comprising the following steps:
[0007] Step 1: Large equipment manufacturers have sufficient computing power and data set resources to design large-scale detection models and use massive datasets from different scenarios. Train sequentially to obtain a strong generalization group model f col (:,Θ col );
[0008] Step 2: Large equipment manufacturers analyze the gradient information of each layer of the group model According to the gradient significance Select and extract some structures and corresponding parameters to obtain a plug-and-play embeddable module - the learning gene Θ lg ,This module contains sufficient detection meta-knowledge for large manufacturers to conduct subsequent R&D or sell it to small manufacturers;
[0009] Step 3: Equipment manufacturers can inherit the learning genes to the lightweight individual model f according to the equipment and scenario requirements of the target task. ind (:,Θ ind ), thereby realizing the transfer of meta-knowledge between heterogeneous NNs. Among them, the individual model consists of learning genes Θ lg and any expansion Θ expIt consists of two parts. The structure of the latter can be designed according to the input and output dimensions, device computing power, etc., and the corresponding parameters are randomly initialized.
[0010] Step 4: Equipment manufacturers use the target signal detection task dataset Train individual models and achieve rapid convergence of individual models with the assistance of learning genes rich in meta-knowledge.
[0011] Optionally, in one embodiment of the present invention, in step 1, the input of the large-scale population model is the result of the least squares (LS) channel estimation. and the results of the Zero Force (ZF) signal detection scheme The output is the estimated signal detection result NN consists of convolutional layers, batch normalization layers, and fully connected layers.
[0012] Optionally, in one embodiment of the present invention, in step 1, the population model is trained using a gradient descent method, and a cost function is minimized by iterating NN parameters. The cost function is a mean square error (MSE), which is described as follows:
[0013]
[0014] Among them, (n) is the sample index, N is the number of samples, x is the real transmitted signal, ||·|| 2 is the Euclidean norm.
[0015] Optionally, in one embodiment of the present invention, in step 2, the gradient of the population model Described as:
[0016]
[0017] in, is the cost function of the group model, Θ col represents the parameters of the population model, k∈{1,…,K} represents the kth training data set, l∈{1,…,L col} represents the layer index of the population model, Represents the parameter index of the lth layer of the population model.
[0018] Optionally, in one embodiment of the present invention, in step 2, the gradient significance of the lth layer on the kth task It is defined as the proportion of parameters with larger gradients in a layer, as described below:
[0019]
[0020] Where σ represents the threshold value. Φ=1, otherwise 0. As the training task number k increases, when the “ When the phenomenon of "gradually decreasing from large to stable" occurs, it means that this layer can be selected as a learning gene. The set of all NN layers that meet the above requirements is the learning gene Θ lg .
[0021] Optionally, in one embodiment of the present invention, in step 3, the input and output of the individual model are consistent with the group model, and the data set is NN consists of convolutional layers, batch normalization layers, and fully connected layers. It is lighter in scale than the swarm model, and any extended part can be freely designed according to the antenna dimension, computing power and capacity of the device carrying the detection model, etc.
[0022] Optionally, in one embodiment of the present invention, in step 4, the individual model is trained using a gradient descent method, and the cost function is described as follows:
[0023]
[0024] Among them, M represents the number of layers contained in the learning gene, Θ m represents the NN parameters at the inherited position of the learning gene in the individual model, trained by the target dataset, ||·|| 2 is the Euclidean norm, and λ represents the weight of the L2 regularization term. The cost function consists of two parts: the first part is the MSE between the emission signal reconstructed by the NN and the true emission signal, and the second part is to limit the update of the learning gene unit parameters so as to retain meta-knowledge while adapting to new tasks.
[0025] The low-cost signal detection model deployment method based on learning genes in the embodiment of the present invention, while ensuring detection performance, extracts and utilizes flexibly expandable, plug-and-play learning gene modules to assist in the design and deployment of intelligent signal detectors, supports meta-knowledge transfer across networks, and realizes rapid training of signal detection models, effectively reducing manufacturers' training costs while protecting their intellectual property rights.
[0026] Additional aspects and advantages of the present invention will be given in part in the following description and in part will be obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] The above and / or additional aspects and advantages of the present invention will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:
[0028] Figure 1 A flowchart of a low-cost signal detection model deployment method based on learning genes provided according to an embodiment of the present invention;
[0029] Figure 2 A framework diagram of a MIMO signal detection system based on deep learning according to an embodiment of the present invention;
[0030] Figure 3 A schematic diagram of extracting and expanding learning genes according to an embodiment of the present invention;
[0031] Figure 4 To learn the inheritance mode of genes in the individual model according to the embodiment of the present invention;
[0032] Figure 5 A low-cost signal detection model deployment method based on learning genes according to an embodiment of the present invention. DETAILED DESCRIPTION
[0033] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present invention, and should not be construed as limiting the present invention.
[0034] The DL-based MIMO detection model has shown very superior performance in many scenarios, but its data-driven characteristics lead to limited generalization capabilities. In a dynamically changing and diverse wireless communication environment, the actual deployment of intelligent signal detectors is still limited. If you want to obtain a high-performance, highly generalized DL-based signal detection network, equipment manufacturers must pay extremely high training costs. In addition, the use of knowledge transfer technologies such as transfer learning and meta-learning requires a lot of AI expertise and sophisticated design of NN, and there are also challenges in adapting to devices with different computing power and capacity. Therefore, the present invention is designed to use learning gene technology to assist the design and deployment of DL-based signal detectors in MIMO systems.
[0035] Figure 1 The present invention provides a flowchart of a low-cost signal detection model deployment method based on learning genes according to an embodiment of the present invention.
[0036] like Figure 1 As shown, the low-cost signal detection model deployment method based on learning genes includes the following steps:
[0037] Step 1: Large equipment manufacturers have sufficient computing power and data set resources to design large-scale detection models and use massive datasets from different scenarios. Train sequentially to obtain a strong generalization group model f col (:,Θ col ).
[0038] In an embodiment of the present invention, the input of the large-scale population model is the result of LS channel estimation. and the results of the ZF signal detection scheme The output is the estimated signal detection result NN consists of convolutional layer, batch normalization layer, and fully connected layer, and the activation function is Tanh. The population model is trained using the gradient descent method to minimize the cost function by iterating the NN parameters. The cost function MSE is described as follows:
[0039]
[0040] Among them, (n) is the sample index, N is the number of samples, x is the real transmitted signal, ||·|| 2 is the Euclidean norm.
[0041] Step 2: Large equipment manufacturers analyze the gradient information of each NN layer of the group model According to the gradient significance Select and extract some structures and corresponding parameters to obtain a plug-and-play embeddable module - the learning gene Θ lg ,This module contains sufficient detection meta-knowledge for large manufacturers to conduct subsequent R&D or sell it to small manufacturers.
[0042] In an embodiment of the present invention, the gradient of the population model Described as:
[0043]
[0044] in, is the cost function of the group model, Θ col represents the parameters of the population model, k∈{1,…,K} represents the kth training data set, l∈{1,…,L col} represents the layer index of the population model, Represents the parameter index of the lth layer of the group model. The gradient significance of the lth layer on the kth task It is defined as the proportion of parameters with larger gradients in a layer, as described below:
[0045]
[0046] Where σ represents the threshold value. Φ=1, otherwise 0. As the training task number k increases, when the “ When the phenomenon of "gradually decreasing from large to stable" occurs, it means that the layer can be extracted as a learning gene. The set of all NN layers that meet the above requirements is the learning gene Θ lg .
[0047] Step 3: Equipment manufacturers can inherit the learning genes to the lightweight individual model f according to the equipment and scenario requirements of the target task. ind (:,Θ ind ), thereby realizing the transfer of meta-knowledge between heterogeneous NNs. Among them, the individual model consists of learning genes Θ lg and any expansion Θ exp It consists of two parts. The structure of the latter can be designed according to the input and output dimensions, device computing power, etc., and the corresponding parameters are randomly initialized.
[0048] In the embodiment of the present invention, the input and output types of the individual model are consistent with those of the group model. The data set is The neural network consists of convolutional layers, batch normalization layers, and fully connected layers, and the activation function is Tanh. The scale of the individual model is lighter than that of the group model, and any extension part can be freely designed according to the antenna dimension, computing power and capacity of the device carrying the detection model, etc.
[0049] Step 4: Equipment manufacturers use the target signal detection task dataset Train individual models and achieve rapid convergence of individual models with the assistance of learning genes rich in meta-knowledge.
[0050] In an embodiment of the present invention, the individual model is trained using the gradient descent method, and the cost function is described as follows:
[0051]
[0052] Among them, M represents the number of layers contained in the learning gene, Θ m represents the NN parameters at the inherited position of the learning gene in the individual model, trained by the target dataset, ||·|| 2 is the Euclidean norm, and λ represents the weight of the L2 regularization term. The cost function consists of two parts: the first part is the MSE between the output of the signal detection NN and the true emission signal, and the second part is to limit the update of the learning gene unit parameters in order to retain meta-knowledge while adapting to new tasks.
[0053] The invention specifically comprises the following steps:
[0054] (1) In the uplink MIMO system, the base station is configured with N t = 32 transmitting antennas, user configuration N r= 8 receiving antennas, carrier frequency is 6GHz, the base station is located at the origin, the height is 15 meters, the user is randomly initialized in a circular area of 50-100 meters around the base station, the height is 1.5 meters, and the channel model is the 3GPP urban microcell scenario without line of sight. According to the above settings, the data set is generated. The single scene data set contains 10,000 CSI samples, 8,100 samples are used for training, 900 for verification, and 1,000 for testing. A total of 55 scene data sets are generated, of which 50 are used to train the group model and 5 are used to train the individual model.
[0055] (2) Figure 2 The DL-based MIMO signal detection framework is shown. The user end sends data and pilot signals. During the coherence time, the base station uses the LS algorithm to estimate the channel based on the received pilot signal. And use the estimated channel and received data signal to perform ZF detection to obtain a rough reconstruction signal Then, and splicing, input into the signal detection NN, Perform refinement to reconstruct a more accurate transmission signal.
[0056] (3) Figure 3 The figure shows the extraction and expansion of learning genes. First, a large-scale population model is designed and trained sequentially on 50 training data sets, with each data set trained for 50 rounds. Then, the gradient significance of each convolutional layer in the population model is analyzed. If as the task k increases, The value of changes from large to small and gradually stabilizes, which means that the gradient gradually tends from large value to zero during the entire training process. This shows that this layer is less sensitive to scene changes and is more likely to contain meta-knowledge. Therefore, this layer can be extracted as a learning gene. The structure and parameter set of this type of neural network layer is the learning gene Θ lg .
[0057] (4) According to the actual situation of the target task, the learning gene is expanded into a lightweight individual model to achieve meta-knowledge transfer. The learning gene can be regarded as a plug-and-play embeddable module. Figure 4 The main extension modes of learning genes are shown: embedding and inheritance. The extension modes can be further divided into upper, middle and lower embedding / inheritance according to where the learning gene is embedded in the individual model or inherited from the group model. Any extension part Θ of the individual model exp The design can be customized according to the actual conditions of the target task, such as the antenna configuration and equipment computing power, and the parameters of this part are randomly initialized.
[0058] (5) After expansion, the individual model is trained using the dataset of the target task. The cost function consists of two parts: the first part is the MSE between the output of the individual model and the true emission signal, and the second part is the L2 regularization term, which is used to limit the update of the parameters of the learning gene unit so that the individual model can retain meta-knowledge while adapting to the new task.
[0059] (6) Figure 5 A low-cost signal detection model deployment method based on learning genes is demonstrated. Equipment manufacturers can be divided into two categories: one is large equipment manufacturers with sufficient computing resources and rich CSI datasets, and the other is small equipment manufacturers with relatively limited computing resources and CSI datasets. Large manufacturers have abundant resources and are able to design large-scale signal detection models and train them sequentially using massive datasets; then they extract learning genes rich in detection meta-knowledge and reuse them when deploying new detectors in different scenarios and devices, so as to quickly adapt to the target tasks. Large equipment manufacturers only need to train the population model once to obtain learning genes, which can greatly reduce the cost of subsequent training of signal detection models. From a long-term perspective, it can reduce the total training overhead.
[0060] In contrast, small equipment manufacturers usually lack the resources to train group models and independently develop learning genes. They can purchase packaged learning gene units from large equipment manufacturers to train their own intelligent signal detectors to cope with various deployment situations. The use of learning genes can also greatly reduce the subsequent design and training costs of small equipment manufacturers. Compared with large manufacturers directly selling CSI data sets, this method can better protect data privacy; in addition, by only selling learning gene units instead of the entire signal detection model, large equipment manufacturers can also effectively protect their intellectual property rights.
[0061] According to the low-cost signal detection model deployment method based on learning genes proposed in the embodiment of the present invention, large equipment manufacturers build large-scale signal detection models, which are sequentially trained by massive data sets of different scenarios to obtain a highly generalized group model; according to the gradient significance of each layer in the group model, partial structures and corresponding parameters are selected and extracted to extract reusable learning gene modules; learning genes can be regarded as a plug-and-play embeddable module that can be flexibly integrated into various lightweight signal detection models to adapt to different application scenarios and devices. The present invention realizes meta-knowledge transfer across network structures, supports rapid training of high-performance, highly generalized MIMO signal detection models, and can effectively reduce the training costs of equipment manufacturers and protect their intellectual property rights.
[0062] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or N embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.
[0063] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In the description of the present invention, the meaning of "N" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined.
[0064] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, fragment or portion of code comprising one or N executable instructions for implementing the steps of a custom logical function or process, and the scope of the preferred embodiments of the present invention includes alternative implementations in which functions may not be performed in the order shown or discussed, including performing functions in a substantially simultaneous manner or in reverse order depending on the functions involved, which should be understood by technicians in the technical field to which the embodiments of the present invention belong.
Claims
1. A low-cost signal detection model deployment method based on learning genes, characterized in that: The following steps are involved: Step 1: Design a large-scale detection model based on a large number of different scene data sets Train sequentially to obtain a strong generalization group model f col (:,Θ col ), K represents the number of training data sets, Θ col represents the parameters of the population model; Step 2: Analyze the gradient information of each neural network layer of the population model According to the gradient significance Select and extract some structures and corresponding parameters to obtain a plug-and-play embeddable module - the learning gene Θ lg , k∈{1,…,K} represents the kth training data set, l∈{1,…,L col } represents the layer index of the group model, L col represents the level of the layer index, Represents the parameter index of the lth layer of the population model; Step 3: According to the equipment and scenario requirements of the target task, inherit the learning gene to the lightweight individual model f ind (:,Θ ind ), the individual model consists of the learning gene Θ lg and customizable arbitrary extension Θ exp It consists of two parts; Step 4: Use the dataset for target signal detection task Train individual models and achieve rapid convergence of individual models with the assistance of learning genes rich in meta-knowledge.
2. The method according to claim 1, characterized in that In step 1, the input of the swarm model is the result of the least squares channel estimation And the zero-forcing signal detection result The output is the estimated signal detection result The neural network consists of convolutional layers, batch normalization layers, and fully connected layers.
3. The method according to claim 1, characterized in that In step 1, the population model is trained using the gradient descent method to minimize the cost function by iterating the neural network parameters. The cost function is described as follows: Where (n) is the sample index, N is the number of samples, x is the real transmitted signal, and ||·||2 is the Euclidean norm.
4. The method according to claim 1, characterized in that: In step 2, the gradient of the population model Described as: in, is the cost function of the group model.
5. The method according to claim 1, characterized in that In step 2, the gradient significance of the lth layer on the kth task is The description is as follows: Where σ represents the threshold value. Φ=1, otherwise 0; with the increase of training task number k, when the lth layer appears in the group model" When the phenomenon of "gradually decreasing from large to stable" occurs, the corresponding layer is selected as the learning gene; the set of all the population model neural network layers that meet the above phenomenon is the learning gene Θ lg .
6. The method according to claim 1, characterized in that In step 3, the input and output of the individual model are consistent with the group model, and any extended part is customized according to the antenna dimension, computing power and capacity of the device carrying the detection model.
7. The method according to claim 1, characterized in that In step 4, the individual model is trained using the gradient descent method, and the cost function is described as follows: Among them, M represents the number of layers contained in the learning gene, Θ m represents the neural network parameters trained by the target data set at the inherited position of the learning gene in the individual model, ||·||2 is the Euclidean norm, and λ represents the weight of the L2 regularization term. The cost function consists of two parts: the first part is the mean square error between the emission signal reconstructed by the neural network and the true emission signal, and the second part is to limit the update of the learning gene unit parameters so as to retain meta-knowledge while adapting to new tasks.
8. The method according to claim 1, characterized in that The learning gene module Θ lg It supports meta-knowledge migration across heterogeneous network structures and can adapt to the needs of equipment manufacturers of different sizes through sales or integration.
9. The method according to claim 1, characterized in that: The training data set of the group model includes CSI samples of multiple wireless communication scenarios, and the data set of the individual model Specific mission data corresponding to the target deployment scenario.
10. The method according to claim 1, characterized in that The neural network scale of the lightweight individual model is smaller than that of the group model, and its extended part Θ exp The parameters are initialized randomly.