Channel modeling method based on meta-learning
Through the channel modeling method based on meta-learning, the generative adversarial network is used to train on the support set and query set, which solves the problem of insufficient generalization capabilities of traditional channel modeling methods in complex dynamic environments, and realizes efficient and flexible channel modeling to meet the needs of 6G systems.
Patent Information
- Application Number
- CN202510398969.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-11
AI Technical Summary
Traditional channel modeling methods lack generalization capabilities in complex dynamic environments, and the data demand is large, so they cannot fully consider the nonlinear characteristics of wireless channels, resulting in inflexible and effective enough in the face of 6G network requirements.
Using a channel modeling method based on meta-learning, a meta-learner is established through a meta-learning framework, a meta-learning machine is used to conduct internal loop training on the support set, and a generalization performance is evaluated in combination with the query set, and fine-tuning training is performed on the target channel data set to generate pseudo-channel data with the same distribution as the target channel data.
It improves the adaptability and generalization capabilities of channel modeling, reduces data acquisition costs, and can quickly adapt to new tasks in the face of scarcity of data, meeting the high requirements of 6G systems for channel modeling.
Smart Images

Figure CN120301541A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of channel modeling, and more specifically, to a channel modeling method based on meta-learning. Background Art
[0002] With the rapid development of mobile communication technology, the construction of the sixth-generation mobile communication (6G) system has become the focus of the global scientific research and industrial communities. The 6G system is expected to provide higher transmission rates, lower latency than the current 5G, and support a wider range of application scenarios, covering multiple fields such as immersive extended reality (XR), autonomous driving, telemedicine, intelligent manufacturing, and smart cities. With the support of these applications, 6G not only needs to meet the requirements of ultra-large bandwidth and low latency, but also needs to handle a large number of terminal connections, support multiple communication modes, including scenarios with extremely high mobility and complex wireless channel environments. Therefore, the successful implementation of the 6G system depends on more accurate and efficient channel modeling technology.
[0003] Channel modeling is a fundamental task in wireless communication systems, and its purpose is to simulate various effects that wireless signals are subjected to during transmission, such as multipath effects, shadow fading, frequency offset, and interference. For 6G, channel modeling not only needs to reflect the performance of signals in different propagation environments, but also needs to be able to dynamically adapt to complex scenarios such as high mobile speeds and ultra-high frequency bands. Therefore, accurate channel modeling is crucial for improving the performance of communication systems, can provide key support for resource optimization allocation, network planning, scheduling algorithm design, etc., and also lays a solid foundation for achieving the performance indicators of 6G (such as ultra-high data rate, ultra-low latency, and large-scale connection).
[0004] Traditional channel modeling methods usually rely on physical models and empirical data, and predict the transmission behavior of signals by considering the attenuation, reflection, scattering and other characteristics of signals in different propagation media. However, these methods face several obvious limitations: (1) Physical models often rely on assumptions and are difficult to handle complex dynamic changes in the real world; (2) The traditional empirical data collection process often requires a large number of field tests, which not only increases the cost but also limits the availability and timeliness of the data: (3) These methods usually cannot fully consider the non-linear characteristics of wireless channels, resulting in them being insufficiently flexible and effective in the face of the rapidly developing 6G network requirements.
[0005] Therefore, how to implement an efficient, flexible, and adaptable channel modeling method has become a key issue in current research.
[0006] In recent years, machine learning (ML)-based channel modeling techniques have received extensive attention. Machine learning methods, especially deep learning techniques, can capture the non-linear relationships in the signal propagation process by automatically extracting features from a large amount of data. This gives machine learning unparalleled advantages in dealing with complex and dynamic wireless environments compared to traditional methods. Machine learning models can learn channel characteristics from historical channel data and can be generalized under different propagation conditions, thus enhancing the adaptability of the model in diverse environments. However, these methods also have certain limitations, especially in cases where data is scarce or of low quality. Machine learning models usually rely on a large amount of reliable training data, but in practical applications, channel data may not fully cover all possible environments, resulting in insufficient generalization ability of the model and thus affecting the performance in actual applications. Summary of the Invention
[0007] To overcome the defects of strong data dependence and weak generalization ability in the above-mentioned prior art, the present invention provides a channel modeling method based on meta-learning.
[0008] To solve the above technical problems, the technical solution of the present invention is as follows: In a first aspect, a channel modeling method based on meta-learning includes: Obtain channel data and establish a channel data set; Based on the meta-learning framework, establish a meta-learner, and divide the channel data set into a support set and a corresponding query set for different specific channel tasks as training samples for meta-learning; wherein, the meta-learner is used to provide initial network parameters for a generative adversarial network with WGAN-GP as the backbone network and capture the commonalities and differences of the generative adversarial network trained on different specific channel tasks, and the generative adversarial network includes a generator and a discriminator; Let the generative adversarial network perform in-loop training on the support set to adapt to the specific data pattern of the specific channel task, and use the query set to evaluate the generalization performance of the generative adversarial network after task adaptation to perform cross-task update on the meta-learner until meta-learning ends; wherein, the task of the generator is to generate synthetic data with the same distribution as the training samples, and the task of the discriminator is to evaluate the quality of the synthetic data output by the generator; Initialize the generative adversarial network with the meta-learner that has completed meta-learning, and perform fine-tuning training on the generative adversarial network on the target channel data set; After completing the fine-tuning training, use the generator to synthesize pseudo-channel data with the same distribution as the target channel data to achieve channel modeling.
[0009] In a second aspect, an electronic device includes: A memory for storing computer-executable instructions or computer programs; A processor, when executing the computer-executable instructions or computer programs stored in the memory, implements the method described in the first aspect.
[0010] In a third aspect, a computer program product includes a computer program or computer-executable instructions, and when the computer program or computer-executable instructions are executed by a processor, the method described in the first aspect is implemented.
[0011] Compared with the prior art, the beneficial effects of the technical solution of the present invention are: This application discloses a channel modeling method based on meta-learning, which uses meta-learning to improve the adaptability of a generative adversarial network in a data-scarce environment and introduces a self-optimizing "meta-model" to quickly adapt to new tasks, thereby effectively improving the learning efficiency, accuracy, and stability of the model and meeting the high requirements for channel modeling in 6G systems. Description of the Drawings
[0012] Figure 1 It is a schematic flowchart of the channel modeling method based on meta-learning in Embodiment 1 of this application.
[0013] Figure 2 It is a block diagram of an OFDM system in Embodiment 1 of this application.
[0014] Figure 3 It is a schematic structural diagram of a generative adversarial network in Embodiment 1 of this application.
[0015] Figure 4 It is a comparison diagram of PDP and FCF of the real channel and the generated channel in the C1-LOS scenario in Embodiment 2 of this application (1400 training data).
[0016] Figure 5 It is a comparison diagram of PDP and FCF of the real channel and the generated channel in the EPA scenario in Embodiment 2 of this application (1400 training data).
[0017] Figure 6 It is a comparison of PDP and FCF of the real channel and the generated channel in the C1-LOS scenario in Embodiment 2 of this application (50 training data).
[0018] Figure 7 It is a comparison of PDP and FCF of the real channel and the generated channel in the EPA scenario in Embodiment 2 of this application (50 training data).
[0019] Figure 8 It is a comparison of PDP and FCF of the real channel and the generated channel in the B2-NLOS scenario in Embodiment 2 of this application (1400 training data).
[0020] Figure 9 For the comparison of the PDP and FCF of the real channel and the generated channel in the C3-NLOS scenario in Embodiment 2 of this application (1400 training data).
[0021] Figure 10 For the curve graph showing the variation of the effects of normal training, transfer learning, and meta-learning with the number of training samples in Embodiment 2 of this application.
[0022] Figure 11 For the curve graph showing the variation of the effects of normal training of ten channels with the number of training samples in Embodiment 2 of this application.
[0023] Figure 12 For the curve graph showing the variation of the effects of meta-learning of ten channels with the number of training samples in Embodiment 2 of this application.
[0024] Figure 13 For the curve graph showing the variation of the effects of transfer learning of ten channels with the number of training samples in Embodiment 2 of this application. Detailed implementation manners
[0025] In the description and claims of this application and the above-mentioned drawings, the terms "first", "second", etc. are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances, which is only a way of distinguishing objects with the same attributes when describing the embodiments of this application. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, so that a process, method, system, product, or device including a series of units does not necessarily have to be limited to those units, but may include other units not clearly listed or inherent to these processes, methods, products, or devices. The term "determine" broadly covers a variety of actions, which may include obtaining, calculating, computing, processing, deriving, researching, searching (for example, searching in a table, database, or other data structure), finding out, and similar actions, and may also include receiving (for example, receiving information), accessing (for example, accessing data in a memory) and similar actions, and may also include generating, creating, establishing and similar actions, as well as parsing, selecting, choosing and similar actions, etc. The relevant definitions of other terms will be given in the following description.
[0026] It should be noted that when an element is considered to be "connected" to another element, it can be directly connected to the other element or connected to the other element through an intermediate element. In addition, in the following embodiments, "connection", if there is a transmission of electrical signals or data between the connected objects, should be understood as "electrical connection", "communication connection", etc.
[0027] The drawings are only for illustrative purposes and should not be construed as a limitation of this patent; To better illustrate this embodiment, some components in the drawings are omitted, enlarged, or reduced, which do not represent the dimensions of the actual product; For those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.
[0028] The technical solution of the present invention will be further described below with reference to the drawings and embodiments.
[0029] Embodiment 1 In a 6G communication system, channel modeling is a crucial task. Traditional channel modeling methods often rely on a large amount of sample data for training. However, in practical applications, obtaining a large amount of sample data may be both difficult and expensive.
[0030] Facing the demand for efficient and accurate channel modeling in 6G mobile communication networks, this embodiment provides a channel modeling method based on meta-learning to address the problems of insufficient generalization ability and large data requirements of traditional channel modeling methods in complex dynamic environments. Referring to FIG. 1, it includes: Obtain channel data and establish a channel data set; Based on the meta-learning framework, establish a meta-learner, and divide the channel data set into a support set and a corresponding query set for different specific channel tasks as training samples for meta-learning. Among them, the meta-learner is used to provide initial network parameters for a generative adversarial network with WGAN-GP as the backbone network and capture the commonalities and differences of the generative adversarial network trained on different specific channel tasks. The generative adversarial network includes a generator and a discriminator; Let the generative adversarial network perform in-loop training on the support set to adapt to the specific data pattern of the specific channel task, and use the query set to evaluate the generalization performance of the generative adversarial network after task adaptation to perform cross-task update on the meta-learner until the meta-learning ends. Among them, the task of the generator is to generate synthetic data with the same distribution as the training samples, and the task of the discriminator is to evaluate the quality of the synthetic data output by the generator; Initialize the generative adversarial network with the meta-learner that has completed meta-learning, and perform fine-tuning training on the generative adversarial network on the target channel data set; After completing the fine-tuning training, use the generator to synthesize pseudo-channel data with the same distribution as the target channel data to achieve channel modeling.
[0031] It should be noted that in this embodiment, meta-learning enables the model to quickly adapt to new tasks in the case of scarce data, so as to achieve the purpose of improving the adaptability of channel modeling, enhancing the generalization ability and reducing the data acquisition cost. Through fast learning and efficient adaptability, it can still maintain a high modeling accuracy when facing complex channel environments and scarce sample data. Therefore, meta-learning has very important application prospects in 6G systems, especially in various different types of wireless channel modeling tasks.
[0032] Those skilled in the art should understand that the meta-learning method can capture the commonalities and differences between tasks by constructing a "meta-learner", thus helping the model to quickly adjust and optimize when facing new tasks. In the channel modeling of 6G systems, especially in dynamic and changing channel environments, meta-learning can not only effectively improve the modeling accuracy, but also enhance the generalization ability of the model in new environments. By reducing the dependence on large-scale data sets, meta-learning also reduces the costs of data acquisition and annotation, thus providing a more efficient solution for the rapid deployment and optimization of next-generation communication systems.
[0033] The core idea of meta-learning is to construct a "meta-model" so that the model can share knowledge between different tasks and, when encountering a new task, quickly optimize and adjust its own learning strategy through a small number of training samples or experiences. The advantage of this method is that even in the case of scarce data, meta-learning can effectively learn and provide high-quality modeling results in different types of channel tasks.
[0034] It should also be noted that in the meta-learning method, the appropriate division of the channel data set is crucial for the training and evaluation of the model. The channel training data set is divided into a support set and a query set , where ; represents the number of specific channel tasks, and each channel task represents a different channel type. Each channel task has a set of independent support sets and query sets , and the channel data of the support set and the query set have the same distribution under the condition that the data volume is large enough.
[0035] Those skilled in the art should understand that this embodiment introduces digital twin technology. Digital twin realizes real-time simulation and prediction of actual channel conditions by creating a virtual copy of the wireless communication environment. Using digital twin, the channel model can be dynamically updated and optimized according to real-time environmental data, thereby improving the accuracy and adaptability of the model. In the 6G network, digital twin can not only effectively simulate different propagation environments, but also help optimize network configuration and resource scheduling, respond to mobility changes and environmental interference in real time, and further improve the stability and reliability of the communication system. Specifically, this embodiment uses a generative adversarial network as the virtual model in the digital twin, and a large amount of channel data with the same distribution as the physical entity can be generated based on a small amount of channel data.
[0036] In some preferred embodiments, the acquisition of the channel data is realized by building a semi-physical simulation platform, which includes two baseband processing units, two universal software radio peripherals (USRPs) and a channel simulator.
[0037] Among them, the first baseband processing unit is responsible for generating the baseband signal and transmitting the signal to USRP1. USRP1 modulates the baseband signal to the carrier frequency and performs channel simulation through the channel simulator. The signal simulated by the channel simulator is received by USRP2 and transmitted to the second baseband processing unit for subsequent signal processing. All devices complete the signal transmission through wired connections. Demonstratively, the baseband processing of the signal is undertaken by two upper computers to realize the functions of the baseband processing unit.
[0038] Referring to Figure 2, the time-varying channel impulse response (CIR) between the transmitter and the receiver in the wireless channel can be modeled as:
[0039] In the formula, represents the complex channel amplitude of the th path at the th moment, represents its corresponding time delay, and g(τ) represents the impulse response of the transmitting and receiving filters.
[0040] Performing Fourier transform on Equation (1), the frequency response of the time-varying channel can be obtained as:
[0041] In the formula, is the Fourier transform of g(τ).
[0042] Since the entire system can be regarded as an orthogonal frequency division multiplexing (OFDM) system, the channel frequency response (CFR) at the nth OFDM symbol of the kth subcarrier can be written as:
[0043] Wherein, , N is the number of OFDM symbols, , K is the number of subcarriers. Without loss of generality, considering one OFDM symbol, Equation (3) can be written as:
[0044] In the OFDM system, the representation of the received signal in the frequency domain can be written as:
[0045] Wherein, , represents the received signal of the receiver, , ) represents the pilot matrix, ) represents the diagonal matrix of the determined elements, , represents the channel vector, represents a Gaussian white noise with a mean of 0 and a variance of .
[0046] Estimate the channel vector using the linear minimum mean square error from Equation (5), and we can obtain:
[0047] Wherein, represents the channel correlation matrix, ) represents the mathematical expectation, represents the conjugate transpose, is the least squares estimate of, which can be written as:
[0048] In some specific implementation processes, the specific process is as follows: S101. Baseband signal generation: The first baseband processing unit (Computer 1) is responsible for generating the baseband signal, setting the number of subcarriers to 512, the transmission time interval to 100, and the number of orthogonal frequency division multiplexing symbols to 14, and sending these signals to USRP1 through a wired connection.
[0049] S102. Signal modulation and transmission: After receiving the baseband signal, USRP1 modulates it to the 2.5 GHz carrier frequency and sends the modulated signal to the channel simulator through a wired connection.
[0050] S103. Channel simulation: The channel simulator receives the modulated signal and performs channel simulation on it to simulate various effects in the actual wireless channel, such as multipath fading, noise, etc. The channel scenarios of the channel simulator are respectively set to the A1-LOS (line of sight in indoor office environment), B1-NLOS (urban small cell), B2-NLOS (severe urban small cell), C1-LOS (suburban macro cell), C2-NLOS (urban macro cell), C3-NLOS (severe urban macro cell) channel scenarios in WINNER II, the EPA (pedestrian mobility scenario), EVA (vehicle mobility scenario), HT (hilly terrain scenario) channel scenarios in 3GPP (3rd Generation Partnership Project) LTE (Long Term Evolution), and the SUI-2 and SUI-6 channel scenarios of Stanford University.
[0051] S104. Signal reception and demodulation: The signal processed by the channel simulator is received by USRP2. USRP2 demodulates the signal back to the baseband and transmits the demodulated signal to the second baseband processing unit (computer 2) through a wired connection.
[0052] S105. Signal processing: The second baseband processing unit receives the demodulated signal and performs subsequent signal processing and analysis.
[0053] Through the above steps, all devices complete the transmission and processing of signals through wired connections, ensuring the accuracy and controllability of the signals, and then establishing a channel data set.
[0054] It should be noted that in the above preferred embodiments, the simulation platform is used as the physical entity in the digital twin, and building the simulation platform can reduce the cost of collecting data.
[0055] In some preferred embodiments, the generator sequentially includes a fully connected layer for expanding the data dimension, a batch normalization layer, a recombination operation layer for adjusting the data dimension, multiple upsampling modules for extracting high-level semantic features of channel data, and a cropping layer for outputting channel impulse response data of a fixed size; wherein, the upsampling module includes an upsampling layer, a convolutional layer, a batch normalization layer, and an activation function layer; The discriminator sequentially includes a padding layer for expanding the data dimension, multiple downsampling modules for extracting core features of channel data, a flattening operation layer for adjusting the data dimension, and a fully connected layer for outputting the authenticity of the synthetic data, wherein the downsampling module includes a convolutional layer and an activation function layer.
[0056] In some specific implementation processes, referring to FIG. 3, the structure of the generator includes a fully connected layer, a restructuring operation layer, a batch normalization layer, and six upsampling modules. Finally, fixed-size channel impulse response data is generated through a cropping operation. The fully connected layer is used to expand the data dimension for subsequent convolution operations; the restructuring operation adjusts the data dimension to match the requirements of the subsequent convolution layer; the batch normalization layer helps to accelerate training and improve stability. Each upsampling module consists of an upsampling layer, a convolution layer, a batch normalization layer, and an activation function layer. The first five activation functions use LeakyReLU, and the last activation function uses tanh. The upsampling module is mainly used to extract high-level semantic features of the channel data.
[0057] In some specific implementation processes, the structure of the discriminator includes a zero-padding operation layer (as the padding layer), seven downsampling modules, a flattening operation layer, and a fully connected layer. The zero-padding operation layer is to expand the data dimension to adapt to the subsequent downsampling modules. The downsampling module consists of a convolution layer and an activation function layer. To reduce the computational amount, 25% of the neuron parameters remain unchanged during each update. The role of the downsampling module is to extract the core features of the channel data for determining whether the input data is real collected data. The flattening operation converts the data into a dimension suitable for the fully connected layer, and finally a value of 0 or 1 is output through the fully connected layer to determine the authenticity of the synthetic data output by the generator.
[0058] It should be understood that when training the generative adversarial network, input training data for training. After the model converges, a large amount of synthetic data similar to the distribution of the channel data can be generated through the generator.
[0059] In some preferred embodiments, making the generative adversarial network perform inner-loop training on the support set to adapt to the specific data pattern of the specific channel task, and using the query set to evaluate the generalization performance of the generative adversarial network after task adaptation to perform cross-task update on the meta-learner includes: For each generative adversarial network, in each inner-loop training cycle of each specific channel task: Using the meta-learner, with the initial network parameters and Initialize the generator and the discriminator. These parameters are the starting points of model training. Train and update the network parameters of the generator and the discriminator on the support set until the corresponding loss function converges or reaches the preset number of training cycles; After completing the inner-loop training, in each outer-loop training cycle: Calculate the loss function of each generative adversarial network on the corresponding query set, and then determine the meta-loss function; The initial network parameters of the meta-learner are updated across tasks with the goal of minimizing the meta-loss function.
[0060] It should be emphasized that the role of the support set is to enable the model to learn and update parameters on each channel task, so that it can exhibit strong adaptability on specific channel tasks. The query set is used to evaluate the generalization performance of the model after task adaptation. It measures the performance of the adapted model on different channel tasks by calculating the meta-loss, thereby verifying the generalization ability of the model and its adaptability to new channel tasks. The calculation of the meta-loss function takes into account the differences between the support set and the query set, and guides the model to adjust during training, so that the model can not only adapt to individual channel tasks, but also have good generalization ability among multiple channel tasks.
[0061] Those skilled in the art should understand that the generator is usually used to generate channel task-specific samples or strategies, while the discriminator is used to evaluate the channel quality of the generator's output. In each inner-loop training cycle of a specific channel task, the support set is used to train the network. Specifically, the generator generates channel data through the support set, and the discriminator provides feedback by comparing the generated channel data (i.e., synthetic data) with the real channel data.
[0062] In some alternative embodiments, the loss function of the generative adversarial network is:
[0063] where and represent the network parameters of the generator and the discriminator respectively; , represents a random variable that follows a uniform distribution on , represents the channel data set, represents the synthetic data (i.e., the generated channel data) obtained after the random noise passes through the generator; and are the probability distributions of Gaussian white noise and respectively; is 's probability distribution, and λ is an adjustable parameter greater than zero; represents the mathematical expectation of the discriminant value obtained after the real channel data (used as the channel data for training) in the corresponding data set passes through the discriminator; represents the mathematical expectation of the discriminant value obtained after the synthetic data passes through the discriminator; represents the loss function penalty term added by WGAN to ensure the stability of GAN training.
[0064] In some alternative embodiments, during the inner-loop training cycle, the network parameters of the generator and the discriminator are updated based on the gradient descent method, where The parameter update process of the generator is expressed as:
[0065] In the formula, represents the network parameters of the generator determined after training using the -th support set; represents the inner-loop learning rate; represents the loss function of the generator calculated on the support set ; represents the gradient when the generator loss function is ; And, the parameter update process of the discriminator is expressed as:
[0066] In the formula, represents the network parameters of the discriminator determined after training using the -th support set ; represents the loss function of the discriminator calculated on the support set ; represents the gradient when the discriminator loss function is ;
[0067] It should be understood that the loss function of the discriminator generally measures the accuracy of the discriminator during the task adaptation process, while the loss function of the generator measures the difference between the data generated by the generator and the real data.
[0068] Furthermore, minimizing the meta-loss function is denoted as the meta-objective function, which includes the sum of the loss functions of multiple channel tasks and incorporates the error on the query set of each channel task, expressed as:
[0069] In the formula, represents the meta-loss function, represents the loss function of the generative adversarial network.
[0070] represents that when the corresponding network parameters are and on the query set The loss calculated above measures the performance of the model on the query set and is usually a task-specific metric that reflects the generalization ability of the model in the channel task. Exemplarily, it can be calculated based on Equation (8).
[0071] It should be emphasized that in the outer loop training cycle, minimizing the meta-objective function is the key goal, which measures the overall performance of the model on all channel tasks.
[0072] In some alternative embodiments, in the outer loop training cycle, the update process of the network parameters of the generator and the discriminator is expressed as:
[0073]
[0074] where represents the meta-learning rate; represents the sum of all loss functions when training the network with M query sets when the discriminator parameters are ; represents the gradient of the update of the discriminator network parameters when the loss function is and the discriminator parameters are ; represents the sum of all loss functions when training the network with M query sets when the generator parameters are ; represents the gradient of the update of the generator network parameters when the loss function is and the generator parameters are ; and the generator parameters are .
[0075] It should be noted that in the outer loop training cycle, in order to make the adaptation ability of the model more generalized between different tasks, the meta-learning rate β is used to control the pace of cross-task updates. The role of the meta-learning rate β is to adjust the parameter update speed in the cross-task update process to ensure that the model can not only quickly adapt to a single channel task but also effectively transfer and generalize between multiple channel tasks. Specifically, the query set is used to verify the generalization ability of the updated model on different channel tasks. By calculating the meta-loss function based on the query set, the model can be further optimized to better handle new channel task scenarios.
[0076] In some preferred embodiments, in the fine-tuning training cycle, the fine-tuning update process adopts an optimization strategy similar to meta-learning, and the parameter update process of the generative adversarial network is expressed as follows:
[0077]
[0078] In the formula, represents the fine-tuning learning rate; represents that the corresponding network parameter is and the loss function on the target channel dataset when represents at the gradient of the model on and represent the network parameters of the fine-tuned discriminator and generator.
[0079] It should be noted that once the meta-learning stage ends, the model will use the parameters learned in the meta-learning stage and and use the target domain data (i.e., the target channel dataset) for fine-tuning for a specific target channel task. At this time, the parameters obtained by meta-learning already have strong cross-task adaptation capabilities and can be used as a relatively general initialization value. Next, the model will be fine-tuned on a dedicated dataset for the target channel task so that the parameters can be further optimized to better adapt to the characteristics and data distribution of the target channel task. In the fine-tuning stage, first, the parameters of the generative adversarial network are initialized to the values obtained in the meta-learning stage, and this initialized parameter set can quickly adapt to the training data of the target channel task and be further adjusted on this basis. Using the target channel dataset as the input of the model in the fine-tuning stage, the model will perform refinement optimization by training on these channel data so that the model can achieve the best performance on the target channel task.
[0080] When the fine-tuning training ends and the generator is used for testing or application, the network parameter is used to generate a pseudo-channel dataset . It should be understood that the final output of the above embodiments includes the network parameters , and the generated data , where is pseudo-channel data (or a pseudo-channel dataset composed of pseudo-channel data) with the same distribution as the training data.
[0081] As a non-limiting example, in the testing stage, is used to generate a pseudo-channel dataset . The goal of this stage is to verify whether the model can effectively generate samples similar to real channel data when facing unseen data, and then test its generalization ability and generation ability. Specifically, the generated dataset will be verified through a series of evaluation criteria.
[0082] This embodiment also provides the pseudocode of the channel modeling method as follows: Table 1 Channel Modeling Algorithm Based on Generative Adversarial Network and Meta-Learning
[0083] Embodiment 2 This embodiment conducts a comparative experiment on the meta-learning-based channel modeling method and the channel modeling method based on the transfer learning framework proposed in Embodiment 1.
[0084] Among them, transfer learning can effectively improve the performance of the model under different conditions by applying the knowledge learned from one channel environment to other channel environments. This method is particularly applicable when the data of a specific channel environment is scarce. By transferring the existing knowledge, it can quickly adapt to the new environment, thereby improving the learning efficiency and accuracy of the model.
[0085] For the channel data of the source task , first, randomly initialize the network parameters of the generator and discriminator of the generative adversarial network , , and use the channel data of the source channel task to perform GAN training. After a period of adversarial training, the parameters of the generator and discriminator will be continuously optimized until the network converges. At this time, the generator can better capture the statistical characteristics of the source task channel data and can generate high-quality channel samples. Once the network converges and obtains the optimized parameters of the generator and discriminator , , when performing the target channel task training, the parameters obtained from the source channel task training can be , used as the initial values. Based on these already trained parameters, the target channel task can be further trained. For the channel data of the target channel task, retraining is performed on the basis of the network parameters of the source channel task, so as to achieve the rapid convergence of the network and quickly obtain the optimized generative adversarial network parameters and .
[0086] Specifically, in the transfer learning framework, the A1-LOS channel scenario data is used as the source domain for model pre-training. Through source domain training, the model can learn the basic feature representation of channel estimation. Subsequently, the pre-trained model is transferred to the target domains of the C1-LOS channel scenario and the EPA channel scenario, and the C1-LOS scenario and EPA scenario data are used for fine-tuning, so that the model can adapt to the multipath effect and signal attenuation characteristics under the C1-LOS scenario and the EPA scenario.
[0087] It should be specifically noted that the above channel modeling method based on transfer learning and generative adversarial networks is not an existing technology in this field.
[0088] Under the meta-learning framework, a multi-task learning environment is constructed, and 11 different channel scenarios are jointly trained as 11 independent channel tasks (i.e., the specific channel tasks). Each channel task contains its specific channel feature distribution, such as path loss, multipath delay, Doppler shift, etc. When facing specific C1-LOS channel scenarios and EPA channel scenarios (i.e., the target channel tasks), a small amount of data of this scenario (i.e., the target channel dataset) is used for fine-tuning, so that the model can converge quickly and reach the ideal performance index.
[0089] To comprehensively evaluate the performance of transfer learning and meta-learning methods in C1-LOS scenarios and EPA scenarios, a comparative analysis is carried out from two dimensions: channel data distribution and system-level performance. First, at the level of channel characteristic analysis, in-depth comparison is made on the channel data of the measured C1-LOS scenario and EPA scenario, the channel data of the C1-LOS scenario and EPA scenario generated by the transfer learning method, and the channel data of the C1-LOS scenario and EPA scenario generated by the meta-learning method. Specifically, two key indicators are extracted: PDP (power delay profile), which is used to analyze the delay spread characteristics and power distribution of multipath signals; FCF (frequency correlation function), which is used to evaluate the frequency selective fading characteristics of the channel. By comparing the statistical distributions of these indicators, the matching degree of the generated data and the measured data in terms of channel characteristics can be intuitively shown. Figures 4 and 5 respectively show the comparison of the collected channel data and the generated channel data in terms of PDP and FCF obtained by training with 1400 training samples in the C1-LOS scenario and EPA scenario. The results shown in the figures indicate that when the amount of training data is sufficient, both transfer learning and meta-learning can significantly improve the generation effect of channel data and achieve relatively excellent performance.
[0090] Figures 6 and 7 respectively show the comparison of the PDP and FCF between the collected channel data and the generated channel data when training with 50 training samples in the C1-LOS scenario and the EPA scenario. It can be seen from the figures that as the amount of training data decreases, the performance of the model shows a certain downward trend. However, in the case of scarce data, meta-learning demonstrates more superior performance compared to transfer learning. Specifically, through its powerful fast adaptation ability, meta-learning can effectively learn the basic characteristics of the channel based on a small amount of data, and is superior to transfer learning in terms of the fitting degree of PDP and FCF. Therefore, meta-learning shows strong robustness in the case of scarce data, can better cope with the challenges brought by insufficient data, and further verifies the advantages of meta-learning under low-resource conditions.
[0091] To further verify the advantages of meta-learning in complex channel modeling, Figures 8 and 9 respectively show the comparison of the PDP and FCF between the collected channel data and the generated channel data when training with 1400 training samples in the B2-NLOS scenario and the C3-NLOS scenario. The results show that even in a complex non-line-of-sight channel environment, the meta-learning method can still effectively model the channel characteristics and generate PDP and FCF that match the actual channel data. Specifically, although the channel environments in the B2-NLOS and C3-NLOS scenarios are relatively complex, with strong multipath effects and large-scale fading, meta-learning can still better capture the changing rules of the channel through its superior fast adaptation ability. Compared with transfer learning, meta-learning shows more superior modeling effects in these two complex channel scenarios. Especially in terms of the fitting accuracy of PDP and FCF, the meta-learning model is significantly better than transfer learning.
[0092] This experimental result further verifies the robustness and efficiency of meta-learning in the face of complex channels. Whether it is a flat channel or a non-flat channel with significant fading and multipath effects, meta-learning can maintain high performance with limited training data support. This shows that meta-learning not only has advantages in simple scenarios, but also can provide more reliable channel data generation ability in complex channel modeling. Compared with transfer learning, meta-learning can achieve better generalization performance in different channel environments through its unique fast adaptation mechanism, thus providing an effective solution for the channel modeling requirements in complex wireless communication systems.
[0093] Figure 10 shows the average effect change curves of normal training, transfer learning, and meta - learning methods in 11 different channel environments. The horizontal axis represents the number of training samples, and the vertical axis is the minimum mean square error (MSE) of the probability density profile (PDP) and frequency - correlation function (FCF) between the real channel data and the generated channel data. It can be seen from the figure that the meta - learning method shows better performance than the transfer - learning method on all channels, and the MSE value is significantly lower than that of transfer learning. This indicates that meta - learning can more effectively extract adaptive knowledge from a small number of samples and maintain better performance in different channel environments. Moreover, when the number of samples is very small, the effect of normal training is worse than that of transfer learning and meta - learning because the training data is scarce and the model cannot learn the data distribution of the channel. When the training data increases, the model can accurately learn the channel data distribution.
[0094] In addition, as the number of samples in the training data set increases, the MSE values of the PDP and FCF both show a gradually decreasing trend whether using transfer learning or meta - learning. This indicates that more training samples help to improve the training accuracy and adaptability of the model. Especially in the process of model fine - tuning, increasing the sample size can better reduce the training error and improve the accuracy of channel modeling. Therefore, the number of samples has an obvious positive impact on the effects of transfer learning and meta - learning, further emphasizing the importance of large - scale data sets in the channel - modeling task.
[0095] To show the variation of model accuracy with the amount of training data for different channels under different methods, Figures 11, 12, and 13 respectively show the trends of the MSE of the PDP between the collected channel data and the generated channel data of different channel types changing with the number of training samples under the three methods of normal training, meta - learning, and transfer learning. As the amount of training data increases, the MSE drops rapidly, indicating that the expansion of the data volume can effectively improve the model's ability to learn the channel data distribution.
[0096] It can be understood that the optional items in the above - mentioned Embodiment 1 are also applicable to this embodiment, so they will not be described repeatedly here.
[0097] Embodiment 3 This embodiment provides a computer - readable storage medium, on which at least one instruction, at least one program, a code set, or an instruction set is stored. The at least one instruction, at least one program, the code set, or the instruction set is loaded and executed by a processor, so that the processor executes some or all of the steps of the method provided in Embodiment 1 of this application.
[0098] It can be understood that the storage medium can be transient or non-transient. Exemplarily, the storage medium includes but is not limited to various media that can store program codes, such as USB flash drives, external hard drives, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs.
[0099] Exemplarily, the processor can be a central processing unit (CPU), a microprocessor unit (MPU), a digital signal processor (DSP), or a field programmable gate array (FPGA), etc.
[0100] Exemplarily, the read-only memory includes but is not limited to MASK ROM, PROM, EPROM, EEPROM, Flash, etc.
[0101] Exemplarily, the random access memory includes but is not limited to DRAM, SRAM, SDRAM, DDR SDRAM, etc.
[0102] In some examples, a computer program product is provided, which can be specifically implemented in a way of hardware, software, or a combination thereof. As a non-limiting example, the computer program product can be embodied as the storage medium, and can also be embodied as a software product, such as an SDK (Software Development Kit), etc.
[0103] As a non-limiting example, a computer program product is provided, which includes a computer program or computer-executable instructions. The computer program or computer-executable instructions are stored in a computer-readable storage medium. The processor of the electronic device reads the computer program or computer-executable instructions from the computer-readable storage medium, and the processor executes the computer-executable instructions, so that the electronic device executes some or all of the steps of the method described in the embodiments of the present application.
[0104] In some examples, a computer program is provided, which includes computer-readable code. When the computer-readable code runs in a computer device, the processor in the computer device executes to implement some or all of the steps in the method.
[0105] This embodiment also provides an electronic device, including a memory and a processor. The memory stores at least one instruction, at least one program, a code set or an instruction set. When the processor executes the at least one instruction, at least one program, the code set or the instruction set, it implements part or all of the steps of the method described in Embodiment 1.
[0106] In some examples, a hardware entity of the electronic device is provided, including: a processor, a memory and a communication interface; wherein, the processor generally controls the overall operation of the electronic device; the communication interface is used to enable the electronic device to communicate with other terminals or servers through a network; the memory is configured to store instructions and applications executable by the processor, and can also cache data to be processed or already processed by the processor and each module in the electronic device (including but not limited to image data, audio data, voice communication data and video communication data), and can be implemented by flash memory (FLASH), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM) or random access memory (RAM).
[0107] The processor may include one or more processing elements. Therefore, the processor may include one or more integrated circuits (ICs) configured to execute the functions of the processor. In addition, each integrated circuit may include circuits (such as a first circuit, a second circuit, and other circuits, etc.) configured to execute the functions of the processor.
[0108] Further, data transmission may be performed between the processor, the communication interface and the memory through a bus. The bus may include any number of interconnected buses and bridges, and the bus connects various circuits of one or more processors and memories together.
[0109] It can be understood that the optional items in the above Embodiment 1 are equally applicable to this embodiment, so they will not be described repeatedly here.
[0110] The same or similar reference numerals correspond to the same or similar components; The terms describing the positional relationship in the drawings are only for illustrative purposes and should not be construed as a limitation to this application; It should be noted that, without conflict, the embodiments and features in the embodiments of this application may be combined with each other.
[0111] In different specific implementations, the method or system described in the present application can be implemented in software, hardware, or a combination thereof. In addition, the order of the steps of the method can be changed, and various elements can be added, reordered, combined, omitted, modified, etc.
[0112] Obviously, the above embodiments of the present application are merely examples for clearly explaining the present application, rather than limitations on the implementation manners of the present application, and are not used to limit the present application. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. Each discrete structural / functional module or unit can be integrated together to form an independent part, or each module can exist alone, or two or more modules can be integrated to form an independent part. The structures and functions of the discrete components can be implemented as a combined structure or component. It is not necessary and impossible to enumerate all the implementation manners here. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included within the protection scope of the claims of the present application.
Claims
1. A channel modeling method based on meta - learning, characterized in that, Including: Obtain channel data and establish a channel data set; Based on the meta-learning framework, establish a meta-learner, and divide the channel data set into a support set and a corresponding query set for different specific channel tasks as training samples for meta-learning; wherein, the meta-learner is used to provide initial network parameters for a generative adversarial network with WGAN-GP as the backbone network and capture the commonalities and differences of the generative adversarial network trained on different specific channel tasks, and the generative adversarial network includes a generator and a discriminator; Let the generative adversarial network perform in-loop training on the support set to adapt to the specific data pattern of the specific channel task, and use the query set to evaluate the generalization performance of the generative adversarial network after task adaptation to perform cross-task update on the meta-learner until the meta-learning ends; wherein, the task of the generator is to generate synthetic data with the same distribution as the training samples, and the task of the discriminator is to evaluate the quality of the synthetic data output by the generator; Initialize the generative adversarial network with the meta-learner that has completed meta-learning, and perform fine-tuning training on the generative adversarial network on the target channel data set; After completing the fine-tuning training, use the generator to synthesize pseudo-channel data with the same distribution as the target channel data to achieve channel modeling.
2. The channel modeling method based on meta-learning according to claim 1, wherein, The step of letting the generative adversarial network perform in-loop training on the support set to adapt to the specific data pattern of the specific channel task, and using the query set to evaluate the generalization performance of the generative adversarial network after task adaptation to perform cross-task update on the meta-learner includes: For each generative adversarial network, in each in-loop training cycle of each specific channel task: Using the meta-learner, with the initial network parameters and Initialize the generator and the discriminator, and train and update the network parameters of the generator and the discriminator on the support set until the corresponding loss function converges or reaches a preset number of training epochs; After completing the in-loop training, in each out-of-loop training cycle: Calculate the loss function of each generative adversarial network on the corresponding query set, and then determine the meta-loss function; With the goal of minimizing the meta-loss function, perform cross-task update on the initial network parameters of the meta-learner.
3. A channel modeling method based on meta-learning according to claim 2, characterized in that The loss function of the generative adversarial network is: Wherein, and respectively represent the network parameters of the generator and the discriminator; , represents a random variable that follows a uniform distribution, represents the channel dataset, represents the synthetic data obtained after the random noise passes through the generator; and are respectively the probability distributions of white Gaussian noise and ; is 's probability distribution, λ is an adjustable parameter greater than zero; represents the mathematical expectation of the discrimination value obtained after the true channel data in the corresponding dataset passes through the discriminator; represents the mathematical expectation of the discrimination value obtained after the synthetic data passes through the discriminator; represents the loss function penalty term added by the WGAN to ensure the stability of GAN training.
4. A channel modeling method based on meta-learning according to claim 3, characterized in that, Denote minimizing the meta-loss function as the meta-objective function, which is expressed as: Among them, represents the meta-loss function, and L represents the loss function of the generative adversarial network.
5. A channel modeling method based on meta-learning according to claim 2, wherein, In the in-loop training cycle, update the network parameters of the generator and the discriminator based on the gradient descent method, where The parameter update process of the generator is expressed as: In the formula, represents the network parameters of the generator determined after training using the -th support set; represents the inner loop learning rate; represents the loss function of the generator calculated on the support set ; represents the gradient when the generator loss function is ; And the parameter update process of the discriminator is expressed as: In the formula, represents the network parameters of the discriminator determined after training using the th support set; represents the loss function of the discriminator calculated on the support set; represents the gradient when the discriminator loss function is .
6. The channel modeling method based on meta-learning according to claim 2, characterized in that In the out-of-loop training cycle, the network parameter update process of the generator and the discriminator is expressed as: In the formula, represents the meta learning rate; represents when the discriminator parameters are at this time, the sum of all loss functions when training the network using M query sets ; represents when the loss function is , and the discriminator parameters are at this time, the gradient of the discriminator network parameter update; represents when the generator parameters are at this time, the sum of all loss functions when training the network using M query sets ; represents when the loss function is , and the generator parameters are at this time, the gradient of the generator network parameter update.
7. A channel modeling method based on meta-learning according to claim 1, characterized in that, In the fine-tuning training cycle, the parameter update process of the generative adversarial network is expressed as follows: In the formula, represents the fine-tuning learning rate; represents that the corresponding network parameter is and the loss function on the target channel dataset when represents on the gradient of the model; and represent the network parameters of the fine-tuned discriminator and generator.
8. A channel modeling method based on meta-learning according to any one of claims 1-7, characterized in that: The generator sequentially includes a fully connected layer for expanding the data dimension, a batch normalization layer, a recombination operation layer for adjusting the data dimension, a plurality of upsampling modules for extracting high-level semantic features of channel data, and a cropping layer for outputting channel impulse response data of a fixed size; wherein, the upsampling module includes an upsampling layer, a convolutional layer, a batch normalization layer, and an activation function layer; The discriminator sequentially includes a padding layer for expanding the data dimension, a plurality of downsampling modules for extracting core features of channel data, a flattening operation layer for adjusting the data dimension, and a fully connected layer for outputting the authenticity of the synthetic data, wherein the downsampling module includes a convolutional layer and an activation function layer.
9. An electronic device, characterized in that, Comprising: a memory for storing computer-executable instructions or a computer program; a processor for implementing the method according to any one of claims 1-8 when executing the computer-executable instructions or the computer program stored in the memory.
10. A computer program product, comprising a computer program or computer-executable instructions, characterized in that, When the computer program or the computer-executable instructions are executed by the processor, the method according to any one of claims 1-8 is implemented.