A Machine Learning-Based Optimization Method for 6G Terahertz Communication
By using a machine learning-based reinforcement learning framework and a two-layer update mechanism, the transmitter control parameters of 6G terahertz communication are optimized, solving the problem of poor adaptability of traditional methods in dynamic environments, and achieving adaptive improvement of system performance and stable guarantee of communication quality.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHENGDU TECH UNIV
- Filing Date
- 2026-01-22
- Publication Date
- 2026-04-17
AI Technical Summary
Traditional optimization methods struggle to achieve real-time, globally optimal control in high-dimensional, complex, and dynamic 6G terahertz communication environments, resulting in system performance failing to reach theoretical limits.
A machine learning-based reinforcement learning framework is adopted to perform end-to-end joint decision-making through a policy network. A two-layer update mechanism combining importance sampling and orthogonal vortex balance optimization algorithm is used to optimize the transmitter control parameters, such as beamforming parameters, transmit power and modulation scheme.
It achieves adaptive optimization of the 6G terahertz communication system, improves decision accuracy and robustness, overcomes path loss and channel time-varying bottlenecks, ensures high-reliability and low-latency communication, and supports integrated communication and sensing.
Smart Images

Figure CN121567229B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of communication technology, specifically to a 6G terahertz communication optimization method based on machine learning. Background Technology
[0002] Sixth-generation (6G) mobile communication technology aims to realize the grand vision of the Internet of Things, digital twins, and immersive experiences, placing unprecedented demands on the peak data rate, latency, connection density, and reliability of communication systems. The terahertz band (0.1~10THz), with its ultra-large bandwidth (supporting Tbit / s-level transmission rates) and high-precision positioning and sensing potential, is widely recognized as one of the key enabling technologies for 6G. However, terahertz communication also faces severe challenges. The extremely short wavelength of terahertz waves leads to severe path loss, poor penetration, and limited communication coverage. Simultaneously, its communication links are susceptible to dynamic environmental factors such as human body obstruction and mobility, resulting in rapidly changing channel states and highly unstable communication quality.
[0003] To overcome these challenges, 6G systems commonly employ very large-scale antenna arrays (VMAs) and beamforming techniques to compensate for path loss. Furthermore, smart metasurfaces, as an emerging technology, can intelligently reshape channels and enhance signal coverage by dynamically adjusting the wireless propagation environment through low-cost reflective elements. However, the introduction of these technologies also dramatically increases the system's control dimensions. For example, base stations need to adjust beamforming parameters, transmit power, and modulation and coding schemes in real time, while simultaneously coordinating the control of the RIS (Radio Reflection System) reflection parameters. Traditional model-based or fixed-rule optimization methods struggle to achieve real-time, globally optimal control in such a high-dimensional, complex, and dynamic environment, resulting in system performance far from its theoretical limits. Summary of the Invention
[0004] The purpose of this application is to provide a machine learning-based optimization method for 6G terahertz communication, which aims to solve the problems of poor adaptability and poor optimization effect of traditional optimization methods due to the high-dimensional, complex and dynamic 6G terahertz communication environment, and to achieve adaptive and continuous improvement of communication system performance.
[0005] This application is achieved through the following technical solution:
[0006] A machine learning-based optimization method for 6G terahertz communication includes:
[0007] During 6G terahertz communication, communication status information is collected; wherein, the communication status information includes channel status information, beam quality index information and / or transmission performance index information;
[0008] Based on the communication status information, a policy network is used to select actions and determine the transmitter control parameters; wherein, the transmitter control parameters include beamforming parameters, transmit power and / or modulation method;
[0009] Based on the aforementioned transmitter control parameters, the transmitter is controlled to transmit communication signals, and communication performance indicators and communication status information at the next moment after executing the transmitter control parameters are obtained.
[0010] Based on communication status information, transmitter control parameters, communication performance indicators, and communication status information at the next moment, a historical playback experience pool is constructed.
[0011] Determine cycle information; wherein, the cycle information is the arrival of an optimization cycle and / or the arrival of a correction cycle; one correction cycle includes multiple optimization cycles;
[0012] If the periodic information is an optimized period, then based on the historical replay experience pool, an importance sampling strategy is used to perform a single-layer update on the policy network to determine the updated policy network.
[0013] If the periodic information is the correction period, then based on the historical replay experience pool, the policy network is updated in two layers using an importance sampling strategy and a hyperparameter optimization strategy based on the orthogonal vortex balance optimization algorithm to determine the updated policy network.
[0014] Based on the updated policy network, the transmitter control parameters in the subsequent 6G terahertz communication process are optimized to achieve 6G terahertz communication optimization.
[0015] In one possible implementation, it also includes:
[0016] After each parameter synchronization cycle, the network parameters of the policy network are synchronized to the target network so that the target network can guide the update of the policy network.
[0017] In one possible implementation, a historical playback experience pool is constructed based on communication status information, transmitter control parameters, communication performance indicators, and communication status information at the next moment, including:
[0018] The reward generated by executing the transmitter control parameters is determined based on the communication performance indicators.
[0019] The communication status information, transmitter control parameters, reward, and next-moment communication status information are constructed into a quadruple and placed into the historical replay experience pool.
[0020] The number of historical playback experience pools is fixed, and the number is limited by a first-in-first-out rule.
[0021] In one possible implementation, based on a historical replay experience pool, a single-layer update of the policy network is performed using an importance sampling strategy to determine the updated policy network, including:
[0022] For any sample in the historical replay experience pool, the priority corresponding to the sample is obtained, and the initial sampling probability is obtained according to the priority. The target sampling probability corresponding to the sample is obtained by performing importance sampling processing and normalization processing on the initial sampling probability.
[0023] Based on the target sampling probability corresponding to the sample, a fixed number of samples are collected from the historical playback experience pool to obtain the target sample.
[0024] Based on the target sample, the policy network is updated in a single layer using a strategy of target network guidance and gradient descent optimization to determine the updated policy network.
[0025] In one possible implementation, for any sample in the historical playback experience pool, the priority corresponding to the sample is obtained, and an initial sampling probability is obtained based on the priority. The target sampling probability corresponding to the sample is obtained by performing importance sampling processing and normalization processing on the initial sampling probability, including:
[0026] For any sample in the historical playback experience pool, determine the first Q value corresponding to the transmitter control parameters output by the policy network based on the communication status information;
[0027] Based on the communication status information at the next moment, the target network is used to select actions, and the maximum Q value generated based on the action selection is determined to obtain the second Q value.
[0028] Based on the first Q value, the second Q value, and the reward in the sample, obtain the priority corresponding to the sample, and obtain the sum of the priorities corresponding to all samples;
[0029] Based on the priority of the sample and the sum of priorities, determine the sampling probability of each sample in the historical replay experience pool;
[0030] Based on the sampling probability corresponding to the sample and the total number of samples in the historical replay experience pool, the importance of the sample is obtained by using an exponential function.
[0031] The importance of the sample is normalized to obtain the target sampling probability of the sample.
[0032] In one possible implementation, based on a historical replay experience pool, the policy network is updated in two layers using an importance sampling strategy and a hyperparameter optimization strategy based on an orthogonal vortex equilibrium optimization algorithm to determine the updated policy network, including:
[0033] An importance sampling strategy is used to determine the target samples, and the update counter is set to 1.
[0034] The hyperparameters involved in the gradient descent update process of the policy network are encoded to obtain multiple encoding vectors; wherein, the hyperparameters include discount factors and learning rates;
[0035] For any given encoding vector, the policy network is updated multiple times using simulated gradient descent with target samples, and the policy network after the simulated gradient descent update is tested with target samples. The average reward obtained from the test is used as the fitness of the encoding vector.
[0036] Based on the fitness of the encoded vector, the optimal encoded vector is determined from all encoded vectors.
[0037] Based on the optimal encoding vector, the encoding vector is updated quickly using a polar vortex search strategy to determine the updated encoding vector.
[0038] An information balancing search strategy is used to balance and update the updated encoding vector to determine the balanced and updated encoding vector.
[0039] An adaptive global update is performed on the balanced updated encoding vector using a Gaussian orthogonal search strategy to determine the globally updated encoding vector.
[0040] If the count value of the update counter is greater than the preset maximum number of updates, the target encoding vector is determined based on the globally updated encoding vector to achieve inner-layer optimization; otherwise, the count value of the update counter is incremented by one, and the fitness of the encoding vector is returned based on the globally updated encoding vector to perform the next update.
[0041] Based on the target encoding vector and target samples, the policy network is updated using a strategy of target network guidance and gradient descent optimization to determine the updated policy network and achieve outer layer optimization.
[0042] In one possible implementation, the policy network is updated multiple times using simulated gradient descent with target samples, and the policy network after the simulated gradient descent updates is tested with target samples. The average reward obtained from the test is used as the fitness of the encoding vector, including:
[0043] Based on the discount factor in the encoded vector, the loss function value is obtained using the target sample;
[0044] The policy network is updated multiple times using the loss function value and the learning rate in the encoding vector;
[0045] The policy network updated by the simulated gradient descent is tested using the communication state information in the target sample to determine the test reward;
[0046] The difference between the test reward and the reward in the target sample is obtained, and the mean of the differences corresponding to all target samples is used as the average reward to obtain the fitness of the encoding vector.
[0047] In one possible implementation, based on the optimal encoding vector, a polar radius vortex search strategy is used to quickly update the encoding vector, and the updated encoding vector is determined, including:
[0048] Initialize the polar diameter parameters and vortex angle parameters, and obtain the first vortex search control factor and the second vortex search control factor based on the polar diameter parameters and vortex angle parameters;
[0049] Obtain the center vector corresponding to all encoded vectors, and obtain the first vortex search quantity based on the center vector and the first vortex search control factor;
[0050] The second vortex search quantity is obtained based on the optimal encoding vector and the second vortex search control factor;
[0051] Based on the optimal encoding vector, the first vortex search quantity, and the second vortex search quantity, the encoding vector is updated quickly to obtain the updated encoding vector.
[0052] In one possible implementation, an information balancing search strategy is used to balance and update the updated encoding vector, determining the balanced and updated encoding vector, including:
[0053] For any updated encoding vector, determine the encoding vector with the closest Euclidean distance among the other updated encoding vectors, and obtain the nearest vector corresponding to the updated encoding vector;
[0054] The updated encoding vectors are arranged in ascending order of fitness to obtain the sorted encoding vectors;
[0055] Based on the sequence number of the sorted encoding vector and its corresponding nearest vector, obtain the reference information corresponding to the sorted encoding vector;
[0056] The update speed in the current update process is determined based on the arranged encoding vector, the optimal parameter vector, the reference information corresponding to the arranged encoding vector, and the update speed of the arranged encoding vector in the previous update process.
[0057] Based on the update speed in the current update process, the updated encoding vector is balanced and updated to determine the balanced encoding vector.
[0058] In one possible implementation, an adaptive global update is performed on the balanced updated encoding vector using a Gaussian orthogonal search strategy to determine the globally updated encoding vector, including:
[0059] Based on the balanced updated encoding vector, an orthogonal diagonal matrix is obtained, and based on the orthogonal diagonal matrix, a Gaussian distribution is used to obtain the global information acquisition vector;
[0060] Obtain the centroid vector corresponding to the balanced updated encoding vector, and obtain the overall information acquisition vector based on the centroid vector using Cauchy distribution;
[0061] Based on the optimal encoding vector, the optimal information collection vector is obtained using the Levy distribution;
[0062] Based on the global information acquisition vector, the overall information acquisition vector, and the optimal information acquisition vector, the encoding vector after the balanced update is adaptively updated globally to determine the encoding vector after the global update.
[0063] Compared with the prior art, this application has the following advantages and beneficial effects:
[0064] This application provides a machine learning-based 6G terahertz communication optimization method. It employs a reinforcement learning framework for communication performance optimization. Through continuous interaction with the environment and trial-and-error learning, the policy network can continuously adapt to channel changes, user movement, and sudden interference. The experience replay mechanism based on importance sampling prioritizes learning high-value experiences, greatly accelerating the learning process. Furthermore, the policy network is updated in two layers using an importance sampling strategy and a hyperparameter optimization strategy based on the orthogonal vortex balance optimization algorithm, further improving decision accuracy and the policy network's robustness in complex scenarios. Attached Figure Description
[0065] To more clearly illustrate the technical solutions of the exemplary embodiments of this application, the accompanying drawings used in the embodiments will be briefly described below. It should be understood that the following drawings only show some embodiments of this application and should not be considered as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort. In the drawings:
[0066] Figure 1 A flowchart illustrating a machine learning-based 6G terahertz communication optimization method provided in this application embodiment. Detailed Implementation
[0067] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the embodiments and accompanying drawings. The illustrative embodiments and descriptions of this application are only for explaining this application and are not intended to limit this application.
[0068] like Figure 1 As shown in the figure, this application provides a 6G terahertz communication optimization method based on machine learning, including:
[0069] S101. During 6G terahertz communication, collect communication status information; wherein, the communication status information includes channel status information, beam quality index information and / or transmission performance index information;
[0070] This application's embodiments integrate channel state information, beam quality indicators, and / or transmission performance indicators into the action space of reinforcement learning, and perform end-to-end joint decision-making through a policy network. This overcomes the local optimum problem caused by traditional step-by-step or independent optimization methods, and can find the optimal combination of parameters from a global perspective, thereby maximizing the overall system performance.
[0071] For example, channel state information may include a channel matrix, channel gain, Doppler shift, angle of arrival, and / or departure angle. For instance, the channel matrix, estimated from pilot signals, contains amplitude and phase information for the terahertz channel. Beam quality metrics may include signal-to-noise ratio, signal-to-interference-plus-noise ratio, and / or reference signal received power; transmission performance metrics may include bit error rate, throughput, and / or transmission delay. However, it is worth noting that this is merely an example illustrating an embodiment of this application; in actual implementation, other information may be used as communication state information to achieve state space acquisition.
[0072] S102. Based on the communication status information, a policy network is used to select actions and determine the transmitter control parameters; wherein, the transmitter control parameters include beamforming parameters, transmit power and / or modulation method;
[0073] For example, the beamforming parameters can be codeword indices for analog beamforming or weight vectors for digital beamforming; the modulation method can be QPSK, 16-QAM, 64-QAM, or other related signal modulation methods; and the transmit power can be the power value allocated for the current transmission.
[0074] Alternatively, the terahertz band (0.1~10THz) possesses abundant untapped spectrum resources, providing communication systems with ultra-large bandwidth and enabling 6G networks to support Tbit / s-level transmission rates. This band lies between light and microwaves, possessing the spectral resolution of light and the penetrability of microwaves. Furthermore, the shorter wavelength of terahertz waves allows for smaller antenna sizes, enabling the integration of very large-scale antenna arrays into wireless communication systems, providing high-precision positioning and high-resolution sensing capabilities.
[0075] The Reconfigurable Intelligent Surface (RIS) consists of numerous low-cost reflective elements, each capable of adjusting the phase and amplitude of the incident signal in real time, thus intelligently reshaping the wireless channel. The RIS can establish an additional communication link between the base station (BS) and the user, achieving precise coherent superposition of the reflected signal with the BS's direct signal. This enhances signal strength in the target area while suppressing signal strength in areas of interference. Therefore, the RIS's control parameters can also be used as transmitter control parameters, enabling comprehensive optimization of 6GHz terahertz communication.
[0076] S103. Based on the transmitter control parameters, control the transmitter to transmit communication signals, and obtain communication performance indicators and communication status information at the next moment after executing the transmitter control parameters.
[0077] S104. Construct a historical playback experience pool based on communication status information, transmitter control parameters, communication performance indicators, and communication status information at the next moment.
[0078] By building a historical replay experience pool, historical communication optimization experience can be stored, which facilitates subsequent optimization of the policy network, gradually improving decision accuracy and ensuring decision adaptability.
[0079] S105. Determine cycle information; wherein, the cycle information is the arrival of the optimization cycle and / or the arrival of the correction cycle; one correction cycle includes multiple optimization cycles;
[0080] For example, a timer can be used to time the process, performing a single-layer optimization once every optimization cycle; similarly, performing a double-layer optimization once every correction cycle, thereby achieving adaptive adjustment.
[0081] S106. If the periodic information is an optimized periodicity, then based on the historical replay experience pool, an importance sampling strategy is used to perform a single-layer update on the policy network to determine the updated policy network.
[0082] 6GHz terahertz channels are highly dynamic and uncertain. This application employs a reinforcement learning framework, enabling the policy network to continuously adapt to channel changes, user movement, and sudden interference through continuous interaction with the environment and trial-and-error learning. An experience replay mechanism based on importance sampling prioritizes learning high-value experiences, significantly accelerating the learning process and improving the model's robustness in complex scenarios.
[0083] By employing an importance sampling strategy to perform single-layer updates on the policy network, the weights of the policy network can be quickly adjusted, enabling it to respond rapidly to short-term changes in the environment and ensuring the real-time performance of communication optimization.
[0084] S107. If the periodic information is the correction period, then based on the historical playback experience pool, the policy network is updated in two layers using an importance sampling strategy and a hyperparameter optimization strategy based on the orthogonal vortex balance optimization algorithm to determine the updated policy network.
[0085] A two-layer update strategy is employed, combining an importance sampling strategy with a hyperparameter optimization strategy based on the orthogonal vortex equilibrium optimization algorithm. This introduces inner-layer hyperparameter optimization and automatically finds the optimal learning rate and discount factor using an advanced orthogonal vortex equilibrium optimization algorithm. This allows the reinforcement learning model to adaptively adjust its learning step size and foresight based on the evolution of the data distribution in the experience pool, effectively solving the problem of learning stagnation or slow convergence caused by fixed hyperparameters. This significantly improves the model's learning efficiency and the upper limit of its final convergence performance.
[0086] S108. Based on the updated strategy network, optimize the transmitter control parameters in the subsequent 6G terahertz communication process to achieve 6G terahertz communication optimization.
[0087] Through the aforementioned intelligent collaborative optimization, this application effectively overcomes bottlenecks such as path loss and channel time-varying in terahertz communication. Specific technical effects include: precise beamforming and RIS channel reshaping efficiently focus signal energy onto the user; combined with adaptive power and modulation selection, spectral efficiency is maximized, achieving near-theoretical transmission rates. In the face of obstructions or rapid movement, the RIS and beam can be quickly adjusted to rebuild or enhance the communication link, significantly reducing the probability of interruption and ensuring the requirements for high-reliability, low-latency communication. Intelligent power control avoids unnecessary energy waste, achieving green and energy-saving communication while ensuring communication quality. The optimized high-precision beam is not only used for communication but can also be used collaboratively for environmental sensing, providing technical support for integrated communication and sensing.
[0088] In one possible implementation, it also includes:
[0089] After each parameter synchronization cycle, the network parameters of the policy network are synchronized to the target network so that the target network can guide the update of the policy network.
[0090] Algorithms such as DQN (Deep Q-Network) make decisions through a policy network and guide updates through a target network. Similarly, algorithms like Deep Reinforcement Learning (DRL) and Deep Deterministic Policy Gradient (DDPG) all involve the policy network and target network working together to make decisions and updates.
[0091] In one possible implementation, a historical playback experience pool is constructed based on communication status information, transmitter control parameters, communication performance indicators, and communication status information at the next moment, including:
[0092] The reward generated by executing the transmitter control parameters is determined based on the communication performance indicators.
[0093] The communication status information, transmitter control parameters, reward, and next-moment communication status information are constructed into a quadruple and placed into the historical replay experience pool.
[0094] The number of historical playback experience pools is fixed, and the number is limited by a first-in-first-out rule.
[0095] In one possible implementation, based on a historical replay experience pool, a single-layer update of the policy network is performed using an importance sampling strategy to determine the updated policy network, including:
[0096] For any sample in the historical replay experience pool, the priority corresponding to the sample is obtained, and the initial sampling probability is obtained according to the priority. The target sampling probability corresponding to the sample is obtained by performing importance sampling processing and normalization processing on the initial sampling probability.
[0097] Based on the target sampling probability corresponding to the sample, a fixed number of samples are collected from the historical playback experience pool to obtain the target sample.
[0098] Based on the target sample, the policy network is updated in a single layer using a strategy of target network guidance and gradient descent optimization to determine the updated policy network.
[0099] The experience replay mechanism based on importance sampling prioritizes learning high-value experiences, which greatly accelerates the learning process and improves the model's decision robustness in complex scenarios.
[0100] In one possible implementation, for any sample in the historical playback experience pool, the priority corresponding to the sample is obtained, and an initial sampling probability is obtained based on the priority. The target sampling probability corresponding to the sample is obtained by performing importance sampling processing and normalization processing on the initial sampling probability, including:
[0101] For any sample in the historical playback experience pool, determine the first Q value corresponding to the transmitter control parameters output by the policy network based on the communication status information;
[0102] Based on the communication status information at the next moment, the target network is used to select actions, and the maximum Q value generated based on the action selection is determined to obtain the second Q value.
[0103] Based on the first Q value, the second Q value, and the reward in the sample, the priority corresponding to the sample is obtained as follows:
[0104]
[0105] in, This represents the communication status information in the i-th sample. This represents the transmitter control parameters in the i-th sample, i.e., the actions that have been executed. This represents the reward in the i-th sample. This represents the communication state information at the next time step in the i-th sample. Indicates will The input policy network generates actions, i = 1, 2, ..., I, where I represents the total number of samples in the historical replay experience pool. Indicates the discount factor. This indicates the priority in the i-th sample. Indicates in Below, the policy network output regarding Q value; Indicates in Below, the policy network output regarding The Q value; max indicates taking the maximum value;
[0106] Obtain the sum of priorities corresponding to all samples, and based on the priorities of the samples and the sum of priorities, determine the sampling probability corresponding to each sample in the historical replay experience pool as follows:
[0107]
[0108] in, This represents the sampling probability corresponding to the i-th sample;
[0109] Based on the sampling probability corresponding to the sample and the total number of samples in the historical replay experience pool, the importance of the sample is obtained using an exponential function:
[0110]
[0111] in, This indicates the importance of the i-th sample. This represents the importance control factor, and is set to 2.
[0112] The importance of the sample is normalized to obtain the target sampling probability of the sample.
[0113] By employing the aforementioned importance sampling strategy, we can avoid the bias that may be introduced by existing priority experience replay, where some experiences are oversampled while others are sampled at low frequency, thus improving the stability of the algorithm.
[0114] In one possible implementation, based on a historical replay experience pool, the policy network is updated in two layers using an importance sampling strategy and a hyperparameter optimization strategy based on an orthogonal vortex equilibrium optimization algorithm to determine the updated policy network, including:
[0115] An importance sampling strategy is used to determine the target samples, and the update counter is set to 1.
[0116] The hyperparameters involved in the gradient descent update process of the policy network are encoded to obtain multiple encoding vectors; wherein, the hyperparameters include discount factors and learning rates;
[0117] For example, we can first set upper and lower limits for hyperparameters, then randomly initialize them between these limits and encode them into vectors, thus obtaining the encoded vector. During the update and optimization process, there is a possibility of optimization exceeding the limits. For hyperparameters that exceed the limits, they should be randomly reset or reset to the limit with the smallest difference; for example, if they exceed the upper limit, they should be reset to the upper limit.
[0118] For any given encoding vector, the policy network is updated multiple times using simulated gradient descent with target samples, and the policy network after the simulated gradient descent update is tested with target samples. The average reward obtained from the test is used as the fitness of the encoding vector.
[0119] Based on the fitness of the encoded vector, the optimal encoded vector is determined from all encoded vectors.
[0120] Based on the optimal encoding vector, the encoding vector is updated quickly using a polar vortex search strategy to determine the updated encoding vector.
[0121] An information balancing search strategy is used to balance and update the updated encoding vector to determine the balanced and updated encoding vector.
[0122] An adaptive global update is performed on the balanced updated encoding vector using a Gaussian orthogonal search strategy to determine the globally updated encoding vector.
[0123] If the count value of the update counter is greater than the preset maximum number of updates, the target encoding vector is determined based on the globally updated encoding vector to achieve inner-layer optimization; otherwise, the count value of the update counter is incremented by one, and the fitness of the encoding vector is returned based on the globally updated encoding vector to perform the next update.
[0124] Based on the target encoding vector and target samples, the policy network is updated using a strategy of target network guidance and gradient descent optimization to determine the updated policy network and achieve outer layer optimization.
[0125] By introducing an inner layer of hyperparameter optimization and employing an advanced orthogonal vortex equilibrium optimization algorithm, the optimal learning rate and discount factor are automatically found. This enables the reinforcement learning model to adaptively adjust its learning step size and foresight based on the evolution of the data distribution in the experience pool, effectively solving the problem of learning stagnation or slow convergence caused by fixed hyperparameters, and significantly improving the model's learning efficiency and the upper limit of the final converged performance.
[0126] This two-layer optimization mechanism can effectively adapt to changes in scenarios and equipment aging, enabling the updated decision network to maintain accurate decision-making behavior in dynamic environments, improving communication optimization capabilities while also enhancing communication stability.
[0127] In one possible implementation, the policy network is updated multiple times using simulated gradient descent with target samples, and the policy network after the simulated gradient descent updates is tested with target samples. The average reward obtained from the test is used as the fitness of the encoding vector, including:
[0128] Based on the discount factor in the encoded vector, the loss function value is obtained using the target sample;
[0129] The policy network is updated multiple times using the loss function value and the learning rate in the encoding vector;
[0130] The policy network updated by the simulated gradient descent is tested using the communication state information in the target sample to determine the test reward;
[0131] It's worth noting that the simulated gradient descent update here is only for testing purposes and does not actually update the policy network, thus ensuring normal 6G terahertz communication. For example, the policy network can be copied, and then, based on the hyperparameters in the encoding vector, a gradient descent optimization strategy can be used to perform simulated gradient descent updates on the copied policy network. The policy network after the simulated gradient descent update can then be tested using the communication state information in the target samples to determine the test reward corresponding to each target sample.
[0132] The difference between the test reward and the reward in the target sample is obtained, and the mean of the differences corresponding to all target samples is used as the average reward to obtain the fitness of the encoding vector.
[0133] A higher average reward indicates better hyperparameters in the encoding vector, so the average reward can be used as the fitness of the encoding vector.
[0134] Optionally, an anti-fluctuation mechanism can be set, which means that when the absolute value of the average reward is less than the preset reward change threshold, the original hyperparameters remain unchanged, that is, the discount factor and learning rate remain unchanged.
[0135] For deep reinforcement learning algorithms that employ a policy network and a target network in cooperation, obtaining the loss function value through a discount factor is a fairly standard operation. Similarly, for gradient descent algorithms, updating the gradient through the learning rate is also a fairly standard operation. Therefore, the embodiments in this application will not be described in detail.
[0136] In one possible implementation, based on the optimal encoding vector, a polar radius vortex search strategy is used to quickly update the encoding vector, and the updated encoding vector is determined, including:
[0137] The initial polar diameter and vortex angle parameters are as follows:
[0138]
[0139]
[0140] in, This represents the polar radius parameter corresponding to the j-th encoded vector. This represents the vortex angle parameter corresponding to the j-th encoded vector. Represents pi (π). Represents the first random number between (0,1). This represents the second random number between (0,1). This represents the first constant term between (5, 10). This represents the second constant term between (5, 10);
[0141] Based on the polar diameter parameter and the vortex angle parameter, the first vortex search control factor and the second vortex search control factor are obtained as follows:
[0142]
[0143]
[0144] in, This represents the first vortex search control factor corresponding to the j-th encoded vector. This represents the second vortex search control factor corresponding to the j-th encoding vector. Represents the sine function. Represents the cosine function. Indicates all Take the maximum value from the middle. Indicates all Take the maximum value from the middle.
[0145] Obtain the center vectors corresponding to all encoded vectors, and based on the center vectors and the first vortex search control factor, obtain the first vortex search quantity as follows:
[0146]
[0147] in, This indicates the search volume for the first vortex. This represents the center vector, meaning that the parameter of each dimension is the mean of the parameters of all encoded vectors in the same dimension.
[0148] Let represent the j-th encoded vector in the t-th update process, where j = 1, 2, ..., NP, and NP represents the total number of encoded vectors;
[0149] Based on the optimal encoding vector and the second vortex search control factor, the second vortex search quantity is obtained as follows:
[0150]
[0151] in, This indicates the search volume for the second vortex. Represents the optimal encoding vector;
[0152] Based on the optimal encoding vector, the first vortex search value, and the second vortex search value, the encoding vector is quickly updated to obtain the quickly updated encoding vector as follows:
[0153]
[0154] in, This represents the encoded vector after the j-th fast update.
[0155] This application provides a polar radius vortex search, simulating vortex motion in fluid mechanics. It possesses powerful global exploration capabilities, effectively preventing the algorithm from getting trapped in local optima. Furthermore, this vortex motion revolves around a known optimal solution; as the algorithm progresses, the convergence accuracy gradually improves, comprehensively enhancing search efficiency.
[0156] In one possible implementation, an information balancing search strategy is used to balance and update the updated encoding vector, determining the balanced and updated encoding vector, including:
[0157] For any updated encoding vector, determine the encoding vector with the closest Euclidean distance among the other updated encoding vectors, and obtain the nearest vector corresponding to the updated encoding vector;
[0158] The updated encoding vectors are arranged in ascending order of fitness to obtain the sorted encoding vectors;
[0159] Based on the sequence number of the sorted encoding vector and its corresponding nearest vector, the reference information corresponding to the sorted encoding vector is obtained as follows:
[0160]
[0161] Where m represents the index of the sorted encoded vector. These are reference coefficients for the constant term, such as 1, 2, etc. This represents the update speed of the nearest vector corresponding to the encoded vector after the m-th permutation during the t-th update process. This represents the reference information corresponding to the encoded vector after the m-th permutation, where m = 1, 2, ..., NP;
[0162] Based on the rearranged encoding vector, the optimal parameter vector, the reference information corresponding to the rearranged encoding vector, and the update speed of the rearranged encoding vector in the previous update process, the update speed in the current update process is determined as follows:
[0163]
[0164] in, Let represent the encoded vector after the m-th permutation during the t-th update process. This represents the update speed of the encoded vector after the m-th permutation in the t-th update process, which is the update speed in the previous update process; This represents the update speed of the encoded vector after the m-th permutation during the (t+1)-th update process, which is the update speed in the current update process; This represents a third random number between (0,1). This represents the fourth random number between (0,1). The information learning coefficient is represented as a constant term and is set to 0.85;
[0165] Based on the update rate in the current update process, the updated encoding vector is balanced and updated to determine the balanced encoding vector as follows:
[0166]
[0167] in, This represents the encoding vector after the m-th balance update.
[0168] The information-balanced search provided in this application draws on swarm intelligence, balancing exploration (finding new solutions) and utilization (deepening the current optimal solution) through information sharing between individuals and the group. This effectively maintains the diversity of encoding vectors in the early and mid-stages of the algorithm, thereby avoiding getting trapped in local optima.
[0169] In one possible implementation, an adaptive global update is performed on the balanced updated encoding vector using a Gaussian orthogonal search strategy to determine the globally updated encoding vector, including:
[0170] Based on the balanced updated encoding vector, an orthogonal diagonal matrix is obtained, and based on the orthogonal diagonal matrix, a Gaussian distribution is used to obtain the global information acquisition vector;
[0171] For example, D encoded vectors can be constructed into a D*D dimensional matrix B, where D represents the total dimension of the parameters in the encoded vectors;
[0172] Decompose matrix B into: B = CEC -1 Where C represents a matrix composed of the eigenvectors of matrix B, and E represents an orthogonal diagonal matrix;
[0173] An orthogonal vector is constructed by taking the diagonal elements of the orthogonal diagonal matrix that have values, and then using a Gaussian distribution based on the orthogonal vector to obtain the global information acquisition vector:
[0174]
[0175] in, This represents the encoded vector after the nth balancing update during the t-th update process. Describing orthogonal vectors, Denotes the Gaussian distribution factor. This represents the fifth random number between (0,1). express The corresponding global information acquisition vector;
[0176] Obtain the centroid vector corresponding to the balanced updated encoding vector, and obtain the overall information acquisition vector based on the centroid vector using Cauchy distribution;
[0177] For example, we can first obtain the fitness of each balanced updated encoding vector, then calculate the sum of fitness, divide the fitness of the balanced updated encoding vector by the sum of fitness as the weighting weight, and use the weighting weight to perform a weighted summation of the balanced updated encoding vectors in the same dimension to obtain the centroid vector. That is, the parameter of each dimension in the centroid vector is obtained by weighted summation of the parameters of all balanced updated encoding vectors in the same dimension.
[0178] Then, based on the centroid vector, the Cauchy distribution is used to obtain the overall information acquisition vector as follows:
[0179]
[0180] in, Represents the overall information collection vector. This represents the fifth random number between (0,1). Let T represent the Cauchy distribution factor, and T represent the preset maximum number of updates. Represents the centroid vector;
[0181] Based on the optimal encoding vector, the optimal information collection vector is obtained using the Levy distribution:
[0182]
[0183] in, This represents the optimal information collection vector. This represents the sixth random number between (0,1). Represents the Lévy distribution factor. Represents the optimal encoding vector;
[0184] Based on the global information acquisition vector, the overall information acquisition vector, and the optimal information acquisition vector, an adaptive global update is performed on the balanced updated encoding vector to determine the globally updated encoding vector as follows:
[0185]
[0186]
[0187] in, This represents the encoded vector after the nth global update. express The update speed during update t+1. express The update speed during the t-th update process. Represents the natural constant.
[0188] The Gaussian orthogonal search provided in this application combines the characteristics of various random distributions to achieve multi-directional and adaptive search, which can effectively improve the global search capability of the algorithm and enable the algorithm to escape local optima during the local search process.
[0189] Optionally, simulated annealing can be used to control the above update strategies, further ensuring the update speed of the algorithm.
[0190] The combination of these three elements ensures the efficiency and globality of hyperparameter optimization, thus laying a solid foundation for the optimal performance of the policy network. This two-layer optimization mechanism can effectively adjust the decision-making ability of the policy network, making its decisions more accurate, thereby guaranteeing continuous and adaptive optimization of communication performance.
[0191] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of this application. It should be understood that the above description is only a specific embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A 6G terahertz communication optimization method based on machine learning, characterized in that, include: During 6G terahertz communication, communication status information is collected; wherein, the communication status information includes channel status information, beam quality index information and / or transmission performance index information; Based on the communication status information, a policy network is used to select actions and determine the transmitter control parameters; wherein, the transmitter control parameters include beamforming parameters, transmit power and / or modulation method; Based on the aforementioned transmitter control parameters, the transmitter is controlled to transmit communication signals, and communication performance indicators and communication status information at the next moment after executing the transmitter control parameters are obtained. Based on communication status information, transmitter control parameters, communication performance indicators, and communication status information at the next moment, a historical playback experience pool is constructed. Determine cycle information; wherein, the cycle information is the arrival of an optimization cycle and / or the arrival of a correction cycle; one correction cycle includes multiple optimization cycles; If the periodic information is an optimized period, then based on the historical replay experience pool, an importance sampling strategy is used to perform a single-layer update on the policy network to determine the updated policy network. If the periodic information is the correction period, then based on the historical replay experience pool, the policy network is updated in two layers using an importance sampling strategy and a hyperparameter optimization strategy based on the orthogonal vortex balance optimization algorithm to determine the updated policy network. Based on the updated policy network, the transmitter control parameters in the subsequent 6G terahertz communication process are optimized to achieve 6G terahertz communication optimization. Based on the historical replay experience pool, an importance sampling strategy is used to perform a single-layer update on the policy network to determine the updated policy network, including: For any sample in the historical replay experience pool, obtain the priority corresponding to the sample, and obtain the initial sampling probability according to the priority. By performing importance sampling processing and normalization processing on the initial sampling probability, the target sampling probability corresponding to the sample is obtained. Based on the target sampling probability corresponding to the sample, a fixed number of samples are collected from the historical playback experience pool to obtain the target sample. Based on the target sample, the policy network is updated in a single layer using a strategy of target network guidance and gradient descent optimization to determine the updated policy network. Based on the historical replay experience pool, a two-layer update is performed on the policy network using an importance sampling strategy and a hyperparameter optimization strategy based on the orthogonal vortex equilibrium optimization algorithm to determine the updated policy network, including: An importance sampling strategy is used to determine the target samples, and the update counter is set to 1. The hyperparameters involved in the gradient descent update process of the policy network are encoded to obtain multiple encoding vectors; wherein, the hyperparameters include discount factors and learning rates; For any given encoding vector, the policy network is updated multiple times using simulated gradient descent with target samples, and the policy network after the simulated gradient descent update is tested with target samples. The average reward obtained from the test is used as the fitness of the encoding vector. Based on the fitness of the encoded vector, the optimal encoded vector is determined from all encoded vectors. Based on the optimal encoding vector, the encoding vector is updated quickly using a polar vortex search strategy to determine the updated encoding vector. An information balancing search strategy is used to balance and update the updated encoding vector to determine the balanced and updated encoding vector. An adaptive global update is performed on the balanced updated encoding vector using a Gaussian orthogonal search strategy to determine the globally updated encoding vector. If the count value of the update counter is greater than the preset maximum number of updates, the target encoding vector is determined based on the globally updated encoding vector to achieve inner-layer optimization; otherwise, the count value of the update counter is incremented by one, and the fitness of the encoding vector is returned based on the globally updated encoding vector to perform the next update. Based on the target encoding vector and target samples, the policy network is updated using a strategy of target network guidance and gradient descent optimization to determine the updated policy network and achieve outer layer optimization.
2. The 6G terahertz communication optimization method based on machine learning according to claim 1, characterized in that, Also includes: After each parameter synchronization cycle, the network parameters of the policy network are synchronized to the target network so that the target network can guide the update of the policy network.
3. The 6G terahertz communication optimization method based on machine learning according to claim 1, characterized in that, Based on communication status information, transmitter control parameters, communication performance indicators, and communication status information at the next moment, a historical playback experience pool is constructed, including: The reward generated by executing the transmitter control parameters is determined based on the communication performance indicators. The communication status information, transmitter control parameters, reward, and next-moment communication status information are constructed into a quadruple and placed into the historical replay experience pool. The number of historical playback experience pools is fixed, and the number is limited by a first-in-first-out rule.
4. The 6G terahertz communication optimization method based on machine learning according to claim 1, characterized in that, For any sample in the historical playback experience pool, the priority corresponding to the sample is obtained, and an initial sampling probability is obtained based on the priority. The target sampling probability corresponding to the sample is obtained by performing importance sampling processing and normalization processing on the initial sampling probability, including: For any sample in the historical playback experience pool, determine the first Q value corresponding to the transmitter control parameters output by the policy network based on the communication status information; Based on the communication status information at the next moment, the target network is used to select actions, and the maximum Q value generated based on the action selection is determined to obtain the second Q value. Based on the first Q value, the second Q value, and the reward in the sample, obtain the priority corresponding to the sample, and obtain the sum of the priorities corresponding to all samples; Based on the priority of the sample and the sum of priorities, determine the sampling probability of each sample in the historical replay experience pool; Based on the sampling probability corresponding to the sample and the total number of samples in the historical replay experience pool, the importance of the sample is obtained by using an exponential function. The importance of the sample is normalized to obtain the target sampling probability of the sample.
5. The 6G terahertz communication optimization method based on machine learning according to claim 4, characterized in that, The policy network is updated multiple times using simulated gradient descent with target samples, and then tested with target samples after the simulated gradient descent update. The average reward obtained from the test is used as the fitness of the encoding vector, including: Based on the discount factor in the encoded vector, the loss function value is obtained using the target sample; The policy network is updated multiple times using the loss function value and the learning rate in the encoding vector; The policy network updated by the simulated gradient descent is tested using the communication state information in the target sample to determine the test reward; The difference between the test reward and the reward in the target sample is obtained, and the mean of the differences corresponding to all target samples is used as the average reward to obtain the fitness of the encoding vector.
6. The 6G terahertz communication optimization method based on machine learning according to claim 4, characterized in that, Based on the optimal encoding vector, a polar vortex search strategy is used to quickly update the encoding vector, and the updated encoding vector is determined, including: Initialize the polar diameter parameters and vortex angle parameters, and obtain the first vortex search control factor and the second vortex search control factor based on the polar diameter parameters and vortex angle parameters; Obtain the center vector corresponding to all encoded vectors, and obtain the first vortex search quantity based on the center vector and the first vortex search control factor; The second vortex search quantity is obtained based on the optimal encoding vector and the second vortex search control factor; Based on the optimal encoding vector, the first vortex search quantity, and the second vortex search quantity, the encoding vector is updated quickly to obtain the updated encoding vector.
7. The 6G terahertz communication optimization method based on machine learning according to claim 4, characterized in that, An information-balanced search strategy is used to balance and update the updated encoding vector, determining the balanced and updated encoding vector, including: For any updated encoding vector, determine the encoding vector with the closest Euclidean distance among the other updated encoding vectors, and obtain the nearest vector corresponding to the updated encoding vector; The updated encoding vectors are arranged in ascending order of fitness to obtain the sorted encoding vectors; Based on the sequence number of the sorted encoding vector and its corresponding nearest vector, obtain the reference information corresponding to the sorted encoding vector; The update speed in the current update process is determined based on the arranged encoding vector, the optimal encoding vector, the reference information corresponding to the arranged encoding vector, and the update speed of the arranged encoding vector in the previous update process. Based on the update speed in the current update process, the updated encoding vector is balanced and updated to determine the balanced encoding vector.
8. The 6G terahertz communication optimization method based on machine learning according to claim 4, characterized in that, An adaptive global update is performed on the balanced updated encoding vector using a Gaussian orthogonal search strategy to determine the globally updated encoding vector, including: Based on the balanced updated encoding vector, an orthogonal diagonal matrix is obtained, and based on the orthogonal diagonal matrix, a Gaussian distribution is used to obtain the global information acquisition vector; Obtain the centroid vector corresponding to the balanced updated encoding vector, and obtain the overall information acquisition vector based on the centroid vector using Cauchy distribution; Based on the optimal encoding vector, the optimal information collection vector is obtained using the Levy distribution; Based on the global information acquisition vector, the overall information acquisition vector, and the optimal information acquisition vector, the encoding vector after the balanced update is adaptively updated globally to determine the encoding vector after the global update.
Citation Information
Patent Citations
RIS auxiliary enhanced communication method based on evolution guidance strategy gradient
CN120223127A
Layered agent-based space-air-ground caching and resource optimization method and system
CN121126539A