A beam control algorithm for wireless communication
By combining deep learning and multi-objective optimization algorithms with reinforcement learning, the beam direction and width are dynamically adjusted, solving the problem of unstable signal transmission in complex environments caused by traditional beam control algorithms, and achieving efficient signal coverage and stable communication quality.
Patent Information
- Application Number
- CN202411700863.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-26
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2044-11-26
AI Technical Summary
Existing beam control algorithms struggle to adapt to changes in user location and network load fluctuations in complex communication environments, leading to unstable signal transmission and reduced communication robustness. This is especially true in 5G and future communication systems, where traditional fixed beamwidths cannot flexibly adapt to multipath reflection paths.
By constructing a channel state feature map using a deep learning model, and combining multi-objective optimization algorithms and reinforcement learning, the beam direction and width are dynamically adjusted to predict future channel change trends and optimize beam control strategies to adapt to complex environments.
It improves the effectiveness and stability of signal coverage, ensures communication quality, balances signal coverage quality and energy consumption, and adapts to high-density users and high-load network environments.
Smart Images

Figure CN119603698B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of communication beam control technology, and in particular to a beam control algorithm for wireless communication. Background Technology
[0002] In wireless communication systems, with the increasing demand for high speed and high capacity from users, traditional fixed beam control technology can no longer meet the requirements of complex communication environments. Especially in 5G and future communication systems, the application of millimeter wave communication technology makes signal propagation more susceptible to environmental obstacles and multipath effects, leading to signal attenuation and reduced communication performance. Therefore, how to improve channel utilization and coverage quality through more intelligent beam control algorithms has become the focus of current research.
[0003] Most existing beam control algorithms are based on real-time channel state information (CSI), which is poorly adaptable to changes in user location and environmental characteristics. Especially when users are moving at high speed or network load fluctuates greatly, the system has difficulty responding quickly and adjusting the beam direction, resulting in unstable signal transmission. In addition, traditional beam control often uses a fixed beamwidth, which cannot flexibly adapt to complex propagation environments and is difficult to effectively cover multipath reflection paths, further reducing the robustness of communication. Summary of the Invention
[0004] This invention provides a beam control algorithm for wireless communication.
[0005] A beam control algorithm for wireless communication includes:
[0006] S1: Initialize the beam set by selecting the initial beam direction set Θ from the preset beam mode library. init , used to cover a preset spatial area;
[0007] S2: Real-time acquisition of historical channel state information (CSI) from user terminals, and construction of a channel state feature map F using a deep learning model. CSI Channel state feature map F CSI The system is constructed based on the user's mobile history, environmental reflection characteristics, and multipath effects to predict the user's future channel change trends.
[0008] S3: Based on the channel state feature map F CSI Based on the future channel change trend, the beam direction set is initially adjusted. The adjustment process adopts a multi-objective optimization algorithm, which comprehensively considers channel gain, interference signal strength and system power consumption, and generates multiple candidate beam directions.
[0009] S4: For each candidate beam direction, the adaptive tuning module based on reinforcement learning adjusts the beamwidth so that the beam is not only focused on the target user direction, but also covers potential multipath reflection paths to improve the robustness of signal transmission.
[0010] S5: Based on changes in user location and real-time network load, the beam control algorithm is self-iteratio optimized. By continuously updating the beam pattern library and deep learning model, the beam control can better adapt to the complex wireless communication environment of the future.
[0011] Optionally, the initial beam set in S1 specifically includes:
[0012] S11: Based on historical data and user distribution characteristics in the wireless communication system, multiple beam pattern libraries are pre-built, and the beam pattern libraries include a variety of preset beam pattern sets with different widths, directions and coverage ranges;
[0013] S12: Based on the current user's geographical location, target spatial area requirements, and coverage target, match a set of beam patterns from the beam pattern library;
[0014] S13: Based on the shape and size of the target spatial region, select an initial beam direction set that can cover the target spatial region from the set of matched beam patterns.
[0015] Optionally, S2 specifically includes:
[0016] S21: Receive historical channel status information of user terminals in real time through base stations or access points in the wireless communication system. The historical channel status information includes channel gain, signal strength, and channel attenuation information of user terminals at different time points.
[0017] S22: Construct the channel state feature matrix H(t). The elements in the matrix represent the channel state of the user at different locations and time points. The channel state feature matrix H(t) is constructed based on the user's movement trajectory and the channel features at different time points. Each element represents the channel state at different times and locations.
[0018] S23: Input the channel state feature matrix H(t) into a pre-trained deep learning model to extract the spatial and temporal features of the channel state;
[0019] S24: The extracted channel state features are processed by a deep learning model to generate a channel state feature map, which reflects the influence of user mobility history, environmental reflection characteristics and multipath effects on the channel state.
[0020] S25: Using the generated channel state feature map, predict the user's future channel change trend for dynamic adjustment and optimization of beam direction.
[0021] Optionally, the channel state is also affected by multipath propagation effects, which are modeled as follows: Where, α k f is the attenuation coefficient for the k-th path. k Let t be the frequency of the k-th path, t be time, and K be the total number of multipath paths.
[0022] Optionally, the deep learning model combines convolutional neural networks (CNNs) and recurrent neural networks (RNNs) to extract spatial and temporal features. The convolutional layers of the CNNs are used to extract spatial features, outputting a feature map F. conv Recurrent neural networks are used to capture time dependencies and output the hidden state features h at time t. t ;
[0023] Channel State Feature Map F CSI The generation is represented as: F CSI =f(F conv ,h t ), where f represents the deep learning model's response to the feature map F. conv and hidden state features h t The fusion operation, F CSI This is the final generated channel state feature map;
[0024] Using the channel state feature map F CSI By learning from historical data, the future channel state is predicted. The prediction model uses a recurrent neural network, and the prediction is expressed as follows:
[0025] in, The predicted channel state at the next time point t+1 includes the time variations of channel gain, interference signal strength, and multipath propagation effects, and f is the prediction function of the recurrent neural network model.
[0026] Optionally, S3 specifically includes:
[0027] S31: From the channel state feature map F CSI Extract future channel change trends, including the time variations of channel gain, interference signal strength, and multipath effects;
[0028] S32: Adjust the initial beam direction set Θ using a multi-objective optimization algorithm based on future channel change trends. init The optimization objectives include:
[0029] Channel gain maximization: The objective function is: Among them, G i (θ i ) indicates the beam direction θ i The corresponding channel gain, where N is the number of beam directions;
[0030] Minimize interference signal strength: The objective function is: Among them, I i (θ i ) indicates the beam direction θ i Interference signal strength at the location;
[0031] System power consumption minimization: The optimization objective function is: Among them, P i (θ i θ represents the beam direction. i Corresponding system power consumption;
[0032] S33: Through a multi-objective optimization algorithm, multiple candidate beam directions θ are generated by comprehensively considering the weights of channel gain, interference signal strength, and system power consumption. cand This meets the optimization requirements for future channel state changes and serves as the basis for the next beam direction adjustment.
[0033] Optionally, S4 specifically includes:
[0034] S41: θ for each candidate beam direction cand As a state input in a reinforcement learning environment, the state includes current channel state information, user location, and multipath reflection path information;
[0035] S42: Define the action space of reinforcement learning as the adjustment range of the beamwidth Δθ, and the action a t This indicates a reduction or expansion of the current beamwidth, with the goal of optimizing channel gain and coverage.
[0036] S43: Define the reward function R t The function that comprehensively considers channel gain, user-direction signal quality, and multipath reflection path coverage is expressed as:
[0037] R t =w4·C(θ) cand ,Δθ)+w5·C(θ cand ,Δθ), where C(θ) cand Δθ) represents the beam direction θ cand And the channel gain under beamwidth Δθ, C(θ) cand , Δθ) represents the beam's ability to cover potential multipath reflection paths, and w4 and w5 are the weighting coefficients for channel gain and multipath coverage effect;
[0038] S44: Based on a reinforcement learning algorithm, the beamwidth is adjusted through multiple training iterations to focus the beam on the target user direction while covering multipath reflection paths to optimize signal quality and coverage. The optimized beamwidth Δθ is output. opt It is used for adjusting the beam direction during subsequent communication.
[0039] Optionally, the reinforcement learning algorithm includes defining the state in reinforcement learning as the current beam direction, beam width, channel state information, and user location information; the action is to adjust the beam width (which can expand or shrink the beam); the reward measures the effect after beam adjustment, including channel gain, user direction signal quality, and coverage of multipath reflection paths.
[0040] The reinforcement learning algorithm begins training by initializing the Q-table or, at each time step, selecting an action (i.e., adjusting the beamwidth) based on the current state and executing the action. After executing the action, it updates the state to the next state and updates the Q-table based on the feedback calculated by the reward function.
[0041] Through repeated iterations, the reinforcement learning algorithm gradually learns how to select the optimal beamwidth adjustment strategy, and finally outputs the optimized beamwidth.
[0042] Optionally, S5 specifically includes:
[0043] S51: Real-time acquisition of user terminal location information and network load, wherein the network load includes the current base station connection count, channel occupancy rate and signal interference intensity;
[0044] S52: Input user location change information and network load data into the deep learning model. The deep learning model identifies user location movement patterns and load distribution characteristics through continuous online training and learning.
[0045] S53: Based on the output of the deep learning model, dynamically adjust the beam direction and beam width to adapt to changes in the user's location and ensure the effectiveness of signal coverage;
[0046] S54: Self-updates the beam pattern library, optimizes the beam direction, beam width and coverage in the beam pattern library based on the actual distribution of users and network load, and generates new beam patterns suitable for the current network status.
[0047] S55: Through error feedback, the parameters of the deep learning model are updated. The convolutional neural network part of the deep learning model updates the weights of the spatial feature extraction layer, and the recurrent neural network part of the deep learning model updates the weights of the time series prediction layer. The updated model will be able to better adapt to changes in user location and dynamic changes in network load.
[0048] Optionally, the user location change data is obtained by collecting user location (x) data. u ,y u The change in position Δd is calculated based on the positional difference. u This refers to the user's displacement between two consecutive moments; the network load data is obtained by collecting the real-time connection count N of the base station. conn Channel occupancy rate U channel Interference signal strength I intf Real-time calculation of network load L net ;
[0049] The deep learning model uses a convolutional neural network to handle changes in user location and a recurrent neural network to handle changes in network load over time.
[0050] Input data: X input =(Δd) u ,L net ), where Δd u L represents the user's movement displacement. net This indicates the real-time network load status;
[0051] Output data: (Δθ) pred ,ΔW pred )=f DL (X input ), where Δθ pred : Predicted beam direction adjustment, ΔW pred : Predicted beamwidth adjustment amount;
[0052] Based on real-time user location changes and network load data, the deep learning model undergoes continuous online training and learning updates. At each time step t, new user location and network load data are input into the model to generate beam adjustment suggestions. These suggestions are then combined with actual beam performance and evaluated using an error function. Calculate the difference between the model's predicted values and the actual values:
[0053] Where, θ actual and W actual For the actual beam direction and width, θ i and W i For the beam direction and width predicted by the model, the parameters of the deep learning model are updated through error feedback and backpropagation. The convolutional neural network part updates the weights of the spatial feature extraction layer, and the recurrent neural network part updates the weights of the time series prediction layer.
[0054] The beneficial effects of this invention are:
[0055] This invention introduces a channel state feature map construction method based on deep learning, which can accurately predict the future channel change trend of users. Compared with traditional algorithms that rely on real-time channel state information (CSI), this invention not only considers the user's historical movement trajectory, but also combines environmental reflection characteristics and multipath effects, which greatly improves the prediction accuracy of channel state. Based on these features, the beam direction can be dynamically adjusted, thereby improving the effectiveness and stability of signal coverage and ensuring communication quality in complex propagation environments.
[0056] This invention, by combining deep learning models, achieves adaptive dynamic beam control based on user location changes and network load. By extracting spatial features using convolutional neural networks (CNNs) and capturing temporal features using recurrent neural networks (RNNs), it can analyze the movement trends of user terminals and changes in network load in real time, accurately predict the adjustment of beam direction and beamwidth. In complex and ever-changing wireless communication environments, it can more flexibly optimize channel gain and effectively avoid signal interference and coverage blind spots, thereby improving the overall communication quality of the system.
[0057] This invention comprehensively considers factors such as channel gain, interference signal strength, and system power consumption through a multi-objective optimization algorithm to dynamically optimize beam direction and width. Through an adaptive tuning module based on reinforcement learning, it can balance signal coverage quality and energy consumption under different user scenarios and communication conditions, and achieve precise adjustment of beam control strategy. It shows significant advantages in high-density user and high-load network environments, and can ensure that the communication system maintains a high energy efficiency ratio and stable network performance without sacrificing user experience. Attached Figure Description
[0058] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only for this invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0059] Figure 1 This is a schematic diagram of the algorithm steps in an embodiment of the present invention;
[0060] Figure 2 This is a schematic diagram of the constructed channel state feature map according to an embodiment of the present invention. Detailed Implementation
[0061] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments. It should also be noted that, to make the embodiments more comprehensive, the following embodiments are the best and preferred embodiments, and those skilled in the art can use other alternative methods to implement some well-known technologies; moreover, the accompanying drawings are only for more specific description of the embodiments and are not intended to specifically limit the present invention.
[0062] It should be noted that the use of terms such as "an embodiment," "an embodiment," "an exemplary embodiment," and "some embodiments" in the specification indicates that the described embodiment may include a specific feature, structure, or characteristic, but not every embodiment necessarily includes that specific feature, structure, or characteristic. Furthermore, when a specific feature, structure, or characteristic is described in connection with an embodiment, implementing such a feature, structure, or characteristic in conjunction with other embodiments (whether explicitly described or not) should be within the knowledge of those skilled in the art.
[0063] Generally, terms can be understood at least partly from their use in context. For example, depending at least partly on the context, the term "one or more" as used herein can be used to describe any feature, structure, or characteristic in a singular sense, or a combination of features, structures, or characteristics in a plural sense. Additionally, the term "based on" can be understood not necessarily to convey an exclusive set of factors, but rather, alternatively, depending at least partly on the context, to allow for the presence of other factors that are not necessarily explicitly described.
[0064] like Figures 1-2 As shown, a beam control algorithm for wireless communication includes:
[0065] S1: Initialize the beam set by selecting the initial beam direction set Θ from the preset beam mode library. init , used to cover a preset spatial area;
[0066] S2: Real-time acquisition of historical channel state information (CSI) from user terminals, and construction of a channel state feature map F using a deep learning model. CSI Channel state feature map F CSI The system is constructed based on the user's mobile history, environmental reflection characteristics, and multipath effects to predict the user's future channel change trends.
[0067] S3: Based on the channel state feature map F CSI Based on the future channel change trend, the beam direction set is initially adjusted. The adjustment process adopts a multi-objective optimization algorithm, which comprehensively considers channel gain, interference signal strength and system power consumption, and generates multiple candidate beam directions.
[0068] S4: For each candidate beam direction, the adaptive tuning module based on reinforcement learning adjusts the beamwidth so that the beam is not only focused on the target user direction, but also covers potential multipath reflection paths to improve the robustness of signal transmission.
[0069] S5: Based on changes in user location and real-time network load, the beam control algorithm is self-iteratio optimized. By continuously updating the beam pattern library and deep learning model, the beam control can better adapt to the complex wireless communication environment of the future.
[0070] The initial beam set in S1 specifically includes:
[0071] S11: Based on historical data and user distribution characteristics in the wireless communication system, multiple beam pattern libraries are pre-built. The beam pattern libraries include a variety of preset beam pattern sets with different widths, directions and coverage ranges.
[0072] The beam pattern library is constructed based on historical data and user distribution characteristics in wireless communication systems. Several preset beam pattern sets are defined, each with the following parameters:
[0073] Beam direction angle θ: Beam propagation direction, range [0°, 360°];
[0074] Beamwidth Δθ: The range of angles covered by the beam, in degrees;
[0075] Coverage radius R: The signal coverage area of the beam, in meters;
[0076] The beam pattern library presets beam patterns of different directions, widths, and coverage areas for later selection.
[0077] S12: Based on the current user's geographical location, target spatial area requirements, and coverage target, match a set of beam patterns from the beam pattern library;
[0078] The matching is based on whether the user's geographical location is within the beam's coverage area, and whether the beam's direction and width are suitable for covering the target spatial area.
[0079] The conditions for the matching process are:
[0080]
[0081] Where, (x u ,y u ) represents the current user's geographic location, A beam The coverage area of each beam is determined by its direction θ, width Δθ, and radius R. tar getIt is the area of the target spatial region; when these conditions are met, the matching is successful, and a set of beam patterns that meet the conditions is selected from the beam pattern library.
[0082] S13: Based on the shape and size of the target spatial region, select an initial beam direction set that can cover the target spatial region from the set of matched beam patterns.
[0083] Based on the shape and size of the target spatial region, an initial beam direction set is selected from the matched beam pattern set. The goal of the selection is to ensure that the beam direction set can cover the maximum portion of the target region. The conditions for selecting the beam direction are as follows:
[0084] The target region is defined as a rectangular or polygonal region with a width and height of W. target and H target The conditions for beam coverage are: Δθ m R is the beamwidth in the beam pattern. m This represents the coverage radius in the beam pattern.
[0085] This formula ensures that the beam angle can cover the width of the target area.
[0086] From the set of matched beam patterns, select a beam direction θ that meets the conditions. m So that it covers the entire A tar get The final initial beam direction set is Θ init ,in:
[0087] Θ init ={θ1,θ2,..,θ n} makes
[0088] A beam pattern set is a subset of beam patterns selected from a beam pattern library based on specific conditions (such as user location and coverage requirements). A beam pattern set is a subset of beam patterns that meet the requirements matched from the beam pattern library for the current communication scenario.
[0089] A beam direction set is a set of specific beam directions selected from the matching beam patterns in a beam pattern set, based on the shape and size of the target spatial region. The beam direction set is the actual beam direction configuration ultimately used to cover the target region.
[0090] The beam pattern library is a collection of all beam patterns.
[0091] A beam pattern set is a subset of beam patterns selected from the beam pattern library based on specific requirements.
[0092] The beam direction set is the specific beam direction configuration selected from the beam pattern set to cover the target area.
[0093] S2 specifically includes:
[0094] S21: Receive historical channel status information of user terminals in real time through base stations or access points in the wireless communication system. The historical channel status information includes channel gain, signal strength and channel attenuation information of user terminals at different time points.
[0095] S22: Construct the channel state feature matrix H(t), where the elements represent the user's channel state at different locations and times. The channel state feature matrix H(t) is constructed based on the user's movement trajectory and channel features at different times. Each element represents the channel state at different times and locations, expressed as:
[0096] in;
[0097] h ij (t) represents the channel state information of the user terminal from base station i to user location j at time t, including channel gain, signal strength and attenuation characteristics. i is the number of base stations, j is the number of sampling points at the location of the user terminal, and t represents time. This matrix records the channel state of the user at different times and locations, and is used to describe the channel characteristics between the user and the base station.
[0098] S23: Input the channel state feature matrix H(t) into the pre-trained deep learning model to extract the spatial and temporal features of the channel state;
[0099] S24: The extracted channel state features are processed by a deep learning model to generate a channel state feature map. The channel state feature map reflects the impact of user mobility history, environmental reflection characteristics and multipath effects on the channel state.
[0100] S25: Using the generated channel state feature map, predict the user's future channel change trend for dynamic adjustment and optimization of beam direction.
[0101] Channel conditions are also affected by multipath propagation effects, which are modeled as follows: Where, α k f is the attenuation coefficient for the k-th path. k Let be the frequency of the k-th path, t be the time, and K be the total number of multipath paths. This represents the superposition of signals from multiple paths from the base station to the user terminal, reflecting the impact of multipath propagation in the channel.
[0102] Deep learning models combine convolutional neural networks (CNNs) and recurrent neural networks (RNNs) to extract spatial and temporal features. The convolutional layers of the CNN are used to extract spatial features, outputting a feature map F.conv Recurrent neural networks are used to capture time dependencies and output the hidden state features h at time t. t ;
[0103] A convolutional layer is represented as: F conv =ReLU(W conv *H(t)+b conv ), where F conv W is the feature map output by the convolutional layer. conv The kernel weight matrix * represents the convolution operation, b conv ReLU is the bias term, and ReLU is the activation function.
[0104] Recurrent neural networks process time series features, represented as follows:
[0105] h t =σ(W h ·x t +U h ·h t-1 +b h );h t Let x be the hidden state at time t. t It is the current input feature, h t-1 It is the hidden state at the previous point in time, W h and U h Let b be the weight matrix. h It is the bias term, and σ is the activation function;
[0106] Channel State Feature Map F CSI The generation is represented as: F CSI =f(F conv ,h t ), where f represents the deep learning model's response to the feature map F. conv and hidden state features h t The fusion operation, F CSI This is the final generated channel state feature map;
[0107] Using the channel state feature map F CSI By learning from historical data, the future channel state is predicted. The prediction model uses a recurrent neural network, and the prediction is expressed as follows:
[0108] in, Let g be the predicted channel state at the next time point t+1, including the time variations of channel gain, interference signal strength, and multipath propagation effects, and g be the prediction function of the recurrent neural network model.
[0109] S3 specifically includes:
[0110] S31: From the channel state feature map FCSI Extract future channel change trends, including the time variations of channel gain, interference signal strength, and multipath effects;
[0111] S32: Adjust the initial beam direction set Θ using a multi-objective optimization algorithm based on future channel change trends. init The optimization objectives include:
[0112] Channel gain maximization: The objective function is: Among them, G i (θ i ) indicates the beam direction θ i The corresponding channel gain, where N is the number of beam directions;
[0113] Minimize interference signal strength: The objective function is: Among them, I i (θ i ) indicates the beam direction θ i Interference signal strength at the location;
[0114] System power consumption minimization: The optimization objective function is: Among them, P i (θ i θ represents the beam direction. i Corresponding system power consumption;
[0115] S33: Through a multi-objective optimization algorithm, multiple candidate beam directions θ are generated by comprehensively considering the weights of channel gain, interference signal strength, and system power consumption. cand This meets the optimization requirements for future channel state changes and serves as the basis for the next beam direction adjustment.
[0116] The steps of the multi-objective optimization algorithm are as follows:
[0117] Step 1: Determine the objective functions. To solve this optimization problem, three objective functions are defined:
[0118] Channel gain maximization: f1(θ) i ) = G i (θ i ), where G i (θ i ) is the beam direction θ i The corresponding channel gain.
[0119] Minimize interference signal strength: f2(θ) i ) = I i (θ i ), where I i (θ i ) is the beam direction θ i The strength of the interference signal at that location.
[0120] System power consumption minimization: f3(θ) i ) = P i (θ i ), where P i (θ i ) is the beam direction θ i The corresponding system power consumption.
[0121] Step 2: Determine the weights. Assign weights w1, w2, and w3 to each objective function, representing the importance of channel gain, interference signal strength, and system power consumption in the optimization process, respectively. The weights must satisfy the following condition: w1 + w2 + w3 = 1. The weight values are set according to the priorities in the actual application. For example, if you want to give priority to channel gain, you can set w1 to be larger.
[0122] Step 3: Construct a comprehensive objective function. Use the weighted sum method to combine the three objective functions into a single objective function:
[0123] F(θ i )=w1·f1(θ i )-w2·f2(θ i )-w3·f3(θ i ), f1(θ i The channel gain that needs to be maximized is w1,f2(θ). i ) and f3(θ i ) needs to be minimized, therefore negative weights w2 and w3 are required.
[0124] Step 4: Initialize the candidate beam direction set, starting from the initial beam direction set θ init One or more initial candidate beam directions are selected as the initial points for optimization.
[0125] Step 5: Iterative optimization, using optimization algorithms to optimize the comprehensive objective function F(θ) i The algorithm iterates on the beam to find the optimal beam direction, taking gradient descent as an example:
[0126] 1. Calculate the gradient: For the combined objective function F(θ) i Calculate θ relative to the beam direction i gradient:
[0127] 2. Update beam direction: Update the beam direction θ according to the negative gradient direction. i :
[0128] Where η is the learning rate, which determines the step size.
[0129] 3. Convergence condition: When the change in gradient... The iteration stops when the number of iterations is less than the set threshold or the maximum number of iterations is reached.
[0130] Step 6: Generate multiple candidate beam directions. During the iterative optimization process, record all beam directions that meet the optimization objective and generate multiple candidate beam direction sets θ. cand This candidate set contains multiple optimization results.
[0131] Step 7: Select the optimal beam direction or combination from the candidate beam direction set based on the weight allocation and the prediction of future channel state.
[0132] S4 specifically includes:
[0133] S41: θ for each candidate beam direction cand As the state input in a reinforcement learning environment, the state includes current channel state information, user location, and multipath reflection path information;
[0134] S42: Define the action space of reinforcement learning as the adjustment range of the beamwidth Δθ, and the action a t This indicates a reduction or expansion of the current beamwidth, with the goal of optimizing channel gain and coverage.
[0135] S43: Define the reward function R t The function that comprehensively considers channel gain, user-direction signal quality, and multipath reflection path coverage is expressed as:
[0136] R t =w4·G(θ) cand ,Δθ)+w5·C(θ cand ,Δθ), where G(θ) cand Δθ) represents the beam direction θ cand And the channel gain under beamwidth Δθ, C(θ) cand , Δθ) represents the beam's ability to cover potential multipath reflection paths, and w4 and w5 are the weighting coefficients for channel gain and multipath coverage effect;
[0137] S44: Based on a reinforcement learning algorithm, the beamwidth is adjusted through multiple training iterations to focus the beam on the target user direction while covering multipath reflection paths to optimize signal quality and coverage. The optimized beamwidth Δθ is output. opt It is used for adjusting the beam direction during subsequent communication.
[0138] The reinforcement learning algorithm includes defining the state in reinforcement learning as the current beam direction, beamwidth, channel state information, and user location information; the action is to adjust the beamwidth (which can expand or shrink the beam); the reward measures the effect after beam adjustment, including channel gain, user direction signal quality, and coverage of multipath reflection paths.
[0139] The reinforcement learning algorithm begins training by initializing the Q-table or, at each time step, selecting an action (i.e., adjusting the beamwidth) based on the current state and executing the action. After executing the action, it updates the state to the next state and updates the Q-table based on the feedback calculated by the reward function.
[0140] Through repeated iterations, the reinforcement learning algorithm gradually learns how to select the optimal beamwidth adjustment strategy, and finally outputs the optimized beamwidth.
[0141] The specific steps of the reinforcement learning algorithm are as follows:
[0142] 1. Define reinforcement learning:
[0143] state s t At time step t, state s t Including the current beam direction θ cand The beamwidth Δθ, channel state information, and user location information are included, with the state indicating the current beam configuration and environmental conditions.
[0144] Action a t The action space includes adjustments to the beamwidth Δθ, and action a t Defined as increasing or decreasing beamwidth, the actions include: a t ∈{Δθ+δ,Δθ-δ}, where δ is the step size for beamwidth adjustment;
[0145] Reward R t The reward function measures the quality of each action. This function should take into account channel gain, user direction signal quality, and multipath reflection path coverage, and is defined as follows:
[0146] R t =w4·G(θ) cand ,Δθ)+w5·C(θ cand ,Δθ);
[0147] 2. Training steps of reinforcement learning algorithms:
[0148] 2.1, Initialize a Q-table Q(s) t a t ), which records the reward value of state-action pairs.
[0149] 2.2 Execute actions and obtain feedback: At each time step t, based on the current state s tAction a is selected using an ∈-greedy strategy. t ,Right now:
[0150]
[0151] In the early stages of training, the environment is explored with high probability. As training progresses, exploration is gradually reduced, and the weight of utilizing existing strategies is increased.
[0152] 2.3 Execute actions and update the state: Execute action a t Adjust the beamwidth Δθ and update to the next state s. t+1 At the same time, you will receive a reward R. t The reward is calculated using the aforementioned reward function.
[0153] 2.4 Update Q-values or DQN model: Update the Q-table using the following formula: Q(s) t a t )←Q(s t a t )+ɑ[R t +γmax a′ Q(s t+1 ,a′)-Q(s t a t )], where α is the learning rate and γ is the discount factor.
[0154] By repeating the above steps and training multiple times, the model will gradually learn to select the optimal beamwidth adjustment strategy to achieve the best results in terms of channel gain and multipath reflection coverage.
[0155] S5 specifically includes:
[0156] S51: Real-time acquisition of user terminal location information and network load, including the current base station connection count, channel occupancy rate and signal interference intensity;
[0157] S52: Input user location change information and network load data into the deep learning model. The deep learning model identifies user location movement patterns and load distribution characteristics through continuous online training and learning.
[0158] S53: Based on the output of the deep learning model, dynamically adjust the beam direction and beam width to adapt to changes in the user's location and ensure the effectiveness of signal coverage;
[0159] S54: Self-updates the beam pattern library, optimizes the beam direction, beam width and coverage in the beam pattern library based on the actual distribution of users and network load, and generates new beam patterns suitable for the current network status.
[0160] S55: Through error feedback, the parameters of the deep learning model are updated. The convolutional neural network part of the deep learning model updates the weights of the spatial feature extraction layer, and the recurrent neural network part of the deep learning model updates the weights of the time series prediction layer. The updated model will be able to better adapt to changes in user location and dynamic changes in network load.
[0161] User location change data is collected by user location (x) u ,y u The change in position Δd is calculated based on the positional difference. u This refers to the user's displacement between two consecutive moments; network load data is obtained by collecting the real-time connection count N of the base station. conn Channel occupancy rate U channel Interference signal strength I intf Real-time calculation of network load L net ;
[0162] Deep learning models use convolutional neural networks to handle changes in user location and recurrent neural networks to handle changes in network load over time.
[0163] Input data: X input =(Δd) u ,L net ), where Δd u L represents the user's movement displacement. net This indicates the real-time network load status;
[0164] Output data: (Δθ) pred ,ΔW pred )=f DL (X input ), where Δθ pred : Predicted beam direction adjustment, ΔW pred : Predicted beamwidth adjustment amount;
[0165] Based on real-time user location changes and network load data, the deep learning model undergoes continuous online training and learning updates. At each time step t, new user location and network load data are input into the model to generate beam adjustment suggestions. These suggestions are then combined with actual beam performance and evaluated using an error function. Calculate the difference between the model's predicted values and the actual values:
[0166] Where, θ actual and W actual For the actual beam direction and width, θ i and W iFor the beam direction and width predicted by the model, the parameters of the deep learning model are updated through error feedback and backpropagation. The convolutional neural network part updates the weights of the spatial feature extraction layer, and the recurrent neural network part updates the weights of the time series prediction layer.
[0167] After each adjustment of the beam direction and width based on the deep learning model, the beam pattern library also needs to be updated accordingly to maintain its adaptability. The conditions for updating the beam pattern library are as follows:
[0168] When the user's location changes significantly (Δdu) or network load fluctuates drastically, the beam direction and width need to be adjusted. Therefore, the beam patterns in the beam pattern library need to be dynamically updated. The adjustment result Δθ from the deep learning model output is then used. pred ΔW pred Adjustments are made by applying them to the beam pattern library.
[0169] This invention encompasses any substitutions, modifications, equivalent methods, and solutions made within the spirit and scope of this invention. To provide the public with a thorough understanding of this invention, specific details are described in detail in the following preferred embodiments; however, those skilled in the art will fully understand the invention even without these details. Furthermore, to avoid unnecessary misunderstanding of the essence of this invention, well-known methods, processes, procedures, components, and circuits are not described in detail.
[0170] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A beam control algorithm for wireless communication, characterized in that, include: S1: Initialize the beam set by selecting the initial beam direction set from the preset beam pattern library. , used to cover a preset spatial area; S2: Real-time acquisition of historical channel state information from user terminals, and construction of channel state feature maps using a deep learning model. Channel state feature map The system is constructed based on the user's mobile history, environmental reflection characteristics, and multipath effects to predict the user's future channel change trends. S3: Based on the channel state feature map Based on the future channel change trend, the beam direction set is initially adjusted. The adjustment process adopts a multi-objective optimization algorithm, which comprehensively considers channel gain, interference signal strength and system power consumption, and generates multiple candidate beam directions. S4: For each candidate beam direction, the adaptive tuning module based on reinforcement learning adjusts the beam width so that the beam is not only focused on the target user direction, but also covers potential multipath reflection paths. S5: Based on changes in user location and real-time network load, the beam control algorithm is iteratively optimized, and the beam pattern library and deep learning model are continuously updated. The initial beam set in S1 specifically includes: S11: Based on historical data and user distribution characteristics in the wireless communication system, multiple beam pattern libraries are pre-built, and the beam pattern libraries include a variety of preset beam pattern sets with different widths, directions and coverage ranges; S12: Based on the current user's geographical location, target spatial area requirements, and coverage target, match a beam pattern set from the beam pattern library; S13: Based on the shape and size of the target spatial region, select an initial beam direction set that can cover the target spatial region from the set of matched beam patterns.
2. The beam control algorithm for wireless communication according to claim 1, characterized in that, S2 specifically includes: S21: Receive historical channel status information of user terminals in real time through base stations or access points in the wireless communication system. The historical channel status information includes channel gain, signal strength, and channel attenuation information of user terminals at different time points. S22: Constructing the channel state feature matrix The elements in the matrix represent the channel state of the user at different locations and time points; the channel state feature matrix. It is constructed based on the user's movement trajectory and channel characteristics at different time points, with each element representing the channel state at different times and locations; S23: Transfer the channel state feature matrix The input is fed into a pre-trained deep learning model for feature extraction, extracting the spatial and temporal features of the channel state; S24: The extracted channel state features are processed by a deep learning model to generate a channel state feature map, which reflects the influence of user mobility history, environmental reflection characteristics and multipath effects on the channel state. S25: Using the generated channel state feature map, predict the user's future channel change trend for dynamic adjustment and optimization of beam direction.
3. The beam control algorithm for wireless communication according to claim 2, characterized in that, The channel state is also affected by multipath propagation effects, which are modeled as follows: ,in, For the first The attenuation coefficient of the path, For the first The frequency of the path, For time, This represents the total number of multipath paths. Indicates the user terminal at time From the base station To user location Channel state information, including channel gain, signal strength, and attenuation characteristics.
4. The beam control algorithm for wireless communication according to claim 3, characterized in that, The deep learning model combines convolutional neural networks and recurrent neural networks to extract spatial and temporal features. The convolutional layers of the convolutional neural network are used to extract spatial features and output feature maps. Recurrent neural networks are used to capture time dependencies and output time. Hidden state features ; Channel state characteristic map The generation is represented as: ,in, This indicates the deep learning model's response to feature maps. and hidden state features The fusion operation This is the final generated channel state feature map; Using channel state feature maps By learning from historical data, the future channel state is predicted. The prediction model uses a recurrent neural network, and the prediction is expressed as follows: ,in, For the next predicted time point The channel state, including the time variations of channel gain, interference signal strength, and multipath propagation effects. This is the prediction function for the recurrent neural network model.
5. The beam control algorithm for wireless communication according to claim 1, characterized in that, S3 specifically includes: S31: From the channel state feature map Extract future channel change trends, including the time variations of channel gain, interference signal strength, and multipath effects; S32: Adjust the initial beam direction set using a multi-objective optimization algorithm based on future channel change trends. The optimization objectives include: Channel gain maximization: The objective function is: ,in, Indicates beam direction The corresponding channel gain, Number of beam directions; Minimize interference signal strength: The objective function is: ,in, Indicates beam direction Interference signal strength at the location; System power consumption minimization: The optimization objective function is: ,in, Beam direction Corresponding system power consumption; S33: Through a multi-objective optimization algorithm, multiple candidate beam directions are generated by comprehensively considering the weights of channel gain, interference signal strength, and system power consumption. This meets the optimization requirements for future channel state changes.
6. The beam control algorithm for wireless communication according to claim 5, characterized in that, S4 specifically includes: S41: For each candidate beam direction As a state input in a reinforcement learning environment, the state includes current channel state information, user location, and multipath reflection path information; S42: Define the action space of reinforcement learning as the beamwidth. Adjustment range, action This indicates a reduction or expansion of the current beamwidth, with the goal of optimizing channel gain and coverage. S43: Define the reward function The function that comprehensively considers channel gain, user-direction signal quality, and multipath reflection path coverage is expressed as: ,in, Indicates beam direction and beamwidth Channel gain below, This indicates the beam's ability to cover potential multipath reflection paths. and These are the weighting coefficients for channel gain and multipath coverage effect; S44: Based on a reinforcement learning algorithm, the beamwidth is adjusted through multiple training iterations to focus the beam on the target user direction while covering multipath reflection paths to optimize signal quality and coverage. The optimized beamwidth is then output. .
7. The beam control algorithm for wireless communication according to claim 6, characterized in that, The reinforcement learning algorithm includes defining the state in reinforcement learning as the current beam direction, beamwidth, channel state information, and user location information; the action is to adjust the beamwidth; and the reward measures the effect after beam adjustment, including channel gain, user direction signal quality, and coverage of multipath reflection paths. The reinforcement learning algorithm begins training by initializing the Q-table or, at each time step, selecting an action based on the current state and executing that action. After executing the action, it updates the state to the next state and updates the Q-table based on the feedback calculated by the reward function. Through repeated iterations, the reinforcement learning algorithm gradually learns how to select the optimal beamwidth adjustment strategy, and finally outputs the optimized beamwidth.
8. The beam control algorithm for wireless communication according to claim 1, characterized in that, S5 specifically includes: S51: Real-time acquisition of user terminal location information and network load, wherein the network load includes the current base station connection count, channel occupancy rate and signal interference intensity; S52: Input user location change information and network load data into the deep learning model. The deep learning model identifies user location movement patterns and load distribution characteristics through continuous online training and learning. S53: Based on the output of the deep learning model, dynamically adjust the beam direction and beam width to adapt to changes in the user's location and ensure the effectiveness of signal coverage; S54: Self-updates the beam pattern library, optimizes the beam direction, beam width and coverage in the beam pattern library based on the actual distribution of users and network load, and generates new beam patterns suitable for the current network status. S55: Through error feedback, update the parameters of the deep learning model. The convolutional neural network part of the deep learning model updates the weights of the spatial feature extraction layer, and the recurrent neural network part of the deep learning model updates the weights of the time series prediction layer.
9. The beam control algorithm for wireless communication according to claim 8, characterized in that, The user location change information is obtained by collecting user location data. Changes, calculate positional differences This refers to the user's displacement between two consecutive moments; the network load data is obtained by collecting the real-time connection count of the base station. Channel occupancy rate Interference signal strength Real-time calculation of network load ; The deep learning model uses a convolutional neural network to handle changes in user location and a recurrent neural network to handle changes in network load over time. Input data: ,in Indicates the user's movement displacement. This indicates the real-time network load status; Output data: ,in, Predicted beam direction adjustment amount : Predicted beamwidth adjustment amount; Based on real-time data on user location changes and network load, the deep learning model undergoes continuous online training and learning updates at each time step. The new user location and network load data are input into the model to generate beam adjustment suggestions. These suggestions are then combined with the actual beam effect and analyzed using an error function. Calculate the difference between the model's predicted values and the actual values: ,in, and For the actual beam direction and width, and For the beam direction and width predicted by the model, the parameters of the deep learning model are updated through error feedback and backpropagation. The convolutional neural network part updates the weights of the spatial feature extraction layer, and the recurrent neural network part updates the weights of the time series prediction layer.
Citation Information
Patent Citations
Power line carrier communication method and system
CN118611707A