Wireless signal coverage enhancement system and method based on reinforcement learning
By integrating multi-source data and employing intelligent closed-loop control, the system can perceive obstacle movement in real time and dynamically optimize the beam and power strategies for millimeter-wave communication. This solves the signal coverage problem in NLOS scenarios and improves positioning accuracy and communication reliability.
Patent Information
- Application Number
- CN202511087085.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-05
- Publication Date
- 2025-12-30
AI Technical Summary
In millimeter-wave communication, the positioning error exceeds 15 meters in NLOS scenarios and requires 200ms to detect obstruction, resulting in a 30ms communication interruption, which cannot meet the wireless signal coverage requirements of high-precision applications.
By fusing multi-source sensing data from lidar and cameras, the system extracts real-time information on terminal motion status and environmental obstacles, calculates relative motion parameters and collision time, constructs a dynamic diffraction topology map, uses graph neural networks to predict path loss, dynamically optimizes beam pointing and power adjustment, and combines reinforcement learning for real-time compensation.
Real-time perception of high-speed dynamic obstacles, optimization of beamforming and power strategies, shortening of interruption recovery time, improvement of positioning accuracy and communication reliability, and provision of stable and continuous millimeter-wave coverage.
Smart Images

Figure CN121240106A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of communication technology, and more specifically, to a wireless signal coverage enhancement system and method based on reinforcement learning. Background Technology
[0002] With the rapid development of 5G / 6G communication technology, the millimeter wave (mm Wave) band (24-100GHz) has become a key carrier for realizing gigabit-level transmission due to its ultra-large bandwidth advantage.
[0003] However, millimeter-wave signals suffer from severe non-line-of-sight (NLOS) attenuation during propagation: penetration loss is as high as 20-40 dB (concrete walls), diffraction capability is weak (wavelength 1-10 mm level), and multipath effect is significant;
[0004] In high-speed dynamic obstacle scenarios, traditional coverage enhancement methods based on Received Signal Strength Indicator (RSSI) have a positioning error of more than 15 meters in NLOS scenarios, and require 200ms to detect obstruction. The compensation lag leads to a 30ms communication interruption, which cannot meet the wireless signal coverage requirements of high-precision applications. Summary of the Invention
[0005] This invention provides a wireless signal coverage enhancement system and method based on reinforcement learning, which solves the technical problems in related technologies where the positioning error exceeds 15 meters in NLOS scenarios, and the need for 200ms to detect occlusion and the compensation lag leads to a 30ms communication interruption.
[0006] This invention provides a method for enhancing wireless signal coverage based on reinforcement learning, comprising the following steps:
[0007] S100, Joint Perception and Threat Recognition: By fusing multi-source perception data such as lidar and cameras, it extracts terminal motion status and environmental obstacle information in real time, calculates relative motion parameters and collision time, comprehensively assesses threat level based on obstacle size and distance, and outputs quantitative threat indicators and potential high-risk targets.
[0008] S200, Dynamic Reconstruction of Diffraction Path: Identify effective reflecting surfaces from point cloud data, construct a time-varying diffraction topology map by combining dynamic obstacle positions, use graph neural networks to predict signal loss of each path, screen low-loss candidate path sets, and provide physical propagation model support for beam optimization;
[0009] S300, Real-time Beam Optimization: Selects the optimal diffraction path based on path loss, terminal movement speed and threat level, calculates beam pointing angle and compensates for deviations caused by high-speed terminal movement, adaptively adjusts beam width to balance coverage and anti-interference capability, and generates three-dimensional beam control commands.
[0010] S400, Power Enhancement Decision: Combines path loss and beamforming gain to predict communication quality, uses reinforcement learning to dynamically decide on power adjustment, ensures power safety through regulatory constraints, thermal protection and interference coordination mechanisms, and outputs transmit power commands that meet real-time requirements.
[0011] S500, closed-loop feedback: verifies the signal strength and quality improvement effect after compensation, diagnoses the root cause of anomalies, incrementally updates the path prediction model parameters and stores reinforcement learning experience data, and realizes system self-optimization.
[0012] Furthermore, the specific steps for joint perception and threat identification are as follows:
[0013] S110, Multi-source data spatiotemporal alignment: Input LiDAR raw point cloud data, camera raw target detection results and terminal motion state vector;
[0014] Downsampling and noise filtering are performed on the raw point cloud data of the lidar, non-maximum suppression is performed on the raw target detection results of the camera, point cloud and image features are extracted for registration, the transformation matrix from lidar to base station is calculated, the transformation matrix from camera to base station is calculated, coordinate transformation is performed using the transformation matrix, data alignment is performed based on timestamp, linear interpolation is used to compensate for transmission delay, and a synchronous data buffer is constructed.
[0015] S120, Motion State Extraction: Based on historical positions, calculate velocity and acceleration, apply Kalman filtering to smooth state estimation, and perform outlier detection and processing; extract obstacle center points and boundaries, establish target association and tracking, and estimate obstacle motion parameters; use motion models to predict the state at the next moment, calculate prediction uncertainty, and update state estimation confidence;
[0016] S130, Relative Motion Analysis: Calculate the relative position of the terminal with respect to each obstacle, calculate the relative velocity and acceleration, and analyze the relative motion trend; calculate the approach / remote velocity, analyze the change in motion direction, and assess the relative attitude relationship; predict the nearest approach point, estimate the collision probability, and calculate the avoidance time margin; calculate the relative position of the terminal with respect to each obstacle, calculate the relative velocity and acceleration, and analyze the relative motion trend.
[0017] S140, Collision Threat Assessment: Calculate collision time, assess collision probability, and analyze the severity of collision consequences; integrate multi-factor scoring, apply threat level mapping, and identify high-threat targets.
[0018] Furthermore, the specific steps for dynamic reconstruction of the diffraction path are as follows:
[0019] S210, Environmental Feature Extraction: Segment and cluster point cloud data, identify continuous planar regions, calculate the normal vector and boundary contour of each plane, extract surface texture features and reflection characteristics, apply a deep learning model for material classification, calculate radar cross-section; filter low reflectivity surfaces based on radar cross-section threshold, remove metallic surfaces, merge adjacent effective reflection areas, calculate the area and orientation of each reflection surface, and establish a reflection surface index structure;
[0020] S220, Dynamic Diffraction Graph Construction: Extract the corner points of the reflecting surface as static vertices, take the positions of dynamic obstacles as dynamic vertices, predict the future positions of obstacles, merge duplicate or too close vertices, assign a unique identifier to each vertex, check the visibility between vertex pairs, calculate the Euclidean distance between vertices, apply the maximum distance constraint, construct the adjacency matrix, and optimize the connectivity of the graph.
[0021] S230, Path Loss Prediction: Initialize node feature vectors, perform multi-layer graph convolution operations, aggregate neighbor node information, update node hidden states, and apply non-linear activation functions; extract path start and end point features, calculate path physical length, apply loss prediction models, consider material influence factors, and output loss prediction values.
[0022] Furthermore, environmental feature extraction includes the following:
[0023] The formula for calculating radar cross section is as follows:
[0024] ;
[0025] in, Let J be the radar cross section of the j-th surface. For millimeter wave wavelength, For the first The attenuation coefficient of a path indicates the degree of signal strength attenuation. For carrier frequency, For the first The delay of the path, For belonging to the surface A subset of point clouds The angular constant of a spherical solid. For modulo squaring operations;
[0026] The formula for calculating material classification is as follows:
[0027] ;
[0028] in, A convolutional neural network for material classification is used to identify surface material types. For the first Point cloud data of a surface, Output material types, including metal, glass, and concrete;
[0029] The formula for calculating the effective reflective surface selection is as follows:
[0030] ;
[0031] in, This is the set of effective reflective surfaces after filtering. For the j-th surface, The scattering cross-sectional area threshold, "This indicates the material types that need to be excluded."
[0032] Furthermore, the construction of dynamic diffraction patterns includes the following:
[0033] The formula for generating vertex sets is as follows:
[0034] ;
[0035] ;
[0036] ;
[0037] in, Let be the set of graph vertices at time t. A static vertex represents a corner point of the reflecting surface. For dynamic vertices, For the first The three-dimensional vector of the position of each obstacle For the first The three-dimensional velocity vector of an obstacle. For the current moment, This is a function for extracting corner points;
[0038] The formula for calculating edge set construction is as follows:
[0039] ;
[0040] in, Let be the set of graph edges at time t. Let be the three-dimensional coordinates of any two vertices. Maximum diffraction distance (default) ), For unobstructed line-of-sight detection functions, Calculate the Euclidean distance.
[0041] Furthermore, path loss prediction includes the following:
[0042] The calculation formula for the GNN propagation model is as follows:
[0043] ;
[0044] in, For the first Layer nodes The hidden state vector has dimension d. For the first The weight matrix of the layer GNN has a dimension of d×2d. As vertex The neighborhood group, This is a vector concatenation operation. To modify the activation function of the linear unit, For GNN layer index, The dimension of the hidden state vector;
[0045] The formula for calculating end-to-end path loss is as follows:
[0046] ;
[0047] in, The predicted diffraction loss for path p. Here, θ is a parameterized path loss prediction function, and θ is a learnable parameter. Let be the final hidden state vector of the starting node. This represents the final hidden state vector of the endpoint node. For path The physical length, This is the path loss coefficient. This represents the total number of layers in the GNN.
[0048] Furthermore, the specific steps for real-time beam optimization are as follows:
[0049] S310, Optimal Path Selection: Calculate the main propagation direction unit vector for each path, evaluate the impact of terminal motion on path stability, generate path stability index, normalize the stability index, and mark paths with stability below a threshold; extract path loss features, calculate the threat level around the path, fuse multi-dimensional evaluation indicators, apply weight coefficients, and generate a comprehensive path score; sort all candidate paths, select the path with the highest score, verify path feasibility, record the selection results, and update path status information;
[0050] S320, Beam pointing calculation: Extract the coordinates of key points on the path, calculate the distance from the point to the terminal, calculate the distance from the point to the base station, find the optimal reflection point, and record the position of the reflection point; calculate the relative position vector, decompose the horizontal / vertical components, apply inverse trigonometric functions, convert the angle unit, and ensure the angle range; decompose the terminal speed, calculate the azimuth compensation and pitch compensation, apply the compensation coefficient, and limit the compensation range;
[0051] S330, Beamwidth Adaptive: Calculate terminal velocity magnitude, apply velocity normalization, calculate base width, limit width range, and record baseline values; assess threat level, check velocity conditions, determine adjustment direction, calculate adjustment amount, and apply adjustment rules; merge base width and adjustment amount, apply upper and lower limit constraints, ensure width rationality, generate final width, and update control parameters.
[0052] Furthermore, the specific steps for power enhancement decision-making are as follows:
[0053] S410, Communication Status Assessment: Estimate atmospheric absorption and scattering loss based on carrier frequency and propagation distance, consider the impact of carrier frequency on propagation attenuation, calculate path loss based on free space propagation model, and superimpose diffraction loss, atmospheric loss and other losses to obtain equivalent total loss; read the current transmit power setting of the base station, calculate the interference level based on the location and power of neighboring base stations, calculate the antenna directivity gain based on beamwidth, and comprehensively consider transmit power, path loss, interference and gain.
[0054] S420, Reinforcement Learning Decision: Normalize the predicted SINR by dividing it by the target threshold, divide the terminal speed by the maximum speed limit, directly use the normalized threat level index, divide the current power by the maximum power limit, and combine the normalized components into a state vector; input the normalized state vector into the policy network, calculate the power adjustment suggestion through the policy network, adjust the exploration level based on the threat level to ensure that the power adjustment amount is within a reasonable range, and generate the actual power adjustment command; evaluate the degree of achievement of the current SINR relative to the target, apply a penalty according to the size of the power adjustment amount, determine whether cell handover should occur, and combine the scores of each item according to the weight to generate the final reward signal;
[0055] S430, Power Safety Constraints: Add the current power and the adjustment amount to ensure that it does not exceed the maximum power limit and does not fall below the minimum power requirement. Limit the power exceeding the range to the legal range and save the power value that meets the regulatory requirements. Read the current temperature of the RF module, check whether it exceeds the temperature threshold, determine whether power reduction is needed based on the temperature status, apply the power setting after thermal protection, and record the thermal protection triggering situation. Detect the signals of surrounding neighboring base stations, calculate the impact of neighboring signals on this cell, set the power backoff amount according to the interference level, ensure that it does not cause excessive interference to neighboring cells, and output the power value that meets the coordination requirements.
[0056] Furthermore, the specific steps of the closed-loop feedback are as follows:
[0057] S510, Compensation Effect Evaluation: Sample and analyze the channel response after compensation, calculate the channel response power spectral density, apply the path loss calculation formula, consider the influence of transmit power, and output the actual loss result; extract the channel state before and after compensation, calculate the channel response norm, evaluate the improvement of compensation effect, normalize the result, and generate performance index value; measure the signal-to-interference-plus-noise ratio before compensation, measure the signal-to-interference-plus-noise ratio after compensation, calculate the SINR improvement, evaluate the improvement effect, and record the verification results;
[0058] S520, Root Cause Analysis: Check compensation effectiveness indicators, assess loss prediction error, apply anomaly judgment rules, mark abnormal states, and record detection results; analyze changes in environmental threats, assess model prediction error, monitor hardware performance indicators, determine fault type, and generate diagnostic report.
[0059] S530, online model update: collect terminal location information, record motion state data, extract environmental features, construct training samples, and update the dataset; calculate prediction error gradient, update network weights, apply regularization constraints, verify model performance, and save updated parameters; record state transition data, calculate instant reward values, store interaction experience, update the experience pool, and execute policy optimization.
[0060] The present invention also proposes a reinforcement learning-based wireless signal coverage enhancement system for performing the steps of the aforementioned reinforcement learning-based wireless signal coverage enhancement method, including:
[0061] Joint perception and threat identification module: It integrates LiDAR point cloud, camera visual data and terminal motion status, and through multi-source spatiotemporal alignment and relative motion analysis, it identifies the collision threat level of dynamic obstacles in real time and outputs quantitative threat indicators and a list of high-risk targets.
[0062] Dynamic Diffraction Path Reconstruction Module: Based on environmental reflective surface features and real-time obstacle positions, a dynamic diffraction topology map is constructed; a graph neural network is used to predict the diffraction path loss of the signal in a complex environment, and a set of low-loss candidate paths is selected.
[0063] Real-time beam optimization module: Combines path loss, terminal speed and threat level to select the optimal diffraction path and calculate the beam pointing angle; dynamically compensates for terminal motion offset, adaptively adjusts beamwidth to balance coverage and anti-interference capability, and outputs three-dimensional beam control commands;
[0064] Power enhancement decision module: Analyzes the impact of path loss and beam gain on communication quality, and dynamically adjusts the transmit power using reinforcement learning; outputs a safe power command through regulatory constraints, thermal protection, and interference coordination mechanisms.
[0065] Closed-loop feedback and learning module: verifies the quality of the compensated signal, diagnoses the root cause of anomalies, and triggers incremental model updates.
[0066] The beneficial effects of this invention are as follows:
[0067] This invention solves the signal coverage problem in millimeter-wave NLOS scenarios through multi-source data fusion and intelligent closed-loop control mechanisms. It can perceive the trajectory of high-speed dynamic obstacles in real time, predict the optimal diffraction path by combining environmental electromagnetic characteristics, and dynamically optimize beamforming and power strategies to suppress the impact of multipath interference and penetration loss on signal quality. In extreme obstruction scenarios, it can shorten the interruption recovery time, improve positioning accuracy and communication reliability, and provide stable and continuous millimeter-wave coverage for high-bandwidth and low-latency applications such as vehicle networking and industrial IoT. Attached Figure Description
[0068] Figure 1 This is a flowchart of a wireless signal coverage enhancement method based on reinforcement learning according to the present invention;
[0069] Figure 2 This is a structural block diagram of a wireless signal coverage enhancement system based on reinforcement learning according to the present invention.
[0070] In the diagram: 101, Joint Perception and Threat Identification Module; 102, Dynamic Reconstruction Module for Diffraction Path; 103, Real-time Beam Optimization Module; 104, Power Enhancement Decision Module; 105, Closed-Loop Feedback and Learning Module. Detailed Implementation
[0071] The subject matter described herein will now be discussed with reference to exemplary embodiments. It should be understood that these embodiments are discussed only to enable those skilled in the art to better understand and implement the subject matter described herein, and changes may be made to the function and arrangement of the elements discussed without departing from the scope of this specification. Various processes or components may be omitted, substituted, or added as needed in the examples. Furthermore, features described in some examples may be combined in other examples.
[0072] like Figure 1 As shown, a wireless signal coverage enhancement method based on reinforcement learning includes the following steps:
[0073] S100, Joint Perception and Threat Recognition: By fusing multi-source perception data such as lidar and cameras, it extracts the terminal's motion status (position, speed) and environmental obstacle information in real time, calculates relative motion parameters (relative displacement, approach speed) and collision time (TTC), comprehensively assesses the threat level based on obstacle size and distance, and outputs quantitative threat indicators and potential high-risk targets;
[0074] In one embodiment of the present invention, the following steps are specifically included:
[0075] S110, Spatiotemporal alignment of multi-source data:
[0076] The input includes raw point cloud data from the LiDAR, containing the three-dimensional coordinates and reflection intensity of the points; raw target detection results from the camera, containing the target bounding box and category information; and the terminal motion state vector, containing the velocity magnitude and direction.
[0077] Data preprocessing: Downsampling and noise filtering are performed on the raw point cloud data of the LiDAR, non-maximum suppression (NMS) is performed on the raw target detection results of the camera, and point cloud and image features are extracted for registration;
[0078] Coordinate system transformation: Calculate the transformation matrix from lidar to base station, calculate the transformation matrix from camera to base station, and apply the transformation matrix to perform coordinate transformation;
[0079] Coordinate transformation formula:
[0080] ;
[0081] ;
[0082] ;
[0083] in, This is a 4×4 homogeneous transformation matrix from the sensor to the base station, containing rotation and translation information. The original sensor coordinates are represented in homogeneous coordinates. These are the point coordinates in the base station coordinate system, and the transformed homogeneous coordinates.
[0084] Time synchronization: Data alignment is performed based on timestamps, linear interpolation is used to compensate for transmission delays, and a synchronization data buffer is constructed;
[0085] Output the aligned point cloud dataset, unified to the base station coordinate system, and output the aligned obstacle list, including location and size information;
[0086] S120, Motion State Extraction:
[0087] Terminal state estimation: Based on historical location, velocity and acceleration are calculated, and Kalman filtering is applied to smooth the state estimation, and outlier detection and processing are performed;
[0088] Formulas for calculating terminal motion parameters:
[0089] speed:
[0090] ;
[0091] Acceleration:
[0092] ;
[0093] in, The sampling time interval is the time difference between two adjacent data samples. This is a three-dimensional vector representing the terminal's location in the base station coordinate system. For the first The three-dimensional vector of the terminal position at time t. For the first The three-dimensional vector of the terminal position at time t. Let be a three-dimensional vector of terminal velocity, representing the instantaneous velocity of the terminal. For the first The three-dimensional vector of the terminal velocity at time t. For the first The three-dimensional vector of the terminal velocity at time t. This is a three-dimensional vector of terminal acceleration, representing the instantaneous acceleration of the terminal;
[0094] Obstacle state tracking: Extract the center point and boundary of obstacles, establish target association and tracking, and estimate obstacle motion parameters;
[0095] Formulas for calculating obstacle motion parameters:
[0096] Location:
[0097] ;
[0098] speed:
[0099] ;
[0100] in, The center position of the i-th obstacle is calculated by averaging the vertices. The velocity vector of the i-th obstacle is estimated using Kalman filtering. For the set of vertices of the obstacle boundary, Historical location sequences are used for velocity estimation. This is a Kalman filter used for smoothing velocity estimation. This is the list of aligned obstacles for the i-th time, containing their position and outline information. This represents the center position of the i-th obstacle in the historical location. A function for calculating the average value, used to determine the center location of the obstacle;
[0101] State prediction: Use the motion model to predict the state at the next moment, calculate the prediction uncertainty, and update the confidence of the state estimate;
[0102] Output:
[0103] ;
[0104] ;
[0105] in, This represents the complete state of the terminal, including its position, velocity, and acceleration. This is a list of obstacle states, including the position and velocity of each obstacle.
[0106] S130, Relative Motion Analysis:
[0107] Calculate the relative position of the terminal with respect to each obstacle, calculate the relative velocity and acceleration, and analyze the relative motion trend; calculate the approach / remote velocity, analyze the change in motion direction, and assess the relative attitude relationship; predict the nearest approach point (CPA), estimate the collision probability, and calculate the avoidance time margin;
[0108] Relative state calculation: Calculate the relative position of the terminal with respect to each obstacle, calculate the relative velocity and acceleration, and analyze the relative motion trend;
[0109] Formulas for calculating relative motion parameters:
[0110] Relative displacement:
[0111] ;
[0112] Relative speed:
[0113] ;
[0114] Approach speed:
[0115] ;
[0116] in, Let be the direction vector from the terminal to the i-th obstacle, representing the relative positional relationship. Let be the velocity vector of the terminal relative to the i-th obstacle, representing the relative motion velocity. This is a scalar approach velocity; a positive value indicates approach, and a negative value indicates moving away. For vectors The Euclidean norm of the distance represents the magnitude of the distance. This is the vector dot product operator;
[0117] Output:
[0118] ;
[0119] in, This is a list of relative motion parameters, including relative position, relative velocity, and approach velocity;
[0120] S140, Collision Threat Assessment:
[0121] Collision risk assessment: Calculate the time to collision (TTC), assess the probability of a collision, and analyze the severity of the consequences of a collision;
[0122] Collision time calculation formula:
[0123] ;
[0124] Threat level calculation: Based on a comprehensive multi-factor scoring system, a threat level mapping is applied to identify high-threat targets;
[0125] Threat level quantification formula:
[0126] ;
[0127] ;
[0128] in, The Sigmoid activation function maps the input to the (0,1) interval. Let i be the collision time with the i-th obstacle. The projected area of the obstacle. The collision time weight reflects the importance of time urgency. Distance weights reflect the importance of spatial proximity. Area weighting reflects the importance of obstacle size. The Euclidean distance to the obstacle. The width of the obstacle. The height of the obstacle;
[0129] Threat determination criteria: When and At that time, it was determined to be a high threat. It generates threat level reports, marks high-risk areas, and provides avoidance recommendations;
[0130] Threat levels for each obstacle (0~1), where 0 represents no threat and 1 represents the highest threat. The highest threat level, representing the maximum threat value among all obstacles;
[0131] S200, Dynamic Reconstruction of Diffraction Path: Identify effective reflective surfaces (non-metallic materials, high reflectivity) from point cloud data, construct a time-varying diffraction topology map by combining the position of dynamic obstacles, use graph neural network (GNN) to predict the signal loss of each path, screen a set of low-loss candidate paths, and provide physical propagation model support for beam optimization;
[0132] In one embodiment of the present invention, the following steps are specifically included:
[0133] S210, Environmental Feature Extraction:
[0134] Feature extraction of reflective surface:
[0135] The point cloud data is segmented and clustered to identify continuous planar regions. The normal vector and boundary contour of each plane are calculated, surface texture features and reflection properties are extracted, a deep learning model is applied for material classification, and the radar cross section (RCS) is calculated.
[0136] ;
[0137] in, Let sm be the radar cross-section of the j-th surface. For millimeter wave wavelengths (28 GHz), ), For the first The attenuation coefficient of a path indicates the degree of signal strength attenuation. For carrier frequency, For the first The delay of the path, For belonging to the surface A subset of point clouds The angular constant of a spherical solid. For modulo squaring operations;
[0138] Material classification:
[0139] ;
[0140] in, A convolutional neural network for material classification is used to identify surface material types. For the first Point cloud data of a surface, Output material types, including metal, glass, concrete, etc.
[0141] Effective reflective surface screening: Filter low reflectivity surfaces based on RCS threshold, eliminate metallic surfaces, merge adjacent effective reflective areas, calculate the area and orientation of each reflective surface, and establish a reflective surface index structure;
[0142] ;
[0143] in, This is the set of effective reflective surfaces after filtering. For the j-th surface, The scattering cross-sectional area threshold, "These are the material types that need to be excluded;
[0144] S220, Dynamic Diffraction Pattern Construction:
[0145] Vertex set generation: Extract the corner points of the reflective surface as static vertices, take the positions of dynamic obstacles as dynamic vertices, predict the future positions of obstacles, merge duplicate or too close vertices, and assign a unique identifier to each vertex.
[0146] ;
[0147] ;
[0148] ;
[0149] in, Let be the set of graph vertices at time t. A static vertex represents a corner point of the reflecting surface. For dynamic vertices, For the first The three-dimensional vector of the position of each obstacle For the first The three-dimensional velocity vector of an obstacle. For the current moment, This is a function for extracting corner points;
[0150] Edge set construction: Check the visibility between vertex pairs, calculate the Euclidean distance between vertices, apply the maximum distance constraint, construct the adjacency matrix, and optimize the connectivity of the graph;
[0151] ;
[0152] in, Let be the set of graph edges at time t. Let be the three-dimensional coordinates of any two vertices. Maximum diffraction distance (default) ), For unobstructed line-of-sight detection functions, Calculate the Euclidean distance;
[0153] S230, Path Loss Prediction:
[0154] GNN propagation model: Initialize node feature vectors, perform multi-layer graph convolution operations, aggregate neighbor node information, update node hidden states, and apply non-linear activation functions;
[0155] ;
[0156] in, For the first Layer nodes The hidden state vector has dimension d. For the first The weight matrix of the layer GNN has a dimension of d×2d. As vertex The neighborhood group, This is a vector concatenation operation. To modify the activation function of the linear unit, For GNN layer index, The dimension of the hidden state vector;
[0157] End-to-end path loss: Extract the features of the start and end points of the path, calculate the physical length of the path, apply the loss prediction model, consider the influence of material, and output the predicted loss value.
[0158] ;
[0159] in, The predicted diffraction loss for path p. Here, θ is a parameterized path loss prediction function, and θ is a learnable parameter. Let be the final hidden state vector of the starting node. This represents the final hidden state vector of the endpoint node. For path The physical length, This is the path loss coefficient. This represents the total number of layers in the GNN.
[0160] S300, Real-time Beam Optimization: Selects the optimal diffraction path based on path loss, terminal movement speed and threat level, calculates beam pointing angle and compensates for deviations caused by high-speed terminal movement, adaptively adjusts beam width to balance coverage and anti-interference capability, and generates three-dimensional beam control commands.
[0161] In one embodiment of the present invention, the following steps are specifically included:
[0162] S310, Optimal Path Selection:
[0163] Path stability assessment: Calculate the main propagation direction unit vector for each path, evaluate the impact of terminal motion on path stability, generate path stability index, normalize the stability index, and mark paths with stability below the threshold.
[0164] ;
[0165] in, This is the stability index for the k-th path, with a value ranging from [0,1]. A larger value indicates a more stable path. For path The principal direction unit vector represents the main propagation direction of the path. The terminal velocity vector, This is the vector norm operator, which calculates the magnitude of a vector.
[0166] Comprehensive scoring: Extract path loss features, calculate the threat level around the path, integrate multi-dimensional evaluation indicators, apply weight coefficients, and generate a comprehensive path score;
[0167] ;
[0168] in, The path loss weight reflects the degree to which loss affects the score. The stability weight reflects the importance of path stability. The threat avoidance weight reflects the importance of avoiding high-threat areas. For path The highest threat level of obstacles within a 3m radius, with a value ranging from [0,1]. The diffraction loss for the k-th path is... This is the overall score for the k-th path; a larger value indicates a better path.
[0169] Optimal path selection: Sort all candidate paths, select the path with the highest score, verify the feasibility of the path, record the selection results, and update the path status information;
[0170] ;
[0171] in, To select the optimal diffraction path, For the candidate path set, This is the overall score for the k-th path;
[0172] S320, Beam pointing calculation:
[0173] Main reflection point localization: Extract the coordinates of key points along the path, calculate the distance from the point to the terminal, calculate the distance from the point to the base station, find the optimal reflection point, and record the position of the reflection point;
[0174] ;
[0175] in, The coordinates of the main reflection point, and its three-dimensional position vector in the base station coordinate system. Here are the coordinates of the base station location, with the origin of the base station coordinate system at (0,0,0). The symbol for Euclidean distance calculation is... Let be the coordinates of any point on the path;
[0176] Beam azimuth / downtilt: Calculate the relative position vector, decompose the horizontal / vertical components, apply inverse trigonometric functions, convert the angle unit, and ensure the angle range;
[0177] ;
[0178] ;
[0179] in, This refers to the beam azimuth angle, the pointing angle in the horizontal plane. This is the beam downtilt angle, the pointing angle in the vertical plane. The coordinate components of the main reflection point, in the three-dimensional coordinates of the base station coordinate system. These are the base station coordinate components, typically originating at (0,0,0). For the arctangent function, consider the quadrant-based arctangent calculation;
[0180] Motion compensation: Decompose terminal velocity, calculate azimuth and pitch compensation, apply compensation coefficients, and limit the compensation range;
[0181] ;
[0182] in, This is the speed compensation coefficient, a proportional factor that converts speed into angular offset. This is the unit vector for azimuth direction, and the pointing unit vector within the horizontal plane. The pitch direction is a unit vector, and the pointing unit vector in the vertical plane is another unit vector. This is the azimuth compensation amount, the angle compensation value in the horizontal plane. This is the downtilt angle compensation amount, the angle compensation value in the vertical plane. This is the vector dot product operator;
[0183] S330, adaptive beamwidth:
[0184] Basic beamwidth: Calculate the terminal velocity, apply velocity normalization, calculate the basic beamwidth, limit the beamwidth range, and record the baseline value;
[0185] ;
[0186] in, The maximum beamwidth is the maximum achievable beam angle. This is a reference speed, a baseline value used to normalize the terminal speed. For terminal speed, The basic beamwidth;
[0187] Threat adaptive adjustment: Assess threat level, check velocity conditions, determine adjustment direction, calculate adjustment amount, and apply adjustment rules;
[0188] ;
[0189] in, This is the beamwidth adjustment amount. The highest threat level, with a value range of [0,1]. The speed of the terminal;
[0190] Final beamwidth: Combine the base beamwidth and adjustment amount, apply upper and lower limit constraints to ensure the beamwidth is reasonable, generate the final beamwidth, and update the control parameters;
[0191] ;
[0192] in, For the final beamwidth, Due to minimum beamwidth limitations, Due to maximum beamwidth limitations, Based on the basic beamwidth, This is the beamwidth adjustment amount.
[0193] Output:
[0194] ;
[0195] The final beamwidth is the actual beam angle after considering various factors. The complete beam control vector includes the compensated azimuth, downtilt angle, and beamwidth.
[0196] S400, Power Enhancement Decision: Combines path loss and beamforming gain to predict communication quality, uses reinforcement learning (DDPG) to dynamically decide the power adjustment amount, ensures power safety through regulatory constraints, thermal protection and interference coordination mechanisms, and outputs transmit power commands that meet real-time requirements;
[0197] In one embodiment of the present invention, the following steps are specifically included:
[0198] S410, Communication Status Assessment:
[0199] Equivalent path loss: Atmospheric absorption and scattering loss are estimated based on carrier frequency and propagation distance. The influence of carrier frequency on propagation attenuation is considered. Path loss is calculated based on free space propagation model. Diffraction loss, atmospheric loss and other losses are superimposed to obtain equivalent total loss.
[0200] ;
[0201] in, Atmospheric loss (default 0.1dB / km at 28GHz), For carrier frequency, This refers to the distance from the terminal to the base station. This is the equivalent total path loss. This is a frequency-dependent loss term. This refers to distance-related loss terms.
[0202] Expected SINR (Signal-to-Interference-plus-Noise Ratio) calculation: Read the current transmit power setting of the base station, calculate the interference level based on the location and power of neighboring base stations, calculate the antenna directivity gain according to the beamwidth, and comprehensively consider transmit power, path loss, interference and gain;
[0203] ;
[0204] ;
[0205] in, This is the current transmission power. Interference from neighboring cells To enhance beamforming gain, Beamwidth, in radians. The predicted signal-to-interference-plus-noise ratio;
[0206] Output:
[0207] ;
[0208] in This is a communication quality state vector, with each component in dB.
[0209] S420, Reinforcement Learning Decision Making:
[0210] DDPG state space construction: Normalize the predicted SINR by dividing it by the target threshold, divide the terminal speed by the maximum speed limit, directly use the normalized threat level index, divide the current power by the maximum power limit, and combine the normalized components into a state vector.
[0211] ;
[0212] in, The target signal-to-interference-plus-noise ratio (SINR) threshold. For maximum speed limit, Due to maximum transmit power limitations, For terminal speed, The normalized state vector;
[0213] Action generation: Input the normalized state vector into the policy network, calculate the power adjustment suggestion through the policy network, adjust the exploration level based on the threat level, ensure that the power adjustment amount is within a reasonable range, and generate the actual power adjustment command;
[0214] ;
[0215] in, The output function of the policy network takes the input state vector as input and outputs a power adjustment suggestion. These are the policy network parameters, including network weights and biases. To explore the noise, it follows a normal distribution with a mean of 0. The noise standard deviation is positively correlated with the threat level. This is the power adjustment amount;
[0216] Reward function design: Evaluate the degree of achievement of the current SINR relative to the target, apply penalties according to the power adjustment amount, determine whether cell handover has occurred, combine the scores of each item according to weights, and generate the final reward signal;
[0217] ;
[0218] in, This is a power adjustment penalty factor used to balance communication quality and power stability. This is a switching indicator function; it takes the value 1 when a switch occurs, and 0 otherwise. This is an instant reward value, dimensionless. It is the square of the power adjustment amount, used to penalize excessive power changes;
[0219] Output: This is the power adjustment amount;
[0220] S430, power safety constraints:
[0221] Regulatory constraints: Add the current power and the adjustment amount to ensure that it does not exceed the maximum power limit, ensure that it is not lower than the minimum power requirement, limit the power that exceeds the range to the legal range, and save the power value that meets the regulatory requirements;
[0222] ;
[0223] in, For minimum power limitation, Due to maximum power limitation, This is a temporary power value and must meet the following requirements. ;
[0224] Thermal limiting protection: Read the current temperature of the RF module, check whether it exceeds the temperature threshold, determine whether power reduction is required based on the temperature status, apply the power setting after thermal protection, and record the thermal protection triggering status;
[0225] ;
[0226] in, For the RF module temperature, The power reduction for thermal protection is a fixed value. This is the power value after thermal protection. Temperature threshold;
[0227] Interference Coordination: Detect signals from neighboring base stations, calculate the impact of neighboring signals on the current cell, set the power backoff amount according to the interference level to ensure that there is no excessive interference to neighboring cells, and output a power value that meets the coordination requirements;
[0228] If neighboring regions exist ,but:
[0229] ;
[0230] in, For the neighboring cell reference signal received power, For neighboring cell transmission power, For interference protection intervals, a fixed value is used. The threshold for neighboring cell signal strength;
[0231] Output: This is the final power control command;
[0232] S500, closed-loop feedback: verifies the signal strength and quality improvement effect after compensation, diagnoses the root cause of abnormalities (environmental changes / model mismatch / hardware failure), incrementally updates the path prediction model parameters and stores reinforcement learning experience data, and realizes system self-optimization;
[0233] In one embodiment of the present invention, the following steps are specifically included:
[0234] S510, Compensation Effect Evaluation:
[0235] Actual path loss calculation: Sampling analysis is performed on the compensated channel response to calculate the channel response power spectral density. The path loss calculation formula is applied, taking into account the influence of transmit power, and the actual loss result is output.
[0236] ;
[0237] in, The path loss (dB) is the actual measured value, representing the amount of attenuation during actual signal propagation. The transmit power command (dBm) indicates the signal power transmitted by the base station. The number of sampling points represents the number of points used to sample the channel response. Let be the channel response at the nth sampling point, and let represent the channel impulse response value at the nth sampling point. The power of the channel response represents the signal power at that point.
[0238] Compensation effectiveness index: Extract the channel state before and after compensation, calculate the channel response norm, evaluate the improvement of compensation effect, normalize the results, and generate the effectiveness index value;
[0239] ;
[0240] in, This is a compensation effectiveness index, with a value range of [0,1], representing the quality of the compensation effect. The Frobenius norm is used to calculate the norm of a matrix. The channel response after compensation at the current moment represents the channel state after compensation. The channel response at the previous time step represents the channel state before compensation. The sampling interval represents the time difference between two adjacent samples.
[0241] SINR Improvement Verification: Measure the signal-to-interference-plus-noise ratio (SINR) before compensation, measure the SINR after compensation, calculate the SINR improvement, evaluate the improvement effect, and record the verification results.
[0242] ;
[0243] in, The signal-to-interference-plus-noise ratio (SIR) improvement (dB) indicates the degree of improvement in SIR before and after compensation. The signal-to-interference-plus-noise ratio (SIR) after compensation represents the signal quality after compensation. The signal-to-interference-plus-noise ratio (SIR / NDR) before compensation is represented by the signal quality before compensation.
[0244] Output: This is the performance evaluation vector, which includes compensation performance, SINR improvement, and loss prediction error.
[0245] S520, root cause analysis of the fault:
[0246] Anomaly detection: Check compensation effectiveness indicators, assess loss prediction error, apply anomaly judgment rules, mark abnormal states, and record detection results;
[0247] ;
[0248] in, This is an abnormality indicator; 1 indicates an abnormality, and 0 indicates normal operation. As a compensation performance index, its value range is [0,1]. This represents the actual path loss. To predict path loss, The loss deviation threshold represents the maximum acceptable prediction error. The compensation effectiveness threshold represents the minimum acceptable compensation effect.
[0249] Root cause classification: Analyze changes in environmental threats, assess model prediction errors, monitor hardware performance indicators, determine fault types, and generate diagnostic reports;
[0250] Environmental mutation: when and ;
[0251] Model mismatch: ;
[0252] Hardware failure: 3 times consecutively ;
[0253] Output: The fault diagnosis report includes the root cause category (environmental mutation / model mismatch / hardware failure) and the corresponding confidence level;
[0254] S530, online model updates:
[0255] Dataset augmentation: Collect terminal location information, record motion state data, extract environmental features, construct training samples, and update the dataset;
[0256] ;
[0257] in, The new training samples include location, velocity, threat level, and actual loss. The terminal position vector contains three-dimensional coordinates (x, y, z). The terminal velocity vector contains velocity components in three directions. This represents the maximum threat level, with a value ranging from [0,1]. This represents the actual path loss.
[0258] GNN parameter update: Calculate the prediction error gradient, update the network weights, apply regularization constraints, verify model performance, and save the updated parameters;
[0259] ;
[0260] in, These are the parameters for the GNN model, including the weights and biases of all layers in the network. The learning rate is a hyperparameter that controls the step size of parameter updates. To predict loss for the model, For actual measurement of loss, An operator for calculating the gradient with respect to the parameter θ. The square of the Euclidean distance is used to calculate the prediction error;
[0261] DDPG Experience Replay: Records state transition data, calculates instant reward values, stores interactive experiences, updates the experience pool, and performs strategy optimization;
[0262] ;
[0263] in, This serves as an experience replay buffer, storing a queue of historical interaction data. This represents the current state, including communication quality and environmental characteristics. For power adjustment amount, For immediate reward, indicating the effectiveness of the action. The next state is the system state after the action is performed.
[0264] Output the updated GNN model parameters, including the new weights and biases, and also output the updated DDPG model parameters, including the parameters of the policy network and the value network.
[0265] like Figure 2 As shown, this invention also proposes a wireless signal coverage enhancement system based on reinforcement learning, comprising the following modules:
[0266] Joint Perception and Threat Recognition Module 101: It integrates LiDAR point cloud, camera visual data and terminal motion status, and identifies the collision threat level of dynamic obstacles (such as vehicles) in real time (TTC calculation) through multi-source spatiotemporal alignment and relative motion analysis. It outputs quantitative threat indicators and a list of high-risk targets, providing an environmental perception basis for subsequent path optimization.
[0267] Diffraction path dynamic reconstruction module 102: Based on the environmental reflective surface characteristics (material, scattering intensity) and real-time obstacle positions, a dynamic diffraction topology map is constructed; a graph neural network (GNN) is used to predict the diffraction path loss of the signal in a complex environment, and a set of low-loss candidate paths is screened to provide physical propagation model support for beam optimization.
[0268] Real-time beam optimization module 103: Combines path loss, terminal speed and threat level to select the optimal diffraction path and calculate the beam pointing angle; dynamically compensates for terminal motion offset, adaptively adjusts beamwidth to balance coverage and anti-interference capability, and outputs three-dimensional beam control commands (azimuth / elevation / beamwidth).
[0269] Power enhancement decision module 104: Analyzes the impact of path loss and beam gain on communication quality, and dynamically adjusts the transmit power using reinforcement learning (DDPG); through regulatory constraints (power upper limit), thermal protection (temperature limit) and interference coordination (neighbor cell avoidance) mechanisms, it outputs a safe power command to maintain a stable SINR.
[0270] Closed-loop feedback and learning module 105: Verify the signal quality after compensation (loss error, SINR improvement rate), diagnose the root causes of anomalies (environmental changes / model mismatch / hardware failure); trigger incremental model updates (GNN parameter calibration, DDPG experience playback) to achieve system self-optimization and long-term reliability improvement.
[0271] Based on the above methods and systems, the following example is given: Millimeter-wave dynamic NLOS compensation monitoring example (highway scenario).
[0272] Step 1: Joint Perception and Threat Identification
[0273] A smart car is traveling at high speed on a highway when its onboard sensors detect a truck on its right suddenly changing lanes.
[0274] The system uses lidar to capture the outline of the truck, cameras to identify the vehicle type, and combines the vehicle's GPS location and speed to calculate that the truck will completely block the base station signal in 0.8 seconds, thus marking it as a high-threat target.
[0275] Step 2: Dynamic reconstruction of diffraction path
[0276] The system scans the environment, identifies the central median metal guardrail and roadside billboards as effective reflective surfaces, and constructs a dynamic diffraction topology map.
[0277] Predict two candidate paths using a graph neural network:
[0278] Radiation path around the top of the truck (higher loss)
[0279] Billboard reflection path (more stable)
[0280] Step 3: Real-time Beam Optimization
[0281] The billboard reflection path is selected as the main link, and the calculated beam needs to be pointed to a specific reflection point of the billboard. The beam is dynamically widened to 25 degrees according to the vehicle speed to compensate for the pointing deviation caused by the high speed of the vehicle and generate three-dimensional beam control commands.
[0282] Step 4: Power Enhancement Decision
[0283] Predicting path loss necessitates increased transmit power, but neighboring cell signals indicate potential interference with ambulance communications. After balancing system requirements, the power is slightly increased and limited to regulatory safety levels to avoid cross-cell interference.
[0284] Step 5: Closed-loop feedback
[0285] After the occlusion occurred, the signal remained stable through the reflection path of the billboard. The system confirmed that the SINR (signal-to-noise ratio) improvement met the target, and stored the successful compensation parameters in the model training library to optimize the response in similar future scenarios.
[0286] In this example, the system yielded the following results: the video call did not experience any interruption due to the truck obstructing the view, and the navigation system maintained centimeter-level positioning accuracy, verifying the effectiveness of the "monitoring-compensation-optimization" closed loop.
[0287] The embodiments of the present invention have been described above, but the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of the present invention, all of which are within the protection scope of the present invention.
Claims
1. A method for wireless signal coverage enhancement based on reinforcement learning, the method comprising: The method comprises the following steps: S100, joint perception and threat identification: by fusing multi-source perception data such as laser radar and camera, the terminal motion state and environmental obstacle information are extracted in real time, the relative motion parameters and collision time are calculated, the threat level is evaluated by comprehensively considering the size and distance of the obstacle, and the quantitative threat index and potential high-risk target are output; S200, dynamic reconstruction of diffraction path: the effective reflecting surface is identified from the point cloud data, the time-varying diffraction topology graph is constructed combined with the dynamic obstacle position, the signal loss of each path is predicted by using the graph neural network, and the low-loss candidate path set is screened out to provide physical propagation model support for beam optimization; S300, real-time beam optimization: the optimal diffraction path is selected based on the path loss, terminal motion speed and threat level, the beam pointing angle is calculated and the deviation caused by the high-speed motion of the terminal is compensated, the beam width is adaptively adjusted to balance the coverage range and anti-interference ability, and the three-dimensional beam control instruction is generated; S400, power reinforcement decision: the communication quality is predicted combined with the path loss and beamforming gain, the power adjustment amount is dynamically decided by using reinforcement learning, the power safety is guaranteed through the regulation constraint, thermal protection and interference coordination mechanism, and the transmission power instruction meeting the real-time requirement is output; S500, closed-loop feedback: the signal strength and quality improvement effect after compensation are verified, the abnormal root cause is diagnosed, the path prediction model parameters are incrementally updated, and the reinforcement learning experience data is stored, so that the system is self-optimized.
2. The method of claim 1, wherein, The specific steps of joint perception and threat identification are as follows: S110, time and space alignment of multi-source data: input laser radar original point cloud data, camera original target detection result and terminal motion state vector; The laser radar original point cloud data is down-sampled and noise filtered, the camera original target detection result is non-maximum suppression, the point cloud and image features are extracted for registration, the transformation matrix of the laser radar to the base station is calculated, the transformation matrix of the camera to the base station is calculated, the coordinate conversion is applied by using the transformation matrix, the data alignment is carried out based on the time stamp, the transmission delay is compensated by using linear interpolation, and the synchronous data buffer is constructed; S120, motion state extraction: the speed and acceleration are calculated based on the historical position, the state estimation is smoothed by applying Kalman filtering, and the abnormal value is detected and processed; The center point and boundary of the obstacle are extracted, the target association and tracking are established, and the motion parameters of the obstacle are estimated; The next time state is predicted by using the motion model, the prediction uncertainty is calculated, and the confidence of the state estimation is updated; S130, relative motion analysis: the relative position of the terminal and each obstacle is calculated, the relative speed and acceleration are calculated, and the relative motion trend is analyzed; The approach / away speed is calculated, the motion direction change is analyzed, and the relative attitude relationship is evaluated; The closest approach point is predicted, the collision probability is estimated, and the avoidance time margin is calculated; The relative position of the terminal and each obstacle is calculated, the relative speed and acceleration are calculated, and the relative motion trend is analyzed; S140, collision threat evaluation: the collision time is calculated, the collision probability is evaluated, and the collision consequence severity is analyzed; The high-threat target is determined by comprehensively scoring multiple factors and applying threat level mapping.
3. The method of claim 1, wherein, The specific steps of dynamic reconstruction of diffraction path are as follows: S210, environment feature extraction: segment and cluster the point cloud data, identify continuous planar regions, calculate the normal vector and boundary contour of each plane, extract surface texture features and reflection characteristics, apply a deep learning model for material classification, and calculate the radar cross section area; filter low reflectivity surfaces according to the threshold of radar cross section area, remove metal material surfaces, merge adjacent effective reflection areas, calculate the area and direction of each reflection surface, and establish a reflection surface index structure; S220, dynamic diffraction graph construction: extract the reflection surface corner points as static vertices, take the dynamic obstacle position as dynamic vertices, predict the future position of the obstacle, merge repeated or close vertices, assign a unique identifier to each vertex, check the visibility between vertex pairs, calculate the Euclidean distance between vertices, apply maximum distance constraints, construct an adjacency matrix, and optimize the connectivity of the graph; S230, path loss prediction: initialize the node feature vector, perform multi-layer graph convolution operations, aggregate neighbor node information, update the node hidden state, and apply a nonlinear activation function; extract the path start and end point features, calculate the path physical length, apply the loss prediction model, consider the material influence factors, and output the loss prediction value. In the environment feature extraction, the following contents are included:
4. The method of claim 3, wherein, The calculation formula of radar cross section area is as follows: The calculation formula of material classification is as follows: ; in, Let J be the radar cross section of the j-th surface. For millimeter wave wavelength, For the first The attenuation coefficient of a path indicates the degree of signal strength attenuation. For carrier frequency, For the first The delay of the path, For belonging to the surface A subset of point clouds The angular constant of a spherical solid. For modulo squaring operations; The calculation formula of effective reflection surface screening is as follows: ; wherein, is a material classification convolutional neural network for identifying surface material types, is a first point cloud data of a surface, is a material type output including metal, glass, concrete; In the dynamic diffraction graph construction, the following contents are included: ; wherein, is the effective reflection surface set after screening, is the jth surface, is the scattering cross-sectional area threshold, is the material type that needs to be excluded.
5. The method of claim 4, wherein, The calculation formula of vertex set generation is as follows: The calculation formula of edge set construction is as follows: ; ; ; wherein, is a set of graph vertices for time t, is a static vertex, representing a reflecting facet corner, is a dynamic vertex, is a position three-dimensional vector of the th obstacle, is a velocity three-dimensional vector of the th obstacle, is a current time, is a function to extract a corner. In the path loss prediction, the following contents are included: ; wherein, is a set of graph edges at time t, is a three-dimensional coordinate of any two vertices, is a maximum diffraction distance, is a line-of-sight unocclusion detection function, is a Euclidean distance calculation.
6. The method of claim 5, wherein the method further comprises: The calculation formula of GNN propagation model is as follows: The calculation formula of end-to-end path loss is as follows: ; wherein, is the layer GNN weight matrix of dimension d x 2d, is the hidden state vector of the layer GNN of dimension d, is the hidden state vector of the layer GNN of dimension d, is the neighbor set of vertex v, is the vector concatenation operation, is the rectified linear unit activation function, is the GNN layer index, is the hidden state vector dimension; The specific steps of real-time beam optimization are as follows: ; wherein, is the predicted diffraction loss for path p, is the parameterized path loss prediction function, and is the final hidden state vector for the start node, is the final hidden state vector for the end node, is the physical length of path , is the path loss coefficient, is the total number of layers of the GNN.
7. The method of claim 1, wherein, S310, optimal path selection: calculate the unit vector of the main propagation direction of each path, evaluate the impact of terminal motion on path stability, generate a path stability index, normalize the stability index, and mark paths with stability below the threshold; extract path loss features, calculate path perimeter threat levels, fuse multi-dimensional evaluation indicators, apply weight coefficients, and generate path comprehensive scores; sort all candidate paths, select the path with the highest score, verify path feasibility, record selection results, and update path state information; S320, beam pointing calculation: extract path key point coordinates, calculate point-to-terminal distance, calculate point-to-base station distance, find the optimal reflection point, and record the reflection point position; calculate the relative position vector, decompose the horizontal / vertical components, apply the inverse trigonometric function, convert the angle unit, and ensure the angle range; decompose the terminal velocity, calculate the azimuth and elevation compensation amounts, apply compensation coefficients, and limit the compensation range; S330, beam width adaptation: calculate the terminal velocity, apply velocity normalization, calculate the basic width, limit the width range, record the reference value; evaluate the threat level, check the velocity condition, determine the adjustment direction, calculate the adjustment amount, apply the adjustment rule; combine the basic width and adjustment amount, apply upper and lower limit constraints, ensure the width rationality, generate the final width, and update the control parameters. The specific steps of power reinforcement decision are as follows:
8. The method of claim 1, wherein, S410, communication state evaluation: estimate atmospheric absorption and scattering loss according to carrier frequency and propagation distance, consider the influence of carrier frequency on propagation attenuation, calculate path loss based on free space propagation model, superimpose diffraction loss, atmospheric loss and other losses to get equivalent total loss; read the current transmission power setting of the base station, calculate the interference level based on the location and power of the adjacent base station, calculate the antenna directivity gain according to the beam width, comprehensively consider the transmission power, path loss, interference and gain; S420, reinforcement learning decision: normalize the predicted SINR by dividing the target threshold, normalize the terminal speed by dividing the maximum speed limit, directly use the normalized threat level index, divide the current power by the maximum power limit, and combine each normalized component into a state vector; input the normalized state vector into the policy network, calculate the power adjustment suggestion through the policy network, adjust the exploration degree based on the threat level, ensure that the power adjustment is within a reasonable range, and generate the actual power adjustment instruction; evaluate the degree of achievement of the current SINR relative to the target, apply a penalty according to the size of the power adjustment, judge whether cell switching occurs, combine each score according to the weight to generate the final reward signal; S430, power safety constraint: add the current power and adjustment amount to ensure that it does not exceed the maximum power limit and does not fall below the minimum power requirement, limit the power that exceeds the range to a legal interval, save the power value that meets the regulatory requirements; read the current temperature of the radio frequency module, check whether it exceeds the temperature threshold, determine whether power reduction is needed according to the temperature state, apply the power setting after thermal protection, and record the thermal protection triggering condition; detect the signals of surrounding adjacent base stations, calculate the influence of adjacent signals on the cell, set the power backoff amount according to the interference level to ensure that it does not cause excessive interference to the adjacent cell, and output the power value that meets the coordination requirements.
9. The method of claim 1, wherein, The specific steps of closed-loop feedback are as follows: S510, compensation effect evaluation: sample and analyze the compensated channel response, calculate the channel response power spectral density, apply the path loss calculation formula, consider the influence of transmission power, and output the actual loss result; extract the channel state before and after compensation, calculate the channel response norm, evaluate the improvement of compensation effect, normalize the result, and generate the performance index value; measure the signal-to-interference-and-noise ratio before compensation, measure the signal-to-interference-and-noise ratio after compensation, calculate the SINR improvement, evaluate the improvement effect, and record the verification result; S520, fault root cause analysis: check the compensation performance index, evaluate the loss prediction error, apply the abnormality judgment rule, mark the abnormal state, and record the detection result; analyze the change of environmental threats, evaluate the model prediction error, monitor the hardware performance index, determine the fault type, and generate a diagnosis report; S530, online model update: collect terminal location information, record motion state data, extract environmental features, build training samples, and update the data set; calculate the prediction error gradient, update the network weight, apply regularization constraint, verify the model performance, and save the updated parameters; record state transition data, calculate the immediate reward value, store interaction experience, update the experience pool, and perform policy optimization.
10. A wireless signal coverage enhancement system based on reinforcement learning, characterized by, Steps for performing a method of wireless signal coverage enhancement based on reinforcement learning as claimed in any one of claims 1-9, comprising: Joint perception and threat identification module: fuse lidar point cloud, camera vision data and terminal motion state, through multi-source spatio-temporal alignment and relative motion analysis, real-time identify the collision threat level of dynamic obstacles, output quantitative threat indicators and high-risk target list; Diffractive path dynamic reconstruction module: based on the characteristics of the environment reflecting surface and the real-time obstacle position, construct a dynamic diffraction topology graph; use graph neural network to predict the diffraction path loss of the signal in the complex environment, and screen a low-loss candidate path set; Real-time beam optimization module: combine path loss, terminal speed and threat level to select the optimal diffraction path and calculate the beam pointing angle; dynamically compensate for terminal motion offset, and adaptively adjust the beam width to balance the coverage range and anti-interference ability, output three-dimensional beam control instructions; Power reinforcement decision module: analyze the influence of path loss and beam gain on communication quality, and dynamically adjust the transmission power by using reinforcement learning; output safe power instructions through regulation constraints, thermal protection and interference coordination mechanism; Closed-loop feedback and learning module: verify the compensated signal quality, diagnose the root cause of the abnormality; trigger incremental model update.
Citation Information
Cited By
Low-altitude communication node deployment method and system based on deep learning and reinforcement learning
CN121586007A
Wireless communication network intelligent interference management system based on deep learning
CN122092994A