Network Optimization Method, Device and Equipment Based on the Fusion of Geometry and Deep Learning

By using the Poisson point process in the wireless network to generate a random topology structure, and combining analytical geometry and deep deterministic strategy gradient algorithms, the problems of the respective limitations of analytical geometry and deep learning in the existing technology are solved, and efficient optimization of wireless network performance and effective utilization of spectrum resources are achieved.

CN120075849BActive Publication Date: 2025-06-27EAST CHINA JIAOTONG UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510525883.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-06-27
Estimated Expiration
2045-04-25

AI Technical Summary

Technical Problem

In the prior art, analytical geometry methods are difficult to achieve accurate tuning for specific scenarios. Deep reinforcement learning ignores the geometric characteristics and global characteristics of the network physical layer, resulting in slow convergence and poor generalization capabilities, and is unable to effectively adapt to dynamically changing network environments.

Method used

The Poisson point process is used to generate the random topology of the wireless network, and combined with the results derived from analytical geometry, the deep deterministic strategy gradient algorithm is used to optimize the network parameters to realize the joint optimization of spectrum resource allocation and power control.

Benefits of technology

By combining analytical geometry and deep learning, the joint optimization of wireless network performance is achieved, the spectrum resource utilization efficiency is improved, the link interrupt probability is reduced, and the adaptability is strong.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120075849B_ABST
    Figure CN120075849B_ABST
Patent Text Reader

Abstract

The present application provides a network optimization method, device and equipment based on the fusion of geometry and deep learning, belonging to the field of network resource allocation. The method includes: generating a random distribution of macro base stations, roadside units and vehicles by using a Poisson point process to construct the topological structure of a communication network; constructing a channel analysis function of the communication network according to the topological structure; deriving the analysis function to obtain the performance analysis result of the communication network; constructing a deep deterministic policy network on the premise of the analysis result; combining the deep deterministic policy network and the proximal policy optimization algorithm to jointly optimize the resource allocation and power control of the communication network to obtain a preliminary communication network; and locally optimizing the preliminary communication network according to an optimization constraint function to obtain a target communication network. The present application effectively improves the performance of a wireless network, reduces the link interruption probability and improves the resource utilization efficiency by combining random geometry, deep learning algorithms and local optimization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of network resource allocation, and particularly relates to a network optimization method, device, and equipment based on the integration of geometry and deep learning. Background Art

[0002] Regarding the problem of increased spectrum interference due to intensified wireless backhaul link resource competition in the integrated access and backhaul solution proposed by 3GPP (3rd Generation Partnership Project), the combination with millimeter-wave technology has become an important solution. The millimeter-wave band provides abundant spectrum resources. Although its propagation loss is high, resulting in limited coverage, it shows significant advantages in short-distance high-speed transmission scenarios. In addition, through the application of multi-antenna beamforming technology and cooperative transmission technology, the link efficiency and system performance of millimeter-wave communication can be further improved.

[0003] When analyzing the performance of heterogeneous millimeter-wave integrated access and backhaul networks, stochastic geometry modeling and analytic geometry theory play important roles. For example, using the Poisson point process to model the distribution of base station nodes, indicators such as the network average interference level and coverage probability can be derived, providing a theoretical basis for network planning.

[0004] However, in the prior art, the analytic geometry method has limitations in practical applications, such as simplified assumptions making it difficult to capture the optimal configuration, and problems of dealing with complex expressions and large computational overhead. Therefore, it becomes crucial to introduce deep learning to jointly optimize network performance. In particular, the deep deterministic policy gradient algorithm shows unique advantages in solving problems with continuous action spaces.

[0005] In the prior art, attempts have also been made to separately improve the analytic geometry model and deep learning algorithms to solve the above problems, including using a more realistic complex node distribution model combined with geographic information system data, and using transfer learning and reinforcement learning improvement methods to enhance the generalization ability. However, these improvements also face challenges, such as increased model complexity and complex reward function design.

[0006] In summary, neither the analytic geometry nor deep learning alone can fully exploit the potential of heterogeneous millimeter-wave integrated access and backhaul networks. Therefore, a method that organically combines the two is needed to make full use of their respective advantages and overcome each other's limitations to meet the strict requirements of future wireless communication networks for high performance and high reliability. Such a method can achieve the joint optimization of network performance and adapt to the changing network environment. Summary of the Invention

[0007] The purpose of the embodiments of this application is to provide a network optimization method, device, and equipment based on the fusion of geometry and deep learning, so as to solve the problem in the prior art that it is difficult for analytical geometry methods to achieve precise tuning for specific scenarios to determine the optimal configuration, while the isolated application of deep reinforcement learning ignores the geometric and global characteristics of the network physical layer, resulting in slow convergence, poor generalization ability, and inability to effectively adapt to the dynamically changing network environment.

[0008] To solve the above technical problems, this application is implemented as follows:

[0009] In the first aspect, the embodiments of this application provide a network optimization method based on the fusion of geometry and deep learning. The method includes:

[0010] Use the Poisson point process to generate the random distribution of macro base stations, roadside units, and vehicles to construct the topological structure of the communication network; among them, the macro base station is modeled by a two-dimensional Poisson point process, and the roadside units and vehicles are respectively modeled by homogeneous spatial Poisson point processes with different densities;

[0011] According to the topological structure, construct the channel analysis function of the communication network;

[0012] Derive the analysis function to obtain the performance analysis result of the communication network;

[0013] Based on the analysis result, construct a deep deterministic policy network; among them, the architecture characteristics of the deep deterministic policy network include: state space, action space, and reward function;

[0014] Combine the deep deterministic policy network and the proximal policy optimization algorithm to jointly optimize the resource allocation and power control of the communication network to obtain a preliminary communication network;

[0015] According to the optimization constraint function, perform local optimization on the preliminary communication network to obtain the target communication network; among them, the optimization constraint function includes link signal-to-noise ratio constraint, spectrum allocation ratio constraint, and power constraint.

[0016] Compared with the prior art, the above technical solution provided by this application at least includes the following beneficial effects:

[0017] This application first generates a random topology of the wireless network through a Poisson point process, and based on the results derived from analytic geometry, further optimizes the network parameters using the Deep Deterministic Policy Gradient (DDPG) algorithm to achieve joint optimization of spectrum resource allocation and power control. This method provides a necessary network environment for subsequent applications of the DDPG network. Analytic geometry is used to calculate and predict network performance, providing data support and constraint conditions for the DDPG network. By using the DDPG algorithm, the network can learn the optimal wireless resource management strategy and power control strategy in the environment to maximize the rate and reduce the link outage probability. After the DDPG network obtains a preliminary solution, further optimization and fine-tuning are performed through the Proximal Policy Optimization (PPO) algorithm to ensure that the solution is more accurate.

[0018] This application effectively improves the performance of the wireless network, reduces the link outage probability, and enhances the spectrum resource utilization efficiency through a method combining stochastic geometry, the DDPG algorithm, the PPO algorithm, and local optimization, and has significant application prospects.

[0019] Preferably, the specific steps for constructing the channel analytic function of the communication network according to the topology include:

[0020] Based on the line-of-sight probability model and the piecewise path loss model, establish the link blockage function and the path loss function;

[0021] Use the single-slope path loss model to describe the signal attenuation function;

[0022] Describe the directional transmission gain function of the multi-antenna array through the sector antenna model;

[0023] Derive the signal-to-interference-plus-noise ratio (SINR) function based on the link blockage function, the path loss function, the signal attenuation function, and the directional transmission gain function.

[0024] Preferably, the specific steps for constructing the DDPG network with the analytic results as the constraint conditions include:

[0025] Construct the state space according to the analytic results; the state space includes the spectrum resource allocation ratio, the macro base station transmit power, the link mode, and the network throughput;

[0026] Formulate the reward function according to the analytic results;

[0027] Form the DDPG network based on the state space and the reward function.

[0028] Preferably, the specific steps for formulating the reward function based on the rate and power consumption include:

[0029] The reward function is:

[0030] ,

[0031] ,

[0032] Among them, represents the reward value, represents the interruption probability, represents the successful coverage rate, is the power consumption, , and are weight coefficients, which respectively control the interruption probability, maximize the rate, and reduce power consumption, represents the throughput.

[0033] Preferably, combining the deep deterministic policy network and the proximal policy optimization algorithm to jointly optimize the resource allocation and power control of the communication network, the specific steps to obtain the preliminary communication network include:

[0034] Optimizing the communication network based on the deep deterministic policy network to obtain a preliminary solution;

[0035] Based on the preliminary solution, using the proximal policy optimization algorithm for secondary solution to obtain the preliminary communication network.

[0036] Preferably, the specific steps to locally optimize the preliminary communication network according to the optimization constraint function to obtain the target communication network include:

[0037] Based on the preliminary communication network, using the grid search function to construct a local search network; among them, the grid search function is used to construct a local search network near the spectrum ratio and power optimal point;

[0038] Using the optimization constraint function to optimize the local search network to form the target communication network.

[0039] Preferably, the specific steps to use the optimization constraint function to optimize the local search network to form the target communication network include:

[0040] Determine the optimization objective function, optimization objective, and optimization constraint function;

[0041] The optimization objective function is:

[0042] ,

[0043] The optimization objective is:

[0044] ,

[0045] The optimization constraint function is:

[0046] ,

[0047] ,

[0048] ,

[0049] wherein, represents the network space throughput, represents the basic data rate, represents the proportion of energy harvesting time, represents the spectrum deviation penalty coefficient, represents the current spectrum allocation ratio, represents the optimal value of spectrum allocation, represents the power deviation penalty coefficient, represents the current transmission power of the macro base station, represents the optimal transmission power; represents the joint optimization of the spectrum allocation ratio and the transmission power of the macro base station to maximize the system throughput; represents the signal-to-interference-plus-noise ratio; represents the transmission power of the macro base station, then and respectively represent the minimum and maximum values of the transmission power of the macro base station.

[0050] Preferably, the specific steps further include:

[0051] Performing local optimization on the preliminary communication network according to the optimization constraint function to obtain target information; the target information includes the algorithm convergence curve, throughput, link outage probability, and distance relationship.

[0052] In a second aspect, an embodiment of the present application provides a network optimization device based on the fusion of geometry and deep learning, including:

[0053] A geometric modeling module, configured to generate random distributions of macro base stations, roadside units, and vehicles using a Poisson point process to construct the topological structure of the communication network; wherein, the macro base stations are modeled using a two-dimensional Poisson point process, and the roadside units and vehicles are respectively modeled using uniform spatial Poisson point processes with different densities;

[0054] A function module, configured to construct a channel analysis function of the communication network according to the topological structure;

[0055] An analysis module, configured to derive the analysis function to obtain the performance analysis result of the communication network;

[0056] A policy network module, configured to construct a deep deterministic policy network with the analysis result as a constraint condition; wherein, the architecture features of the deep deterministic policy network include: a state space, an action space, and a reward function;

[0057] A joint optimization module, which is used to combine the deep deterministic policy network and the proximal policy optimization algorithm to jointly optimize the resource allocation and power control of the communication network, so as to obtain a preliminary communication network;

[0058] A local optimization module, which is used to locally optimize the preliminary communication network according to an optimization constraint function to obtain a target communication network; wherein, the optimization constraint function includes link signal-to-noise ratio constraint, spectrum allocation ratio constraint and power constraint.

[0059] In a third aspect, an embodiment of the present application provides an electronic device, which includes a processor, a memory, and a program or instruction stored in the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of the method described in the first aspect are implemented.

[0060] It can be understood that the beneficial effects of the technical solutions provided in the above second aspect and third aspect can refer to the relevant descriptions in the above first aspect, and will not be repeated here.

[0061] Additional aspects and advantages of the present application will be given in part in the following description, become apparent in part from the following description, or be understood through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] The above and / or additional aspects and advantages of the present application will become apparent and be readily understood from the description of the embodiments in conjunction with the following drawings, where:

[0063] Figure 1 is a schematic flowchart of a network optimization method based on the fusion of geometry and deep learning provided by some embodiments of the present application;

[0064] Figure 2 is a block diagram of a network optimization device based on the fusion of geometry and deep learning shown by some embodiments of the present application;

[0065] Figure 3 is a block diagram of an electronic device shown by some embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0066] The technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments of the present application belong to the scope of protection of the present application.

[0067] The terms "first", "second", etc. in the description and claims of this application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of this application can be implemented in an order other than those illustrated or described herein. In addition, "and / or" in the description and claims means at least one of the connected objects, and the character " / " generally indicates that the related objects before and after are in an "or" relationship.

[0068] The following combines the accompanying drawings and details a network optimization method provided by an embodiment of this application based on the fusion of geometry and deep learning through specific embodiments and their application scenarios.

[0069] Figure 1 It is a schematic flowchart of a network optimization method based on the fusion of geometry and deep learning shown in the first embodiment of this application. Please refer to Figure 1 This method includes:

[0070] Step S101: Use a Poisson point process to generate a random distribution of macro base stations, roadside units, and vehicles to construct the topological structure of the communication network; among them, the macro base station is modeled by a two-dimensional Poisson point process, and the roadside unit and the vehicle are respectively modeled by uniform spatial Poisson point processes with different densities.

[0071] In a possible implementation, first, model the random distribution of network nodes. The macro base station is modeled by a two-dimensional Poisson point process with a density of and is randomly distributed in the two-dimensional plane, denoted as . In addition, the positions of the vehicle and the roadside unit as nodes are modeled by uniform spatial Poisson point processes with densities of and respectively. The position of the vehicle node is denoted as an independent and uniform spatial Poisson point process , and the position of the roadside unit is denoted as . By simulating the spatial randomness of the base station and vehicle user nodes, the uncertainty in the real network environment is captured.

[0072] Next, set the antenna resources. On the basis of the above random distribution modeling, further configure the antenna resources of the macro base station and the roadside unit. Assume that the macro base station is configured with a large-scale MIMO antenna array, and the roadside unit and vehicle users are configured with single antennas.

[0073] Furthermore, it is necessary to establish link connection relationships. Based on the above random distribution modeling, the link connections are described in detail. The wireless backhaul transmission scheme studied in this embodiment adopts a half-duplex mode and is executed in two stages, each stage lasting for the time length of one time slot. The specific process is as follows: First, attempt to establish a backhaul link from the macro base station to the roadside unit. If this connection is successfully established, then continue to construct the link from the roadside unit to the vehicle; conversely, if the backhaul link between the macro base station and the roadside unit fails to be successfully established, then the macro base station directly transmits data to the vehicle.

[0074] It should be noted that this embodiment uses a Poisson point process to construct the wireless network topology to simulate the random distribution of macro base stations, roadside units, and vehicles. Through random geometric modeling, the main purpose is to generate the topology of the wireless network, determine the specific positions of macro base stations, roadside units, and vehicle users, and their link connection relationships, so as to form spatial geometric information. These information provide the necessary network environment and data basis for subsequent analytical geometry derivations.

[0075] For example, consider a downlink transmission IAB network, which consists of a macro base station, a super-densely distributed roadside unit, and vehicles. The vehicles are connected to the roadside unit through the downlink, and the roadside unit backhauls to the macro base station through a millimeter-wave link, or the vehicles directly access the macro base station through the downlink.

[0076] Among them, the macro base station is modeled by a two-dimensional Poisson point process with a density of in a given network area, randomly distributed in the two-dimensional plane, and denoted as . In addition, the positions of vehicle nodes and roadside units are modeled by uniform spatial Poisson point processes with densities of and respectively. The positions of vehicle nodes are denoted as an independent and homogeneous spatial Poisson point process , and the position distribution of roadside units is . By simulating the spatial randomness of base stations and vehicle user nodes, the uncertainties in the real network environment are captured.

[0077] Due to the ultra-dense deployment of roadside units in the network, the number of roadside units far exceeds the number of macro base stations in a limited area. Assume that the spatial positions of roadside units, macro base stations, and vehicles respectively form sets: the roadside unit set , the macro base station set , and the vehicle set . Therefore, this embodiment can use these sets to describe the point processes composed of roadside units, macro base stations, and vehicles:

[0078] , , ,

[0079] wherein, 、 and are all natural numbers, respectively representing the current labels of the roadside unit, the macro base station, and the vehicle.

[0080] According to the Slivnyak-Mecke theorem, consider a typical vehicle user (reference vehicle user) located at the origin of coordinates, that is, let . Let the roadside unit closest to it be: , and the closest macro base station be: . The distances from the vehicle user to the roadside unit and the macro base station are respectively expressed as: , , and these distances are used to model the path loss of wireless communication signals and network coverage analysis.

[0081] Step S102: Construct a channel analysis function for the communication network according to the topological structure.

[0082] Specifically, it includes: based on the line-of-sight probability model and the segmented path loss model, establish a link blocking function and a path loss function; use a single-slope path loss model to describe the signal attenuation function; describe the directional transmission gain function of the multi-antenna array through a sector antenna model; according to the link blocking function, the path loss function, the signal attenuation function, and the directional transmission gain function, derive the signal-to-interference ratio function.

[0083] First, based on the network topological structure established by stochastic geometry, establish a channel analysis function and perform analytic geometry analysis to derive an analytic conclusion of network performance.

[0084] In a possible implementation manner, a millimeter-wave channel analysis function is established. Due to the characteristics of high frequency and large bandwidth, millimeter-wave communication has a wide range of applications in scenarios such as vehicle-to-everything (V2X) networks, 5G / 6G cellular communications, and driverless V2X communications. However, due to the high directivity and high sensitivity to blockage of millimeter waves, its channel modeling is different from traditional cellular communications. To accurately characterize the millimeter-wave channel characteristics, this embodiment uses a line-of-sight (LOS) sphere approximation model to describe the blocking probability of the link, and uses different path loss models to respectively characterize the signal attenuation under LOS and non-line-of-sight (NLOS) propagation conditions.

[0085] To characterize the high sensitivity of millimeter-wave communication to obstacle blockage, a LOS sphere approximation model is used to characterize whether the link is in the LOS state. That is, this embodiment assumes that the LOS probability of the link only depends on the distance between the communication parties and a preset maximum LOS distance threshold 。LOS probability The expression is as follows:

[0086] ,

[0087] wherein, represents the backhaul LOS probability; is an indicator function, when it takes the value of 1, otherwise it takes the value of 0; is the distance from the roadside unit or macro base station to the vehicle; is the maximum coverage distance of the link in the LOS state.

[0088] This formula means that when the link distance is less than or equal to the link must be in the LOS state, that is, there is no blocking effect. When the link distance

[0089] Since the path loss of millimeter wave communication varies significantly due to different propagation environments. In this embodiment, a segmented path loss model is adopted, and different path loss exponents are defined for LOS and NLOS propagation conditions respectively. The path loss model can be expressed as:

[0090] ,

[0091] wherein, represents the path loss of millimeter wave communication; represents the signal propagation distance of millimeter wave communication, and are the LOS and NLOS path loss exponents respectively, represents the adjustment coefficient.

[0092] In this embodiment, a Rayleigh fading model with unit mean is adopted to describe the small-scale fading effect. Therefore, the small-scale fading coefficient between the vehicle and its serving roadside unit follows an exponential distribution with a mean of 1, which can effectively simulate the Rayleigh fading situation between the macro base station and the roadside unit.

[0093] Furthermore, a Sub-6Ghz channel analysis function is established. In the wireless communication environment of the Sub-6 GHz band, the channel gain is mainly affected by path loss and small-scale fading. In order to accurately describe the path loss of the signal, the expression of the single-slope path loss model adopted in this embodiment is as follows:

[0094] ,

[0095] wherein, Represents the path loss in the Sub-6 GHz band; Represents the signal propagation distance in the Sub-6 GHz band, is the path loss exponent.

[0096] In the Sub-6 GHz band, small-scale fading is usually modeled by Rayleigh fading to characterize the random fluctuations of the signal in a multipath propagation environment. Specifically, in this embodiment, it is assumed that the channel fading coefficient is , and it follows an exponential distribution with a unit mean, expressed as:

[0097] ,

[0098] where, represents the location of the roadside unit; represents the location of the vehicle.

[0099] Furthermore, an antenna model is established. In a millimeter-wave communication system, to overcome the severe path loss of high-frequency signals, each macro base station is equipped with a large and compact antenna array to achieve high-gain directional beamforming. This technology enhances communication reliability by concentrating the transmit power in the target direction, increasing the link gain and reducing interference. To simplify the modeling, in this embodiment, a sector antenna model is used to approximately describe the beamforming gain, and its mathematical expression is as follows:

[0100] ,

[0101] where, represents the beamforming gain; represents the main lobe gain, which provides the maximum signal gain when pointing to the target direction; represents the side lobe gain, which reflects the signal loss when the beam is not aligned with the target direction; is the beam width, which determines the angular range covered by the beam; represents the angular deviation of the beam relative to the target direction.

[0102] In millimeter-wave communication, the beam width is usually narrow to ensure that the signal energy is focused in the target direction, thereby improving the signal-to-interference-plus-noise ratio and reducing co-channel interference. During the association process, the macro base station scans the space to obtain the best beam alignment. Therefore, the main lobe of the serving macro base station points to the serving roadside unit, that is, the beamforming gain of the serving macro base station is always , while the main lobes of other macro base stations do not necessarily point to the direction of the serving roadside unit.

[0103] In addition, according to Euclidean principles, the distance of the link can be expressed as:

[0104] ,

[0105] Among them, represents the link distance; , respectively represent the starting and ending coordinates of the link.

[0106] Furthermore, based on the above model, an analytical function of the signal-to-interference ratio received by a typical vehicle and a roadside unit can be established.

[0107] The signal-to-interference ratio of a typical vehicle can be expressed as:

[0108] ,

[0109] Among them, represents the signal-to-interference ratio of the typical vehicle; represents the transmitting power of the roadside unit; represents the small-scale fading coefficient between the typical vehicle and its serving roadside unit, and follows an exponential distribution with a mean of 1; represents the path loss of its link, is the total interference from at , is the total interference from at .

[0110] The signal-to-interference ratio of the serving roadside unit can be expressed as:

[0111] ,

[0112] Among them, represents the signal-to-interference ratio of the roadside unit; represents the transmitting power of the macro base station; represents the main lobe gain; represents the small-scale fading coefficient between the serving roadside unit and the serving macro base station, and follows an exponential distribution with a mean of 1; is the path loss of its link; represents the total interference from at ; represents the total interference from at . However, the main lobe beam directions of other macro base stations may not all be aligned with the serving roadside unit. Therefore, in this embodiment, it is assumed that the main lobe of other macro base stations points to the serving roadside unit with a probability This probability mainly depends on the beam width of the main lobe of the macro base station beam.

[0113] Step S103: Derive an analytical function to obtain the performance analysis result of the communication network;

[0114] This step aims to determine the coverage probability of a typical receiver in the network based on the signal-to-interference-plus-noise ratio (SINR). In addition, the spatial throughput of the network will also be determined to characterize the overall performance of the entire network.

[0115] In the process of determining the coverage probability, first consider the link from the macro base station to the serving roadside unit in the downlink backhaul, and then consider the link from the roadside unit to the typical vehicle in the downlink access. In this embodiment, the joint coverage probability of two consecutive hops will be calculated, and the joint coverage probability can be expressed as:

[0116] ,

[0117] where, , are respectively and SINR threshold values.

[0118] The proof formula for the joint coverage probability is:

[0119] , ,

[0120] ,

[0121] where, represents the joint coverage probability, represents the Laplace transform of the interference generated by other roadside units to the serving vehicle; represents the link coefficient from the macro base station to the roadside unit; represents the link coefficient from the roadside unit to the vehicle; represents the Laplace transform of the interference generated by other roadside units to the serving roadside unit; represents the Laplace transform of the interference generated by other macro base stations to the serving roadside unit.

[0122] The network spatial throughput can be expressed as:

[0123] ,

[0124] where, represents the network throughput; is the rate per unit spectrum; is the overhead factor; represents the successful coverage rate.

[0125] Step S104: Based on the analysis result, construct a deep deterministic policy network; among them, the architecture characteristics of the deep deterministic policy network include: state space, action space, and reward function;

[0126] Specifically, it includes: constructing a state space according to the parsing result; the state space includes the spectrum resource allocation ratio, macro base station transmission power, link mode, and network throughput; formulating a reward function according to the parsing result; and forming a deep deterministic policy network based on the state space and the reward function.

[0127] Specifically, a five-dimensional state space is constructed based on the topology generated by the stochastic geometry model and the network performance results analyzed by analytic geometry, including: spectrum resource allocation ratio, macro base station transmission power, roadside unit transmission power, link mode, and network throughput.

[0128] Based on the analytic geometry analysis, the deep deterministic policy gradient algorithm is used to optimize the network performance. Based on the network topology features extracted from the parsing result, the state space, action space, and reward function are designed to ensure the efficient configuration and optimization of network resources. The design is as follows:

[0129] State space: The network topology features calculated from the parsing result, such as link mode, coverage probability, and power control, are used as the input of the state space of the deep deterministic policy network. The state vector includes the key information of the current network.

[0130] Action space: For multiple resource allocation problems such as power allocation, spectrum reuse factor, and backhaul path selection, the action space is jointly optimized. By reasonably selecting power control and spectrum allocation ratio, the deep deterministic policy network can adaptively adjust the policy during the training process to improve the overall network performance and ensure that the system resources can be effectively allocated according to the real-time demand.

[0131] Defining the reward function and designing the reward function: The reward function cleverly combines multiple key performance indicators, including network rate, delay, and fairness, and introduces stochastic geometry constraints, such as interference threshold, to ensure that while pursuing performance optimization, excessive interference is avoided.

[0132] Rate: The rate of the network is affected by the spectrum allocation and the selection of link mode, and the goal is to maximize the rate.

[0133] Power consumption: Power consumption is an important indicator in optimization. Increasing power may improve the rate, but it will also increase the power consumption. Therefore, there is a term in the reward function to penalize power consumption so that the agent can learn how to reduce power consumption without sacrificing performance.

[0134] In the deep deterministic policy network environment, the reward function is the basis for the agent to optimize the policy. Considering the link outage probability, rate, and power consumption comprehensively, the expression of the reward function is:

[0135] ,

[0136] ,

[0137] wherein, represents the reward value, represents the interruption probability, represents the successful coverage rate, is the power consumption, , and are weight coefficients, which respectively control the interruption probability, maximize the rate and reduce the power consumption, represents the throughput.

[0138] Step S105: Combine the Deep Deterministic Policy Network and the Proximal Policy Optimization algorithm to jointly optimize the resource allocation and power control of the communication network, and obtain a preliminary communication network;

[0139] Specifically, it includes: optimizing the communication network based on the Deep Deterministic Policy Network to obtain a preliminary solution; for the preliminary solution, using the Proximal Policy Optimization algorithm for secondary solution to obtain a preliminary communication network.

[0140] Specifically, on the basis of obtaining a preliminary solution by the Deep Deterministic Policy Gradient algorithm, further use the Proximal Policy Optimization algorithm file for optimization and fine-tuning. The Deep Deterministic Policy Network module is responsible for the generation and optimization of the action policy, and is more suitable for decision-making problems with a continuous action space; the Proximal Policy Optimization module has strong stability in the overall evolution of the policy and the valuation accuracy, and can effectively constrain the amplitude of the policy update; the two realize joint training and cross-optimization by sharing the environmental state, policy results, evaluation information, etc.; this collaborative mechanism takes into account the fineness of decision-making and the stability of policy convergence, and is applicable to communication system scenarios with multiple nodes, dynamic changes, and latency sensitivity.

[0141] Step S106: Perform local optimization on the preliminary communication network according to the optimization constraint function to obtain the target communication network; wherein, the optimization constraint function includes link signal-to-noise ratio constraint, spectrum allocation ratio constraint and power constraint.

[0142] Specifically, it includes: based on the preliminary communication network, using the grid search function to construct a local search network; wherein, the grid search function is used to construct a local search network near the spectrum ratio and power optimal points; using the optimization constraint function to optimize the local search network to form the target communication network.

[0143] For example, perform local grid search near the optimal solution determined by the Deep Deterministic Policy Gradient algorithm and the Proximal Policy Optimization algorithm to accurately optimize the network performance. Among them, the grid search function is defined as constructing a local search grid near the optimal points of the spectrum ratio and power;

[0144] The optimized objective function adopts a V-shaped objective function, and its expression is:

[0145] ,

[0146] The expression of the optimization objective is:

[0147] ,

[0148] Since the link outage probability must be lower than a certain threshold, it is required that the signal-to-noise ratio of the link is greater than a certain threshold, and the signal-to-interference ratio of the link must satisfy the following conditions. The optimization constraint function is:

[0149] ,

[0150] ,

[0151] ,

[0152] Among them, represents the network space throughput, represents the basic data rate, represents the proportion of energy harvesting time, represents the spectrum deviation penalty coefficient, represents the current spectrum allocation ratio, represents the optimal value of spectrum allocation, represents the power deviation penalty coefficient, represents the current macro base station transmission power, represents the optimal transmission power; represents the joint optimization of the spectrum allocation ratio and the macro base station transmission power to maximize the system throughput; represents the signal-to-interference ratio; represents the threshold value; represents the macro base station transmission power, then and represent the minimum and maximum values of the macro base station transmission power respectively.

[0153] According to the optimization constraint function, the preliminary communication network is locally optimized to obtain the target information; the target information includes the algorithm convergence curve, throughput, link outage probability and distance relationship.

[0154] Specifically, finally, a comprehensive visual display and performance verification of the optimization results are carried out, specifically including: the convergence curves of the Deep Deterministic Policy Gradient algorithm and the Proximal Policy Optimization algorithm to show the optimization efficiency of the algorithms; the two-dimensional throughput surface graph of the local grid search optimization results to clearly display the exact optimal solution after optimization; the comparative analysis of the network throughput before and after optimization; the analysis of the relationship between the link outage probability and the distance to clarify the effectiveness of the optimization scheme.

[0155] The network optimization method based on the fusion of geometry and deep learning provided by the above embodiments, on the one hand, combines analytic geometry with continuous control optimization based on the Deep Deterministic Policy Gradient algorithm. The Poisson point process is used to generate the random distribution of base stations, roadside units, and vehicles to simulate the randomness of the wireless network. Through the framework of analytic geometry, preliminary analysis results are obtained, revealing the geometric relationships between devices in the network and their impact on communication performance. The Deep Deterministic Policy Gradient algorithm is introduced, taking the analytic characteristics of analytic geometry as the state input. The Deep Deterministic Policy Network realizes the optimization of wireless network resources and power control through training and learning. Combining the policy-value evaluation structure, it can perform efficient optimization in a non-stationary dynamic environment, improving the convergence speed and policy stability. Therefore, by combining the stochastic geometry method with the Deep Deterministic Policy Gradient, the learned policy can not only provide high performance but also improve the interpretability and transparency of the policy.

[0156] Secondly, an efficient communication network optimization scheme that comprehensively utilizes millimeter-wave technology, multi-antenna beamforming, and cooperative transmission technology to meet the stringent requirements of next-generation communication networks for high capacity, low latency, and large-scale coverage. The millimeter-wave band provides extremely rich spectrum resources, which can significantly improve the capacity and bandwidth of the network. However, its propagation loss is relatively high, restricting the coverage range, especially in large-scale transmission and complex environments, presenting significant challenges. To address this issue, the multi-antenna beamforming technology is introduced. By achieving highly directional beam transmission in high-speed transmission scenarios over short distances, it effectively compensates for the high propagation loss of millimeter waves, improving the reliability and transmission efficiency of the link. At the same time, the integration of cooperative transmission technology enables multiple communication nodes to share resources and information, further enhancing the overall performance and transmission success rate of the system. This scheme can not only achieve flexible optimization in an uncertain environment but also provide theoretical and technical support for the popularization of millimeter-wave technology in practical applications, promoting the efficient deployment and operation of future 5G and 6G networks.

[0157] On the other hand, for the intelligent resource management strategy of multi-objective optimization, a reward function that comprehensively considers rate, link outage probability, and power control is designed, and the spectrum resources and macro base station power configuration are jointly optimized, enabling the reinforcement learning intelligent agent to learn an optimization strategy that balances multiple objectives, rather than simply maximizing the rate or minimizing power consumption. Secondly, for the adaptive selection of link modes, the intelligent adaptive selection of link modes is realized through reinforcement learning, and the optimal link mode is adjusted according to the real-time network state to improve the overall system rate and reduce the outage probability. On the basis of obtaining a preliminary solution by the deep deterministic policy gradient algorithm, the proximal policy algorithm file is further used for optimization and fine-tuning to make the optimization result more accurate. This is different from the traditional network optimization method with a fixed link mode, enabling the network to flexibly adjust the connection strategy in different environments and improving the adaptability of the network.

[0158] By combining stochastic geometry, deep deterministic policy gradient algorithm, proximal policy optimization algorithm and local optimization, the performance of the wireless network is effectively improved, the link outage probability is reduced, and the utilization efficiency of spectrum resources is enhanced, with significant application prospects.

[0159] It should be noted that for the network optimization method based on the fusion of geometry and deep learning provided in the embodiments of the present application, the execution subject can be a network optimization device based on the fusion of geometry and deep learning, or a control module in the network optimization device based on the fusion of geometry and deep learning for executing the loading of the network optimization method based on the fusion of geometry and deep learning. In the embodiments of the present application, taking the network optimization device based on the fusion of geometry and deep learning executing the loading of the network optimization method based on the fusion of geometry and deep learning as an example, the network optimization device based on the fusion of geometry and deep learning provided in the embodiments of the present application is described.

[0160] Figure 2 is a schematic diagram of a network optimization device based on the fusion of geometry and deep learning shown in the second embodiment of the present application. Please refer to Figure 2 , the network optimization device 200 based on the fusion of geometry and deep learning includes:

[0161] Geometric modeling module 201: used to generate the random distribution of macro base stations, roadside units and vehicles by using the Poisson point process to construct the topological structure of the communication network; among them, the macro base stations are modeled by two-dimensional Poisson point process, and the roadside units and vehicles are respectively modeled by uniform spatial Poisson point processes with different densities;

[0162] Function module 202: used to construct the channel analysis function of the communication network according to the topological structure;

[0163] Specifically, it includes: establishing a link blocking function and a path loss function based on a line-of-sight probability model and a segmented path loss model; using a single-slope path loss model to describe the signal attenuation function; describing the directional transmission gain function of a multi-antenna array through a sector antenna model; and deriving a signal-to-interference ratio function according to the link blocking function, the path loss function, the signal attenuation function, and the directional transmission gain function.

[0164] Analysis module 203: used to derive an analysis function to obtain the performance analysis result of the communication network;

[0165] Policy network module 204: used to construct a deep deterministic policy network with the analysis result as a constraint condition; among them, the architecture characteristics of the deep deterministic policy network include: a state space, an action space, and a reward function;

[0166] Specifically, it includes: constructing a state space according to the analysis result; the state space includes the spectrum resource allocation ratio, the macro base station transmission power, the link mode, and the network throughput; formulating a reward function according to the analysis result; and forming a deep deterministic policy network based on the state space and the reward function.

[0167] The reward function is:

[0168] ,

[0169] ,

[0170] Among them, represents the reward value, represents the outage probability, represents the successful coverage rate, is the power consumption, , and are weight coefficients, which respectively control the outage probability, maximize the rate, and reduce the power consumption, represents the throughput.

[0171] Joint optimization module 205: used to combine the deep deterministic policy network and the proximal policy optimization algorithm to jointly optimize the resource allocation and power control of the communication network to obtain a preliminary communication network;

[0172] Specifically, it includes: optimizing the communication network based on the deep deterministic policy network to obtain a preliminary solution; based on the preliminary solution, using the proximal policy optimization algorithm for secondary solution to obtain a preliminary communication network.

[0173] Local optimization module 206: used to locally optimize the preliminary communication network according to the optimization constraint function to obtain the target communication network; among them, the optimization constraint function includes link signal-to-noise ratio constraint, spectrum allocation ratio constraint, and power constraint.

[0174] Specifically, based on the preliminary communication network, a local search network is constructed by using a grid search function; wherein, the grid search function is used to construct a local search network near the spectrum ratio and power optimal points; and the local search network is optimized by using an optimization constraint function to form a target communication network.

[0175] Specifically, an optimization objective function, an optimization objective, and an optimization constraint function are determined;

[0176] The optimization objective function is:

[0177] ,

[0178] The optimization objective is:

[0179] ,

[0180] The optimization constraint function is:

[0181] ,

[0182] ,

[0183] ,

[0184] Among them, represents the network space throughput; represents the basic data rate; represents the proportion of energy harvesting time; represents the spectrum deviation penalty coefficient; represents the current spectrum allocation ratio; represents the optimal value of spectrum allocation; represents the power deviation penalty coefficient; represents the current macro base station transmission power; represents the optimal transmission power; represents the joint optimization of the spectrum allocation ratio and the maximum system throughput of the macro base station transmission power; represents the signal-to-interference ratio; represents the threshold value; represents the macro base station transmission power, then and represent the minimum and maximum values of the macro base station transmission power respectively.

[0185] The specific steps of this embodiment further include: locally optimizing the preliminary communication network according to the optimization constraint function to obtain target information; the target information includes an algorithm convergence curve, throughput, link interruption probability, and distance relationship.

[0186] The network optimization device based on the fusion of geometry and deep learning in the embodiments of the present application may be a device, or a component, an integrated circuit, or a chip in a terminal. The device may be a mobile electronic device or a non-mobile electronic device. Exemplarily, the mobile electronic device may be a mobile phone, a tablet computer, a laptop computer, a handheld computer, a vehicle-mounted electronic device, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc., and the non-mobile electronic device may be a server, a Network Attached Storage (NAS), a personal computer (PC), a television (TV), a teller machine, or a self-service machine, etc. The embodiments of the present application do not make specific limitations.

[0187] The network optimization device based on the fusion of geometry and deep learning in the embodiments of the present application may be a device with an operating system. The operating system may be the Android operating system, the iOS operating system, or other possible operating systems. The embodiments of the present application do not make specific limitations.

[0188] The network optimization device based on the fusion of geometry and deep learning provided in the embodiments of the present application can implement Figure 1 each process implemented by the network optimization method based on the fusion of geometry and deep learning in the method embodiments. To avoid repetition, it will not be elaborated here.

[0189] Optionally, please refer to Figure 3 , the embodiments of the present application further provide an electronic device 300, including a processor 301, a memory 302, and a computer program 303 stored on the memory 302 and executable on the processor 301. When the computer program 303 is executed by the processor 301, it implements each process of the above-mentioned network optimization method embodiment based on the fusion of geometry and deep learning, and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0190] The embodiments of the present application further provide a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, it implements each process of the above-mentioned network optimization method embodiment based on the fusion of geometry and deep learning, and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0191] Among them, the processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs, etc.

[0192] Another embodiment of the present application provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement each process of the above embodiments of the network optimization method based on the fusion of geometry and deep learning, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.

[0193] It should be understood that the chip mentioned in the embodiments of the present application may also be referred to as a system-on-chip, system chip, chip system, or system-on-chip, etc.

[0194] It should be noted that in this article, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including that element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in a reverse order according to the functions involved. For example, the methods described may be performed in an order different from that described, and various steps may be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.

[0195] Through the description of the above embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc), and includes several instructions for causing a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of the present application.

[0196] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative rather than restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them fall within the protection scope of the present application.

Claims

1. A network optimization method based on the fusion of geometry and deep learning, characterized in that: include: The Poisson point process is used to generate the random distribution of macro base stations, roadside units and vehicles to construct the topological structure of the communication network; the macro base stations are modeled by a two-dimensional Poisson point process, and the roadside units and vehicles are modeled by uniform spatial Poisson point processes with different densities. Constructing a channel analysis function of the communication network according to the topological structure; Derivation of the analytical function to obtain a performance analysis result of the communication network; Based on the analysis result, a deep deterministic policy network is constructed; wherein the architecture features of the deep deterministic policy network include: state space, action space and reward function; Combining the deep deterministic policy network and the proximal policy optimization algorithm, jointly optimizing resource allocation and power control of the communication network, and obtaining a preliminary communication network; The preliminary communication network is locally optimized according to the optimization restriction function to obtain a target communication network; wherein the optimization restriction function includes link signal-to-noise ratio restriction, spectrum allocation ratio restriction and power restriction.

2. The network optimization method based on the fusion of geometry and deep learning according to claim 1 is characterized in that: The specific steps of constructing the channel resolution function of the communication network according to the topological structure include: Based on the line-of-sight probability model and the segmented path loss model, the link blocking function and path loss function are established; A single slope path loss model is used to describe the signal attenuation function; The directional transmission gain function of multi-antenna array is described by sector antenna model; The signal-to-interference ratio function is derived based on the link blocking function, path loss function, signal attenuation function and directional transmission gain function.

3. The network optimization method based on the fusion of geometry and deep learning according to claim 1 is characterized in that: The specific steps of constructing a deep deterministic strategy network using the analysis results as constraints include: Constructing a state space according to the analysis results; the state space includes spectrum resource allocation ratio, macro base station transmission power, link mode and network throughput; Formulate a reward function according to the analysis result; The deep deterministic policy network is formed based on the state space and the reward function.

4. The network optimization method based on the fusion of geometry and deep learning according to claim 3 is characterized in that: The reward function is: , , in, Represents the reward value, represents the interruption probability, represents the successful coverage rate, To consume power, , and are weight coefficients, which respectively control the interruption probability, maximize the rate and reduce the power consumption, Indicates throughput.

5. The network optimization method based on the fusion of geometry and deep learning according to claim 1 is characterized in that: The specific steps of combining the deep deterministic policy network and the proximal policy optimization algorithm to jointly optimize the resource allocation and power control of the communication network to obtain a preliminary communication network include: Optimizing the communication network based on the deep deterministic strategy network to obtain a preliminary solution; Based on the preliminary solution, a proximal strategy optimization algorithm is used to perform a secondary solution to obtain a preliminary communication network.

6. The network optimization method based on the fusion of geometry and deep learning according to claim 1 is characterized in that: The specific steps of locally optimizing the preliminary communication network according to the optimization restriction function to obtain the target communication network include: Based on the preliminary communication network, a local search network is constructed using a grid search function; wherein the grid search function is used to construct a local search network near the optimal point of spectrum ratio and power; The local search network is optimized using the optimization restriction function to form the target communication network.

7. The network optimization method based on the fusion of geometry and deep learning according to claim 6 is characterized in that: The specific steps of optimizing the local search network by using the optimization restriction function to form a target communication network include: Determining an optimization objective function, an optimization target, and the optimization restriction function; The optimization objective function is: , The optimization goal is: , The optimization restriction function is: , , , in, represents the network space throughput; Indicates the basic data rate; Indicates the proportion of energy harvesting time; represents the spectrum deviation penalty coefficient; Indicates the current spectrum allocation ratio; Indicates the optimal value of spectrum allocation; represents the power deviation penalty coefficient; Indicates the current macro base station transmit power; represents the optimal transmit power; Represents the joint optimization spectrum allocation ratio Transmit power of macro base station Maximizing system throughput; It represents the signal-to-interference ratio; Indicates the threshold value; represents the macro base station transmit power, then and They respectively represent the minimum and maximum values ​​of the macro base station transmit power.

8. The network optimization method based on the fusion of geometry and deep learning according to claim 1 is characterized in that: The specific steps also include: The preliminary communication network is locally optimized according to the optimization restriction function to obtain target information; the target information includes an algorithm convergence curve, throughput, link interruption probability and distance relationship.

9. A network optimization device based on the fusion of geometry and deep learning, used to execute the network optimization method based on the fusion of geometry and deep learning as described in any one of claims 1 to 8, characterized in that: include: A geometric modeling module is used to generate a random distribution of macro base stations, roadside units and vehicles using a Poisson point process to construct a topological structure of a communication network; wherein the macro base stations are modeled using a two-dimensional Poisson point process, and the roadside units and vehicles are modeled using uniform spatial Poisson point processes with different densities; A function module, used for constructing a channel analysis function of the communication network according to the topological structure; An analysis module, used for deriving the analysis function to obtain a performance analysis result of the communication network; A policy network module, used to construct a deep deterministic policy network with the analysis results as constraints; wherein the architecture features of the deep deterministic policy network include: state space, action space and reward function; A joint optimization module, used to combine the deep deterministic policy network and the proximal policy optimization algorithm to jointly optimize the resource allocation and power control of the communication network to obtain a preliminary communication network; A local optimization module is used to locally optimize the preliminary communication network according to an optimization restriction function to obtain a target communication network; wherein the optimization restriction function includes link signal-to-noise ratio restriction, spectrum allocation ratio restriction and power restriction.

10. An electronic device, characterized in that: include: A memory, a processor, and a program or instruction stored in the memory and executable on the processor, wherein the program or instruction, when executed by the processor, implements the steps of the network optimization method based on the fusion of geometry and deep learning as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Deep reinforcement learning communication interference resource allocation method fused with noise network

    CN115866760A

  • Heterogeneous network resource energy efficiency optimization method based on deep reinforcement learning

    CN116567667A