Network optimization method, device and equipment based on geometric and deep learning fusion

By combining Poisson point process and deep deterministic strategy gradient algorithm in wireless networks, the problems of the respective limitations of analytical geometry and deep learning in the prior art are solved, and the joint optimization of wireless network performance is achieved, and the efficiency and adaptability of spectrum resource utilization are improved.

CN120075849AActive Publication Date: 2025-05-30EAST CHINA JIAOTONG UNIVERSITY
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510525883.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-05-30
Estimated Expiration
2045-04-25

AI Technical Summary

Technical Problem

In the prior art, analytical geometry methods are difficult to achieve accurate tuning for specific scenarios. Deep reinforcement learning ignores the geometric characteristics and global characteristics of the network physical layer, resulting in slow convergence and poor generalization capabilities, and is unable to effectively adapt to dynamically changing network environments.

Method used

The Poisson point process is used to generate the random topology of the wireless network, and combined with the results derived from analytical geometry, the deep deterministic strategy gradient algorithm is used to optimize the network parameters to realize the joint optimization of spectrum resource allocation and power control.

Benefits of technology

By combining analytical geometry and deep learning, the joint optimization of wireless network performance is achieved, the spectrum resource utilization efficiency is improved, the link interrupt probability is reduced, and the adaptability is strong.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120075849A_ABST
    Figure CN120075849A_ABST
Patent Text Reader

Abstract

The invention provides a network optimization method, device and equipment based on geometry and deep learning fusion, and belongs to the field of network resource allocation. The method comprises the following steps: generating random distribution of a macro base station, a road side unit and a vehicle by adopting a Poisson point process so as to construct a topological structure of a communication network; constructing a channel analytic function of the communication network according to the topological structure; deriving an analytic function to obtain a performance analysis result of the communication network; constructing a deep deterministic policy network by taking an analysis result as a precondition; in combination with a deep deterministic policy network and a near-end policy optimization algorithm, resource allocation and power control of the communication network are jointly optimized to obtain a preliminary communication network; and performing local optimization on the preliminary communication network according to the optimization limiting function to obtain a target communication network. By combining stochastic geometry, a deep learning algorithm and local optimization, the wireless network performance is effectively improved, the link outage probability is reduced, and the resource utilization efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of network resource allocation, and particularly relates to a network optimization method, device, and equipment based on the fusion of geometry and deep learning. Background Art

[0002] Regarding the problem of increased spectrum interference due to intensified wireless backhaul link resource competition in the integrated access and backhaul solution proposed by 3GPP (3rd Generation Partnership Project), combining millimeter-wave technology has become an important solution. The millimeter-wave band provides abundant spectrum resources. Although its propagation loss is high, resulting in limited coverage, it shows significant advantages in short-distance high-speed transmission scenarios. In addition, through the application of multi-antenna beamforming technology and cooperative transmission technology, the link efficiency and system performance of millimeter-wave communication can be further improved.

[0003] When analyzing the performance of heterogeneous millimeter-wave integrated access and backhaul networks, stochastic geometry modeling and analytic geometry theory play important roles. For example, using the Poisson point process to model the distribution of base station nodes, indicators such as the network average interference level and coverage probability can be derived, providing a theoretical basis for network planning.

[0004] However, in the prior art, the analytic geometry method has limitations in practical applications, such as simplified assumptions making it difficult to capture the optimal configuration, as well as problems of dealing with complex expressions and large computational overhead. Therefore, it becomes crucial to introduce deep learning to jointly optimize network performance. In particular, the deep deterministic policy gradient algorithm shows unique advantages in solving continuous action space problems.

[0005] In the prior art, attempts have also been made to separately improve the analytic geometry model and deep learning algorithms to solve the above problems, including adopting a more realistic complex node distribution model combined with geographic information system data, and using transfer learning and reinforcement learning improvement methods to enhance the generalization ability. However, these improvements also face challenges, such as increased model complexity and complex reward function design.

[0006] In summary, neither using analytic geometry nor deep learning alone can fully exploit the potential of heterogeneous millimeter-wave integrated access and backhaul networks. Therefore, a method that organically combines the two is needed to make full use of their respective advantages and overcome their respective limitations to meet the strict requirements of future wireless communication networks for high performance and high reliability. Such a method can achieve the joint optimization of network performance and adapt to the changing network environment. Summary of the Invention

[0007] The purpose of the embodiments of this application is to provide a network optimization method, device, and equipment based on the fusion of geometry and deep learning, so as to solve the problem in the prior art that it is difficult for analytical geometry methods to achieve precise tuning for specific scenarios to determine the optimal configuration, while the isolated application of deep reinforcement learning ignores the geometric and global characteristics of the network physical layer, resulting in slow convergence and poor generalization ability, and being unable to effectively adapt to the dynamically changing network environment.

[0008] To solve the above technical problems, this application is implemented as follows: In a first aspect, the embodiments of this application provide a network optimization method based on the fusion of geometry and deep learning. The method includes: Using a Poisson point process to generate a random distribution of macro base stations, roadside units, and vehicles to construct the topological structure of the communication network; among them, the macro base stations are modeled by a two-dimensional Poisson point process, and the roadside units and vehicles are respectively modeled by uniform spatial Poisson point processes with different densities; According to the topological structure, construct a channel analysis function for the communication network; Derive the analysis function to obtain the performance analysis results of the communication network; Taking the analysis results as preconditions, construct a deep deterministic policy network; among them, the architecture characteristics of the deep deterministic policy network include: state space, action space, and reward function; Combining the deep deterministic policy network and the proximal policy optimization algorithm, jointly optimize the resource allocation and power control of the communication network to obtain a preliminary communication network; According to the optimization constraint function, perform local optimization on the preliminary communication network to obtain the target communication network; among them, the optimization constraint function includes link signal-to-noise ratio constraint, spectrum allocation ratio constraint, and power constraint.

[0009] Compared with the prior art, the above technical solutions provided by this application at least include the following beneficial effects: This application first generates a random topological structure of the wireless network through a Poisson point process, and based on the results derived from analytical geometry, uses the deep deterministic policy gradient algorithm to further optimize the network parameters to achieve the joint optimization of spectrum resource allocation and power control. This method provides a necessary network environment for the subsequent application of the deep deterministic policy gradient network. Analytical geometry is used to calculate and predict network performance, providing data support and constraint conditions for the deep deterministic policy network. By using the deep deterministic policy gradient algorithm, the network can learn the optimal wireless resource management strategy and power control strategy in the environment to maximize the rate and reduce the link outage probability. After the deep deterministic policy network obtains a preliminary solution, further optimization and fine-tuning are performed through the proximal policy algorithm to ensure that the solution is more accurate.

[0010] This application effectively improves the performance of wireless networks, reduces the probability of link interruption, and enhances the efficiency of spectrum resource utilization by combining random geometry, the deep deterministic policy gradient algorithm, the proximal policy optimization algorithm, and local optimization, showing significant application prospects.

[0011] Preferably, the specific steps for constructing the channel analysis function of the communication network according to the topological structure include: Based on the line-of-sight probability model and the piecewise path loss model, establish the link blocking function and the path loss function; Use the single-slope path loss model to describe the signal attenuation function; Describe the directional transmission gain function of the multi-antenna array through the sector antenna model; Derive the signal-to-interference-plus-noise ratio function based on the link blocking function, the path loss function, the signal attenuation function, and the directional transmission gain function.

[0012] Preferably, the specific steps for constructing the deep deterministic policy network with the analysis result as the constraint condition include: Construct the state space according to the analysis result; the state space includes the spectrum resource allocation ratio, the macro base station transmit power, the link mode, and the network throughput; Formulate the reward function according to the analysis result; Form a deep deterministic policy network based on the state space and the reward function.

[0013] Preferably, the specific steps for formulating the reward function based on rate and power consumption include: The reward function is: , , where, represents the reward value, represents the interruption probability, represents the successful coverage rate, is the consumed power, , and are weight coefficients, controlling the interruption probability, maximizing the rate, and reducing the power consumption respectively, represents the throughput.

[0014] Preferably, the specific steps for jointly optimizing the resource allocation and power control of the communication network by combining the deep deterministic policy network and the proximal policy optimization algorithm to obtain the preliminary communication network include: Optimize the communication network based on the deep deterministic policy network to obtain a preliminary solution; Based on the preliminary solution, use the proximal policy optimization algorithm for secondary solution to obtain the preliminary communication network.

[0015] Preferably, the specific steps of locally optimizing the preliminary communication network according to the optimization constraint function to obtain the target communication network include: Based on the preliminary communication network, a local search network is constructed using a grid search function; wherein, the grid search function is used to construct a local search network near the spectrum ratio and power optimal points; The local search network is optimized using the optimization constraint function to form the target communication network.

[0016] Preferably, the specific steps of optimizing the local search network using the optimization constraint function to form the target communication network include: Determine the optimization objective function, optimization objective, and optimization constraint function; The optimization objective function is: , The optimization objective is: , The optimization constraint function is: , , , Wherein, represents the network space throughput, represents the basic data rate, represents the proportion of energy harvesting time, represents the spectrum deviation penalty coefficient, represents the current spectrum allocation ratio, represents the optimal value of spectrum allocation, represents the power deviation penalty coefficient, represents the current macro base station transmission power, represents the optimal transmission power; represents jointly optimizing the spectrum allocation ratio and the macro base station transmission power to maximize the system throughput; represents the signal-to-interference-plus-noise ratio; represents the threshold value; represents the macro base station transmission power, then and respectively represent the minimum and maximum values of the macro base station transmission power.

[0017] Preferably, the specific steps further include: Locally optimize the preliminary communication network according to the optimization constraint function to obtain target information; the target information includes the algorithm convergence curve, throughput, link outage probability, and distance relationship.

[0018] In a second aspect, an embodiment of the present application provides a network optimization device based on the fusion of geometry and deep learning, including: A geometric modeling module, configured to generate random distributions of macro base stations, roadside units, and vehicles using a Poisson point process to construct a topological structure of a communication network; wherein, the macro base stations are modeled using a two-dimensional Poisson point process, and the roadside units and vehicles are respectively modeled using uniform spatial Poisson point processes with different densities; A function module, configured to construct a channel analysis function of the communication network according to the topological structure; An analysis module, configured to derive the analysis function to obtain a performance analysis result of the communication network; A policy network module, configured to construct a deep deterministic policy network with the analysis result as a constraint condition; wherein, the architecture features of the deep deterministic policy network include: a state space, an action space, and a reward function; A joint optimization module, configured to jointly optimize resource allocation and power control of the communication network by combining the deep deterministic policy network and the proximal policy optimization algorithm to obtain a preliminary communication network; A local optimization module, configured to perform local optimization on the preliminary communication network according to an optimization constraint function to obtain a target communication network; wherein, the optimization constraint function includes link signal-to-noise ratio constraints, spectrum allocation ratio constraints, and power constraints.

[0019] In a third aspect, an embodiment of the present application provides an electronic device, which includes a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of the method described in the first aspect are implemented.

[0020] It can be understood that the beneficial effects of the technical solutions provided in the above second aspect and third aspect can refer to the relevant descriptions in the above first aspect, and will not be repeated here.

[0021] Additional aspects and advantages of the present application will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] The above and / or additional aspects and advantages of the present application will become apparent and easy to understand from the description of the embodiments in conjunction with the following drawings, wherein: Figure 1 is a schematic flowchart of a network optimization method based on the fusion of geometry and deep learning provided by some embodiments of the present application; Figure 2 is a block diagram of a network optimization device based on the fusion of geometry and deep learning shown by some embodiments of the present application; Figure 3It is a block diagram of an electronic device shown in some embodiments of the present application. Detailed implementation manners

[0023] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0024] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order different from those illustrated or described herein. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / " generally indicates an "or" relationship between the associated objects before and after.

[0025] Next, a network optimization method based on the fusion of geometry and deep learning provided in the embodiments of the present application will be described in detail in conjunction with the accompanying drawings, through specific embodiments and their application scenarios.

[0026] Figure 1 It is a schematic flowchart of a network optimization method based on the fusion of geometry and deep learning shown in the first embodiment of the present application. Please refer to Figure 1 , and the method includes: Step S101: Generate a random distribution of macro base stations, roadside units, and vehicles using a Poisson point process to construct the topological structure of the communication network; among them, the macro base station is modeled by a two-dimensional Poisson point process, and the roadside unit and the vehicle are respectively modeled by uniform spatial Poisson point processes with different densities.

[0027] In a possible implementation manner, first, a random distribution model of network nodes is performed. The macro base station is modeled by a two-dimensional Poisson point process with a density of , randomly distributed in the two-dimensional plane, and is denoted as . In addition, the positions of the vehicle and the roadside unit as nodes are modeled by uniform spatial Poisson point processes with densities of and respectively. The position of the vehicle node is denoted as an independent and uniform spatial Poisson point process , and the position of the roadside unit is denoted as . By simulating the spatial randomness of the base station and the vehicle user nodes, the uncertainty in the real network environment is captured.

[0028] Next, set the antenna resources. Based on the above random distribution modeling, further configure the antenna resources of the macro base station and the roadside unit. Assume that the macro base station is configured with a large-scale MIMO antenna array, while the roadside unit and vehicle users are configured with single antennas.

[0029] Furthermore, it is necessary to establish link connection relationships. Based on the above random distribution modeling, describe the link connections in detail. The wireless backhaul transmission schemes studied in this embodiment all adopt the half-duplex mode and are executed in two stages, with each stage lasting for the time length of one time slot. The specific process is as follows: First, attempt to establish a backhaul link from the macro base station to the roadside unit. If this connection is successfully established, then continue to construct the link from the roadside unit to the vehicle; conversely, if the backhaul link between the macro base station and the roadside unit fails to be successfully established, then the macro base station directly transmits data to the vehicle.

[0030] It should be noted that this embodiment uses the Poisson point process to construct the wireless network topology to simulate the random distribution of macro base stations, roadside units, and vehicles. Through random geometric modeling, the main purpose is to generate the topology of the wireless network, determine the specific positions of macro base stations, roadside units, and vehicle users and their link connection relationships, so as to form spatial geometric information. These information provide the necessary network environment and data basis for subsequent analytical geometry derivations.

[0031] For example, consider a downlink transmission IAB network, which consists of a macro base station, a super-densely distributed roadside unit, and vehicles. The vehicles are connected to the roadside unit through the downlink, and the roadside unit backhauls to the macro base station through the millimeter-wave link, or the vehicles directly access the macro base station through the downlink.

[0032] Among them, the macro base station is modeled by a two-dimensional Poisson point process with a density of in a given network area, randomly distributed in the two-dimensional plane, and denoted as . In addition, the positions of vehicle nodes and roadside units are modeled by uniform spatial Poisson point processes with densities of and respectively. The positions of vehicle nodes are denoted as an independent and homogeneous spatial Poisson point process , and the position distribution of roadside units is . By simulating the spatial randomness of base stations and vehicle user nodes, the uncertainties in the real network environment are captured.

[0033] Due to the ultra-dense deployment of roadside units in the network, the number of roadside units far exceeds the number of macro base stations in a limited area. Assume that the spatial positions of roadside units, macro base stations, and vehicles respectively form sets: the roadside unit set , the macro base station set , and the vehicle set Therefore, this embodiment can use these sets to describe the point processes composed of roadside units, macro base stations, and vehicles: , , , wherein, 、 and are all natural numbers, respectively representing the current labels of roadside units, macro base stations, and vehicles.

[0034] According to the Slivnyak-Mecke theorem, consider a typical vehicle user (reference vehicle user) located at the origin of coordinates, that is, let . Let the roadside unit closest to it be set as: , and the closest macro base station be set as: . The distances from the vehicle user to the roadside unit and the macro base station are respectively expressed as: , , and these distances are used to model the path loss of wireless communication signals and network coverage analysis.

[0035] Step S102: Construct a channel analysis function for the communication network according to the topological structure.

[0036] Specifically, it includes: based on the line-of-sight probability model and the segmented path loss model, establishing a link blocking function and a path loss function; using a single-slope path loss model to describe the signal attenuation function; describing the directional transmission gain function of the multi-antenna array through a sector antenna model; and deriving the signal-to-interference ratio function according to the link blocking function, path loss function, signal attenuation function, and directional transmission gain function.

[0037] First, based on the network topological structure established by random geometry, establish a channel analysis function and perform analytic geometry analysis to derive the analytic conclusion of network performance.

[0038] In a possible implementation manner, a millimeter-wave channel analysis function is established. Due to the characteristics of high frequency and large bandwidth, millimeter-wave communication has wide applications in scenarios such as vehicle-to-everything (V2X), 5G / 6G cellular communication, and driverless V2X communication. However, due to the high directivity and high sensitivity to blockage of millimeter waves, its channel modeling is different from traditional cellular communication. To accurately characterize the millimeter-wave channel characteristics, this embodiment uses the line-of-sight (LOS) sphere approximation model to describe the blocking probability of the link, and uses different path loss models to respectively characterize the signal attenuation under LOS and non-line-of-sight (NLOS) propagation conditions.

[0039] In order to characterize the high sensitivity of millimeter wave communication to obstacles, the line-of-sight sphere approximation model is used to characterize whether the link is in the LOS state. That is, this embodiment assumes that the LOS probability of the link depends only on the distance between the two communicating parties. And the preset maximum LOS distance threshold . LOS probability The expression is: , in, represents the return LOS probability; is the indicator function, when hour The value is 1, otherwise the value is 0; is the distance from the roadside unit or macro base station to the vehicle; It is the maximum coverage distance of the link in the LOS state.

[0040] This formula means that when the link distance When the link distance is When the link is in the NLOS state, the signal needs to propagate through scattering or reflection, and the path loss increases significantly.

[0041] Since the path loss of millimeter wave communication varies significantly due to different propagation environments, this embodiment adopts a segmented path loss model to define different path loss exponents for LOS and NLOS propagation conditions. The path loss model can be expressed as: , in, represents the path loss of mmWave communication; Indicates the signal propagation distance of millimeter communication, and are the LOS and NLOS path loss indices, respectively, Represents the adjustment factor.

[0042] This embodiment uses a unit mean Rayleigh fading model to describe the small-scale fading effect. Therefore, the small-scale fading coefficient between the vehicle and its serving roadside unit follows an exponential distribution with a mean of 1, which can effectively simulate the Rayleigh fading between the macro base station and the roadside unit.

[0043] Furthermore, a Sub-6 GHz channel analysis function is established. In the wireless communication environment of the Sub-6 GHz frequency band, the channel gain is mainly affected by path loss and small-scale fading. In order to accurately describe the path loss of the signal, this embodiment uses the expression of the single slope path loss modulus as follows: , in, Represents the path loss in the Sub-6 GHz band; Represents the signal propagation distance in the Sub-6 GHz band, is the path loss exponent. In the Sub-6 GHz band, small-scale fading is usually modeled using Rayleigh fading to characterize the random fluctuations of the signal in a multipath propagation environment. Specifically, in this embodiment, it is assumed that the channel fading coefficient is , and it follows an exponential distribution with a unit mean, expressed as: , where, represents the position of the roadside unit; represents the position of the vehicle.

[0044] Furthermore, an antenna model is established. In a millimeter-wave communication system, to overcome the severe path loss of high-frequency signals, each macro base station is equipped with a large and compact antenna array to achieve high-gain directional beamforming. This technology enhances communication reliability by concentrating the transmission power in the target direction, increasing the link gain and reducing interference. To simplify the modeling, in this embodiment, a sector antenna model is used to approximately describe the beamforming gain, and its mathematical expression is as follows: , where, represents the beamforming gain; represents the main lobe gain, which provides the maximum signal gain when pointing to the target direction; represents the side lobe gain, which reflects the signal loss when the beam is not aligned with the target direction; is the beam width, which determines the angular range covered by the beam; represents the angular deviation of the beam relative to the target direction.

[0045] In millimeter-wave communication, the beam width is usually narrow to ensure that the signal energy is focused in the target direction, thereby improving the signal-to-interference-plus-noise ratio and reducing co-channel interference. During the association process, the macro base station scans the space to obtain the best beam alignment. Therefore, the main lobe of the serving macro base station points to the serving roadside unit, that is, the beamforming gain of the serving macro base station is always , while the main lobes of other macro base stations do not necessarily point to the direction of the serving roadside unit.

[0046] In addition, according to Euclidean principles, the distance of the link can be expressed as: , where, represents the link distance; , respectively represent the starting and ending coordinates of the link.

[0047] Furthermore, based on the above model, an analytical function of the signal-to-interference-plus-noise ratio (SINR) received by a typical vehicle and a roadside unit can be established.

[0048] The SINR of a typical vehicle can be expressed as: , where represents the signal-to-interference ratio (SIR) of a typical vehicle; represents the transmit power of the roadside unit; represents the small-scale fading coefficient between a typical vehicle and its serving roadside unit, and follows an exponential distribution with a mean of 1; represents the path loss of its link, is the total interference from at , is the total interference from at .

[0049] The SINR of the serving roadside unit can be expressed as: , where represents the SIR of the roadside unit; represents the transmit power of the macro base station; represents the main lobe gain; represents the small-scale fading coefficient between the serving roadside unit and the serving macro base station, and follows an exponential distribution with a mean of 1; is the path loss of its link; represents the total interference from at ; represents the total interference from at . However, the main lobe beam directions of other macro base stations may not all be aligned with the serving roadside unit. Therefore, in this embodiment, it is assumed that the main lobe of other macro base stations points to the serving roadside unit with a probability , and this probability mainly depends on the beam width of the main lobe of the macro base station beam.

[0050] Step S103: Derive the analytical function to obtain the performance analysis result of the communication network; This step aims to determine the coverage probability of a typical receiver in the network based on the SINR. In addition, the spatial throughput of the network will also be determined to characterize the overall performance of the entire network.

[0051] In the process of determining the coverage probability, first consider the link from the macro base station to the serving roadside unit in the downlink backhaul, and then consider the link from the roadside unit to the typical vehicle in the downlink access. In this embodiment, the joint coverage probability of two consecutive hops will be calculated, and the joint coverage probability can be expressed as: , where, , are respectively and signal-to-interference ratio thresholds.

[0052] The proof formula for the joint coverage probability is: , , , where, represents the joint coverage probability, represents the Laplace transform of the interference generated by other roadside units to the serving vehicle; represents the link coefficient from the macro base station to the roadside unit; represents the link coefficient from the roadside unit to the vehicle; represents the Laplace transform of the interference generated by other roadside units to the serving roadside unit; represents the Laplace transform of the interference generated by other macro base stations to the serving roadside unit.

[0053] The network space throughput can be expressed as: , where, represents the network throughput; is the rate per unit spectrum; is the overhead factor; represents the successful coverage rate.

[0054] Step S104: Based on the parsing result, construct a deep deterministic policy network; among them, the architecture characteristics of the deep deterministic policy network include: state space, action space, and reward function; Specifically, construct the state space according to the parsing result; the state space includes the spectrum resource allocation ratio, macro base station transmission power, link mode, and network throughput; formulate the reward function according to the parsing result; form the deep deterministic policy network based on the state space and the reward function.

[0055] Specifically, construct a five-dimensional state space with the topology generated by the random geometry model and the network performance results analyzed by analytic geometry, including: spectrum resource allocation ratio, macro base station transmission power, roadside unit transmission power, link mode, network throughput.

[0056] Based on the analysis of analytic geometry, the deep deterministic policy gradient algorithm is used to optimize the network performance. Based on the network topology features extracted from the analysis results, the state space, action space, and reward function are designed to ensure the efficient configuration and optimization of network resources. The design is as follows: State space: The network topology features calculated from the analysis results, such as link mode, coverage probability, and power control, are used as the input of the state space of the deep deterministic policy network. The state vector includes the key information of the current network.

[0057] Action space: For multiple resource allocation problems such as power allocation, spectrum reuse factor, and backhaul path selection, the action space is jointly optimized. By reasonably selecting the power control and spectrum allocation ratio, the deep deterministic policy network can adaptively adjust the policy during the training process to improve the overall network performance and ensure that the system resources can be effectively allocated according to real-time requirements.

[0058] Define the reward function and design the reward function: The reward function cleverly integrates multiple key performance indicators, including network rate, latency, and fairness, and introduces stochastic geometry constraints, such as interference thresholds, to ensure that while pursuing performance optimization, excessive interference is avoided.

[0059] Rate: The rate of the network is affected by spectrum allocation and link mode selection, and the goal is to maximize the rate.

[0060] Power consumption: Power consumption is an important indicator in optimization. Increasing power may increase the rate but also increase power consumption. Therefore, there is a term in the reward function to penalize power consumption so that the agent learns how to reduce power consumption without sacrificing performance.

[0061] In the deep deterministic policy network environment, the reward function is the basis for the agent to optimize the policy. Considering the link outage probability, rate, and power consumption comprehensively, the expression of the reward function is: , , where, represents the reward value, represents the outage probability, represents the successful coverage rate, is the power consumption, , and are weight coefficients, which control the outage probability, maximize the rate, and reduce power consumption respectively, represents the throughput.

[0062] Step S105: Combine the Deep Deterministic Policy Network and the Proximal Policy Optimization algorithm to jointly optimize the resource allocation and power control of the communication network, and obtain a preliminary communication network; Specifically, it includes: optimizing the communication network based on the Deep Deterministic Policy Network to obtain a preliminary solution; and performing secondary solution using the Proximal Policy Optimization algorithm on the preliminary solution to obtain a preliminary communication network.

[0063] Specifically, on the basis of obtaining a preliminary solution by the Deep Deterministic Policy Gradient algorithm, the Proximal Policy Optimization algorithm file is further used for optimization and fine-tuning. The Deep Deterministic Policy Network module is responsible for the generation and optimization of action policies, and is more suitable for decision-making problems with a continuous action space; the Proximal Policy Optimization module has strong stability in the overall evolution of the policy and the valuation accuracy, and can effectively constrain the amplitude of policy updates; the two share the environmental state, policy results, evaluation information, etc. to achieve joint training and cross-optimization; this collaborative mechanism takes into account the fineness of decision-making and the stability of policy convergence, and is suitable for multi-node, dynamically changing, and delay-sensitive communication system scenarios.

[0064] Step S106: Perform local optimization on the preliminary communication network according to the optimization constraint function to obtain the target communication network; where the optimization constraint function includes link signal-to-noise ratio constraint, spectrum allocation ratio constraint, and power constraint.

[0065] Specifically, it includes: based on the preliminary communication network, using a grid search function to construct a local search network; where the grid search function is used to construct a local search network near the spectrum ratio and power optimal points; using the optimization constraint function to optimize the local search network to form the target communication network.

[0066] For example, further perform local grid search near the optimal solution determined by the Deep Deterministic Policy Gradient algorithm and the Proximal Policy Optimization algorithm to accurately optimize the network performance. Among them, the grid search function is defined as constructing a local search grid near the optimal points of the spectrum ratio and power; The optimization objective function uses a V-shaped objective function, and the expression is: , The expression of the optimization objective is: , Since the link outage probability must be lower than a certain threshold, it is required that the signal-to-noise ratio of the link is greater than a certain threshold, and the signal-to-interference ratio of the link must satisfy the following conditions. The optimization constraint function is: , , , Among them, represents the network space throughput, represents the basic data rate, represents the proportion of energy harvesting time, represents the spectrum deviation penalty coefficient, represents the current spectrum allocation ratio, represents the optimal value of spectrum allocation, represents the power deviation penalty coefficient, represents the current macro base station transmission power, represents the optimal transmission power; represents the jointly optimized spectrum allocation ratio and the macro base station transmission power to maximize the system throughput; represents the signal-to-interference-plus-noise ratio; represents the macro base station transmission power, then and respectively represent the minimum and maximum values of the macro base station transmission power. According to the optimization constraint function, perform local optimization on the preliminary communication network to obtain the target information; the target information includes the algorithm convergence curve, throughput, link outage probability, and distance relationship.

[0067] Specifically, finally, conduct a comprehensive visual display and performance verification on the optimization results, specifically including: the convergence curves of the deep deterministic policy gradient algorithm and the proximal policy optimization algorithm to show the algorithm optimization efficiency; the two-dimensional throughput surface map of the local grid search optimization results to clearly display the accurate optimal solution after optimization; the comparative analysis of the network throughput before and after optimization; the analysis of the link outage probability and distance relationship to clarify the effectiveness of the optimization scheme.

[0068] The network optimization method based on the fusion of geometry and deep learning provided by the above embodiments, on the one hand, combines analytic geometry with continuous control optimization based on the deep deterministic policy gradient algorithm. Use the Poisson point process to generate the random distribution of base stations, roadside units, and vehicles to simulate the randomness of the wireless network. Through the framework of analytic geometry, obtain the preliminary analysis results, reveal the geometric relationships between devices in the network and their impacts on communication performance, introduce the deep deterministic policy gradient algorithm, use the analytic characteristics of analytic geometry as the state input, and the deep deterministic policy network realizes the optimization of wireless network resources and power control through training and learning. Combining the policy-value evaluation structure, it can perform efficient optimization in a non-stationary dynamic environment, improve the convergence speed and policy stability. Therefore, by combining the stochastic geometry method with the deep deterministic policy gradient, the learned policy can not only provide high performance but also improve the interpretability and transparency of the policy.

[0069] Secondly, an efficient communication network optimization solution that comprehensively utilizes millimeter-wave technology, multi-antenna beamforming, and cooperative transmission technology is proposed to meet the stringent requirements of next-generation communication networks for high capacity, low latency, and large-scale coverage. The millimeter-wave band provides extremely rich spectrum resources, which can significantly improve the capacity and bandwidth of the network. However, its propagation loss is relatively high, restricting the coverage range, especially presenting significant challenges in large-scale transmission and complex environments. To address this issue, multi-antenna beamforming technology is introduced. By achieving highly directional beam transmission in high-speed transmission scenarios over short distances, it effectively compensates for the high propagation loss of millimeter waves, enhancing the reliability and transmission efficiency of the link. Meanwhile, the integration of cooperative transmission technology enables multiple communication nodes to share resources and information, further enhancing the overall performance and transmission success rate of the system. This solution can not only achieve flexible optimization in uncertain environments but also provide theoretical and technical support for the popularization of millimeter-wave technology in practical applications, promoting the efficient deployment and operation of future 5G and 6G networks.

[0070] On the other hand, a multi-objective optimization intelligent resource management strategy is designed. A reward function that comprehensively considers rate, link outage probability, and power control is designed to jointly optimize spectrum resources and macro base station power configuration, enabling the reinforcement learning intelligent agent to learn an optimization strategy that balances multiple objectives, rather than simply maximizing the rate or minimizing power consumption. Secondly, link mode adaptive selection is achieved through reinforcement learning. The optimal link mode is adaptively selected according to the real-time network state, improving the overall system rate and reducing the outage probability. Based on the initial solution obtained by the deep deterministic policy gradient algorithm, the proximal policy algorithm is further used for optimization and fine-tuning to make the optimization result more accurate. This is different from traditional network optimization methods with fixed link modes, enabling the network to flexibly adjust connection strategies in different environments and improving the adaptability of the network.

[0071] By combining stochastic geometry, deep deterministic policy gradient algorithm, proximal policy optimization algorithm, and local optimization, the performance of wireless networks can be effectively improved, the link outage probability can be reduced, and the utilization efficiency of spectrum resources can be enhanced, showing significant application prospects.

[0072] It should be noted that for the network optimization method based on the fusion of geometry and deep learning provided in the embodiments of this application, the execution subject can be a network optimization device based on the fusion of geometry and deep learning, or a control module in the network optimization device based on the fusion of geometry and deep learning for executing the loading of the network optimization method based on the fusion of geometry and deep learning. In the embodiments of this application, taking the network optimization device based on the fusion of geometry and deep learning executing the loading of the network optimization method based on the fusion of geometry and deep learning as an example, the network optimization device based on the fusion of geometry and deep learning provided in the embodiments of this application is described.

[0073] Figure 2 is a schematic diagram of a network optimization device based on the fusion of geometry and deep learning shown in the second embodiment of the present application. Please refer to Figure 2 . The network optimization device 200 based on the fusion of geometry and deep learning includes: Geometry modeling module 201: used to generate the random distribution of macro base stations, roadside units, and vehicles by using the Poisson point process to construct the topological structure of the communication network; among them, the macro base station is modeled by a two-dimensional Poisson point process, and the roadside unit and the vehicle are respectively modeled by a homogeneous spatial Poisson point process with different densities; Function module 202: used to construct the channel analysis function of the communication network according to the topological structure; Specifically include: based on the line-of-sight probability model and the segmented path loss model, establish the link blocking function and the path loss function; use the single-slope path loss model to describe the signal attenuation function; use the sector antenna model to describe the directional transmission gain function of the multi-antenna array; according to the link blocking function, the path loss function, the signal attenuation function, and the directional transmission gain function, deduce the signal-to-interference ratio function.

[0074] Analysis module 203: used to deduce the analysis function to obtain the performance analysis result of the communication network; Policy network module 204: used to construct a deep deterministic policy network with the analysis result as a constraint condition; among them, the architecture characteristics of the deep deterministic policy network include: state space, action space, and reward function; Specifically include: constructing the state space according to the analysis result; the state space includes the spectrum resource allocation ratio, the macro base station transmission power, the link mode, and the network throughput; formulating the reward function according to the analysis result; forming a deep deterministic policy network based on the state space and the reward function.

[0075] The reward function is: , , wherein, represents the reward value, represents the outage probability, represents the successful coverage rate, is the power consumption, , and are weight coefficients, which respectively control the outage probability, maximize the rate, and reduce the power consumption, represents the throughput.

[0076] Joint optimization module 205: used to combine the deep deterministic policy network and the proximal policy optimization algorithm to jointly optimize the resource allocation and power control of the communication network to obtain a preliminary communication network; Specifically, it includes: optimizing the communication network based on the deep deterministic policy network to obtain a preliminary solution; and based on the preliminary solution, using the proximal policy optimization algorithm for secondary solution to obtain a preliminary communication network.

[0077] Local optimization module 206: used to locally optimize the preliminary communication network according to the optimization constraint function to obtain the target communication network; where the optimization constraint function includes link signal-to-noise ratio constraint, spectrum allocation ratio constraint, and power constraint.

[0078] Specifically, it includes: constructing a local search network based on the preliminary communication network using a grid search function; where the grid search function is used to construct a local search network near the spectrum ratio and power optimal points; and optimizing the local search network using the optimization constraint function to form the target communication network.

[0079] Specifically, determine the optimization objective function, optimization objective, and optimization constraint function; The optimization objective function is: , The optimization objective is: , The optimization constraint function is: , , , Among them, represents the network space throughput; represents the basic data rate; represents the proportion of energy harvesting time; represents the spectrum deviation penalty coefficient; represents the current spectrum allocation ratio; represents the optimal spectrum allocation value; represents the power deviation penalty coefficient; represents the current macro base station transmission power; represents the optimal transmission power; represents the joint optimization of the spectrum allocation ratio and the macro base station transmission power to maximize the system throughput; represents the signal-to-interference-plus-noise ratio; represents the threshold value; represents the macro base station transmission power, then and respectively represent the minimum and maximum values of the macro base station transmission power.

[0080] The specific steps of this embodiment further include: locally optimizing the preliminary communication network according to the optimized constraint function to obtain target information; the target information includes an algorithm convergence curve, throughput, link interruption probability, and distance relationship.

[0081] The network optimization device based on the fusion of geometry and deep learning in the embodiments of the present application can be a device, or a component, integrated circuit, or chip in a terminal. This device can be a mobile electronic device or a non-mobile electronic device. Exemplarily, the mobile electronic device can be a mobile phone, a tablet computer, a laptop computer, a handheld computer, a vehicle-mounted electronic device, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc., and the non-mobile electronic device can be a server, a Network Attached Storage (NAS), a personal computer (PC), a television (TV), a teller machine, or a self-service machine, etc. The embodiments of the present application do not make specific limitations.

[0082] The network optimization device based on the fusion of geometry and deep learning in the embodiments of the present application can be a device with an operating system. This operating system can be the Android operating system, the iOS operating system, or other possible operating systems. The embodiments of the present application do not make specific limitations.

[0083] The network optimization device based on the fusion of geometry and deep learning provided in the embodiments of the present application can implement Figure 1 each process implemented by the network optimization method based on the fusion of geometry and deep learning in the method embodiment. To avoid repetition, it will not be elaborated here.

[0084] Optionally, please refer to Figure 3 , the embodiments of the present application further provide an electronic device 300, including a processor 301, a memory 302, and a computer program 303 stored on the memory 302 and executable on the processor 301. When the computer program 303 is executed by the processor 301, it implements each process of the above-mentioned network optimization method embodiment based on the fusion of geometry and deep learning, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.

[0085] An embodiment of the present application also provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, each process of the above embodiment of the network optimization method based on the fusion of geometry and deep learning is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be elaborated here.

[0086] Among them, the processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes a computer-readable storage medium, such as a computer read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc.

[0087] Another embodiment of the present application provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run a program or instruction to implement each process of the above embodiment of the network optimization method based on the fusion of geometry and deep learning, and the same technical effect can be achieved. To avoid repetition, it will not be elaborated here.

[0088] It should be understood that the chip mentioned in the embodiment of the present application may also be referred to as a system-on-chip, a system chip, a chip system, or a system-on-chip.

[0089] It should be noted that in this article, the term "including", "comprising", or any other variant thereof is intended to cover a non-exclusive inclusion, such that a process, method, article, or device including a series of elements includes not only those elements but also other elements not expressly listed, or further includes elements inherent to such process, method, article, or device. Without further limitation, an element defined by the phrase "including a..." does not exclude the presence of additional identical elements in the process, method, article, or device including that element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in a reverse order according to the functions involved. For example, the described methods may be performed in an order different from that described, and various steps may be added, omitted, or combined. Additionally, the features described with reference to certain examples may be combined in other examples.

[0090] Through the description of the above embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in various embodiments of the present application.

[0091] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative rather than restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them belong to the protection scope of the present application.

Claims

1. A network optimization method based on the fusion of geometry and deep learning, characterized in that: include: The Poisson point process is used to generate the random distribution of macro base stations, roadside units and vehicles to construct the topological structure of the communication network; the macro base stations are modeled by a two-dimensional Poisson point process, and the roadside units and vehicles are modeled by uniform spatial Poisson point processes with different densities. Constructing a channel analysis function of the communication network according to the topological structure; Derivation of the analytical function to obtain a performance analysis result of the communication network; Based on the analysis result, a deep deterministic policy network is constructed; wherein the architecture features of the deep deterministic policy network include: state space, action space and reward function; Combining the deep deterministic policy network and the proximal policy optimization algorithm, jointly optimizing resource allocation and power control of the communication network, and obtaining a preliminary communication network; The preliminary communication network is locally optimized according to the optimization restriction function to obtain a target communication network; wherein the optimization restriction function includes link signal-to-noise ratio restriction, spectrum allocation ratio restriction and power restriction.

2. The network optimization method based on the fusion of geometry and deep learning according to claim 1 is characterized in that: The specific steps of constructing the channel resolution function of the communication network according to the topological structure include: Based on the line-of-sight probability model and the segmented path loss model, the link blocking function and path loss function are established; A single slope path loss model is used to describe the signal attenuation function; The directional transmission gain function of multi-antenna array is described by sector antenna model; The signal-to-interference ratio function is derived based on the link blocking function, path loss function, signal attenuation function and directional transmission gain function.

3. The network optimization method based on the fusion of geometry and deep learning according to claim 1 is characterized in that: The specific steps of constructing a deep deterministic strategy network using the analysis results as constraints include: Constructing a state space according to the analysis results; the state space includes spectrum resource allocation ratio, macro base station transmission power, link mode and network throughput; Formulate a reward function according to the analysis result; The deep deterministic policy network is formed based on the state space and the reward function.

4. The network optimization method based on the fusion of geometry and deep learning according to claim 3 is characterized in that: The reward function is: , , in, Represents the reward value, represents the interruption probability, represents the successful coverage rate, To consume power, , and are weight coefficients, which respectively control the interruption probability, maximize the rate and reduce the power consumption, Indicates throughput.

5. The network optimization method based on the fusion of geometry and deep learning according to claim 1 is characterized in that: The specific steps of combining the deep deterministic policy network and the proximal policy optimization algorithm to jointly optimize the resource allocation and power control of the communication network to obtain a preliminary communication network include: Optimizing the communication network based on the deep deterministic strategy network to obtain a preliminary solution; Based on the preliminary solution, a proximal strategy optimization algorithm is used to perform a secondary solution to obtain a preliminary communication network.

6. The network optimization method based on the fusion of geometry and deep learning according to claim 1 is characterized in that: The specific steps of locally optimizing the preliminary communication network according to the optimization restriction function to obtain the target communication network include: Based on the preliminary communication network, a local search network is constructed using a grid search function; wherein the grid search function is used to construct a local search network near the optimal point of spectrum ratio and power; The local search network is optimized using the optimization restriction function to form the target communication network.

7. The network optimization method based on the fusion of geometry and deep learning according to claim 6 is characterized in that: The specific steps of optimizing the local search network by using the optimization restriction function to form a target communication network include: Determining an optimization objective function, an optimization target, and the optimization restriction function; The optimization objective function is: , The optimization goal is: , The optimization restriction function is: , , , in, represents the network space throughput; Indicates the basic data rate; Indicates the proportion of energy harvesting time; represents the spectrum deviation penalty coefficient; Indicates the current spectrum allocation ratio; Indicates the optimal value of spectrum allocation; represents the power deviation penalty coefficient; Indicates the current macro base station transmit power; represents the optimal transmit power; Represents the joint optimization spectrum allocation ratio Transmit power of macro base station Maximizing system throughput; It represents the signal-to-interference ratio; Indicates the threshold value; represents the macro base station transmit power, then and They respectively represent the minimum and maximum values ​​of the macro base station transmit power.

8. The network optimization method based on the fusion of geometry and deep learning according to claim 1 is characterized in that: The specific steps also include: The preliminary communication network is locally optimized according to the optimization restriction function to obtain target information; the target information includes an algorithm convergence curve, throughput, link interruption probability and distance relationship.

9. A network optimization device based on the fusion of geometry and deep learning, used to execute the network optimization method based on the fusion of geometry and deep learning as described in any one of claims 1 to 8, characterized in that: include: A geometric modeling module is used to generate a random distribution of macro base stations, roadside units and vehicles using a Poisson point process to construct a topological structure of a communication network; wherein the macro base stations are modeled using a two-dimensional Poisson point process, and the roadside units and vehicles are modeled using uniform spatial Poisson point processes with different densities; A function module, used for constructing a channel analysis function of the communication network according to the topological structure; An analysis module, used for deriving the analysis function to obtain a performance analysis result of the communication network; A policy network module, used to construct a deep deterministic policy network with the analysis results as constraints; wherein the architecture features of the deep deterministic policy network include: state space, action space and reward function; A joint optimization module, used to combine the deep deterministic policy network and the proximal policy optimization algorithm to jointly optimize the resource allocation and power control of the communication network to obtain a preliminary communication network; A local optimization module is used to locally optimize the preliminary communication network according to an optimization restriction function to obtain a target communication network; wherein the optimization restriction function includes link signal-to-noise ratio restriction, spectrum allocation ratio restriction and power restriction.

10. An electronic device, characterized in that: include: A memory, a processor, and a program or instruction stored in the memory and executable on the processor, wherein the program or instruction, when executed by the processor, implements the steps of the network optimization method based on the fusion of geometry and deep learning as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Deep reinforcement learning communication interference resource allocation method fused with noise network

    CN115866760A

  • Heterogeneous network resource energy efficiency optimization method based on deep reinforcement learning

    CN116567667A

  • Internet-of-vehicles wireless resource allocation method based on enhanced dual-depth Q network

    CN117412391A

  • Heterogeneous wireless network transparent coexistence method based on multi-agent reinforcement learning

    CN118102345A

  • Resource allocation method based on internal-external circulation and hierarchical reinforcement learning

    CN119854866A