Deployment method, system and equipment of 5G mobile access network of power distribution network and medium
Through multi-agent interactive learning, the optimal deployment strategy is obtained, which solves the problem of deployment cost and transmission delay of 5G mobile access network in the distribution network, and realizes the low latency and high reliability requirements of distribution network services.
Patent Information
- Application Number
- CN202510083206.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-06-06
AI Technical Summary
The existing distribution network communication mechanism cannot meet the high requirements of distribution network protection services for low latency, reliability and cost-effectiveness. The transmission delay and reliability of 5G public networks need to be improved, and the operating costs are relatively high.
Through interactive learning between multiple agents and the environment of the distribution network 5G mobile access network, the optimal deployment strategy is obtained, including the optimization goal of minimizing deployment costs and transmission delay, configuring the deployment of the distribution network 5G mobile access network in the distribution network in the agent, finding work paths and protecting backup paths.
It has achieved the real-time requirement of distribution network services for real-time.
Smart Images

Figure CN120111527A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of network deployment technology, and in particular to a deployment method, system, device and medium for a 5G mobile access network of a distribution network. Background Art
[0002] With the acceleration of the construction of new power systems and the upgrading of power grids, higher requirements will be placed on distribution network protection and communication carrying capacity. The existing distribution network communication mechanism can no longer meet the high requirements of distribution network protection services for low latency and reliability. For existing optical technologies in distribution network protection, optical communication costs are high and network deployment is complex. For 5G communication technology, the transmission latency and reliability of 5G public networks need to be improved, and the operating costs of 5G are high, which limits its promotion and application in actual projects.
[0003] However, the distribution network involves a large number of power equipment, data transmission and monitoring needs, and has high requirements for network reliability, security and real-time performance. When deploying 5G mobile access networks, it is necessary to consider both cost-effectiveness and low latency. Cost-effectiveness requires that the network deployment be economical, while low latency requires that the network be able to quickly respond to the real-time needs of distribution network protection services. Under the joint goals of cost-effectiveness and low latency, the reliable deployment of 5G mobile access networks in distribution networks involves multiple key factors such as cost optimization, network reliability, and deployment efficiency. Therefore, existing methods are difficult to solve the problem of reliable deployment with cost-effectiveness and low latency as the joint optimization goals. Summary of the invention
[0004] Based on the technical problem to be solved, the present invention provides a deployment method, system, equipment and medium for a 5G mobile access network of a distribution network, which can reduce network construction costs and network transmission delays, improve network response speed, and meet the real-time requirements of distribution network services.
[0005] In order to solve the above technical problems, an embodiment of the present invention provides a method for deploying a 5G mobile access network of a distribution network, including:
[0006] Taking minimizing the deployment cost and transmission delay of the distribution network 5G mobile access network as the optimization goal, the optimal deployment strategy is obtained by interactive learning between multiple intelligent agents and the environment of the distribution network 5G mobile access network;
[0007] According to the optimal deployment strategy, the deployment of the distribution network 5G mobile access network is configured in each intelligent agent; wherein, the configuration of the deployment of the distribution network 5G mobile access network includes finding the working path and the protection backup path of the distribution network 5G mobile access network.
[0008] As an improvement of the above scheme, the optimization goal is to minimize the deployment cost and transmission delay of the distribution network 5G mobile access network, and to obtain the optimal deployment strategy through interactive learning between multiple intelligent agents and the environment of the distribution network 5G mobile access network, including:
[0009] For each agent, the optimization goal of minimizing the deployment cost and transmission delay of the 5G mobile access network of the distribution network is expressed as a Markov decision process, which is represented by a four-tuple entry of <current state, action, reward, next state>; where,
[0010] Determine the state of the agent based on the remaining computing resource capacity of the equipment in the 5G mobile access network of the distribution network and the connection status of all candidate sites;
[0011] Determine the action of the agent according to the candidate site locations of the 5G mobile access network of the distribution network, the order of connection of active antenna units, spectrum resource allocation, and the isolation between the working path and the backup protection path;
[0012] Determine a reward function of the agent based on the deployment cost and transmission delay of the 5G mobile access network of the distribution network;
[0013] The Q value set of all actions of the agent is obtained through the deep Q network, and the action with the largest value is selected from the Q value set to obtain the optimal deployment strategy.
[0014] As an improvement of the above scheme, the reward function of the agent is:
[0015]
[0016] in, represents the expected reward at the current distribution network protection service request i and the current time t; represents the deployment cost at the current distribution network protection service request i and the current time t, and α is the weight coefficient of the deployment cost; represents the transmission delay between the current distribution network protection service request i and the current time t, β is the weight coefficient of the transmission delay; if isvalid means that under the current distribution network protection service request i and the current time t, the action a is valid. is notvalid means that under the current distribution network protection service request i and the current time t, the action a is invalid; A represents the device set of the distribution network 5G mobile access network; A Boolean variable representing the connection status of device m and device n at time t. If device m and device n are connected, then otherwise d m,nIndicates the distance between device m and device n; c m,n represents the cost per unit length of optical fiber between device m and device n; c m It represents the average cost per unit length of optical fiber from device m to all other devices; A Boolean variable indicating whether device m is occupied at time t. If device m is occupied, then otherwise τ m,n represents the propagation delay per kilometer of optical fiber between device m and device n; τ m Indicates the processing delay of device m.
[0017] As an improvement of the above solution, the Q value set of all actions of the agent is obtained through the deep Q network, and the action with the largest value is selected from the Q value set to obtain the optimal deployment strategy, including:
[0018] Input the current distribution network protection service request and the current network status into the deep Q network, calculate the Q values corresponding to all actions under the current network status, and obtain the Q value set of all actions;
[0019] The action with the largest value is selected from the Q value set to obtain the optimal deployment strategy.
[0020] As an improvement of the above scheme, the deep Q network includes a multi-level attention-enhanced convolutional neural network and a multi-layer fully connected network;
[0021] The multi-layer fully connected network includes a multi-head attention layer.
[0022] As an improvement of the above solution, after configuring the deployment of the distribution network 5G mobile access network in each intelligent agent according to the optimal deployment strategy, it also includes:
[0023] Observe the network status at the next moment and obtain corresponding rewards;
[0024] Store the four-tuple entry of <current state, action, reward, next state> into the experience replay cache;
[0025] Data is randomly sampled periodically from the experience replay buffer to train the multi-level attention-enhanced convolutional neural network to update parameters of the multi-level attention-enhanced convolutional neural network.
[0026] As an improvement of the above scheme, the Bellman optimal equation is used to update the parameters of the multi-level attention-enhanced convolutional neural network.
[0027] In order to solve the above technical problems, an embodiment of the present invention further provides a deployment system of a 5G mobile access network of a distribution network, including:
[0028] A deployment strategy acquisition module is used to minimize the deployment cost and transmission delay of the distribution network 5G mobile access network as the optimization goal, and obtain the optimal deployment strategy through interactive learning between multiple intelligent agents and the environment of the distribution network 5G mobile access network;
[0029] A configuration deployment module is used to configure the deployment of the distribution network 5G mobile access network in each intelligent agent according to the optimal deployment strategy; wherein, the configuration of the deployment of the distribution network 5G mobile access network includes finding the working path and protection backup path of the distribution network 5G mobile access network.
[0030] As an improvement of the above solution, the deployment strategy acquisition module includes:
[0031] The optimization target expression unit is used to express the optimization target of minimizing the deployment cost and transmission delay of the 5G mobile access network of the distribution network as a Markov decision process for each intelligent agent, which is expressed as a four-tuple entry of <current state, action, reward, next state>; wherein,
[0032] An intelligent agent state determination unit, configured to determine the state of the intelligent agent according to the remaining computing resource capacity of the equipment in the 5G mobile access network of the distribution network and the connection state of all candidate sites;
[0033] An intelligent agent action determination unit, used to determine the action of the intelligent agent according to the candidate site location of the distribution network 5G mobile access network, the order of connection of the active antenna unit, the spectrum resource allocation, and the isolation between the working path and the backup protection path;
[0034] An agent reward function determination unit, used to determine the agent's reward function according to the deployment cost and transmission delay of the distribution network 5G mobile access network;
[0035] The deployment strategy acquisition unit is used to obtain the Q value set of all actions of the intelligent agent through a deep Q network, select the action with the largest value from the Q value set, and obtain the optimal deployment strategy.
[0036] To solve the above technical problems, an embodiment of the present invention further provides a terminal device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein when the processor executes the computer program, it implements the deployment method of the distribution network 5G mobile access network as described in any one of the above items.
[0037] To solve the above technical problems, an embodiment of the present invention further provides a computer-readable storage medium, wherein the computer-readable storage medium includes a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute the deployment method of the distribution network 5G mobile access network as described in any one of the above.
[0038] Compared with the prior art, the embodiments of the present invention provide a method, system, device and medium for deploying a 5G mobile access network of a distribution network, wherein the method takes minimizing the deployment cost and transmission delay of the 5G mobile access network of the distribution network as the optimization goal, and obtains the optimal deployment strategy through interactive learning between multiple intelligent agents and the environment of the 5G mobile access network of the distribution network, and then configures the deployment of the 5G mobile access network of the distribution network in each intelligent agent according to the optimal deployment strategy; wherein the configuration of the deployment of the 5G mobile access network of the distribution network includes finding the working path and the protection backup path of the 5G mobile access network of the distribution network. The present invention can search for the optimal solution in a global scope through the collaboration and competition between multiple intelligent agents, not only considering the current state, but also optimizing future decisions through the long-term accumulated reward value, thereby achieving global optimization, and showing higher flexibility. With the help of the reinforcement learning optimization method enhanced by the multi-head attention mechanism, it can intelligently and efficiently solve the network deployment optimization problem of deployment cost and transmission delay. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] In order to more clearly illustrate the technical solution of the present invention, the drawings used in the implementation mode will be briefly introduced below. Obviously, the drawings described below are only some implementation modes of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0040] Figure 1 It is a flow chart of a method for deploying a 5G mobile access network of a distribution network provided by an embodiment of the present invention;
[0041] Figure 2 This is a schematic diagram of a 5G mobile access network architecture for a distribution network provided by an embodiment of the present invention;
[0042] Figure 3 It is another flow chart of a method for deploying a 5G mobile access network of a distribution network provided by an embodiment of the present invention;
[0043] Figure 4 It is a structural block diagram of a deployment system of a 5G mobile access network for a distribution network provided by an embodiment of the present invention;
[0044] Figure 5 It is a structural block diagram of a terminal device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0045] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0046] See also Figure 1 , Figure 1 It is a flow chart of a method for deploying a 5G mobile access network of a distribution network provided by an embodiment of the present invention, and the method for deploying a 5G mobile access network of a distribution network comprises steps S1 to S2:
[0047] S1. Taking minimizing the deployment cost and transmission delay of the distribution network 5G mobile access network as the optimization goal, the optimal deployment strategy is obtained by interactive learning between multiple intelligent agents and the environment of the distribution network 5G mobile access network;
[0048] S2. According to the optimal deployment strategy, configure the deployment of the distribution network 5G mobile access network in each intelligent agent; wherein, the configuration of the deployment of the distribution network 5G mobile access network includes finding the working path and the protection backup path of the distribution network 5G mobile access network.
[0049] It is worth noting that in order to minimize the deployment cost and transmission delay of the 5G mobile access network of the distribution network, the embodiment of the present invention adopts multi-agent reinforcement learning technology. Through multiple agents, interactive learning is carried out between the environment of the 5G mobile access network of the distribution network. Each agent can be understood as a decision-making unit to deploy a part of the 5G network. Through repeated attempts and environmental feedback, the agent continuously learns and adjusts its deployment strategy, and finally finds a globally optimal deployment plan, which can minimize the deployment cost and transmission delay while meeting the network coverage and performance requirements.
[0050] In an optional implementation, the optimization goal is to minimize the deployment cost and transmission delay of the distribution network 5G mobile access network, and to obtain the optimal deployment strategy by interactive learning between multiple intelligent agents and the environment of the distribution network 5G mobile access network, including:
[0051] For each agent, the optimization goal of minimizing the deployment cost and transmission delay of the 5G mobile access network of the distribution network is expressed as a Markov decision process, which is represented by a four-tuple entry of <current state, action, reward, next state>; where,
[0052] Determine the state of the agent based on the remaining computing resource capacity of the equipment in the 5G mobile access network of the distribution network and the connection status of all candidate sites;
[0053] Determine the action of the agent according to the candidate site locations of the 5G mobile access network of the distribution network, the order of connection of active antenna units, spectrum resource allocation, and the isolation between the working path and the backup protection path;
[0054] Determine a reward function of the agent based on the deployment cost and transmission delay of the 5G mobile access network of the distribution network;
[0055] The Q value set of all actions of the agent is obtained through the deep Q network, and the action with the largest value is selected from the Q value set to obtain the optimal deployment strategy.
[0056] It should be noted that the process of multiple intelligent agents interacting and learning with the environment of the distribution network 5G mobile access network to obtain the optimal deployment strategy adopts the Markov Decision Process (MDP) framework in reinforcement learning. For each intelligent agent, its learning process is modeled as an MDP, which consists of the following four elements:
[0057] State: The state of the agent is used to reflect the current environmental conditions of the 5G mobile access network of the distribution network. It can be based on the remaining computing resource capacity of the equipment, such as the existing equipment in the distribution network, the remaining computing resource capacity on the corresponding multi-access edge computing server, etc., as well as the connection status of all candidate sites that can be used to deploy 5G devices, including whether the devices are connected.
[0058] Action: The action of the agent represents the deployment operation it takes in the environment. The components of the action include selecting candidate site locations; the order of connecting active antenna units (AAUs). If multiple AAUs are deployed, their connection order is determined to optimize signal coverage and interference control; appropriate spectrum resources are allocated to the network to avoid interference and maximize network capacity; the optimal working path and redundant backup protection paths are selected for the network to improve network reliability. Among them, path selection needs to consider path length, bandwidth, reliability, and isolation between paths (for example, physical distance or isolation in network topology) to avoid single point failures.
[0059] Reward: The reward function is used to evaluate the quality of the actions taken by the agent. The design goal of the reward function in the embodiment of the present invention is to guide the agent to learn the strategy of minimizing the deployment cost and transmission delay. Among them, the deployment can take into account the hard costs of MEC servers, central units, distribution units, optical line terminals, splitters and optical cables, and the soft costs of labor, trenching and laying optical cables; the transmission delay can take into account the data transmission delay, waiting delay, propagation delay on the optical path, and processing delay on the MEC server and central unit, distribution unit, and optical line terminal in the active antenna unit and optical network unit.
[0060] State at the next moment: After the intelligent agent takes action, the environmental state of the distribution network 5G mobile access network will change and form a new state.
[0061] Then, through the Deep Q-Network (DQN) algorithm, each agent learns and constructs a Q-value function, which maps states and actions to corresponding Q-values, representing the expected cumulative reward for taking the action in that state. At each time step, the agent selects the action with the largest Q-value, which is the greedy strategy. By continuously interacting with the environment, obtaining rewards and updating the Q-value function, the agent will eventually learn an optimal deployment strategy that minimizes the deployment cost and transmission delay of the 5G mobile access network of the distribution network. Finally, multiple agents work together to complete the optimized deployment of the 5G mobile access network of the entire distribution network.
[0062] Specifically, the reward function of the agent is:
[0063]
[0064] in, represents the expected reward at the current distribution network protection service request i and the current time t; represents the deployment cost at the current distribution network protection service request i and the current time t, and α is the weight coefficient of the deployment cost; represents the transmission delay between the current distribution network protection service request i and the current time t, β is the weight coefficient of the transmission delay; if isvalid means that under the current distribution network protection service request i and the current time t, the action a is valid. is notvalid means that under the current distribution network protection service request i and the current time t, the action a is invalid; A represents the device set of the distribution network 5G mobile access network; A Boolean variable representing the connection status of device m and device n at time t. If device m and device n are connected, then otherwise dm,n Indicates the distance between device m and device n; c m,n represents the cost per unit length of optical fiber between device m and device n; c m It represents the average cost per unit length of optical fiber from device m to all other devices; A Boolean variable indicating whether device m is occupied at time t. If device m is occupied, then otherwise τ m,n represents the propagation delay per kilometer of optical fiber between device m and device n; τ m Indicates the processing delay of device m.
[0065] Furthermore, the Q value set of all actions of the agent is obtained through the deep Q network, and the action with the largest value is selected from the Q value set to obtain the optimal deployment strategy, including:
[0066] Input the current distribution network protection service request and the current network status into the deep Q network, calculate the Q values corresponding to all actions under the current network status, and obtain the Q value set of all actions;
[0067] The action with the largest value is selected from the Q value set to obtain the optimal deployment strategy.
[0068] Preferably, the deep Q network includes a multi-level attention-enhanced convolutional neural network and a multi-layer fully connected network; the multi-layer fully connected network includes a multi-head attention layer.
[0069] Exemplarily, the protection service request information of the current distribution network and the network status information at the current moment are used as input data and sent to the deep Q network. The current distribution network protection service request information includes but is not limited to: service type, service level, service demand, service area, etc.; the network status information at the current moment includes the agent status information mentioned in the above steps, such as: the connection status of all candidate sites, the remaining computing resource capacity of the equipment, etc., and other parameters reflecting the real-time operation status of the network. After receiving the input data, the deep Q network calculates the Q value of all possible actions in the current network state through its internal neural network structure. The output of the deep Q network is a Q value set, which contains the Q values corresponding to all possible actions. The action with the largest Q value is selected from the Q value set as the optimal action, which corresponds to the action with the highest expected benefit in the current state. The process of selecting the optimal action can be regarded as a greedy strategy, that is, selecting the action that seems to be the best at the moment. In order to balance exploration and utilization, the ε-greedy strategy or other more complex strategies can also be adopted to select non-optimal actions with a certain probability to avoid falling into the local optimal solution.
[0070] It is worth noting that in order to improve the performance and learning efficiency of the deep Q network, the deep Q network adopts a combination structure of a multi-level attention-enhanced convolutional neural network and a multi-layer fully connected network. The convolutional neural network (CNN) is good at processing images and spatial data, and can effectively extract spatial features from the input data, such as the location relationship of candidate sites and network topology information. By adding a multi-level attention mechanism, important network areas and key information can be highlighted, and the learning efficiency and expression ability of the network can be improved. Especially when processing high-dimensional data, the multi-level attention mechanism can focus on different levels of features, for example, the low level focuses on local connection information, and the high level focuses on global network topology information. The fully connected network (FCN) is good at processing abstract features, and can map the spatial features extracted by the convolutional neural network to the action space, and finally output the Q value corresponding to each action. The multi-layer fully connected network can learn more complex nonlinear relationships and improve the expression ability of the model. After using the convolutional neural network to extract local features and spatial features, the fully connected layer is responsible for mapping these features to the final Q value output. However, the fully connected layer is usually a simple linear transformation, which may not fully capture the complex interactive relationship between different features. Inserting a multi-head attention layer between the fully connected layers can solve this problem. This combined structure makes full use of the respective advantages of convolutional neural networks and fully connected networks, enabling the deep Q network to more effectively handle complex distribution network 5G mobile access network deployment problems, thereby obtaining a better deployment strategy.
[0071] In an optional implementation, after configuring the deployment of the distribution network 5G mobile access network in each intelligent agent according to the optimal deployment strategy, it also includes:
[0072] Observe the network status at the next moment and obtain corresponding rewards;
[0073] Store the four-tuple entry of <current state, action, reward, next state> into the experience replay cache;
[0074] Data is randomly sampled periodically from the experience replay buffer to train the multi-level attention-enhanced convolutional neural network to update parameters of the multi-level attention-enhanced convolutional neural network.
[0075] Preferably, the Bellman optimal equation is used to update the parameters of the multi-level attention-enhanced convolutional neural network.
[0076] It is worth noting that in order to continuously learn and optimize the deployment strategy, after the agent executes the selected action (deployment strategy), it can continue to observe the network state at the next moment, and calculate the corresponding reward based on the network state at the next moment, and then form a four-tuple <current state, action, reward, next state> with the current state, the action executed, the reward obtained, and the state at the next moment, and store it in the experience replay buffer. The experience replay buffer is a queue or buffer for storing historical experience data, which is used to randomly sample data during the training process to break the correlation between data and improve training efficiency and stability. A certain number of four-tuple data are randomly sampled from the experience replay buffer regularly to train the multi-level attention-enhanced convolutional neural network. This part of the process uses the deep reinforcement learning algorithm. The goal of the training process is to update the parameters of the deep Q network so that it can better estimate the Q value under different states and actions. During the training process, the parameter update uses the Bellman Optimality Equation to calculate the target Q value. The core idea of the Bellman Optimality Equation is to backpropagate the value of the future state to the current state, thereby guiding the action selection in the current state. This process is iterated continuously until the deep Q network converges and achieves the expected performance. This enables the agent to continuously learn and improve its deployment strategy, ultimately achieving intelligent and optimized deployment of the 5G mobile access network for the distribution network.
[0077] In order to enable those skilled in the art to more clearly understand the deployment method of a 5G mobile access network for a distribution network provided in an embodiment of the present invention, a 5G mobile access network architecture for a distribution network is provided below to implement the deployment method of a 5G mobile access network for a distribution network provided in an embodiment of the present invention.
[0078] See also Figure 2 , Figure 2 is a schematic diagram of a 5G mobile access network architecture of a distribution network provided by an embodiment of the present invention, such as Figure 2 As shown, the distribution network 5G mobile access network architecture includes a multi-access edge computing server MEC, a central unit CU, a distribution unit DU, an optical line terminal OLT, a splitter Splitter, a number of optical network units ONU1~ONU3 and a number of active antenna units AAU1~AAU3;
[0079] The central unit CU and the distribution unit DU are located at the same position of the distribution network 5G mobile access network, and a central unit-distribution unit CU-DU is obtained. The optical line terminal OLT and the central unit-distribution unit CU-DU are located at the same position of the distribution network 5G mobile access network, and a central unit-distribution unit-optical line terminal CUDU-OLT is obtained;
[0080] The optical line terminal OLT is connected to the splitter Splitter through optical fiber, and the splitter Splitter is connected to each optical network unit ONU through optical fiber;
[0081] Each of the optical network units ONU and each of the active antenna units AAU are respectively located at the same position of the distribution network 5G mobile access network, and a plurality of optical network units-active antenna units ONU-AAU (ONU1-AAU1~ONU3-AAU3) are obtained;
[0082] The central unit CU-distributed unit DU is connected to each active antenna unit AAU through a fronthaul network, and the central unit-distributed unit CU-DU is connected to the multi-access edge computing server MEC through a backhaul network.
[0083] It should be noted that Multi-access Edge Computing (MEC) refers to a server that deploys computing, storage, and network functions at the edge of the wireless access network, closer to the user device. Compared with traditional cloud computing, MEC transfers computing resources from data centers to edge nodes closer to users, thereby significantly reducing latency, improving bandwidth efficiency, and supporting new applications and services. "Multi-access" means that MEC servers can support multiple wireless access technologies, including 5G.
[0084] The Centralized Unit (CU), also called the central unit, is responsible for handling centralized control and management functions so that it can be centrally placed in the core network.
[0085] The distributed unit (DU) is responsible for handling functions directly related to wireless access, so it can be placed at the edge close to the wireless access point, helping to reduce latency and increase network capacity.
[0086] The Optical Line Terminal (OLT) is a crucial device in the fiber optic access network. It is located in the central office or access point of the network and is responsible for connecting and managing multiple optical network units (ONUs) through the fiber optic network.
[0087] An Optical Network Unit (ONU) is a fiber-optic access network device located at the user side, connected to a splitter and communicating with an optical line terminal (OLT) to provide users with various network services.
[0088] A splitter, also often called an optical splitter in a fiber optic network, is a passive optical device used to split an optical signal on an optical fiber into multiple paths, or to combine multiple optical signals into one path.
[0089] An active antenna unit (AAU) is an antenna unit that integrates a radio frequency (RF) transceiver, a power amplifier, and other electronic components. It integrates the radio frequency unit (RRU) and antenna in a traditional radio base station, which can effectively improve network coverage and capacity.
[0090] It should be noted that a 5G mobile access network architecture for a distribution network provided in an embodiment of the present invention is a distribution communication network architecture of a 5G mobile access network carried by a time-division multiplexing passive optical network (TDM-PON), that is, a 5G RAN (Radio Access Network) based on TDM-PON. The embodiment of the present invention assumes that the CU and DU are centrally deployed in the same candidate site of the network. This deployment method can reduce the propagation delay on the link; then the CU-DU is interconnected with the AAU and MEC server through the fronthaul and backhaul networks respectively, the ONU and the AAU are located at the same location, and the OLT and the CU-DU are located at the same location. Then, the end-to-end transmission of the working path and the backup protection path of the distribution network protection service facing the network is, starting from ONU-AAU, passing through the splitter Splitter, CUDU-OLT, and finally transmitted to the MEC server, wherein the working path and the backup protection path of the distribution network protection service are link-incompatible. This two-stage RAN architecture is the preferred network deployment scenario, which not only ensures the actual cost-effectiveness of the distribution network, but also ensures low-latency transmission of the distribution network.
[0091] In the specific implementation, an intelligent agent is deployed at each ONU-AAU. Each intelligent agent learns to improve its own strategy to obtain rewards by interacting with the environment of the distribution network 5G mobile access network, so as to obtain the optimal strategy in the environment. Among them, multiple agents execute actions in parallel to achieve the global optimization goal of minimizing the joint reliable deployment cost and transmission delay. That is, each intelligent agent tries to find the working path from ONU-AAU to the MEC server and its backup protection path.
[0092] Exemplarily, for each agent, cost-effectiveness and low latency are used as guidance, and the deployment problem is formulated as a Markov decision process, expressed as a four-tuple of <current state, reward, action, next state>. For the state of each agent, it contains the remaining computing resource capacity on the corresponding MEC server and CUDU-OLT, as well as the connection status of all candidate sites; for the action of each agent, since the deployment on each ONU-AAU comes from the protection service of the distribution network, the reliable deployment of the distribution network protection service can be composed of several sub-problems: (1) How to select the candidate site locations of the splitter, CUDU-OLT and MEC server in the order of the functional service chain; (2) How to determine the order of connection of multiple AAUs; (3) How to allocate spectrum resources for the link; (4) How to find a balance between maximizing the isolation between the working path and the backup protection path and minimizing the network deployment cost and transmission delay. Each sub-problem is contained in the action of the agent, and the specific action selection is determined by the maximum reward expectation obtained by selecting different actions for different states, as shown in the formula:
[0093]
[0094] in, represents the expected reward at the current distribution network protection service request i and the current time t; represents the deployment cost at the current distribution network protection service request i and the current time t, and α is the weight coefficient of the deployment cost; represents the transmission delay between the current distribution network protection service request i and the current time t, β is the weight coefficient of the transmission delay; if isvalid means that under the current distribution network protection service request i and the current time t, the action a is valid. is notvalid means that under the current distribution network protection service request i and the current time t, the action a is invalid; A AAU ,B SPlitter ,C cudu ,D mec Represents the device set of AAU, Splitter, CUDU-OLT and MEC in the 5G mobile access network of the distribution network; A Boolean variable representing the connection status between device m and device n at the current time t. If device m and device n are connected, then otherwise d m,n Indicates the distance between device m and device n; c m,n represents the cost per unit length of optical fiber between device m and device n; c m It represents the average cost per unit length of optical fiber from device m to all other devices; A Boolean variable indicating whether device m is occupied at the current time t. If device m is occupied, then otherwise τ m,n represents the propagation delay per kilometer of optical fiber between device m and device n; τ m Indicates the processing delay of device m.
[0095] The strategies for these sub-problems are solved by a deep reinforcement learning algorithm enhanced by a multi-head attention mechanism. The Q value of each agent represents the state-action function obtained through deep Q-Network (DQN) training.
[0096] For example, see Figure 3 , Figure 3 Another flow chart of a method for deploying a 5G mobile access network in a distribution network provided by an embodiment of the present invention is as follows: Figure 3 As shown, the deployment method includes the following steps:
[0097] It should be noted that the parameters of the 5G mobile access network of the distribution network can be initialized first, including the link occupancy status, the occupancy status of AAU, Splitter, CUDU-OLT, and MEC devices, the deployment cost storage parameters, and the transmission delay storage parameters;
[0098] Step 1: Collect the distribution network protection service requests on all ONU-AAUs and the network status S at the current time t t Input into the multi-head attention-enhanced deep Q network (DQN), which consists of a multi-level attention-enhanced convolutional neural network (MHA-CNN) and a multi-layer fully connected network (FC);
[0099] Step 2: CNN calculates the current distribution network protection service request i and the network status at the current time t. Calculate each action in this The corresponding Q value is obtained, and the Q value set of all actions is output Here represents the expected reward obtained by selecting action a under the current distribution network protection service request i and the current time t;
[0100] Step 3: From the Q value set Select the action with the largest value Execute, and then perform reliable deployment in multi-agents, that is, find the working path of each AAU through Splitter, CUDU-OLT and MEC and the backup protection path of each AAU through Splitter, CUDU-OLT and MEC that is incompatible with the working path link;
[0101] Step 4: Execute After the action, the new network state is observed And get the corresponding reward value
[0102] Step 5: The four-tuple entry is stored in the experience replay cache;
[0103] Step 6: Periodically randomly sample a batch of data from the experience replay buffer;
[0104] Step 7: Use the sampled data to train the CNN. During training, use the Bellman optimization equation to iteratively update the parameters of the CNN so that it can more accurately predict the expected rewards of various actions. If the Bellman error converges after updating the DNN parameters, it means that the DNN can fit the value of the action well. Repeat the above steps until the Q value function converges.
[0105] It should be noted that when using the DQN model, the current connection status of the splitter, CU-DU-OLT and MEC server, and the distribution network protection service on the ONU-AAU can be input into the model. The whole process is a continuous cycle of reinforcement learning. The intelligent agent continuously interacts with the environment, observes feedback and optimizes the DQN model, ultimately achieving the goal of efficient and reliable deployment of the 5G mobile access network in the distribution network.
[0106] In summary, the embodiments of the present invention provide a method for deploying a 5G mobile access network for a distribution network, aiming at the limitations of traditional heuristic algorithms in solving the deployment problem of distribution network protection services in large-scale distribution network communication networks. Through the deep reinforcement learning framework, the network deployment optimization problem of deployment cost and transmission delay can be solved intelligently and efficiently; the neural network is used to approximate the Q function, thereby solving the problem of traditional Q learning in high-dimensional state space. In contrast, heuristic algorithms often rely on intuition or experience and may not be able to find the optimal solution or approximate optimal solution within a given computational cost. In addition, the deployment of multiple multi-agents can search for the optimal solution in a global scope through the collaboration and competition between agents, not only considering the current state, but also optimizing future decisions through long-term accumulated reward values, thereby achieving global optimization and showing higher flexibility; the embodiments of the present invention can not only effectively reduce the deployment cost, but also ensure the low-latency transmission performance of the distribution network protection service, and meet the needs of the distribution network protection service. This integrated communication network solution and reliable deployment method provide a practical technical solution for the digital transformation and intelligent upgrading of the distribution network for the distribution network protection service.
[0107] On the basis of the above-mentioned method items, the present invention provides corresponding embodiments of the system items.
[0108] See also Figure 4 , Figure 4 This is a structural block diagram of a deployment system for a 5G mobile access network of a distribution network provided by an embodiment of the present invention. The deployment system for a 5G mobile access network of a distribution network includes:
[0109] A deployment strategy acquisition module 21 is used to minimize the deployment cost and transmission delay of the distribution network 5G mobile access network as the optimization goal, and obtain the optimal deployment strategy through interactive learning between multiple intelligent agents and the environment of the distribution network 5G mobile access network;
[0110] A configuration deployment module 22 is used to configure the deployment of the distribution network 5G mobile access network in each intelligent body according to the optimal deployment strategy; wherein the configuration of the deployment of the distribution network 5G mobile access network includes finding the working path and protection backup path of the distribution network 5G mobile access network.
[0111] In an optional implementation, the deployment strategy acquisition module 21 includes:
[0112] The optimization target expression unit is used to express the optimization target of minimizing the deployment cost and transmission delay of the 5G mobile access network of the distribution network as a Markov decision process for each intelligent agent, which is expressed as a four-tuple entry of <current state, action, reward, next state>; wherein,
[0113] An intelligent agent state determination unit, configured to determine the state of the intelligent agent according to the remaining computing resource capacity of the equipment in the 5G mobile access network of the distribution network and the connection state of all candidate sites;
[0114] An intelligent agent action determination unit, used to determine the action of the intelligent agent according to the candidate site location of the distribution network 5G mobile access network, the order of connection of the active antenna unit, the spectrum resource allocation, and the isolation between the working path and the backup protection path;
[0115] An agent reward function determination unit, used to determine the agent's reward function according to the deployment cost and transmission delay of the distribution network 5G mobile access network;
[0116] The deployment strategy acquisition unit is used to obtain the Q value set of all actions of the intelligent agent through a deep Q network, select the action with the largest value from the Q value set, and obtain the optimal deployment strategy.
[0117] Furthermore, the reward function of the agent is:
[0118]
[0119]
[0120] in, represents the expected reward at the current distribution network protection service request i and the current time t; represents the deployment cost at the current distribution network protection service request i and the current time t, and α is the weight coefficient of the deployment cost; represents the transmission delay between the current distribution network protection service request i and the current time t, β is the weight coefficient of the transmission delay; if isvalid means that under the current distribution network protection service request i and the current time t, the action a is valid. is notvalid means that under the current distribution network protection service request i and the current time t, the action a is invalid; A represents the device set of the distribution network 5G mobile access network; A Boolean variable representing the connection status of device m and device n at time t. If device m and device n are connected, then otherwise d m,n Indicates the distance between device m and device n; c m,n represents the cost per unit length of optical fiber between device m and device n; c m It represents the average cost per unit length of optical fiber from device m to all other devices; A Boolean variable indicating whether device m is occupied at time t. If device m is occupied, then otherwise τ m,n represents the propagation delay per kilometer of optical fiber between device m and device n; τ m Indicates the processing delay of device m.
[0121] In an optional embodiment, the deployment strategy acquisition unit is specifically used to:
[0122] Input the current distribution network protection service request and the current network status into the deep Q network, calculate the Q values corresponding to all actions under the current network status, and obtain the Q value set of all actions;
[0123] The action with the largest value is selected from the Q value set to obtain the optimal deployment strategy.
[0124] Furthermore, the deployment system of the distribution network 5G mobile access network also includes a convolutional neural network update training module, which is used to:
[0125] Observe the network status at the next moment and obtain corresponding rewards;
[0126] Store the four-tuple entry of <current state, action, reward, next state> into the experience replay cache;
[0127] Data is randomly sampled periodically from the experience replay buffer to train the multi-level attention-enhanced convolutional neural network to update parameters of the multi-level attention-enhanced convolutional neural network.
[0128] It should be noted that a deployment system for a 5G mobile access network for a distribution network provided in an embodiment of the present invention is used to execute all the process steps of a deployment method for a 5G mobile access network for a distribution network in the above embodiment. The working principles and beneficial effects of the two correspond one to one, and therefore will not be repeated herein.
[0129] The embodiment of the present invention also provides a terminal device, such as Figure 5 As shown, it is a structural block diagram of a preferred embodiment of a terminal device provided by the present invention. The terminal device includes a processor 31, a memory 32, and a computer program stored in the memory 32 and configured to be executed by the processor 31. When the processor 31 executes the computer program, the deployment method of the 5G mobile access network of the distribution network as described in any of the above embodiments is implemented.
[0130] In addition, an embodiment of the present invention also provides a computer-readable storage medium, which includes a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute the deployment method of the distribution network 5G mobile access network as described in any of the above embodiments.
[0131] When the processor 31 executes the computer program, the steps in the above-mentioned method for deploying a 5G mobile access network in the distribution network are implemented, for example Figure 1 Alternatively, when the processor 31 executes the computer program, the functions of each module in the above-mentioned embodiment of the deployment system of the 5G mobile access network of the distribution network are realized, for example Figure 4 The functions of each module of the distribution network 5G mobile access network deployment system are shown.
[0132] Preferably, the computer program can be divided into one or more modules / units, which are stored in the memory 32 and executed by the processor 31 to implement the present invention. The one or more modules / units can be a series of computer program instruction segments that can implement specific functions, and the instruction segments are used to describe the execution process of the computer program in the terminal device.
[0133] The processor 31 can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor, or the processor 31 can be any conventional processor. The processor 31 is the control center of the terminal device, and various parts of the terminal device are connected using various interfaces and lines.
[0134] The memory 32 mainly includes a program storage area and a data storage area, wherein the program storage area can store an operating system, an application required for at least one function, etc., and the data storage area can store related data, etc. In addition, the memory 32 can be a high-speed random access memory, or a non-volatile memory, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, and a flash card (Flash Card), etc., or the memory 32 can also be other volatile solid-state storage devices.
[0135] It should be noted that the above terminal device may include, but is not limited to, a processor and a memory. Those skilled in the art will understand that Figure 5 The structural block diagram shown is only a structural example of the above-mentioned terminal device and does not constitute a structural limitation of the above-mentioned terminal device. The above-mentioned terminal device may include more or fewer components than shown in the figure, or combine certain components, or different components.
[0136] The above is a preferred embodiment of the present invention. It should be pointed out that a person skilled in the art can make several improvements and modifications without departing from the principle of the present invention. These improvements and modifications are also considered to be within the scope of protection of the present invention.
Claims
1. A method for deploying a 5G mobile access network in a distribution network, characterized in that: include: Taking minimizing the deployment cost and transmission delay of the distribution network 5G mobile access network as the optimization goal, the optimal deployment strategy is obtained by interactive learning between multiple intelligent agents and the environment of the distribution network 5G mobile access network; According to the optimal deployment strategy, the deployment of the distribution network 5G mobile access network is configured in each intelligent agent; wherein, the configuration of the deployment of the distribution network 5G mobile access network includes finding the working path and the protection backup path of the distribution network 5G mobile access network.
2. The method for deploying a 5G mobile access network for a distribution network according to claim 1, characterized in that: The optimization goal is to minimize the deployment cost and transmission delay of the distribution network 5G mobile access network, and to obtain the optimal deployment strategy by interactively learning between multiple intelligent agents and the environment of the distribution network 5G mobile access network, including: For each agent, the optimization goal of minimizing the deployment cost and transmission delay of the 5G mobile access network of the distribution network is expressed as a Markov decision process, which is represented by a four-tuple entry of <current state, action, reward, next state>; where, Determine the state of the agent based on the remaining computing resource capacity of the equipment in the 5G mobile access network of the distribution network and the connection status of all candidate sites; Determine the action of the agent according to the candidate site locations of the 5G mobile access network of the distribution network, the order of connection of active antenna units, spectrum resource allocation, and the isolation between the working path and the backup protection path; Determining a reward function of the agent based on the deployment cost and transmission delay of the 5G mobile access network of the distribution network; The Q value set of all actions of the agent is obtained through the deep Q network, and the action with the largest value is selected from the Q value set to obtain the optimal deployment strategy.
3. The method for deploying a 5G mobile access network for a distribution network according to claim 2, characterized in that: The reward function of the agent is: in, represents the expected reward at the current distribution network protection service request i and the current time t; represents the deployment cost at the current distribution network protection service request i and the current time t, and α is the weight coefficient of the deployment cost; represents the transmission delay between the current distribution network protection service request i and the current time t, and β is the weight coefficient of the transmission delay; Indicates that under the current distribution network protection service request i and the current time t, action a is effective. Indicates that under the current distribution network protection service request i and the current time t, action a is invalid; A represents the device set of the distribution network 5G mobile access network; A Boolean variable representing the connection status of device m and device n at time t. If device m and device n are connected, then otherwise d m,n Indicates the distance between device m and device n; c m,n represents the cost per unit length of optical fiber between device m and device n; c m It represents the average cost per unit length of optical fiber from device m to all other devices; A Boolean variable indicating whether device m is occupied at time t. If device m is occupied, then otherwise τ m,n represents the propagation delay per kilometer of optical fiber between device m and device n; τ m Indicates the processing delay of device m.
4. The method for deploying a 5G mobile access network for a distribution network as claimed in claim 3, characterized in that: The method of obtaining the Q value set of all actions of the agent through the deep Q network, selecting the action with the largest value from the Q value set, and obtaining the optimal deployment strategy includes: Input the current distribution network protection service request and the current network status into the deep Q network, calculate the Q values corresponding to all actions under the current network status, and obtain the Q value set of all actions; The action with the largest value is selected from the Q value set to obtain the optimal deployment strategy.
5. The method for deploying a 5G mobile access network for a distribution network according to claim 4, characterized in that: The deep Q network includes a multi-level attention-enhanced convolutional neural network and a multi-layer fully connected network; The multi-layer fully connected network includes a multi-head attention layer.
6. The method for deploying a 5G mobile access network for a distribution network according to claim 4, characterized in that: After configuring the deployment of the distribution network 5G mobile access network in each intelligent agent according to the optimal deployment strategy, the method further includes: Observe the network status at the next moment and obtain corresponding rewards; Store the four-tuple entry of <current state, action, reward, next state> into the experience replay cache; Data is randomly sampled periodically from the experience replay buffer to train the multi-level attention-enhanced convolutional neural network to update parameters of the multi-level attention-enhanced convolutional neural network.
7. The method for deploying a 5G mobile access network for a distribution network according to claim 6, characterized in that: The Bellman optimal equation is used to update the parameters of the multi-level attention-enhanced convolutional neural network.
8. A deployment system for a 5G mobile access network in a distribution network, characterized in that: include: A deployment strategy acquisition module is used to minimize the deployment cost and transmission delay of the distribution network 5G mobile access network as the optimization goal, and obtain the optimal deployment strategy through interactive learning between multiple intelligent agents and the environment of the distribution network 5G mobile access network; A configuration deployment module is used to configure the deployment of the distribution network 5G mobile access network in each intelligent agent according to the optimal deployment strategy; wherein, the configuration of the deployment of the distribution network 5G mobile access network includes finding the working path and protection backup path of the distribution network 5G mobile access network.
9. The deployment system of the 5G mobile access network of the distribution network as claimed in claim 8, characterized in that: The deployment strategy acquisition module includes: The optimization target expression unit is used to express the optimization target of minimizing the deployment cost and transmission delay of the 5G mobile access network of the distribution network as a Markov decision process for each intelligent agent, which is expressed as a four-tuple entry of <current state, action, reward, next state>; wherein, An intelligent agent state determination unit, configured to determine the state of the intelligent agent according to the remaining computing resource capacity of the equipment in the 5G mobile access network of the distribution network and the connection state of all candidate sites; An intelligent agent action determination unit, used to determine the action of the intelligent agent according to the candidate site location of the distribution network 5G mobile access network, the order of connection of the active antenna unit, the spectrum resource allocation, and the isolation between the working path and the backup protection path; An agent reward function determination unit, used to determine the agent's reward function according to the deployment cost and transmission delay of the distribution network 5G mobile access network; The deployment strategy acquisition unit is used to obtain the Q value set of all actions of the intelligent agent through a deep Q network, select the action with the largest value from the Q value set, and obtain the optimal deployment strategy.
10. A terminal device, characterized in that: It includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, and when the processor executes the computer program, it implements the deployment method of the distribution network 5G mobile access network as described in any one of claims 1 to 7.
11. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute the method for deploying a 5G mobile access network of a distribution network as described in any one of claims 1 to 7.