Personalized intelligent software-based network resource allocation method based on deep reinforcement learning
Through the personalized intelligent network resource allocation method based on deep reinforcement learning, the problem of personalized and differentiated needs in 6G communication networks is solved, accurate and efficient resource allocation is achieved, and network resource utilization is improved.
Patent Information
- Application Number
- CN202410736211.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-07
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2044-06-07
AI Technical Summary
Existing technologies have failed to effectively address the network resource allocation needs of personalized differentiation in 6G communication networks, resulting in long-term low network resource utilization. In addition, resource allocation is limited to wired resources and lacks the introduction of wireless resources.
A personalized intelligent network resource allocation method based on deep reinforcement learning is adopted. By establishing a network demand classification method, constructing a feature matrix and a policy neural network, accurate and efficient software-based network resource allocation is carried out. This includes resource demand inspection, classification, feature matrix construction, training and testing, and ultimately personalized intelligent resource allocation.
It achieves accurate and efficient resource allocation for users' personalized and differentiated needs in 6G networks, improves network resource utilization, and enhances the quality of resource allocation plans.
Smart Images

Figure CN118714011B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a personalized intelligent network resource allocation method based on deep reinforcement learning, belonging to the technical field of 6G network resource management and network virtualization. Background Art
[0002] Currently, one of the challenges facing future 6G communication networks is designing highly elastic networks that meet differentiated new service demands through intelligent, flexible, and efficient dynamic configuration of network resources and full utilization of existing network resources. Virtualization and software-based approaches are key technical approaches to achieving highly elastic 6G communication networks, primarily through Network Function Virtualization (NFV) and Software Defined Networking (SDN). Virtualization and software-based approaches are complementary and complementary, decoupling hardware and functional resources while enabling programmability of networks (functions and resources). This makes future 6G communication networks highly elastic, facilitating flexible service deployment, lowering construction and maintenance costs, and improving resource utilization.
[0003] While current network resource allocation methods based on partial topology attributes can effectively address software-based resource allocation, they suffer from two major issues: 1. These methods are designed to address uniform network service requirements, without employing different topology attribute-assisted methods to address personalized and differentiated needs. This results in long-term low network resource utilization; 2. These methods consider and allocate resources limited to wired resources, necessitating the inclusion of wireless resources and a wider range of resource types. Therefore, it is necessary to propose a personalized intelligent network resource allocation method based on deep reinforcement learning to address these issues. Summary of the Invention
[0004] The purpose of the present invention is to address the defects and shortcomings of the above-mentioned prior art and provide an intelligent network resource allocation method based on deep reinforcement learning, which can realize efficient and intelligent network resource allocation for personalized network service needs.
[0005] The technical solution adopted by the present invention to solve its technical problems is: a personalized intelligent software-based network resource allocation method based on deep reinforcement learning. The method is applied to perform accurate and efficient software-based network resource allocation in medium and large-scale underlying communication networks. An accurate and reliable network demand classification method is established. The underlying network topology attributes corresponding to the classified services are selected to construct a feature matrix and a policy neural network and train them, thereby improving the quality of the network resource allocation solution. The method includes the following steps:
[0006] Step 1: Check the resource requirements of the personalized network service;
[0007] Step 2: Calculate the average resource requirement value of the service, compare it with the threshold, and classify it;
[0008] Step 3: Construct a feature matrix using the selected underlying 6G network attributes as input to the policy neural network;
[0009] Step 4: Train and test the agent on the training set and test set;
[0010] Step 5: Perform personalized intelligent resource allocation for the classified services.
[0011] Optionally, in step 1, the resource requirements include: wired resource requirements, wireless resource requirements, and both wireless and wired resource requirements.
[0012] Optionally, if the network demand only has wired resource requirements, then the software-based resource allocation and deployment will be carried out in the 6G core network; if the network service only has wireless resource requirements, then the resource allocation and deployment will be carried out in the 6G access network; if the network service has both wireless and wired resource requirements, it will be carried out in the access-transmission-core network.
[0013] Optionally, in step 2, the computing requirements to be processed include: calculating the total number of nodes of the network service |TotalElement(VN)|, the average wired resource requirement, the average wireless resource requirement, and the average network processing transmission delay;
[0014] Optionally, in step 2, the types of network service requirements include: multi-connectivity network, high resource demand network service, strict delay request network service, and ordinary network service.
[0015] Optionally, the average wired resource demand of the network service is:
[0016]
[0017] Where a and b represent wired nodes of personalized network services, CPU(a) represents the computing resources of node a, Stor(a) represents the storage resources of node a, Capa(a) represents the capacity resources of node a, Band(ab) represents the bandwidth resources of link ab, VNWiredNode represents the set of wired nodes of network services, |VNWiredNode| represents the total number of wired nodes of network services, VNLink represents the set of links of network services, and |VNLink| represents the total number of links of network services.
[0018] The average wireless resource demand is:
[0019]
[0020] Where c represents the network service wireless node, Spec(c) represents the spectrum resources of node c, VNWireleNode represents the set of network service wireless nodes, and |VNWireleNode| represents the total number of network service wireless nodes.
[0021] The average network processing transmission delay is:
[0022]
[0023] Where, e and f represent network service nodes, VNNode represents the set of network service nodes, |VNNode| represents the total number of network service nodes, VNLink represents the set of network service links, |VNLink| represents the total number of network service links, ProDelay(e) represents the processing delay of node e, and ProDelay(ef) represents the transmission delay of link ef.
[0024] Optionally, if the total number of nodes |TotalElement(VN)| of the network service is greater than a corresponding threshold, then the network service belongs to the multi-connectivity network demand category; if either the average wired resource demand AveWired(VN) or the average wireless resource demand AveWirele(VN) is greater than a corresponding resource threshold, then the network service belongs to the high resource demand category; if the average delay |AveDelay(VN| of the network service is less than a corresponding threshold, then the network service belongs to the strict delay request category. If none of these conditions are met, then the network service is classified as ordinary.
[0025] Optionally, if the network service requirement described in step 3 falls into the multi-connectivity category, network topology attributes related to connectivity need to be introduced to quantify and calculate the underlying 6G network. Currently, candidate node and link attributes include node degree, node center closeness, and link interference. If the network service requirement falls into the high resource demand category, topology attributes related to resource values need to be selected to participate in the quantification of nodes in the 6G network. Possible resource-related attributes include node strength and link strength. If the network service requirement falls into the high latency requirement category, topology attributes related to network connectivity, such as node center closeness and link interference, need to be selected to participate in the quantification of physical nodes in the 6G network. If the network service requirement falls into the general demand category, the high resource demand method can be directly selected to process general network services.
[0026] Optionally, in step 4, a five-layer policy neural network is selected as the learning agent. The training set consists of 100 network services. After training, it is tested on a test set consisting of 100 network services.
[0027] Optionally, in step 5, after completing the training and testing of the above-mentioned intelligent agent, a candidate set of 6G physical nodes is constructed by sorting them in descending order of probability, and personalized software resource allocation is performed for the network service requirements proposed by the user.
[0028] Beneficial effects:
[0029] 1. The present invention classifies the needs raised by each user and allocates network resources using a deep reinforcement learning method, which not only ensures the personalized and differentiated needs of users, but also realizes accurate and efficient network resource allocation, thereby fully improving the utilization rate of 6G network resources.
[0030] 2. The present invention is applied to accurate and efficient software-based network resource allocation in medium and large-scale underlying communication networks. It establishes an accurate and reliable network demand classification method, selects the underlying network topology attributes corresponding to the classified services, constructs a feature matrix and a policy neural network, and trains them, thereby improving the quality of the network resource allocation plan. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 This is a flow chart of the intelligent network resource allocation method based on deep reinforcement learning of the present invention.
[0032] Figure 2 It is a schematic diagram of constructing a set of candidate physical nodes with strong resource allocation capabilities in the present invention. DETAILED DESCRIPTION
[0033] In order to make the objectives, technical solutions and advantages of the present invention more clear, the present invention is described in detail below with reference to the accompanying drawings and specific embodiments.
[0034] like Figure 1 As shown, the present invention provides a personalized intelligent network resource allocation method based on deep reinforcement learning. The method is used to ensure the personalized and differentiated needs of users and achieve accurate and efficient network resource allocation. The method mainly includes the following steps:
[0035] Step 1: Check the resource requirements of the personalized network service.
[0036] Step 2: Calculate the average resource demand value of the service, compare it with the threshold, and classify it;
[0037] Step 3: Use the selected underlying 6G network attributes to construct a feature matrix as the input of the policy neural network;
[0038] Step 4: Train and test the agent in the training set and test set;
[0039] Step 5: Perform personalized intelligent resource allocation for the classified services.
[0040] Steps 1 to 5 are described in detail below.
[0041] In step 1, when a personalized network service request is received, the resource type of the request is first checked. If the network service only requires wired resources, resource allocation and deployment for this network request are performed in the 6G core network. If the network service only requires wireless resources, resource allocation and deployment for this network request are performed in the 6G access network. If the network service requires both wireless and wired resources, resource allocation and deployment are performed across the entire 6G network access-transmission-core network.
[0042] In step 2, the resource requirements of the personalized network service are initially processed and prepared for classification. The following items need to be processed: the total number of nodes of the network service |TotalElement(VN)|, the average wired resource requirements, the average wireless resource requirements, and the average network processing transmission delay. The formula for the average wired resource requirement is:
[0043]
[0044] Where a and b represent wired nodes of the personalized network service, CPU(a) represents the computing resources of node a, Stor(a) represents the storage resources of node a, Capa(a) represents the capacity resources of node a, Band(ab) represents the bandwidth resources of network service link ab, VNWiredNode represents the set of wired nodes of the network service, |VNWiredNode| represents the total number of wired nodes of the network service, VNLink represents the link set of the network service, and |VNLink| represents the total number of links. The formula for average network processing transmission delay is:
[0045]
[0046] Where c represents the network service wireless node, Spec(c) represents the spectrum resource of node c, VNWireleNode represents the set of network service wireless nodes, and |VNWireleNode| represents the number of network service wireless nodes. The formula for average network processing transmission delay is:
[0047]
[0048] Here, e and f represent network service nodes, VNNode represents the set of network service nodes, |VNNode| represents the total number of network service nodes, ProDelay(e) represents the processing delay of node e, and ProDelay(ef) represents the transmission delay of link ef. The four values obtained are compared with the resource thresholds for the 6G network. It is important to note that these thresholds can be set independently, selected from existing 6G research reports and white papers, or through collaboration with telecommunications operators or network service providers, using operator data as a reference. Generally speaking, operators can estimate service thresholds by measuring and calculating the indicators and scale of service requests received over a certain period of time. This threshold setting based on actual demand is valuable as a reference. If the total number of nodes (|TotalElement(VN)|) of a network service is greater than the corresponding threshold, then the network service belongs to the Multi-Connectivity category. If either the average wired resource demand (AveWired(VN)) or the average wireless resource demand (AveWirele(VN)) is greater than the corresponding resource threshold, then the network service belongs to the High Resource Demand category. If the average delay (|AveDelay(VN|)) of the network service is less than the corresponding threshold, then the network service belongs to the Strict Delay Request category. If none of these conditions are met, then the network service is classified as an Ordinary category.
[0049] In step 3, after completing the classification, if the network service requirement falls into the multi-connection category, relevant network topology attributes need to be introduced to quantify and calculate the underlying 6G network. Currently, candidate node and link attributes include node degree, node center closeness, and link interference. More attributes can be found in graph theory references.
[0050] Node degree generally represents the number of directly connected neighboring nodes. A higher number indicates a higher probability that the node serves as a connection hub within the entire network. Node closeness is calculated by summing the distances required to establish a loop-free connection with the remaining nodes in the network. It is generally expressed by counting the number of intermediate nodes. A lower number indicates a higher status for the node within the network, facilitating interconnection. Link interference describes the number of intermediate nodes a link crosses. A higher number indicates a higher likelihood of interference during transmission. After selecting these attributes, appropriate quantitative calculation methods are needed to quantify the resource allocation capabilities of all nodes in the 6G network. Existing research has shown that directly multiplying attributes and resources is inefficient and unreliable. Therefore, a Markov stochastic model approach can be employed. Based on the product of resource attributes, a connection probability matrix for directly connected neighboring nodes and a probability matrix for directly connected global nodes can be constructed. After multiple iterations, when the calculated node resource values stabilize, they can be used as a basis for allocating and deploying network service resources. Next, the physical node with the largest value is selected and compared with the service node with the highest resource and demand. If the service node's requirements (resources, network functions, and processing latency) are all met, that physical node is selected for resource allocation. The remaining service nodes are allocated resources in the same manner. After resource deployment for the service node is complete, the shortest path method is used to assess the resource requirements and transmission latency of the service link. This completes the personalized resource allocation for the network service. Note that if any of these requirements are not met during this process, resource allocation for the network service fails. If the network service demand is high, topological attributes related to network node and link weights are selected to quantify the nodes in the 6G network. Recommended weight-related attributes include node strength and link strength. Both attributes represent the sum of values on a node (link). In resource allocation method research, the sum of node resources (CPU, capacity, storage) can represent strength, and the bandwidth of a link can represent link strength.
[0051] Next, the aforementioned Markov method is applied to the node's location attributes (degree and centrality) within the network. Repeatedly, this method generates stable resource values. Software-defined resource allocation for this network service is then performed using the aforementioned processing model, achieving personalized and efficient resource allocation for high-resource-demand network services. If the network service demand is high-latency, topological attributes related to direct and indirect connections between network nodes and links, such as node-centrality closeness and link interference, are selected. This also contributes to the quantification of physical nodes in the 6G network. Consistent with the quantification methods for the two aforementioned node types, a Markov method is used to calculate stable node values before personalized resource allocation is performed for this network service. During the allocation process, resource allocation for a node is prioritized, with consideration given to the remaining physical nodes surrounding the assigned physical node, which helps reduce link transmission latency. The remaining method is consistent with the aforementioned method. If the network service demand is general, the high-resource-demand method can be directly applied to general network services.
[0052] In step 4, if Figure 2 As shown in the figure, a five-layer policy neural network is selected as the learning agent. The policy neural network consists of the basic elements of a neural network: extraction layer, convolution layer, probabilistic layer, filtering layer, and output layer. When training the agent, the function of the extraction layer is to extract the topological attributes of each node in the 6G physical network to form a feature vector. The feature vectors of all nodes form a final feature matrix. The function of the convolution layer is to perform a convolution operation on the generated feature matrix. The convolution formula is:
[0053] ARV(A)=λ·Vector(A)+deviation
[0054] λ refers to the weight vector of a convolution kernel, and deviation is a constant used to adjust the convolution bias. This allows the feature vector of each node to obtain an available resource vector. Afterwards, in the probability layer, a probability generation operation is performed on the available resource vector of each node, and finally a probability value is obtained. Then, it is filtered through the final output layer and sorted in descending order according to the probability of the physical node. Then, based on the probability vector P of each node, any |A|th node is selected to construct a one-hot encoding vector, that is, the dimension of the vector is equal to the number of network nodes, and only the |A|th position is 1, and the rest are 0. A loss function is constructed as:
[0055]
[0056] y |A| and p |A| They refer to the randomly selected one-hot vector and the value of the |A|th dimension of the predicted probability vector P, respectively. The gradient, g, can then be calculated using reverse derivation. Because we hope that the randomly selected node A will produce better results, model parameter training tends to favor similar decisions. However, if this random node does not produce good results, model parameter training is discouraged and discouraged. Therefore, the following formula is used to determine and update the gradient:
[0057] g:=α·r·g
[0058] a represents the learning rate, which controls the speed of model training. If the rate is too high, training will not converge and the optimal solution will be missed. If the rate is too low, training will be slow, so it is important to choose an appropriate learning rate. Multiplying the reward by the gradient g makes it easier to obtain decisions with larger rewards, which has a greater impact on the agent and makes it more likely to make similar decisions. If the reward is too small or even negative, it will have a smaller impact on agent training. Therefore, during actual training, the rate must be adjusted manually. The training set is set to consist of 100 network services. After training, it is then tested on a test set consisting of 100 multi-connection network services.
[0059] In step 5, if Figure 2 As shown in the figure, after completing the above training and testing, the candidate set of 6G physical nodes is constructed based on descending probability, and personalized software-based resource allocation is performed based on the user's network service requirements. It is important to note that training and testing are completed before the final service is provided. In other words, operators or service providers need to train the 6G network in advance using the above deep reinforcement learning method. In this way, when providing services, intelligent resource allocation and deployment can be directly performed based on the personalized and differentiated network services requested by users.
[0060] In summary, the present invention classifies network service requirements into individual categories and adopts a personalized network resource allocation method based on deep reinforcement learning, thereby ensuring the personalized and differentiated needs of users in a highly flexible 6G network while achieving efficient network resource allocation, thereby fully improving the utilization rate of 6G network resources.
[0061] The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, it should be understood by those skilled in the art that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A personalized intelligent network resource allocation method based on deep reinforcement learning, characterized in that: The method is applied to accurate and efficient software-based network resource allocation in medium-to-large-scale underlying communication networks. An accurate and reliable network demand classification method is established. The underlying network topology attributes corresponding to the classified services are selected to construct a feature matrix and a policy neural network, and then trained. The method includes the following steps: Step 1: Check the types of resource requirements; Step 2: Compare and classify the computing resource requirements with the thresholds. The computing resource requirements that need to be processed include: the total number of nodes of the computing network service |TotalElement(VN)|, the average wired resource requirement, the average wireless resource requirement, and the average network processing transmission delay. The average wired resource requirement of the network service is: Where a and b represent wired nodes of personalized network services, CPU(a) represents the computing resources of node a, Stor(a) represents the storage resources of node a, Capa(a) represents the capacity resources of node a, Band(ab) represents the bandwidth resources of link ab, VNWiredNode represents the set of wired nodes of network services, |VNWiredNode| represents the total number of wired nodes of network services, VNLink represents the set of links of network services, and |VNLink| represents the total number of links of network services. The average wireless resource demand is: Wherein, c represents the personalized network service wireless node, Spec(c) represents the spectrum resource of node c, VNWireleNode represents the set of network service wireless nodes, and |VNWireleNode| represents the total number of personalized network service wireless nodes; The average network processing transmission delay is: Where, e and f represent the personalized network service nodes, VNNode represents the network service node set, |VNNode| represents the total number of network service nodes, VNLink represents the network service link set, |VNLink| represents the total number of network service links, ProDelay(e) represents the processing delay of node e, and ProDelay(ef) represents the transmission delay of link ef; Network service demand categories include: multi-connection network service, high resource demand network service, high latency demand network service, and general (Ordinary) network service. If the total number of nodes of the network service is greater than the corresponding threshold, then the network service belongs to the multi-connection network service; if either the average wired resource demand or the average wireless resource demand is greater than the corresponding resource threshold, then the network service belongs to the high resource demand network service; if the average latency of the network service is less than the corresponding threshold, then the network service belongs to the high latency demand network service. If none of these conditions are met, then the network service is classified as general (Ordinary) network service. Step 3: Use the selected attributes to construct a feature matrix as the input of the policy neural network. If the personalized network service belongs to the multi-connection class, then it is necessary to introduce network topology attributes related to the connection class to quantify and calculate the underlying 6G network. Currently, the candidate node and link attributes are node degree, node center intimacy, and link interference. If the network service demand belongs to the high resource demand class, it is necessary to select topology attributes related to network resources to participate in the quantification of nodes in the 6G network. The relevant attributes that can be referenced are: node strength and link strength. If the network service demand belongs to the high latency demand class, it is necessary to select topology attributes related to the network node center class, namely node center intimacy and link interference, and use them to quantify the physical nodes in the 6G network. If the network service demand belongs to the general demand class, directly select the high resource demand class method to directly process the general class network service. Step 4: Train and test the agent on the training set and test set; Step 5: Perform intelligent resource allocation for the classified services.
2. A personalized intelligent network resource allocation method based on deep reinforcement learning according to claim 1, characterized in that: In step 1, the resource requirements include: wired resource requirements, wireless resource requirements, and both wireless and wired resource requirements.
3. A personalized intelligent network resource allocation method based on deep reinforcement learning according to claim 2, characterized in that: If the network demand only has wired resource requirements, then resource allocation and deployment will be carried out in the 6G core network; if the network service only has wireless resource requirements, then resource allocation and deployment will be carried out in the 6G access network; if the network service has both wireless and wired resource requirements, then resource allocation and deployment will be carried out in the 6G access-transmission-core network.
4. The personalized intelligent network resource allocation method based on deep reinforcement learning according to claim 1, characterized in that: In step 4, a five-layer policy neural network is selected as the learning agent. The training set consists of 100 network services. After training, it is tested in a test set consisting of 100 network services.
5. The personalized intelligent network resource allocation method based on deep reinforcement learning according to claim 1, characterized in that: In step 5, after completing the training and testing of the above-mentioned intelligent agent, a candidate set of 6G physical nodes is constructed based on descending probability, and personalized software resource allocation is performed for the network service requirements proposed by the user.
Citation Information
Patent Citations
Multi-media fusion cooperative communication method for radio resource management network
CN109495361A
Service function chain reconfiguration method based on load balancing
CN111538587A