A computing power network resource management method and device based on a virtual network and reinforcement learning

By combining virtual networks with reinforcement learning, global dynamic scheduling and on-demand allocation of computing resources are achieved, solving the problem of unreasonable resource allocation in dynamic network environments caused by traditional computing resource management methods, and improving resource utilization and service quality.

CN122496477APending Publication Date: 2026-07-31ANHUI NANRUI JIYUAN POWER GRID TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ANHUI NANRUI JIYUAN POWER GRID TECH CO LTD
Filing Date
2026-05-14
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Traditional computing resource management methods suffer from insufficient efficiency, flexibility, and intelligence when facing dynamically changing network environments and diverse business needs. This leads to unreasonable resource allocation, low utilization, difficulty in achieving globally optimal configuration, and inability to meet the dynamic resource scheduling requirements of new computing scenarios such as edge computing and cloud-edge collaboration.

Method used

A computing power network resource management architecture based on the fusion of virtual networks and reinforcement learning is adopted. By virtualizing physical computing power resources and combining them with an improved deep reinforcement learning model, the reward function is reconstructed to achieve global dynamic scheduling and on-demand allocation of computing power resources, thereby improving the efficiency, flexibility and intelligence of resource management.

Benefits of technology

It significantly improves the overall resource utilization of the computing network, avoids resource idleness and waste, ensures the service quality of high-priority services, and fully releases the potential value of the computing network.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122496477A_ABST
    Figure CN122496477A_ABST
Patent Text Reader

Abstract

This invention relates to a computing power network resource management method based on virtual networks and reinforcement learning, comprising: real-time collection of key indicators and information of the computing power network; mapping computing power nodes of the computing power network to virtual network computing power nodes; calculating the available computing power resources and available bandwidth resources of each computing power node in the virtual network; obtaining a state matrix based on the normalized data; obtaining the optimal computing power nodes and optimal computing power links of the virtual network; and generating the final computing power network node and link configuration scheme. This invention abstracts and pools physical computing power resources through virtual networks, combined with the intelligent decision-making capabilities of reinforcement learning, resulting in global dynamic scheduling and on-demand allocation of computing power resources; autonomously outputting optimal virtual network computing power node selection and resource allocation strategies, significantly improving the overall resource utilization of the computing power network, while ensuring the service quality of high-priority services, and fully releasing the potential value of the computing power network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computing power resource management technology, and in particular to a computing power network resource management method and device based on virtual networks and reinforcement learning. Background Technology

[0002] With the rapid development of information technology, traditional computing resource management methods are increasingly revealing their shortcomings in efficiency, flexibility, and intelligence when facing ever-growing computing demands and complex, ever-changing network environments. These methods rely on static configuration and manual decision-making, making it difficult to adapt to the dynamic changes in computing resources and the diverse application needs. In the current field of computing resource management, although network monitoring tools can collect key indicators such as network traffic, latency, and packet loss rate in real time, this data is often analyzed in isolation, failing to fully realize its potential in resource management. At the same time, the management of physical network computing nodes also lacks intelligent means, leading to frequent problems such as unreasonable resource allocation and low utilization.

[0003] Furthermore, existing computing power scheduling mechanisms are mostly based on preset rules or simple thresholds for decision-making, failing to adapt to real-time fluctuations in network status, dynamic changes in business load, and performance differences among computing power nodes. This easily leads to a mismatch between computing power resources and business needs: on the one hand, some high-priority services experience response delays and decreased service quality due to insufficient computing power supply; on the other hand, some idle computing power nodes waste resources due to untimely scheduling, making it difficult to achieve optimal global computing power resource allocation. Simultaneously, traditional management architectures lack unified awareness and collaborative scheduling capabilities for virtual network computing power nodes, failing to effectively integrate distributed computing power resources. This makes it difficult to meet the dynamic resource scheduling needs of emerging computing power scenarios such as edge computing and cloud-edge collaboration, and also fails to provide end-to-end computing power service guarantees for users, severely restricting the efficient operation and value release of the computing power network. Summary of the Invention

[0004] To address the problem that traditional computing power resource management methods cannot meet the requirements of efficiency, flexibility, and intelligence, the primary objective of this invention is to provide a computing power network resource management method based on virtual networks and reinforcement learning that significantly improves the efficiency, flexibility, and intelligence of resource management and substantially enhances the overall resource utilization of computing power networks.

[0005] This invention employs a computing power network resource management architecture based on the fusion of virtual networks and reinforcement learning. Addressing the pain points of traditional computing power resource management, which relies on static configuration and manual decision-making and struggles to adapt to dynamic network environments and diverse business needs, this invention abstracts and pools physical computing power resources through virtual networks. Combined with the intelligent decision-making capabilities of reinforcement learning, it enables global dynamic scheduling and on-demand allocation of computing power resources, completely overcoming the limitations of traditional methods and significantly improving the efficiency, flexibility, and intelligence of resource management. This provides core technical support for the large-scale deployment and efficient operation of computing power networks. Secondly, this invention uses a deep reinforcement learning model based on an improved reward function, addressing the limitations of existing deep reinforcement learning methods. The Xi model suffers from shortcomings in optimizing computing network resources, including insufficient guidance, slow convergence speed, and limited improvement in resource utilization. A reconstructed reward function integrates multiple objectives such as computing node performance, resource utilization, and service quality, enhancing the model's ability to learn the globally optimal allocation of computing resources. This accelerates model convergence and enables the model to autonomously output optimal virtual network computing node selection and resource allocation strategies based on real-time network conditions and service load. This significantly improves the overall resource utilization of the computing network, avoids resource idleness and waste, and ensures the service quality of high-priority services. It brings the advantages of refined and intelligent management of computing resources, fully releasing the potential value of the computing network.

[0006] To achieve the above objectives, the present invention adopts the following technical solution: a method for managing computing network resources based on virtual networks and reinforcement learning, the method comprising the following sequential steps:

[0007] (1) Collect key indicators and information of computing power network in real time. The key indicators include traffic, latency and packet loss rate. The information includes node location, hardware configuration and software version. The key indicators and information together constitute historical data.

[0008] (2) Using virtualization technology, the computing power nodes of the computing power network are mapped to virtual computing power nodes; the configuration parameters of the virtual computing power nodes are set, and the performance of the virtual computing power nodes is evaluated;

[0009] (3) Analyze the current load and resource configuration of the computing power nodes in the virtual network, and calculate the available computing power resources and available bandwidth resources of each computing power node in the virtual network;

[0010] (4) Calculate the available bandwidth of the connected computing power links of the virtual network computing power nodes and the average distance from the virtual network computing power nodes to other virtual network computing power nodes. Perform maximum and minimum normalization processing on the available computing power resources, available bandwidth of the connected computing power links, and average distance from the virtual network computing power nodes to other virtual network computing power nodes for each computing power node in the virtual network to obtain the normalized data. Obtain the state matrix based on the normalized data. ;

[0011] (5) Improve the reward function in the deep reinforcement learning model to obtain the improved model, and then convert the state matrix. Input the improved model and output the available resources of each computing node in the virtual network. The available resources include available computing power resources and available bandwidth resources. Calculate the embedding probability of each computing node in the virtual network, determine the computing nodes and computing links of the relevant virtual networks that meet the computing power task request, and obtain the optimal computing nodes and optimal computing links of the virtual network.

[0012] (6) Based on the optimal virtual network computing power nodes, generate the final computing power network node and link configuration scheme and deploy it into the actual computing power network.

[0013] Step (3) specifically includes the following steps in sequence:

[0014] (3a) Calculate the available computing resources of each computing node in the virtual network. :

[0015] ;

[0016] ;

[0017] in, This represents the computing power request for the i-th task; This indicates the number of computing power task requests in the virtual network; Represents virtual network computing power nodes Computing resources; Represents virtual network computing power nodes Whether it was successfully mapped to a virtual network computing node If successful, then It is 1 if it is true, otherwise it is 0. Indicates the first A collection of virtual network computing nodes Indicates the first Number of virtual network computing nodes; Represents virtual network computing power nodes computing resources This represents the set of computing nodes in virtual network C;

[0018] (3b) Calculate the available bandwidth resources of each computing node in the virtual network. ;

[0019] ;

[0020] ;

[0021] in, Represents virtual network computing power links bandwidth resources; Indicates virtual network computing power network link Whether it was successfully mapped to the virtual network computing power link If successful, then It is 1 if it is true, otherwise it is 0. Indicates the first Number of virtual network computing power links; Represents virtual network computing power links bandwidth resources; This represents a set of virtual network computing power links.

[0022] Step (4) specifically includes the following steps in sequence:

[0023] (4a) Calculate the current virtual network computing power nodes The sum of the bandwidths of all connected virtual network computing links :

[0024] ;

[0025] in, The available bandwidth resources for each computing node in the virtual network; Indicates the virtual network computing power network link; Represents a set of virtual network computing power links;

[0026] (4b) Calculate the computing power of virtual network nodes Average distance to other virtual network computing nodes :

[0027] ;

[0028] in, Represents virtual network computing power nodes Location; Represents virtual network computing power nodes Location Indicates the number of virtual network path hops required for virtual network computing power links;

[0029] (4c) Available computing resources for each computing node in the virtual network , and Max-min normalization is performed separately to accelerate model convergence speed and improve accuracy:

[0030] ;

[0031] ;

[0032] ;

[0033] in, express The normalized state vector; express The minimum value; express The maximum value; express The normalized state vector; express The minimum value; express The maximum value; express The normalized state vector; express The minimum value; express The maximum value;

[0034] (4d) combination , and The state matrix is ​​obtained. :

[0035] ;

[0036] In the formula, express The total number of computing nodes in the medium power range; This represents the set of computing nodes in virtual network C;

[0037] Step (5) specifically includes the following steps in sequence:

[0038] (5a) Improve the deep reinforcement learning model to obtain the improved model: normalize the resource cost and resource benefit in the immediate reward function of the deep reinforcement learning model to obtain the normalized resource cost. With normalized resource revenue :

[0039] ;

[0040] ;

[0041] In the formula, For the start time, End time; Indicates the first A collection of virtual network computing nodes express middle Available computing resources This represents the computing power request for the i-th task. Represents a virtual network computing node; express The maximum available computing power resources of all computing nodes in the system; express The minimum available computing power resources of all computing nodes in the system; express middle Available bandwidth resources; express The maximum available bandwidth resources for all computing power links in the system; Indicate The minimum available bandwidth resources for all computing power links in the system; express Required virtual network path hops; Indicates the virtual network computing power network link; Represents the set of computing power links of the i-th virtual network;

[0042] The state matrix Input the improved model to calculate the available resources of each computing node in the virtual network. :

[0043] ;

[0044] in, Represents the weight matrix. Represents the deviation vector;

[0045] The improved model of the virtual network automatically learns available resources. Each corresponding computing node Resource adaptability score ;

[0046] (5b) Calculate the embedding probability of each virtual network computing node using the softmax function. Scale the resource vector to Inside:

[0047] ;

[0048] In the formula, This represents the set of computing nodes in virtual network C;

[0049] (5c) Based on embedding probability All computing network nodes are sorted in descending order, and the improved model of all virtual network computing links is searched according to the path algorithm to determine all combinations of computing nodes and computing links of virtual networks that meet the computing task request.

[0050] (5d) Calculate the normalized resource cost of virtual network computing nodes and computing links in each combination. and normalized resource income The benefit-cost ratio for each combination of data is calculated. :

[0051] ;

[0052] The calculated benefit-cost ratio for each group Sort the data in descending order and select the virtual network with the largest benefit-cost ratio. The computing nodes and computing links corresponding to the largest benefit-cost ratio are the optimal computing nodes and optimal computing links.

[0053] Another object of the present invention is to provide an electronic device comprising:

[0054] Processor; and

[0055] The memory stores computer program instructions that, when executed by the processor, cause the processor to perform the computing power network resource management method based on virtual networks and reinforcement learning as described above.

[0056] The present invention also provides a computer-readable storage medium having stored thereon computer program instructions, which, when executed by a processor, cause the processor to perform the computing power network resource management method based on virtual networks and reinforcement learning as described above.

[0057] As can be seen from the above technical solution, the beneficial effects of the present invention are as follows: First, the present invention adopts a computing power network resource management architecture based on the fusion of virtual networks and reinforcement learning. Addressing the pain point that traditional computing power resource management relies on static configuration and manual decision-making, making it difficult to adapt to dynamic network environments and diverse business needs, the present invention abstracts and pools physical computing power resources through virtual networks. Combined with the intelligent decision-making capabilities of reinforcement learning, it brings about global dynamic scheduling and on-demand allocation of computing power resources, completely overcoming the limitations of traditional methods and significantly improving the efficiency, flexibility, and intelligence of resource management, providing core technical support for the large-scale deployment and efficient operation of computing power networks; Second, the present invention adopts a deep reinforcement learning architecture based on an improved reward function. The Xi model addresses the shortcomings of existing deep reinforcement learning models in optimizing computing network resources, such as insufficient guidance, slow convergence speed, and limited improvement in resource utilization. It reconstructs a reward function that integrates multiple objectives, including computing node performance, resource utilization, and service quality, thereby enhancing the model's ability to learn the globally optimal allocation of computing resources. This accelerates model convergence and enables the model to autonomously output optimal virtual network computing node selection and resource allocation strategies based on real-time network conditions and service load. This significantly improves the overall resource utilization of the computing network, avoids resource idleness and waste, and ensures the service quality of high-priority services. It brings the advantages of refined and intelligent management of computing resources, fully releasing the potential value of the computing network. Attached Figure Description

[0058] Figure 1 This is a flowchart of the method of the present invention;

[0059] Figure 2 This is a diagram of the deep reinforcement learning model architecture based on the improved reward function in this invention. Detailed Implementation

[0060] like Figure 1 As shown, a method for managing computing network resources based on virtual networks and reinforcement learning includes the following sequential steps:

[0061] (1) Collect key indicators and information of computing power network in real time. The key indicators include traffic, latency and packet loss rate. The information includes node location, hardware configuration and software version. The key indicators and information together constitute historical data.

[0062] (2) Using virtualization technology, the computing power nodes of the computing power network are mapped to virtual computing power nodes; the configuration parameters of the virtual computing power nodes are set, and the performance of the virtual computing power nodes is evaluated;

[0063] (3) Analyze the current load and resource configuration of the computing power nodes in the virtual network, and calculate the available computing power resources and available bandwidth resources of each computing power node in the virtual network;

[0064] (4) Calculate the available bandwidth of the connected computing power links of the virtual network computing power nodes, the average distance from the virtual network computing power nodes to other virtual network computing power nodes, and the available computing power resources of each computing power node in the virtual network. The average distances from virtual network computing power nodes to other virtual network computing power nodes are subjected to maximum and minimum normalization processes to obtain normalized data. The state matrix is ​​then obtained based on the normalized data. ;

[0065] (5) Improve the reward function in the deep reinforcement learning model to obtain the improved model, and then convert the state matrix. Input the improved model and output the available resources of each computing node in the virtual network. The available resources include available computing power resources and available bandwidth resources. Calculate the embedding probability of each computing node in the virtual network, determine the computing nodes and computing links of the relevant virtual networks that meet the computing power task request, and obtain the optimal computing nodes and optimal computing links of the virtual network.

[0066] (6) Based on the optimal virtual network computing power nodes, generate the final computing power network node and link configuration scheme and deploy it into the actual computing power network.

[0067] Step (3) specifically includes the following steps in sequence:

[0068] (3a) Calculate the available computing resources of each computing node in the virtual network. :

[0069] ;

[0070] ;

[0071] in, This represents the computing power request for the i-th task; This indicates the number of computing power task requests in the virtual network; Represents virtual network computing power nodes Computing resources; Represents virtual network computing power nodes Whether it was successfully mapped to a virtual network computing node If successful, then It is 1 if it is true, otherwise it is 0. Indicates the first A collection of virtual network computing nodes Indicates the first Number of virtual network computing nodes; Represents virtual network computing power nodes computing resources This represents the set of computing nodes in virtual network C;

[0072] (3b) Calculate the available bandwidth resources of each computing node in the virtual network. ;

[0073] ;

[0074] ;

[0075] in, Represents virtual network computing power links bandwidth resources; Indicates virtual network computing power network link Whether it was successfully mapped to the virtual network computing power link If successful, then It is 1 if it is true, otherwise it is 0. Indicates the first Number of virtual network computing power links; Represents virtual network computing power links bandwidth resources; This represents a set of virtual network computing power links.

[0076] Step (4) specifically includes the following steps in sequence:

[0077] (4a) Calculate the current virtual network computing power nodes The sum of the bandwidths of all connected virtual network computing links :

[0078] ;

[0079] in, The available bandwidth resources for each computing node in the virtual network; Indicates the virtual network computing power network link; Represents a set of virtual network computing power links;

[0080] (4b) Calculate the computing power of virtual network nodes Average distance to other virtual network computing nodes :

[0081] ;

[0082] in, Represents virtual network computing power nodes Location; Represents virtual network computing power nodes Location Indicates the number of virtual network path hops required for virtual network computing power links;

[0083] (4c) Available computing resources for each computing node in the virtual network , and Max-min normalization is performed separately to accelerate model convergence speed and improve accuracy:

[0084] ;

[0085] ;

[0086] ;

[0087] in, express The normalized state vector; express The minimum value; express The maximum value; express The normalized state vector; express The minimum value; express The maximum value; express The normalized state vector; express The minimum value; express The maximum value;

[0088] (4d) combination , and The state matrix is ​​obtained. :

[0089] ;

[0090] In the formula, express The total number of computing nodes in the medium power range; This represents the set of computing nodes in virtual network C;

[0091] like Figure 2 As shown, step (5) specifically includes the following steps in sequence:

[0092] (5a) Improve the deep reinforcement learning model to obtain the improved model: normalize the resource cost and resource benefit in the immediate reward function of the deep reinforcement learning model to obtain the normalized resource cost. With normalized resource revenue :

[0093] ;

[0094] ;

[0095] In the formula, For the start time, End time; Indicates the first A collection of virtual network computing nodes express middle Available computing resources This represents the computing power request for the i-th task. Represents a virtual network computing node; express The maximum available computing power resources of all computing nodes in the system; express The minimum available computing power resources of all computing nodes in the system; express middle Available bandwidth resources; express The maximum available bandwidth resources for all computing power links in the system; Indicate The minimum available bandwidth resources for all computing power links in the system; express Required virtual network path hops; Indicates the virtual network computing power network link; Represents the set of computing power links of the i-th virtual network;

[0096] Therefore, normalized resource costs are adopted. and resource benefits Using indicators to measure resource costs and resource benefits can eliminate the problem of inconsistent dimensions.

[0097] The state matrix Input the improved model to calculate the available resources of each computing node in the virtual network. :

[0098] ;

[0099] in, Represents the weight matrix. Represents the deviation vector;

[0100] The improved model of the virtual network automatically learns available resources. Each corresponding computing node Resource adaptability score ;

[0101] (5b) Calculate the embedding probability of each virtual network computing node using the softmax function. Scale the resource vector to Inside:

[0102] ;

[0103] In the formula, This represents the set of computing nodes in virtual network C;

[0104] (5c) Based on embedding probability All computing network nodes are sorted in descending order, and the improved model of all virtual network computing links is searched according to the path algorithm to determine all combinations of computing nodes and computing links of virtual networks that meet the computing task request.

[0105] (5d) Calculate the normalized resource cost of virtual network computing nodes and computing links in each combination. and normalized resource income The benefit-cost ratio for each combination of data is calculated. :

[0106] ;

[0107] The calculated benefit-cost ratio for each group Sort the data in descending order and select the virtual network with the largest benefit-cost ratio. The computing nodes and computing links corresponding to the largest benefit-cost ratio are the optimal computing nodes and optimal computing links.

[0108] In summary, this invention adopts a computing power network resource management architecture based on the fusion of virtual networks and reinforcement learning. Addressing the pain points of traditional computing power resource management, which relies on static configuration and manual decision-making and struggles to adapt to dynamic network environments and diverse business needs, this invention abstracts and pools physical computing power resources through virtual networks. Combined with the intelligent decision-making capabilities of reinforcement learning, it enables global dynamic scheduling and on-demand allocation of computing power resources, completely overcoming the limitations of traditional methods and significantly improving the efficiency, flexibility, and intelligence of resource management. This provides core technical support for the large-scale deployment and efficient operation of computing power networks. Furthermore, this invention employs a deep reinforcement learning model based on an improved reward function, addressing the limitations of existing deep learning models. Reinforcement learning models suffer from shortcomings in optimizing computing network resources, such as insufficient guidance, slow convergence speed, and limited improvement in resource utilization. This paper reconstructs a reward function that integrates multiple objectives, including computing node performance, resource utilization, and service quality, enhancing the model's ability to learn the globally optimal allocation of computing resources. This accelerates model convergence and enables the model to autonomously output optimal virtual network computing node selection and resource allocation strategies based on real-time network conditions and service load. This significantly improves the overall resource utilization of the computing network, avoids resource idleness and waste, and ensures the service quality of high-priority services. It brings the advantages of refined and intelligent management of computing resources, fully releasing the potential value of the computing network.

[0109] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the claimed invention. The scope of protection claimed by the appended claims and their equivalents is defined.

Claims

1. A method for managing computing network resources based on virtual networks and reinforcement learning, characterized in that: The method includes the following steps in sequence: (1) Collect key indicators and information of computing power network in real time. The key indicators include traffic, latency and packet loss rate. The information includes node location, hardware configuration and software version. The key indicators and information together constitute historical data. (2) Using virtualization technology, the computing power nodes of the computing power network are mapped to virtual computing power nodes; the configuration parameters of the virtual computing power nodes are set, and the performance of the virtual computing power nodes is evaluated; (3) Analyze the current load and resource configuration of the computing power nodes in the virtual network, and calculate the available computing power resources and available bandwidth resources of each computing power node in the virtual network; (4) Calculate the available bandwidth of the connected computing power links of the virtual network computing power nodes and the average distance from the virtual network computing power nodes to other virtual network computing power nodes. Perform maximum and minimum normalization processing on the available computing power resources, available bandwidth of the connected computing power links, and average distance from the virtual network computing power nodes to other virtual network computing power nodes for each computing power node in the virtual network to obtain the normalized data. Obtain the state matrix based on the normalized data. ; (5) Improve the reward function in the deep reinforcement learning model to obtain the improved model, and then convert the state matrix. Input the improved model and output the available resources of each computing node in the virtual network. The available resources include available computing power resources and available bandwidth resources. Calculate the embedding probability of each computing node in the virtual network, determine the computing nodes and computing links of the relevant virtual networks that meet the computing power task request, and obtain the optimal computing nodes and optimal computing links of the virtual network. (6) Based on the optimal virtual network computing power nodes, generate the final computing power network node and link configuration scheme and deploy it into the actual computing power network.

2. The computing power network resource management method based on virtual networks and reinforcement learning according to claim 1, characterized in that: Step (3) specifically includes the following steps in sequence: (3a) Calculate the available computing resources of each computing node in the virtual network. : ; ; in, This represents the computing power request for the i-th task; This indicates the number of computing power task requests in the virtual network; Represents virtual network computing power nodes Computing resources; Represents virtual network computing power nodes Whether it was successfully mapped to a virtual network computing node If successful, then It is 1 if it is true, otherwise it is 0; Indicates the first A collection of virtual network computing nodes Indicates the first Number of virtual network computing nodes; Represents virtual network computing power nodes computing resources This represents the set of computing nodes in virtual network C; (3b) Calculate the available bandwidth resources of each computing node in the virtual network. ; ; ; in, Represents virtual network computing power links bandwidth resources; Represents virtual network computing power network links Whether it was successfully mapped to the virtual network computing power link If successful, then It is 1 if it is true, otherwise it is 0; Indicates the first Number of virtual network computing power links; Represents virtual network computing power links bandwidth resources; This represents a set of virtual network computing power links.

3. The computing power network resource management method based on virtual networks and reinforcement learning according to claim 1, characterized in that: Step (4) specifically includes the following steps in sequence: (4a) Calculate the current virtual network computing power nodes The sum of the bandwidths of all connected virtual network computing links : ; in, The available bandwidth resources for each computing node in the virtual network; Indicates the virtual network computing power network link; Represents a set of virtual network computing power links; (4b) Calculate the computing power of virtual network nodes Average distance to other virtual network computing nodes : ; in, Represents virtual network computing power nodes Location; Represents virtual network computing power nodes Location Indicates the number of virtual network path hops required for virtual network computing power links; (4c) Available computing resources for each computing node in the virtual network , and Max-min normalization is performed separately to accelerate model convergence speed and improve accuracy: ; ; ; in, express The normalized state vector; express The minimum value; express The maximum value; express The normalized state vector; express The minimum value; express The maximum value; express The normalized state vector; express The minimum value; express The maximum value; (4d) combination , and The state matrix is ​​obtained. : ; In the formula, express The total number of computing nodes in the medium power range; This represents the set of computing nodes in virtual network C.

4. The computing power network resource management method based on virtual networks and reinforcement learning according to claim 1, characterized in that: Step (5) specifically includes the following steps in sequence: (5a) Improve the deep reinforcement learning model to obtain the improved model: normalize the resource cost and resource benefit in the immediate reward function of the deep reinforcement learning model to obtain the normalized resource cost. With normalized resource revenue : ; ; In the formula, For the start time, End time; Indicates the first A collection of virtual network computing nodes express middle Available computing resources This represents the computing power request for the i-th task. Represents a virtual network computing node; express The maximum available computing power resources of all computing nodes in the system; express The minimum available computing power resources of all computing nodes in the system; express middle Available bandwidth resources; express The maximum available bandwidth resources for all computing power links in the system; Indicate The minimum available bandwidth resources for all computing power links in the system; express Required virtual network path hops; Indicates the virtual network computing power network link; Represents the set of computing power links of the i-th virtual network; The state matrix Input the improved model to calculate the available resources of each computing node in the virtual network. : ; in, Represents the weight matrix. Represents the deviation vector; The improved model of the virtual network automatically learns available resources. Each corresponding computing node Resource adaptability score ; (5b) Calculate the embedding probability of each virtual network computing node using the softmax function. Scale the resource vector to Inside: ; In the formula, This represents the set of computing nodes in virtual network C; (5c) Based on embedding probability All computing network nodes are sorted in descending order, and the improved model of all virtual network computing links is searched according to the path algorithm to determine all combinations of computing nodes and computing links of virtual networks that meet the computing task request. (5d) Calculate the normalized resource cost of virtual network computing nodes and computing links in each combination. and normalized resource income The benefit-cost ratio for each combination of data is calculated. : ; The calculated benefit-cost ratio for each group Sort the data in descending order and select the virtual network with the largest benefit-cost ratio. The computing nodes and computing links corresponding to the largest benefit-cost ratio are the optimal computing nodes and optimal computing links.

5. An electronic device, comprising: processor; as well as A memory storing computer program instructions, which, when executed by the processor, cause the processor to perform the computing power network resource management method based on virtual networks and reinforcement learning as described in any one of claims 1-4.

6. A computer-readable storage medium having stored thereon computer program instructions, which, when executed by a processor, cause the processor to perform the computing power network resource management method based on virtual networks and reinforcement learning as described in any one of claims 1-4.