A method and system for power network resource allocation
By combining federated reinforcement learning and game theory, this method optimizes network slice resource allocation, solving the problems of unreasonable resource allocation and insufficient data privacy protection in traditional methods, and achieving secure and efficient network slice resource allocation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ELECTRIC POWER RESEARCH INSTITUTE OF STATE GRID SHANDONG ELECTRIC POWER COMPANY
- Filing Date
- 2023-08-29
- Publication Date
- 2026-07-24
AI Technical Summary
In network slicing scenarios, existing technologies and traditional resource allocation methods cannot meet different performance requirements, and they also suffer from problems such as high communication overhead, insufficient data security and privacy protection, and low user satisfaction.
By combining federated reinforcement learning and game theory, global slice selection strategy parameters are generated through the analysis and training of user agents. The Nash equilibrium point is calculated based on game theory algorithms to optimize the allocation of network slice resources, ensuring data security and user interests.
It achieves secure and efficient allocation of network slice resources, meets diverse user performance needs, ensures data privacy protection, and optimizes resource allocation results.
Smart Images

Figure CN117278555B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of power system technology, and in particular to a method and system for allocating power network resources. Background Technology
[0002] As 5G networks evolve towards massive connectivity, security, and efficiency, network slicing creates multiple isolated, customized virtual networks on a public physical network to meet diverse service needs. This provides customized services to users and addresses the challenge of 5G networks failing to meet diverse performance requirements. However, in network slicing scenarios, traditional fixed resource allocation methods, with their fixed and unreasonable resource allocation, cannot satisfy varying performance demands. Existing optimization solutions mainly include the following: First, using machine learning algorithms for network slice resource allocation, combining random forest algorithms with reinforcement learning's IPPO algorithm to solve the resource allocation technical problem in network slicing. Second, based on the 5G network, dividing limited network resources into slices, and using game theory to transform user optimization strategies into a benefit maximization problem, obtaining the maximum user utility function and achieving the optimal slice resource allocation result. Third, utilizing a federated learning framework, achieving multi-party collaborative training of base station load prediction models between slices while ensuring user privacy and preventing data sharing, thus realizing slice-level distributed base station resource prediction. This also avoids the problem of insufficient base station resource prediction performance in network slicing, which can easily affect user service progress.
[0003] However, existing solutions still have problems and shortcomings: The first solution aggregates performance metrics data observed by all slice user agents to train the network slice resource optimization model, which not only leads to huge communication overhead but also violates data security and privacy protection requirements. The second solution, while effectively meeting the need for reasonable resource allocation, improving the rationality of network resources, and enhancing user experience quality, suffers from suboptimal allocation results and fails to consider user data security and privacy protection. The third solution, while considering user data security and privacy protection, neglects user satisfaction and interests. Summary of the Invention
[0004] This invention provides a method and system for allocating power network resources to solve the aforementioned technical problems in the prior art.
[0005] To provide a basic understanding of some aspects of the disclosed embodiments, a brief summary is given below. This summary is not intended as a general commentary, nor is it intended to identify key / important components or to describe the scope of protection of these embodiments. Its sole purpose is to present some concepts in a simple form as a prelude to the detailed description that follows.
[0006] According to a first aspect of the present invention, a method for allocating power network resources is provided, wherein the power network has a plurality of user agents.
[0007] In one embodiment, the power network resource allocation method includes:
[0008] The analysis of each user agent is performed to obtain the environmental state before the slice task is executed, the environmental state after the slice task is executed, the slice task reward, and the network slice scheme for executing the slice task. The data is then merged and organized to obtain the training dataset.
[0009] Based on the training dataset, each user agent is trained to obtain new slice selection strategy parameters for each user agent; and the new slice selection strategy parameters for each user agent are uploaded to the cloud server so that the cloud server can perform federated aggregation based on the received new slice selection strategy parameters to obtain global slice selection strategy parameters.
[0010] Each user agent trains itself according to the global slice selection strategy parameters. After the user agent training is completed, it selects the corresponding network slice based on the game theory algorithm and the global slice selection strategy parameters to generate the best new network slice scheme.
[0011] In one embodiment, the power network resource allocation method further includes:
[0012] Before obtaining the training dataset, each user agent acquires its own task information and slice selection strategy parameters based on the current environmental state of the power network.
[0013] Each user agent calculates a decision action based on the current environment state and slice selection strategy parameters; and selects the corresponding network slice to execute the task information based on the decision action, thereby obtaining a new environment state and task reward.
[0014] In one embodiment, selecting the corresponding network slice to generate the optimal new network slice scheme based on the game theory algorithm and the global slice selection strategy parameters includes:
[0015] A revenue matrix is generated based on pre-acquired network slice usage statistics and the global slice selection strategy parameters.
[0016] Based on the payoff matrix, the Nash equilibrium point is calculated using a game theory algorithm, and based on the Nash equilibrium point, the corresponding network slice is selected to generate the optimal new network slice scheme.
[0017] In one embodiment, generating a revenue matrix based on pre-acquired network slice usage statistics and the global slice selection strategy parameters includes:
[0018] Based on the statistical data on the practical use of network slices obtained in advance, the user group data characteristics are analyzed.
[0019] Based on the pre-configured slice selection strategy, the power network is analyzed to generate a set of network slice schemes. The set of network slice schemes includes slice selection strategies that users can choose and network slice schemes that slice managers can choose.
[0020] Based on the user group data characteristics, the user goals and the slice management goals are substituted into the user and slice management revenue functions to obtain the corresponding revenue values;
[0021] A revenue matrix is generated based on the slice selection strategies available to users, the network slicing schemes available to slice managers, and the revenue values of users and slice managers.
[0022] In one embodiment, the user group data characteristics include the number of users and the distribution of user types; the user objectives include average user satisfaction; and the slice management objectives include system resource balancing.
[0023] In one embodiment, the formula for calculating the user revenue function is:
[0024]
[0025] ,
[0026]
[0027] In the formula, User revenue value; Average user satisfaction; End-to-end delay for power network communication; Network slicing latency; For the bandwidth of the power grid; N represents the network slice bandwidth; N is the number of users. The penalty for system crashes.
[0028] In one embodiment, the formula for calculating the slice manager's revenue function is as follows:
[0029]
[0030] In the formula, This represents the revenue value for the slice management party. For the task In the corresponding network slice Bandwidth used; For the task In the corresponding network slice Bandwidth used; For the task In the corresponding network slice The bandwidth used.
[0031] In one embodiment, the formula for calculating the Nash equilibrium point is:
[0032]
[0033]
[0034] In the formula, This is the Nash equilibrium point, i.e., the optimal new slicing scheme; User revenue value; This represents the revenue value for the slice management party. Selectable slice selection strategy; Choose a strategy for slicing; Available slicing schemes; For the task The corresponding network slice.
[0035] According to a second aspect of the present invention, a power network resource allocation system is provided, wherein the power network has a plurality of user agents.
[0036] In one embodiment, the power network resource allocation system includes:
[0037] The training dataset generation module is used to analyze each user agent, obtain the environmental state before the user agent executes the slicing task, the environmental state after the slicing task, the slicing task reward, and the network slicing scheme for executing the slicing task, and merge and organize the data to obtain the training dataset.
[0038] The agent federated training module is used to train each user agent according to the training dataset to obtain new slice selection policy parameters for each user agent; and to upload the new slice selection policy parameters of each user agent to the cloud server so that the cloud server can perform federated aggregation according to the received new slice selection policy parameters to obtain global slice selection policy parameters.
[0039] The network slice allocation module is used for each user agent to train the user agent according to the global slice selection strategy parameters, and after the user agent training is completed, to select the corresponding network slice based on the game theory algorithm and the global slice selection strategy parameters to generate the best new network slice scheme.
[0040] In one embodiment, the power network resource allocation system further includes:
[0041] The data acquisition module is used to enable each user agent to obtain its own task information and slice selection strategy parameters based on the current environmental state of the power network before obtaining the training dataset.
[0042] The data calculation module is used by each user agent to calculate decision actions based on the current environment state and slice selection strategy parameters; and to select the corresponding network slice to execute the task information based on the decision actions, thereby obtaining a new environment state and task reward.
[0043] In one embodiment, the network slice allocation module includes a revenue matrix generation module and an equilibrium point allocation module, wherein,
[0044] The revenue matrix generation module is used to generate a revenue matrix based on the pre-acquired network slice usage statistics and the global slice selection strategy parameters.
[0045] The equilibrium point allocation module is used to calculate the Nash equilibrium point using a game theory algorithm based on the payoff matrix, and to select the corresponding network slice to generate the optimal new network slice scheme based on the Nash equilibrium point.
[0046] In one embodiment, the revenue matrix generation module includes: a user analysis module, a set generation module, a revenue calculation module, and a matrix generation module, wherein,
[0047] The user analysis module is used to analyze users based on pre-acquired statistical data on the usage of network slices, and to obtain user group data characteristics.
[0048] The set generation module is used to analyze the power network according to the pre-configured slice selection strategy and generate a set of network slice schemes. The set of network slice schemes includes slice selection strategies that users can choose and network slice schemes that slice managers can choose.
[0049] The revenue calculation module is used to substitute the user's target and the slice manager's target into the user's and slice manager's revenue functions based on the user group data characteristics to obtain the corresponding revenue value.
[0050] The matrix generation module is used to generate a revenue matrix based on the slice selection strategy that the user can choose, the network slicing scheme that the slice manager can choose, as well as the user's revenue value and the slice manager's revenue value.
[0051] In one embodiment, the user group data characteristics include the number of users and the distribution of user types; the user objectives include average user satisfaction; and the slice management objectives include system resource balancing.
[0052] In one embodiment, the formula for calculating the user revenue function is:
[0053]
[0054] ,
[0055]
[0056] In the formula, User revenue value; Average user satisfaction; End-to-end delay for power network communication; Network slicing latency; For the bandwidth of the power grid; N represents the network slice bandwidth; N is the number of users. The penalty for system crashes.
[0057] In one embodiment, the formula for calculating the slice manager's revenue function is as follows:
[0058]
[0059] In the formula, This represents the revenue value for the slice management party. For the task In the corresponding network slice Bandwidth used; For the task In the corresponding network slice Bandwidth used; For the task In the corresponding network slice The bandwidth used.
[0060] In one embodiment, the formula for calculating the Nash equilibrium point is:
[0061]
[0062]
[0063] In the formula, This is the Nash equilibrium point, i.e., the optimal new slicing scheme; User revenue value; This represents the revenue value for the slice management party. Selectable slice selection strategy; Choose a strategy for slicing; Available slicing schemes; For the task The corresponding network slice.
[0064] According to a third aspect of the present invention, a computer device is provided.
[0065] In one embodiment, the computer device includes a memory and a processor, the memory storing a computer program, wherein the processor executes the computer program to implement the steps of the above-described method.
[0066] According to a fourth aspect of the present invention, a computer-readable storage medium is provided.
[0067] In one embodiment, a computer program is stored on the computer-readable storage medium, which, when executed by a processor, implements the steps of the above method.
[0068] The technical solutions provided by the embodiments of the present invention may include the following beneficial effects:
[0069] This invention combines network slicing, federated reinforcement learning, and game theory to achieve network slice allocation based on federated reinforcement learning game theory. The introduction of federated reinforcement learning ensures data security among users and more accurate slice resource allocation. The introduction of game theory guarantees the interests of both the slice manager and the user group and meets the needs of both parties, thereby providing users with safe and efficient services and meeting their diverse performance requirements.
[0070] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit the invention. Attached Figure Description
[0071] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.
[0072] Figure 1 This is a flowchart illustrating a power network resource allocation method according to an exemplary embodiment;
[0073] Figure 2 This is a structural block diagram of a power network resource allocation system according to an exemplary embodiment;
[0074] Figure 3 This is an architecture diagram of a power network resource allocation system according to an exemplary embodiment;
[0075] Figure 4This is a flowchart illustrating the allocation of power grid slice resources according to an exemplary embodiment;
[0076] Figure 5 This is a schematic diagram illustrating the principle of a power network resource allocation process according to an exemplary embodiment;
[0077] Figure 6 This is a slice implementation of a power service topology diagram according to an exemplary embodiment;
[0078] Figure 7 This is a schematic diagram of the structure of a computer device according to an exemplary embodiment. Detailed Implementation
[0079] The following description and accompanying drawings fully illustrate specific embodiments described herein to enable those skilled in the art to practice them. Some embodiments may include or substitute parts and features of other embodiments. The scope of the embodiments herein encompasses the entire scope of the claims and all available equivalents thereof. Throughout this document, the terms “first,” “second,” etc., are used only to distinguish one element from another without requiring or implying any actual relationship or order between the elements. Indeed, a first element can also be referred to as a second element, and vice versa. Furthermore, the terms “comprising,” “including,” or any other variations thereof are intended to cover non-exclusive inclusion, such that a structure, apparatus, or device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a structure, apparatus, or device. Without further limitation, an element defined by the phrase “comprising one…” does not exclude the presence of other identical elements in the structure, apparatus, or device that includes said element. The various embodiments described herein are presented in a progressive manner, with each embodiment focusing on its differences from other embodiments; similar or identical parts between embodiments can be referred to interchangeably.
[0080] The terms "longitudinal," "lateral," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," and "outer" used in this document to indicate orientations or positional relationships are based on the orientations or positional relationships shown in the accompanying drawings. They are used solely for the convenience of describing the document and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. In the description herein, unless otherwise specified and limited, the terms "installed," "connected," and "linked" should be interpreted broadly. For example, they can refer to mechanical or electrical connections, or internal connections between two elements; they can be direct connections or indirect connections through an intermediate medium. Those skilled in the art can understand the specific meaning of the above terms according to the specific circumstances.
[0081] In this document, unless otherwise stated, the term "multiple" means two or more.
[0082] In this article, the character " / " indicates that the objects before and after it are in an "or" relationship. For example, A / B means: A or B.
[0083] In this article, the term "and / or" describes an association between objects, indicating that three relationships can exist. For example, A and / or B means: A or B, or A and B.
[0084] It should be understood that although the steps in the flowchart are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order constraint on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the diagram may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.
[0085] The modules in the apparatus or system of this application can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.
[0086] Where there is no conflict, the embodiments and features in the embodiments of the present invention can be combined with each other.
[0087] Figure 1An embodiment of a power network resource allocation method according to the present invention is shown.
[0088] In this optional embodiment, the power network resource allocation method includes:
[0089] Step S101: Analyze each user agent to obtain the environmental state before the slice task, the environmental state after the slice task, the slice task reward, and the network slice scheme for executing the slice task. Then, merge and organize the data to obtain the training dataset.
[0090] Step S103: Train each user agent according to the training dataset to obtain new slice selection strategy parameters for each user agent; and upload the new slice selection strategy parameters of each user agent to the cloud server so that the cloud server can perform federated aggregation based on the received new slice selection strategy parameters to obtain global slice selection strategy parameters.
[0091] In step S105, each user agent trains itself according to the global slice selection strategy parameters. After the user agent training is completed, the corresponding network slice is selected based on the game theory algorithm and the global slice selection strategy parameters to generate the best new network slice scheme.
[0092] In this optional embodiment, the power network resource allocation method further includes: before obtaining the training dataset, each user agent obtains its own task information and slice selection strategy parameters according to the current environmental state of the power network; each user agent calculates a decision action according to the current environmental state and slice selection strategy parameters; and selects the corresponding network slice to execute the task information according to the decision action to obtain a new environmental state and task reward.
[0093] In this optional embodiment, when selecting the corresponding network slice to generate the optimal new network slice scheme based on the game theory algorithm and the global slice selection strategy parameters, a payoff matrix can be generated according to the pre-acquired network slice usage statistics and the global slice selection strategy parameters; based on the payoff matrix, the Nash equilibrium point is calculated using the game theory algorithm, and based on the Nash equilibrium point, the corresponding network slice is selected to generate the optimal new network slice scheme.
[0094] In this optional embodiment, when generating the revenue matrix based on pre-acquired network slice usage statistics and the global slice selection strategy parameters, the user group data characteristics are obtained by analyzing the pre-acquired network slice usage statistics; the power network is analyzed according to the pre-configured slice selection strategy to generate a set of network slice schemes, which includes slice selection strategies that users can choose and network slice schemes that slice managers can choose; based on the user group data characteristics, the user objectives and slice manager objectives are substituted into the revenue functions of the user and slice manager to obtain the corresponding revenue values; and a revenue matrix is generated based on the slice selection strategies that users can choose, the network slice schemes that slice managers can choose, the user revenue values, and the slice manager revenue values.
[0095] Figure 2 An embodiment of a power network resource allocation system according to the present invention is shown.
[0096] In this optional embodiment, the power network resource allocation system includes:
[0097] The training dataset generation module 201 is used to analyze each user agent, obtain the environmental state before each user agent executes the slicing task, the environmental state after the slicing task, the slicing task reward, and the network slicing scheme for executing the slicing task, and merge and organize the data to obtain the training dataset.
[0098] The agent federated training module 203 is used to train each user agent according to the training dataset to obtain new slice selection policy parameters for each user agent; and upload the new slice selection policy parameters of each user agent to the cloud server so that the cloud server can perform federated aggregation according to the received new slice selection policy parameters to obtain global slice selection policy parameters.
[0099] The network slice allocation module 205 is used for each user agent to train the user agent according to the global slice selection strategy parameters, and after the user agent training is completed, to select the corresponding network slice based on the game theory algorithm and the global slice selection strategy parameters to generate the best new network slice scheme.
[0100] In this optional embodiment, the power network resource allocation system further includes: a data acquisition module (not shown in the figure), used to obtain the task information and slice selection strategy parameters of each user agent based on the current environmental state of the power network before obtaining the training dataset; and a data calculation module (not shown in the figure), used to calculate the decision action of each user agent based on the current environmental state and slice selection strategy parameters; and to select the corresponding network slice to execute the task information based on the decision action, thereby obtaining a new environmental state and task reward.
[0101] In this optional embodiment, the network slice allocation module 205 includes a payoff matrix generation module (not shown in the figure) and an equilibrium point allocation module (not shown in the figure). The payoff matrix generation module is used to generate a payoff matrix based on pre-acquired network slice usage statistics and the global slice selection strategy parameters. The equilibrium point allocation module is used to calculate the Nash equilibrium point using a game theory algorithm based on the payoff matrix, and select the corresponding network slice to generate the optimal new network slice scheme based on the Nash equilibrium point.
[0102] In this optional embodiment, the revenue matrix generation module includes: a user analysis module (not shown in the figure), a set generation module (not shown in the figure), a revenue calculation module (not shown in the figure), and a matrix generation module (not shown in the figure). The user analysis module is used to analyze users based on pre-acquired statistical data on network slice usage to obtain user group data characteristics. The set generation module is used to analyze the power network according to a pre-configured slice selection strategy to generate a set of network slice schemes, including slice selection strategies selectable by users and network slice schemes selectable by slice managers. The revenue calculation module is used to substitute user objectives and slice manager objectives into the revenue functions of users and slice managers based on the user group data characteristics to obtain corresponding revenue values. The matrix generation module is used to generate a revenue matrix based on the slice selection strategies selectable by users, the network slice schemes selectable by slice managers, and the user and slice manager revenue values.
[0103] To facilitate understanding of the above technical solutions of the present invention, the following further describes the above technical solutions of the present invention from the perspectives of architecture and principle, as follows:
[0104] In practical applications, this invention is implemented based on the 5G SA system architecture of the 3GPP standard. The network slicing function end-to-end includes: terminal, radio access network, transport network, edge cloud, and central cloud. In 3GPP TS 38.300 V15.6.0, the network slicing function can divide a user's physical network into multiple logical networks, enabling one network to serve multiple purposes. Network slicing allows users to build multiple end-to-end, virtual, isolated, and on-demand customized dedicated logical networks on a single physical network to meet the different requirements of various industry customers for network capabilities (latency, bandwidth, number of connections, reliability, etc.).
[0105] The architecture of the power network resource allocation system described in this invention is as follows: Figure 3 As shown, its allocation process is as follows: Figure 4 As shown, the power network resource allocation system consists of three main parts: 1. End-to-end power grid slicing architecture: The physical network infrastructure is virtualized into end-to-end connection paths. These paths are divided into network slices with different attributes (such as bandwidth, latency, and cost) according to the slicing scheme. Users of the network select appropriate slices to deploy tasks based on their task characteristics (bandwidth requirements, computational load, latency constraints, etc.) through the network slicing decision system. The actual deployment of tasks will change the state of the entire network environment. 2. Slicing selection function based on federated reinforcement learning: Each user is treated as an agent, and the decision network is trained through federated reinforcement learning to make slice selection decisions that maximize user task execution satisfaction. 3. Slicing scheme generation function based on game theory: The user group and the slice management party are the two sides in a game. Based on the user decision network and feasible slicing schemes, a payoff matrix for both sides is generated. Through game theory algorithms, a Nash equilibrium point is found, thereby obtaining the slicing scheme that maximizes the payoff for the slice management party (such as system network stability, resource balance, and profit).
[0106] Figure 5 The diagram illustrates the principle of power network resource allocation. In this invention, each user in the network is an intelligent agent, labeled as... ; These are the task attributes and requirements that the user needs to perform; yes Real-time monitoring of key statuses of devices related to intelligent network slicing, such as available bandwidth, latency, and reliability; It is the environmental state observed by agent i, which includes user task information and system slice information. It is a slice selection policy model based on reinforcement learning, where wi is the model parameter for user i. The model calculates the action the agent should take by inputting the current environmental state; It refers to the action taken by agent i at time t, which means selecting a slice. Where C represents the slicing scheme. This is the j-th slice. The detailed steps of the solution are as follows:
[0107] (1) User (agent) i observes the current environment and obtains its own task information. Slicing scheme and current network system status Slice selection strategy model parameter w;
[0108] (2) User (agent) i, based on the current environmental state and slice selection strategy Calculate decision-making actions Select slice Assign tasks;
[0109] (3) Each task is executed on the selected slice, and user (agent) i obtains the new environmental state. and their respective returns The reward is calculated based on factors such as the customer satisfaction of user (agent) i and the balance of system resources.
[0110] (4) Based on the training data Train the agents, calculate the loss function value, and update the policy parameters of each agent. ;
[0111] (5) Set the edge cloud / central cloud as the federated learning coordinator and each user as a federated learning participant;
[0112] (6) Each participant uploads its local model parameters and local gradient updates to the coordinating node. For example, local model parameters... and updating gradients .
[0113] (7) The coordinator updates the global model parameters. The parameters are then distributed to all participating intelligent agents to ensure uniformity. ;
[0114] (8) Determine whether the model parameters of each participant have converged or whether the number of training iterations has reached the threshold. If converged or the threshold has been reached, stop the training loop of the federated reinforcement learning model and then execute the slicing scheme based on game theory. If it does not converge, use the agent again to allocate slices and carry out the next round of training.
[0115] Since the slice allocation scheme needs to maximize business benefits while meeting customer needs, the problem is constructed as a game theory-based slice allocation, where the user group and the slice management party are the two sides in the game. Let represent the utility function of a game participant in the state space and strategy / action space. Specifically:
[0116] The slice management team analyzes user group data characteristics (such as the number of users, user type distribution, etc.) based on slice usage statistics, and collects slice selection strategies trained through federated reinforcement learning. Other feasible slice selection strategies, such as expert experience, greedy algorithms, and historical parameters. generated And so on, and generate a set of optional slicing schemes.
[0117] The user group's strategic action space is the selectable slice selection strategy P, and the slice manager's strategic action space is the selectable slice scheme C.
[0118] By combining statistical data characteristics, user objectives (such as average user satisfaction) and slice management objectives (such as system resource balance or revenue) are incorporated into the revenue functions of users and slice management. The profit values for both parties are obtained;
[0119] Based on the slicing schemes available to the administrator and the slicing selection strategies of the users, a payoff matrix is generated; to maximize the payoffs of each party, the optimal equilibrium point is calculated using game theory algorithms. Based on the calculated equilibrium point, the slice management entity generates a new slice plan. .
[0120] Furthermore, in practical applications, the power grid industry has two main categories of typical 5G slicing service requirements: ultra-reliable low-latency communication (uRLLC) slices, belonging to the production control area slices; and enhanced mobile bandwidth (eMBB) and massive machine-type communication (mMTC) slices, belonging to the management information area slices. This invention targets 5G high-capacity, high-bandwidth (eMBB) application scenarios, selecting precise load control services. These services are primarily used to prioritize the disconnection of interruptible loads (such as HVAC and indoor lighting) during power grid faults to ensure load balance in the power grid system, thereby guaranteeing the normal operation of uninterrupted loads (such as those used in hospitals and factories). This service has high real-time and reliability requirements. Based on network bandwidth and latency analysis, a slicing scheme is proposed to develop services for the 5G eMBB scenario. The power service topology diagram for slicing implementation is shown below. Figure 6 As shown.
[0121] When using this method, assume that the task requires a communication end-to-end latency of less than [a certain value]. Bandwidth requirement greater than There are two eligible slices; the slice information is time delay. ,bandwidth .
[0122] User (agent) i, based on the current system state And slice selection strategy, calculate decision actions Select a slice and assign tasks; the intelligent inspection task is executed on the selected slice, and the user obtains the new system status. and their respective returns According to < Train the agents and update the policy parameters of each agent. The edge cloud / central cloud is designated as the coordinator of the federated learning process, and each terminal is a participant in the federated learning process, resulting in a slice selection strategy. .
[0123] The slice management team uses slice usage statistics, such as user type distribution and slice selection strategies. and a set of optional slicing schemes ,in Scheme C i There are two slices. The user's goal and the slice manager's goal are respectively substituted into the user's and slice manager's reward functions. The reward function is constructed from user satisfaction, where user i's satisfaction is... , The average revenue function for the user group is:
[0124]
[0125] In the formula User revenue value; Average user satisfaction; End-to-end delay for power network communication; Network slicing latency; For the bandwidth of the power grid; N represents the network slice bandwidth; N is the number of users. The penalty for system crashes.
[0126] The slice management revenue function is:
[0127]
[0128] In the formula, This represents the revenue value for the slice management team. For the task In the corresponding network slice Bandwidth used; For the task In the corresponding network slice Bandwidth used; For the task In the corresponding network slice The bandwidth used.
[0129] Calculate the revenue for the slice manager and users under each slicing scheme, and generate a revenue matrix. The required revenue functions for each must satisfy the Nash equilibrium condition. The slice manager calculates the Nash equilibrium point, aiming to maximize revenue, if the revenue matrix satisfies the following condition:
[0130]
[0131]
[0132] In the formula, This is the Nash equilibrium point, i.e., the optimal new slicing scheme; User revenue value; This represents the revenue value for the slice management team. Selectable slice selection strategy; Choose a strategy for slicing; Available slicing schemes; For the task The corresponding network slice.
[0133] Figure 7 An embodiment of a computer device according to the present invention is shown. This computer device may be a server and includes a processor, memory, and a network interface connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database stores static and dynamic information data. The network interface communicates with external terminals via a network connection. When the computer program is executed by the processor, it implements the steps in the above-described method embodiment.
[0134] Those skilled in the art will understand that Figure 7 The structure shown is merely a block diagram of a portion of the structure related to the present invention and does not constitute a limitation on the computer device to which the present invention is applied. A specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0135] In one embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.
[0136] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the steps in the method embodiments described above.
[0137] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the methods described above. Any references to memory, storage, databases, or other media used in the embodiments provided by this invention can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical storage, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0138] This invention is not limited to the structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this invention is limited only by the appended claims.
Claims
1. A method for allocating power network resources, characterized in that, The power network has several user agents, and the power network resource allocation method includes: The analysis of each user agent is performed to obtain the environmental state before the slice task is executed, the environmental state after the slice task is executed, the slice task reward, and the network slice scheme for executing the slice task. The data is then merged and organized to obtain the training dataset. Based on the training dataset, each user agent is trained to obtain new slice selection strategy parameters for each user agent; and the new slice selection strategy parameters for each user agent are uploaded to the cloud server so that the cloud server can perform federated aggregation based on the received new slice selection strategy parameters to obtain global slice selection strategy parameters. Each user agent trains itself according to the global slice selection strategy parameters. After the user agent training is completed, it selects the corresponding network slice based on the game theory algorithm and the global slice selection strategy parameters to generate the best new network slice scheme. The process of selecting the corresponding network slice to generate the optimal new network slice scheme based on the game theory algorithm and the global slice selection strategy parameters includes: generating a payoff matrix based on the pre-acquired network slice usage statistics and the global slice selection strategy parameters; calculating the Nash equilibrium point using the game theory algorithm based on the payoff matrix; and selecting the corresponding network slice to generate the optimal new network slice scheme based on the Nash equilibrium point. The formula for calculating the Nash equilibrium point is: In the formula, This is the Nash equilibrium point, i.e., the optimal new slicing scheme; User revenue value; This represents the revenue value for the slice management team. Choose a strategy for slicing; Selectable slice selection strategy; Available slicing schemes; For the task Corresponding network slices; The formula for calculating the user's revenue function is: , In the formula, User revenue value; Average user satisfaction; End-to-end delay for power network communication; Network slicing latency; For the bandwidth of the power grid; N represents the network slice bandwidth; N is the number of users. The penalty for system crash; The formula for calculating the slice management's revenue function is as follows: In the formula, This represents the revenue value for the slice management party. For the task In the corresponding network slice Bandwidth used; For the task In the corresponding network slice Bandwidth used; For the task In the corresponding network slice The bandwidth used.
2. The power network resource allocation method according to claim 1, characterized in that, Also includes: Before obtaining the training dataset, each user agent acquires its own task information and slice selection strategy parameters based on the current environmental state of the power network. Each user agent calculates a decision action based on the current environment state and slice selection strategy parameters; and selects the corresponding network slice to execute the task information based on the decision action, thereby obtaining a new environment state and task reward.
3. The power network resource allocation method according to claim 1, characterized in that, Based on pre-acquired network slice usage statistics and the global slice selection strategy parameters, a revenue matrix is generated, including: Based on the statistical data on the practical use of network slices obtained in advance, the user group data characteristics are analyzed. Based on the pre-configured slice selection strategy, the power network is analyzed to generate a set of network slice schemes. The set of network slice schemes includes slice selection strategies that users can choose and network slice schemes that slice managers can choose. Based on the user group data characteristics, the user goals and the slice management goals are substituted into the user and slice management revenue functions to obtain the corresponding revenue values; A revenue matrix is generated based on the slice selection strategies available to users, the network slicing schemes available to slice managers, and the revenue values of users and slice managers.
4. The power network resource allocation method according to claim 3, characterized in that, The user group data characteristics include the number of users and the distribution of user types; the user goals include average user satisfaction; and the slice management goals include system resource balancing.
5. A power network resource allocation system, characterized in that, The power network has several user intelligent agents, and the power network resource allocation system includes: The training dataset generation module is used to analyze each user agent, obtain the environmental state before the user agent executes the slicing task, the environmental state after the slicing task, the slicing task reward, and the network slicing scheme for executing the slicing task, and merge and organize the data to obtain the training dataset. The agent federated training module is used to train each user agent according to the training dataset to obtain new slice selection policy parameters for each user agent; and to upload the new slice selection policy parameters of each user agent to the cloud server so that the cloud server can perform federated aggregation according to the received new slice selection policy parameters to obtain global slice selection policy parameters. The network slice allocation module is used for each user agent to train the user agent according to the global slice selection strategy parameters, and after the user agent training is completed, to select the corresponding network slice based on the game theory algorithm and the global slice selection strategy parameters to generate the best new network slice scheme. The network slice allocation module includes a payoff matrix generation module and an equilibrium point allocation module. The payoff matrix generation module is used to generate a payoff matrix based on pre-acquired network slice usage statistics and global slice selection strategy parameters. The equilibrium point allocation module is used to calculate the Nash equilibrium point using a game theory algorithm based on the payoff matrix, and select the corresponding network slice based on the Nash equilibrium point to generate the optimal new network slice scheme. The formula for calculating the Nash equilibrium point is: In the formula, This is the Nash equilibrium point, i.e., the optimal new slicing scheme; User revenue value; This represents the revenue value for the slice management party. Selectable slice selection strategy; Choose a strategy for slicing; Available slicing schemes; For the task Corresponding network slices; The formula for calculating the user's revenue function is: , In the formula, User revenue value; Average user satisfaction; End-to-end delay for power network communication; Network slicing latency; For the bandwidth of the power grid; N represents the network slice bandwidth; N is the number of users. The penalty for system crashes; The formula for calculating the user's revenue function is: , In the formula, User revenue value; Average user satisfaction; End-to-end delay for power network communication; Network slicing latency; For the bandwidth of the power grid; N represents the network slice bandwidth; N is the number of users. The penalty for system crashes.
6. The power network resource allocation system according to claim 5, characterized in that, Also includes: The data acquisition module is used to enable each user agent to obtain its own task information and slice selection strategy parameters based on the current environmental state of the power network before obtaining the training dataset. The data calculation module is used by each user agent to calculate decision actions based on the current environment state and slice selection strategy parameters; and to select the corresponding network slice to execute the task information based on the decision actions, thereby obtaining a new environment state and task reward.
7. The power network resource allocation system according to claim 6, characterized in that, The revenue matrix generation module includes: a user analysis module, a set generation module, a revenue calculation module, and a matrix generation module, wherein, The user analysis module is used to analyze users based on pre-acquired statistical data on the usage of network slices, and to obtain user group data characteristics. The set generation module is used to analyze the power network according to the pre-configured slice selection strategy and generate a set of network slice schemes. The set of network slice schemes includes slice selection strategies that users can choose and network slice schemes that slice managers can choose. The revenue calculation module is used to substitute the user's target and the slice manager's target into the user's and slice manager's revenue functions based on the user group data characteristics to obtain the corresponding revenue value. The matrix generation module is used to generate a revenue matrix based on the slice selection strategy that the user can choose, the network slicing scheme that the slice manager can choose, as well as the user's revenue value and the slice manager's revenue value.
8. The power network resource allocation system according to claim 7, characterized in that, The user group data characteristics include the number of users and the distribution of user types; the user goals include average user satisfaction; and the slice management goals include system resource balancing.