Community service network resource scheduling method and device based on hierarchical reinforcement learning

Through the hierarchical reinforcement learning method, the industrial system is divided into communities, high- and low-level strategy models are constructed, and reward functions are designed. This solves the problems of low efficiency and poor stability of traditional reinforcement learning in high-dimensional state space, and realizes an efficient and robust resource scheduling strategy that adapts to dynamic environmental changes.

CN120672044AActive Publication Date: 2025-09-19BEIHANG UNIV
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510752843.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-09-19
Estimated Expiration
2045-06-06

AI Technical Summary

Technical Problem

Traditional reinforcement learning strategies are prone to the "curse of dimensionality" when faced with high-dimensional state spaces and continuous action spaces. The strategy training process converges slowly, has low efficiency, and poor stability. It is difficult to combine the community structure and task hierarchy characteristics in industrial systems, resulting in local optimal resource scheduling strategies and poor global performance. In addition, they are unstable when faced with changes in topology or dynamic adjustments to task requirements, limiting their flexibility and robustness in practical applications.

Method used

A hierarchical reinforcement learning method is used to divide the industrial system into multiple non-overlapping communities. High-level and low-level strategy models are constructed. Global and local reward functions are designed with global benefits and local scheduling satisfaction as the goals respectively. By alternately training high-level and low-level strategy models, global guidance and local resource coordinated optimization driven by the dominant node are achieved.

Benefits of technology

It improves the training efficiency, reliability and generalization ability of resource scheduling strategies, reduces computational complexity, enhances the structural perception and adaptability of strategies, and enables efficient resource allocation in dynamic environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120672044A_ABST
    Figure CN120672044A_ABST
Patent Text Reader

Abstract

The invention provides a community service network resource scheduling method and device based on hierarchical reinforcement learning, and the method comprises the steps: dividing a sample community service network, and obtaining a plurality of non-overlapping sample communities; constructing a high-level strategy model and a low-level strategy model, constructing a global reward function of the high-level strategy model by taking the global benefit as a target, and constructing a local reward function of the low-level strategy model by taking the local scheduling satisfaction as a target; and performing parameter updating on the high-level policy model and the low-level policy model by using the award value of each layer until each model is converged to obtain a high-level policy model and a low-level policy model, dividing the community service network to be scheduled to obtain a plurality of non-overlapped communities to be scheduled, scheduling each community to be scheduled based on the trained high-level policy model and the trained low-level policy model, and scheduling the community to be scheduled. The resource allocation strategy of each community to be scheduled is obtained, the scheduling efficiency is improved, the calculation complexity is reduced, and the structural perception and generalization ability of the strategy is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intersectional technologies of artificial intelligence and industrial Internet, and in particular to a community service network resource scheduling method and device based on hierarchical reinforcement learning. Background Art

[0002] Amid the rapid development of the Industrial Internet, industrial systems are increasingly characterized by high interconnectivity, dynamic heterogeneity, and diverse tasks. Industrial Internet networks typically consist of a large number of heterogeneous devices, sensors, and service nodes, with complex resource collaboration and data exchange relationships between them. In such systems, efficient resource scheduling and optimal allocation directly impact the production efficiency and economic benefits of the entire industrial process. However, due to the complexity of the network structure (such as strong community characteristics, node heterogeneity, and economic constraints between edges), resource scheduling often presents itself as a high-dimensional, non-convex, and tightly coupled optimization problem. Traditional centralized optimization methods struggle to meet the real-time and robustness requirements of large-scale environments, while static strategies struggle to adapt to dynamic load and environmental changes. This results in low overall resource utilization, slow scheduling response, and poor coordination efficiency.

[0003] Reinforcement learning, as an intelligent decision-making method based on interactive learning, has shown great potential in solving dynamic resource management problems in recent years, especially in the face of unknown system dynamics and high-dimensional state space, showing strong adaptability and optimization capabilities.

[0004] However, as industrial systems scale, traditional reinforcement learning strategies are prone to the "curse of dimensionality" when faced with high-dimensional state spaces and continuous action spaces. This leads to slow convergence, low efficiency, and poor stability during policy training. Most reinforcement learning models employ a single policy structure, failing to incorporate the community structure and task-level characteristics of industrial systems. This makes it difficult to achieve effective coordination between the global and local aspects, resulting in locally optimal scheduling strategies and poor global performance. Traditional strategies are often optimized within fixed environments, making them unstable when faced with topological changes or dynamic adjustments to task requirements. This makes policy migration and rapid reuse difficult, limiting their flexibility and robustness in practical applications.

[0005] Therefore, how to improve the training efficiency, reliability, and generalization ability of the resource scheduling strategy model is a technical problem that needs to be solved urgently by those skilled in the art. Summary of the Invention

[0006] The present invention provides a community service network resource scheduling method and device based on hierarchical reinforcement learning, which are used to solve the defects of low training efficiency, poor reliability and poor generalization ability of resource scheduling strategy models in the prior art.

[0007] In one aspect, the present invention provides a community service network resource scheduling method based on hierarchical reinforcement learning, comprising: Construct communities and divide the sample community service network to obtain multiple non-overlapping sample communities; Constructing a high-level policy model, wherein the input of the high-level policy model is the global resource state of the sample community, and the output is the dominant node of the sample community; Constructing a low-level strategy model, wherein the input of the high-level strategy model is the leading node of the sample community, and the output is the resource allocation strategy of each node of the sample community; Constructing a reward function, taking global benefit as the goal, to construct a global reward function of the high-level policy model, and taking local scheduling satisfaction as the goal, to construct a local reward function of the low-level policy model; Model training, using the global reward value of the global reward function and the local reward value of the local reward function to update the parameters of the high-level policy model and the low-level policy model until the high-level policy model and the low-level policy model converge to obtain a high-low-level policy model; The model is applied to divide the service network of the communities to be scheduled into multiple non-overlapping communities to be scheduled. Based on the trained high- and low-level strategy models, each community to be scheduled is scheduled to obtain the resource allocation strategy of each community to be scheduled.

[0008] According to a community service network resource scheduling method based on hierarchical reinforcement learning provided by the present invention, each node in the sample community includes the current resource holdings, unit resource storage cost and unit resource benefits; With the goal of maximizing global benefits, the global reward function of the high-level policy model is constructed, including: A global reward function of the high-level policy model is constructed using a Markov decision process based on the current resource holdings, the unit resource storage cost, and the revenue brought by the unit resource.

[0009] According to a community service network resource scheduling method based on hierarchical reinforcement learning provided by the present invention, the global reward function of the high-level strategy model is:

[0010] in, represents the global reward value, Indicates the The current resource holdings of nodes, Indicates the The current resource holdings of nodes, Indicates the The benefits brought by the unit resources of each node, It represents the measurement of node output efficiency. Indicates the The unit resource storage cost of each node, represents the storage cost adjustment cost, represents the transmission cost, represents the resource coordination cost across nodes, represents the number of nodes in the sample community service network, represents the adjustment factor.

[0011] According to a community service network resource scheduling method based on hierarchical reinforcement learning provided by the present invention, each node in the sample community also includes a task demand; With the goal of maximizing local scheduling satisfaction, the local reward function of the low-level strategy model is constructed, including: A local reward function of the low-level policy model is constructed based on the current resource holdings and the task requirements using a Markov decision process.

[0012] According to a community service network resource scheduling method based on hierarchical reinforcement learning provided by the present invention, the local reward function of the low-level strategy model is:

[0013] Among them, the represents the local reward value, Indicates the The current resource holdings of nodes, Indicates the The task demand of each node.

[0014] According to a community service network resource scheduling method based on hierarchical reinforcement learning provided by the present invention, the state space of each sample community is:

[0015] Among them, the Indicates the The state space of sample communities, Indicates the The number of nodes in a sample community, Indicates the The feature vector of each node includes the current resource holding amount, task demand amount, unit resource storage cost and unit resource benefit.

[0016] According to a community service network resource scheduling method based on hierarchical reinforcement learning provided by the present invention, the action space of the low-level strategy model is:

[0017] Among them, the Represents the resource allocation strategy set for each node, represents the resource allocation strategy of the dominant node, Indicates the resource allocation policy for non-dominant nodes.

[0018] According to a community service network resource scheduling method based on hierarchical reinforcement learning provided by the present invention, the reward value of the global reward function and the global reward function are used to update the parameters of the high-level strategy model and the low-level strategy model until the high-level strategy model and the low-level strategy model converge to obtain the high-level and low-level strategy models, including: Determining an objective function that maximizes device resource utilization and coordination efficiency based on the global reward function and the local reward function; Utilizing the reward value of the global reward function and the global reward function, the high-level strategy model and the low-level strategy model are alternately updated with parameters until the value change of the objective function is less than a preset change, thereby obtaining the high- and low-level strategy models.

[0019] On the other hand, the present invention also provides a community service network resource scheduling device based on hierarchical reinforcement learning, which includes: The community construction module is used to divide the sample community service network into multiple non-overlapping sample communities; A high-level policy model construction module is used to construct a high-level policy model, wherein the input of the high-level policy model is the global resource state of the sample community, and the output is the dominant node of the sample community; A low-level policy model construction module is used to construct a low-level policy model, wherein the input of the high-level policy model is the leading node of the sample community, and the output is the resource allocation strategy of each node of the sample community; A reward function construction module, configured to construct a global reward function of the high-level policy model with global benefit as the goal, and to construct a local reward function of the low-level policy model with local scheduling satisfaction as the goal; A training module, configured to update parameters of the high-level policy model and the low-level policy model using the reward value of the global reward function and the global reward function until both the high-level policy model and the low-level policy model converge to obtain a high-low-level policy model; The resource scheduling module is used to divide the service network of the to-be-scheduled communities into multiple non-overlapping to-be-scheduled communities, schedule each to-be-scheduled community based on the trained high- and low-level strategy models, and obtain the resource allocation strategy for each to-be-scheduled community.

[0020] On the other hand, the present invention also provides an electronic device, which includes a memory, a processor, and a computer program stored in the memory and runnable on the processor. When the processor executes the program, it implements the community service network resource scheduling method based on hierarchical reinforcement learning as described in any one of the above.

[0021] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the community service network resource scheduling methods based on hierarchical reinforcement learning as described above.

[0022] On the other hand, the present invention also provides a computer program product, including a computer program, which, when executed by a processor, implements any of the community service network resource scheduling methods based on hierarchical reinforcement learning as described above.

[0023] The community service network resource scheduling method and device based on hierarchical reinforcement learning provided by the present invention divides the sample community service network to obtain multiple non-overlapping sample communities; constructs a high-level strategy model and a low-level strategy model, constructs a global reward function of the high-level strategy model with global benefits as the goal, and constructs a local reward function of the low-level strategy model with local scheduling satisfaction as the goal; uses the reward values ​​of each layer to update the parameters of the high-level strategy model and the low-level strategy model until each model converges. After obtaining the high- and low-level strategy models, the community service network to be scheduled is divided to obtain multiple non-overlapping communities to be scheduled. Based on the trained high- and low-level strategy models, each community to be scheduled is scheduled to obtain the resource allocation strategy of each community to be scheduled. In this way, by constructing a two-layer reinforcement learning architecture, global guidance and local resource collaborative optimization driven by the dominant node can be achieved. This method models the service network as a graph structure with community divisions, selects the dominant node within the community in the high-level strategy to simplify the decision space, and uses the dominant node as a guide in the low-level strategy to schedule resources for the nodes within the community, thereby improving scheduling efficiency, reducing computational complexity, and enhancing the structural perception and generalization capabilities of the strategy. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0025] Figure 1 Schematic diagram of a process for scheduling community service network resources based on hierarchical reinforcement learning provided by an embodiment of the present invention; Figure 2 Schematic diagram of the structure of a community service network resource scheduling device based on hierarchical reinforcement learning provided by an embodiment of the present invention; Figure 3 It is a structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0026] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0027] Figure 1 It is a flowchart of a community service network resource scheduling method based on hierarchical reinforcement learning provided by an embodiment of the present invention.

[0028] like Figure 1 As shown, the execution subject of the community service network resource scheduling method based on hierarchical reinforcement learning provided by the embodiment of the present invention can be an electronic device, and the method mainly includes the following steps: 101. Construct a community and divide the sample community service network into multiple non-overlapping sample communities; In a specific implementation process, the sample community service network can be modeled as an undirected graph ,in, A collection of nodes, representing devices, sensors or service units, is an edge set, which represents the resource collaboration or communication relationship between nodes. The nodes in the sample community service network can be divided into K non-overlapping communities according to their physical distribution or functional collaboration. .

[0029] For each node It has the following resource feature vector:

[0030] in, Indicates the The current resource holdings of nodes, Indicates the The task demand of each node, Indicates the The benefits brought by the unit resources of each node, Indicates the The unit resource storage cost of each node.

[0031] side Represents a slave node Towards The connection relationship of transmission resources is determined by the transmission cost matrix express, , which represents the unit resource transmission cost.

[0032] 102. Construct a high-level policy model, wherein the input of the high-level policy model is the global resource state of the sample community, and the output is the dominant node of the sample community; In a specific implementation process, the task of the high-level policy model is to Select one or more dominant nodes to form a set , the leading node is responsible for guiding the resource allocation strategy within the sample community. The complete action space of the high-level strategy model is defined as:

[0033] Among them, the represents the complete action space of the high-level policy model, that is, the combined space of the dominant node selection actions of all sample communities, Indicates that from Society A method for selecting one or more dominant nodes from among the nodes.

[0034] In a specific implementation process, the state space of each sample community is:

[0035] Among them, the Indicates the The state space of sample communities.

[0036] The complete state space of the high-level policy model is defined as:

[0037] 103. Construct a low-level strategy model, wherein the input of the high-level strategy model is the leading node of the sample community, and the output is the resource allocation strategy of each node of the sample community; In a specific implementation process, the high-level strategy determines the dominant node Afterwards, the low-level strategy model performs specific resource allocation within each sample community. The action space of the low-level strategy model is:

[0038] Among them, the Represents the resource allocation strategy set for each node, represents the resource allocation strategy of the dominant node, Represents the resource allocation strategy for non-dominant nodes. To ensure the feasibility of resource allocation, the resource amount of non-dominant nodes is realized through standard simple projection:

[0039] in, Represents the standard simple projection.

[0040] 104. Constructing a reward function, taking global benefit as a goal, constructing a global reward function of the high-level policy model, and constructing a local reward function of the low-level policy model with local scheduling satisfaction as a goal; In a specific implementation process, in order to achieve effective coordination of multi-layer strategies, the present invention designs a structured reward function, which includes a global reward function of the high-level strategy model constructed with global benefits as the goal, and a local reward function of the low-level strategy model constructed with local scheduling satisfaction as the goal.

[0041] Specifically, a Markov decision process may be used to construct a global reward function of the high-level policy model based on the current resource holdings, the unit resource storage cost, and the revenue brought by the unit resource.

[0042] The global reward function of the high-level policy model is:

[0043] in, represents the global reward value, Indicates the The current resource holdings of nodes, Indicates the The current resource holdings of nodes, Indicates the The benefits brought by the unit resources of each node, It represents the measurement of node output efficiency. Indicates the The unit resource storage cost of each node, represents the storage cost adjustment cost, represents the transmission cost, represents the resource coordination cost across nodes, represents the number of nodes in the sample community service network, represents the adjustment factor.

[0044] In a specific implementation process, for the global reward function of the above-mentioned high-level strategy model, the quadratic term is used to constrain excessive accumulation of resources, avoid global inefficiency caused by "resource hoarding", reflect the law of diminishing marginal benefits, and be closer to industrial reality. The sigmoid function is used to characterize the nonlinear growth of storage costs. When the resource holdings When it is small, it is cost-sensitive (encourages resource flow); when When it is large, the cost growth rate slows down (allowing moderate caching), which is in line with the actual characteristics of storage resources in industrial systems. The square term of (transmission cost) and resource difference can enhance the balance between adjacent nodes and avoid the "resource island" problem caused by the topological structure. In this way, the reward function can have "structure perception" capabilities, which can guide the high-level policy model to select the dominant node that matches the topological structure and optimize the global resource flow path.

[0045] In a specific implementation process, a Markov decision process can be used to construct a local reward function of the low-level strategy model based on the current resource holdings and the task requirements. The local reward function of the low-level strategy model is:

[0046] Among them, the represents the global reward value, Indicates the The current resource holdings of nodes, Indicates the The task demand of each node.

[0047] 105. Model training: using the global reward value of the global reward function and the local reward value of the local reward function to update the parameters of the high-level policy model and the low-level policy model until both the high-level policy model and the low-level policy model converge to obtain a high-level and low-level policy model; In a specific implementation, interaction with a simulation environment can be performed to obtain a global reward value of a global reward function and a local reward value of a local reward function. Based on the global reward function and the local reward function, an objective function that maximizes the resource utilization and collaborative efficiency of the device can be determined. The global reward value of the global reward function and the local reward value of the local reward function can be used to alternately update the parameters of the high-level policy model and the low-level policy model until both the high-level policy model and the low-level policy model converge, thereby obtaining a high-level and low-level policy model.

[0048] That is, in each training cycle, the low-level policy model is first fixed and the parameters of the high-level policy model are updated; then the high-level policy model is fixed and the parameters of the low-level policy model are updated. Through alternating training and joint optimization of the two-layer policy, the ultimate goal is to maximize the resource utilization and collaborative efficiency of the device.

[0049] Among them, the objective function of maximizing the resource utilization benefit and collaborative efficiency of the device is as follows:

[0050] Among them, the Expressed as a discount factor.

[0051] 106. Apply the model to divide the service network of the to-be-scheduled communities into multiple non-overlapping to-be-scheduled communities, schedule each to-be-scheduled community based on the trained high- and low-level strategy models, and obtain a resource allocation strategy for each to-be-scheduled community.

[0052] In a specific implementation process, after obtaining the high- and low-level strategy models using steps 101-105, the service network of the community to be scheduled can be divided to obtain multiple non-overlapping communities to be scheduled. Based on the trained high- and low-level strategy models, each community to be scheduled is scheduled to obtain the resource allocation strategy of each community to be scheduled.

[0053] The community service network resource scheduling method based on hierarchical reinforcement learning in this embodiment divides the sample community service network to obtain multiple non-overlapping sample communities; constructs a high-level strategy model and a low-level strategy model, and constructs a global reward function of the high-level strategy model with global benefits as the goal, and constructs a local reward function of the low-level strategy model with local scheduling satisfaction as the goal; uses the reward values ​​of each layer to update the parameters of the high-level strategy model and the low-level strategy model until all models converge and the high- and low-level strategy models are obtained. Then, the community service network to be scheduled is divided to obtain multiple non-overlapping communities to be scheduled. Based on the trained high- and low-level strategy models, each community to be scheduled is scheduled to obtain the resource allocation strategy of each community to be scheduled. In this way, by constructing a two-layer reinforcement learning architecture, global guidance and local resource collaborative optimization driven by the dominant node can be achieved. This method models the service network as a graph structure with community divisions, selects the dominant node within the community in the high-level strategy to simplify the decision space, and uses the dominant node as a guide in the low-level strategy to schedule resources for the nodes within the community, thereby improving scheduling efficiency, reducing computational complexity, and enhancing the structural perception and generalization capabilities of the strategy.

[0054] In a specific implementation, dynamic changes in the service network of the scheduled community, such as device startup and shutdown, task migration, etc., may cause changes in the relationships between nodes within the divided qth scheduled community. Therefore, in order to adapt to the dynamic changes in the service network of the scheduled community, in an embodiment of the present invention, dynamic indicators such as the communication frequency and resource interaction intensity of each node in the qth scheduled community can be monitored in real time, and the qth scheduled community can be updated based on these indicators. For example, the communication frequency can be measured by counting the number of inter-node communications per unit time, and the resource interaction intensity can be quantified based on the scale or rate of resource transmission between nodes, thereby constructing a time sequence diagram reflecting the real-time status of the network.

[0055] Specifically, the activity of the qth community to be scheduled can be calculated. If the activity of the qth community to be scheduled is greater than the first preset activity, it means that the nodes in the qth community to be scheduled interact more frequently, its internal structure is complex or the task types are diverse, and it needs to be split to optimize management and resource scheduling. Among them, grouping can be based on communication mode: nodes with high-frequency communication and close connections are divided into a sub-community to ensure high communication efficiency and low latency within the sub-community. For example, on an automated production line, equipment nodes responsible for the same process are divided into a sub-community due to frequent interaction of process data. Grouping can be based on resource type: nodes that process the same type of resources or have strong resource complementarity are grouped into a sub-community. For example, in a data processing community, nodes with rich storage resources are combined with nodes with strong computing power to form a sub-community that focuses on specific data processing tasks. I will not give examples one by one here.

[0056] In a specific implementation, if the activity of the qth to-be-scheduled community is less than a second preset activity, it indicates that the nodes in the qth to-be-scheduled community are insufficiently interacting and need to be merged. The second preset activity is less than the first preset activity. The qth to-be-scheduled community can be merged with other to-be-scheduled communities whose activity is less than the second preset activity into one.

[0057] Based on the same general inventive concept, the present invention also protects a community service network resource scheduling device based on hierarchical reinforcement learning. The community service network resource scheduling device based on hierarchical reinforcement learning provided by the present invention is described below. The community service network resource scheduling device based on hierarchical reinforcement learning described below and the community service network resource scheduling method based on hierarchical reinforcement learning described above can be referenced to each other.

[0058] Figure 2 is a structural diagram of a community service network resource scheduling device based on hierarchical reinforcement learning provided by an embodiment of the present invention. Figure 2As shown, the community service network resource scheduling device based on hierarchical reinforcement learning in this embodiment includes a community construction module 21, a high-level strategy model construction module 22, a low-level strategy model construction module 23, a reward function construction module 24, a training module 25 and a resource scheduling module 26.

[0059] The community construction module 21 is used to divide the sample community service network into multiple non-overlapping sample communities; A high-level policy model construction module 22 is used to construct a high-level policy model, wherein the input of the high-level policy model is the global resource state of the sample community, and the output is the dominant node of the sample community; A low-level policy model construction module 23 is used to construct a low-level policy model, wherein the input of the high-level policy model is the leading node of the sample community, and the output is the resource allocation strategy of each node of the sample community; A reward function construction module 24 is used to construct a global reward function of the high-level policy model with global benefits as the goal, and to construct a local reward function of the low-level policy model with local scheduling satisfaction as the goal; A training module 25 is configured to update parameters of the high-level policy model and the low-level policy model using the reward value of the global reward function and the global reward function until both the high-level policy model and the low-level policy model converge to obtain a high-level and low-level policy model; The resource scheduling module 26 is used to divide the service network of the to-be-scheduled communities into multiple non-overlapping to-be-scheduled communities, schedule each to-be-scheduled community based on the trained high- and low-level strategy models, and obtain a resource allocation strategy for each to-be-scheduled community.

[0060] Figure 3 is a schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. The community service network resource scheduling device based on hierarchical reinforcement learning may include: a processor (processor) 310, a communication interface (Communications Interface) 320, a memory (memory) 330, and a communication bus 340. The processor 310, the communication interface 320, and the memory 330 communicate with each other via the communication bus 340. The processor 310 can call the logic instructions in the memory 330 to execute the community service network resource scheduling method based on hierarchical reinforcement learning.

[0061] Furthermore, the logic instructions in the aforementioned memory 330 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product, stored in a storage medium, includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0062] On the other hand, the present invention also provides a computer program product, which includes a computer program, which can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the community service network resource scheduling method based on hierarchical reinforcement learning provided by the above methods.

[0063] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the community service network resource scheduling method based on hierarchical reinforcement learning provided by the above methods.

[0064] It should be noted that the relevant information that may be involved in the various embodiments of this application are all strictly in accordance with the requirements of laws and regulations, follow the principles of legality, legitimacy and necessity, and are based on the reasonable purposes of business scenarios to process information that users actively provide during the use of products / services or generated due to the use of products / services, as well as information obtained with user authorization.

[0065] The information processed by this application will vary depending on the specific product / service scenario and should be based on the specific scenario in which the user uses the product / service. This information may involve the user's account information, device information, or other related information. This application will treat the relevant information and its processing with a high degree of diligence.

[0066] This application attaches great importance to the security of relevant information and has taken reasonable and feasible security protection measures that comply with industry standards to protect relevant information and prevent unauthorized access, public disclosure, use, modification, damage or loss of relevant information.

[0067] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.

[0068] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.

[0069] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A community service network resource scheduling method based on hierarchical reinforcement learning, characterized in that: include: Construct communities and divide the sample community service network to obtain multiple non-overlapping sample communities; Constructing a high-level policy model, wherein the input of the high-level policy model is the global resource state of the sample community, and the output is the dominant node of the sample community; Constructing a low-level strategy model, wherein the input of the high-level strategy model is the leading node of the sample community, and the output is the resource allocation strategy of each node of the sample community; Constructing a reward function, taking global benefit as the goal, to construct a global reward function of the high-level policy model, and taking local scheduling satisfaction as the goal, to construct a local reward function of the low-level policy model; Model training, using the global reward value of the global reward function and the local reward value of the local reward function to update the parameters of the high-level policy model and the low-level policy model until the high-level policy model and the low-level policy model converge to obtain a high-low-level policy model; The model is applied to divide the service network of the communities to be scheduled into multiple non-overlapping communities to be scheduled. Based on the trained high- and low-level strategy models, each community to be scheduled is scheduled to obtain the resource allocation strategy of each community to be scheduled.

2. The community service network resource scheduling method based on hierarchical reinforcement learning according to claim 1 is characterized in that: Each node in the sample community includes the current resource holdings, unit resource storage cost and unit resource benefits; With the goal of maximizing global benefits, the global reward function of the high-level policy model is constructed, including: A global reward function of the high-level policy model is constructed using a Markov decision process based on the current resource holdings, the unit resource storage cost, and the revenue brought by the unit resource.

3. The community service network resource scheduling method based on hierarchical reinforcement learning according to claim 2 is characterized in that: The global reward function of the high-level policy model is: in, represents the global reward value, Indicates the The current resource holdings of nodes, Indicates the The current resource holdings of nodes, Indicates the The benefits brought by the unit resources of each node, It represents the measurement of node output efficiency. Indicates the The unit resource storage cost of each node, represents the storage cost adjustment cost, represents the transmission cost, represents the resource coordination cost across nodes, represents the number of nodes in the sample community service network, represents the adjustment factor.

4. The community service network resource scheduling method based on hierarchical reinforcement learning according to claim 2 is characterized in that: Each node in the sample community also includes task demand; With the goal of maximizing local scheduling satisfaction, the local reward function of the low-level strategy model is constructed, including: A Markov decision process is used to construct a local reward function of the low-level policy model based on the current resource holdings and the task requirements.

5. The community service network resource scheduling method based on hierarchical reinforcement learning according to claim 4 is characterized in that: The local reward function of the low-level policy model is: Among them, the represents the local reward value, Indicates the The current resource holdings of nodes, Indicates the The task demand of each node.

6. The community service network resource scheduling method based on hierarchical reinforcement learning according to any one of claims 1 to 5, characterized in that: The state space of each sample community is: Among them, the Indicates the The state space of sample communities, Indicates the The number of nodes in the sample community, Indicates the The feature vector of each node includes the current resource holding amount, task demand amount, unit resource storage cost and unit resource benefit.

7. The community service network resource scheduling method based on hierarchical reinforcement learning according to any one of claims 1 to 5, characterized in that: The action space of the low-level strategy model is: Among them, the Represents the resource allocation strategy set for each node, represents the resource allocation strategy of the dominant node, Indicates the resource allocation policy for non-dominant nodes.

8. The community service network resource scheduling method based on hierarchical reinforcement learning according to any one of claims 1 to 5, characterized in that: Using the reward value of the global reward function and the global reward function, updating the parameters of the high-level policy model and the low-level policy model until both the high-level policy model and the low-level policy model converge to obtain a high-level and low-level policy model, including: Determining an objective function that maximizes device resource utilization and coordination efficiency based on the global reward function and the local reward function; Utilizing the reward value of the global reward function and the global reward function, the high-level strategy model and the low-level strategy model are alternately updated with parameters until the value change of the objective function is less than a preset change, thereby obtaining the high- and low-level strategy models.

9. A community service network resource scheduling device based on hierarchical reinforcement learning, characterized in that: include: The community construction module is used to divide the sample community service network into multiple non-overlapping sample communities; A high-level policy model construction module is used to construct a high-level policy model, wherein the input of the high-level policy model is the global resource state of the sample community, and the output is the dominant node of the sample community; A low-level policy model construction module is used to construct a low-level policy model, wherein the input of the high-level policy model is the leading node of the sample community, and the output is the resource allocation strategy of each node of the sample community; A reward function construction module, configured to construct a global reward function of the high-level policy model with global benefit as the goal, and to construct a local reward function of the low-level policy model with local scheduling satisfaction as the goal; A training module, configured to update parameters of the high-level policy model and the low-level policy model using the reward value of the global reward function and the global reward function until both the high-level policy model and the low-level policy model converge to obtain a high-low-level policy model; The resource scheduling module is used to divide the service network of the to-be-scheduled communities into multiple non-overlapping to-be-scheduled communities, schedule each to-be-scheduled community based on the trained high- and low-level strategy models, and obtain the resource allocation strategy for each to-be-scheduled community.

10. An electronic device, characterized in that: The method comprises a memory, a processor and a computer program stored in the memory and executable on the processor. When the processor executes the program, the method for scheduling community service network resources based on hierarchical reinforcement learning as described in any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Big data dynamic allocation and optimal scheduling method based on reinforcement learning

    CN119311407A

  • Logistics operation multi-target supervision method based on complex logistics field group

    CN119338356A

  • Service network scheduling method based on deep reinforcement learning

    CN119342102A

  • Artificial intelligence-based employee post matching and deploying method and system

    CN119494522A

  • Industrial internet multi-service real-time management method for industrial intelligence

    CN119729632A