Multi-Agent Deep Reinforcement Learning Method and System

Through the analysis of the operation logs of the multi-agent system and the construction of the interactive map, the resource competition strategy is optimized, and the problems of unbalanced and non-stable state of resource competition among multi-agents are solved, and more efficient resource utilization and system stability are achieved.

CN119561829BActive Publication Date: 2025-07-18NO 15 INST OF CHINA ELECTRONICS TECH GRP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510125216.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-27
Publication Date
2025-07-18
Estimated Expiration
2045-01-27

AI Technical Summary

Technical Problem

In the traditional multi-agent deep reinforcement learning method, it is difficult to balance resource competition among multi-agents and are prone to fall into a non-stable state.

Method used

By analyzing the operation log data of multiple agents, a multi-agent interaction map is built, load nodes are marked, resource release amount and limited intervals are calculated, resource competition strategies are formulated, and layered deep learning processing is carried out, resource competition strategies are optimized, and resource competition balance model is built.

Benefits of technology

It improves the resource competition balance between multiple agents, reduces the possibility of the system falling into a non-stable state, and improves the stability and resource utilization efficiency of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119561829B_ABST
    Figure CN119561829B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of reinforcement learning, and in particular to a multi-agent deep reinforcement learning method and system. The method includes the following steps: extracting the operation logs of multi-agents and constructing a multi-agent interaction graph to obtain a multi-agent interaction graph; calculating the limited interval of resource release between nodes with different loads in the multi-agent interaction graph to obtain the limited interval of node resource release; formulating a multi-agent resource competition strategy based on the limited interval of node resource release and coupling the hierarchical convergence behaviors of multi-agents to obtain hierarchical convergence behavior coupling data; constructing a multi-agent resource competition balance model based on the hierarchical convergence behavior coupling data to obtain a multi-agent resource competition balance model, and sending the multi-agent resource competition balance model to the cloud platform to perform multi-agent deep reinforcement learning. The present invention makes the cooperation of multi-agents more perfect through the optimization of multi-agent reinforcement learning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of reinforcement learning, and in particular, to a multi-agent deep reinforcement learning method and system. Background Art

[0002] Agents interact in a shared environment and learn how to optimize their respective behavior strategies to achieve the overall system goal through continuous attempts and feedback. Deep reinforcement learning uses deep neural networks to process large state and action spaces, enhancing the agents' perception and decision-making abilities in complex environments. In a multi-agent system, each agent not only needs to consider the impact of its own actions on the environment but also the behavior and influence of other agents. To achieve cooperation, agents need to share information and learn how to cooperate with each other; while in a competitive scenario, agents need to learn adversarial strategies. MADRL is widely applied in fields such as robot coordination control, autonomous driving, smart grid optimization, financial market trading, etc., demonstrating its powerful capabilities and potential in solving multi-agent problems through simulations and practical applications. Through multi-agent deep reinforcement learning, the adaptability and intelligence of the system can be effectively improved, laying a foundation for realizing more intelligent and automated complex systems. However, there are problems in a traditional multi-agent deep reinforcement learning method that it is difficult to balance resource competition among multi-agents and multi-agents are prone to fall into an unstable state. Summary of the Invention

[0003] Based on this, it is necessary to provide a multi-agent deep reinforcement learning method and system to solve at least one of the above technical problems.

[0004] To achieve the above object, a multi-agent deep reinforcement learning method, the method includes the following steps:

[0005] Step S1: Extract the operation logs of multi-agents to obtain multi-agent operation log data; analyze the multi-agent interaction structure of the multi-agent operation log data to obtain multi-agent interaction structure data; construct a multi-agent interaction graph based on the multi-agent interaction structure data to obtain a multi-agent interaction graph;

[0006] Step S2: Label the multi-agent load nodes of the multi-agent interaction graph to obtain multi-agent load node labeling data; calculate the resource release amounts between different nodes of the multi-agent load node labeling data to obtain load node resource release amount data; calculate the limited resource release intervals between different load nodes of the load node resource release amount data to obtain node resource release limited intervals; formulate a multi-agent resource competition strategy based on the node resource release limited intervals to obtain a multi-agent resource competition strategy;

[0007] Step S3: Perform hierarchical deep learning processing according to the multi-agent resource competition strategy to obtain resource competition hierarchical deep learning data; perform multi-agent hierarchical convergence behavior coupling based on the resource competition hierarchical deep learning data to obtain hierarchical convergence behavior coupling data;

[0008] Step S4: Construct a multi-agent resource competition balance model according to the multi-agent resource competition strategy and the hierarchical convergence behavior coupling data to obtain a multi-agent resource competition balance model, and send the multi-agent resource competition balance model to the cloud platform to perform multi-agent deep reinforcement learning.

[0009] By analyzing the operation log data of multi-agents, the present invention can deeply understand the interaction structure and pattern among each agent. This helps to identify key interaction nodes, frequently occurring interaction events, and the time and space distribution rules of interactions. The interaction graph constructed based on the interaction structure data can help to discover bottlenecks and optimization points in the system. By identifying and optimizing the key interactions among multi-agents, the overall efficiency and response speed of the system can be improved. Analyzing the interaction graph helps to identify potential system fault points or unstable factors. By better managing and regulating the interactions among multi-agents, the possibility of system crashes or failures can be reduced, and the stability and reliability of the system can be enhanced. According to the load node annotation data and resource release amount data, the resource usage situation between different nodes can be accurately calculated. This helps to optimize the allocation and utilization of resources, avoid waste and over-consumption of resources, and improve the overall resource utilization efficiency of the system. By calculating the limited interval of resource release and formulating the multi-agent resource competition strategy, the resource competition between different nodes can be effectively managed and mediated. This helps to reduce conflicts and efficiency losses caused by resource competition, and maintain the stability and predictability of system operation. Formulating the resource competition strategy according to the limited interval of node resource release can make the system more flexible and adaptable to different operating environments and workloads. This helps the system to quickly adjust and respond when facing changes and challenges, and maintain efficient operation. Therefore, the present invention provides an optimization process for a traditional multi-agent deep reinforcement learning method, solves the problems that it is difficult to balance resource competition among multi-agents and multi-agents are prone to fall into an unstable state in the traditional multi-agent deep reinforcement learning method, improves the ability of multi-agents to balance resource competition, and reduces the problem that multi-agents are prone to fall into an unstable state.

[0010] Preferably, step S1 includes the following steps:

[0011] Step S11: Extract the operation logs of multi-agents to obtain multi-agent operation log data;

[0012] Step S12: Perform task processing type marking among different agents on the multi-agent operation log data to obtain task processing type marked data;

[0013] Step S13: Perform multi-agent interaction structure analysis on the task processing type marked data according to the multi-agent operation log data to obtain multi-agent interaction structure data;

[0014] Step S14: Perform multi-scale characteristic analysis on the multi-agent interaction structure data to obtain interaction structure multi-scale characteristic data;

[0015] Step S15: Construct a multi-agent interaction graph according to the multi-agent interaction structure data and the interaction structure multi-scale characteristic data to obtain a multi-agent interaction graph.

[0016] Through the extraction of the operation log, the present invention obtains detailed operation data of the multi-agent system. These data include activity records, operation behaviors, state changes, etc. of each agent, providing comprehensive basic data for the overall operation of the system. Marking the task processing type of the operation log data can accurately identify and distinguish the task types executed by different agents in the system. This helps to understand the impact of different tasks on system resources and interaction patterns in a targeted manner during subsequent analysis and optimization. Based on the operation log and the task processing type marked data, the multi-agent interaction structure is analyzed. This step can reveal the relationship network between agents, the information flow path, and common interaction patterns. By deeply analyzing the interaction structure, potential optimization points and system bottlenecks can be discovered. On the basis of the interaction structure data, multi-scale characteristics are analyzed. This includes multi-dimensional characteristic evaluations at multiple levels from micro to macro, such as time scale, space scale, etc. This comprehensive analysis can help identify the dynamic characteristics and laws of agent interactions at different scales. According to the interaction structure data and the multi-scale characteristic data, a multi-agent interaction graph is constructed. This graph not only shows the relationships and interaction patterns between agents, but also reflects their interaction characteristics at different time and space scales. By establishing the graph, an intuitive understanding and analysis basis of the overall operation of the system can be formed.

[0017] Preferably, step S2 includes the following steps:

[0018] Step S21: Perform multi-agent load node annotation on the multi-agent interaction graph according to the multi-agent operation log data to obtain multi-agent load node annotation data;

[0019] Step S22: Perform resource scarcity ranking among different load nodes on the multi-agent load node annotation data to obtain load node resource scarcity ranking data;

[0020] Step S23: Calculate the resource release amounts between different nodes for the multi-agent load node annotation data based on the load node resource shortage sorting data and the multi-agent operation log data, to obtain the load node resource release amount data;

[0021] Step S24: Calculate the limited intervals of resource release between different load nodes for the load node resource release amount data, to obtain the node resource release limited intervals;

[0022] Step S25: Evaluate the performance differences between different load nodes based on the load node resource release amount data and the node resource release limited intervals, to obtain the node performance difference evaluation data;

[0023] Step S26: Formulate the multi-agent resource competition strategy based on the node performance difference evaluation data and the node resource release limited intervals, to obtain the multi-agent resource competition strategy.

[0024] The present invention annotates the load nodes in the interaction graph through the multi-agent operation log data. This step can accurately identify the key nodes in the system, that is, those nodes that undertake important tasks or resource loads during the system operation. According to the load node annotation data, sort the resource shortage among different load nodes. This sorting is based on the analysis of the resource demand and supply capacity of the system, which can help determine which nodes are more critical and tense in resource allocation. Combining the resource shortage sorting data and the operation log, calculate the resource amount released by each load node. This step reveals the resource utilization efficiency of each node in the system and their contribution to the overall system resource release. Based on the load node resource release amount, calculate the limited intervals of resource release between different load nodes. This calculation considers the time and space limitations of node resource release, helping to optimize the timing and space performance of resource allocation. Based on the resource release amount and limited interval data, evaluate the performance differences between different load nodes. This includes the performance differences of each node in the system in resource competition and allocation, thus providing specific data support and basis for optimizing the system performance. According to the performance difference evaluation data and the node resource release limited intervals, formulate the multi-agent resource competition strategy. These strategies can cover aspects such as resource priority allocation, dynamic adjustment, and load balancing to ensure the stable operation and high efficiency of the system under different workloads and competition conditions.

[0025] Preferably, step S24 includes the following steps:

[0026] Step S241: Conduct a fluctuation analysis of the resource release amounts between different nodes based on the load node resource release amount data, to obtain the node resource release amount fluctuation data;

[0027] Step S242: Calculate the extreme difference of resource release fluctuations between nodes with different loads for the node resource release volume fluctuation data to obtain the extreme difference of resource release fluctuations;

[0028] Step S243: Conduct a local similarity fluctuation structure analysis of the node resource release volume fluctuation data based on the extreme difference of resource release fluctuations to obtain local similarity fluctuation structure data;

[0029] Step S244: Perform a non - linear iterative numerical approximation on the node resource release volume fluctuation data based on the local similarity fluctuation structure data to obtain release volume iterative numerical approximation data;

[0030] Step S245: Conduct a local convex optimization process on the local similarity fluctuation structure data according to the release volume iterative numerical approximation data to obtain local fluctuation convex optimization data;

[0031] Step S246: Calculate the finite interval of resource release between nodes with different loads for the node resource release volume fluctuation data according to the release volume iterative numerical approximation data and the local fluctuation convex optimization data to obtain the finite interval of node resource release.

[0032] By analyzing the resource release volume fluctuations of load nodes, the present invention can reveal the dynamic changes and fluctuations in resource utilization among different nodes. This helps to identify the change patterns of resource requirements of nodes in different time periods, and then adjust the resource allocation strategy to cope with the changing requirements. Calculating the extreme difference of resource release fluctuations can quantify the volatility differences in resource release between nodes with different loads. This kind of analysis can help to determine which nodes are the most unstable in terms of resource release and may require additional resource management and adjustment. Based on the extreme difference of resource release fluctuations, conducting a local similarity fluctuation structure analysis can discover the similarity patterns or characteristics in resource release among nodes. This helps to identify groups of nodes with similar resource release patterns, so as to adopt a more precise resource allocation strategy. Using a non - linear iterative numerical approximation method based on the local similarity fluctuation structure can more accurately predict the change trends and patterns of node resource release volumes. This method can effectively process complex resource release data and provide more accurate numerical prediction results. On the basis of the release volume iterative numerical approximation, conduct a convex optimization process on the local similarity fluctuation structure. This step can further optimize the data model to make it more in line with the actual resource release situation and improve the accuracy and reliability of the prediction. Combining the release volume iterative numerical approximation data and the local fluctuation convex optimization data, calculate the finite interval of resource release between nodes with different loads. These interval data can help to determine the limitations and possible optimal release times of each node in terms of resource release, so as to optimize the resource utilization and management strategy.

[0033] Preferably, step S26 includes the following steps:

[0034] Step S261: Calculate the processing efficiency of the same task volume between different nodes based on the node performance difference evaluation data to obtain the node processing efficiency difference data;

[0035] Step S262: Determine the resource allocation priority between different nodes according to the node processing efficiency difference data and the limited interval of node resource release to obtain the node resource allocation priority data;

[0036] Step S263: Design a node interaction polling mechanism according to the node resource allocation priority data and the limited interval of node resource release to obtain the node interaction polling mechanism;

[0037] Step S264: Simulate the operation scenario of the node interaction polling mechanism to obtain the polling operation scenario simulation data; Extract the network congestion characteristics according to the polling operation scenario simulation data to obtain the polling scenario network congestion data;

[0038] Step S265: Control and optimize the node interaction polling mechanism according to the polling scenario network congestion data to obtain the node interaction polling optimization mechanism;

[0039] Step S266: Develop a multi-agent resource competition strategy according to the node interaction polling optimization mechanism and the node resource allocation priority data to obtain the multi-agent resource competition strategy.

[0040] The present invention calculates the difference in processing efficiency for the same task volume based on the evaluation data of node performance differences, and can clarify the efficiency differences of different nodes when processing tasks. This data helps to identify nodes with higher and lower performance, providing a basis for resource allocation strategies. Combining the data on node processing efficiency differences and the limited range of resource release, the resource allocation priorities between different nodes are determined. These priority data can ensure that more resources are allocated to nodes with higher efficiency or better performance within the limited range of resource release, thereby improving the overall efficiency and performance of the system. Based on the node resource allocation priority data and the limited range of resource release, an interactive polling mechanism between nodes is designed. This mechanism can ensure the fair allocation and reasonable utilization of resources in a multi-agent system, avoiding problems such as overly intense resource competition or unbalanced resource utilization. A simulation of the operating scenario of the node interactive polling mechanism is carried out to obtain the simulation data of the polling operating scenario. Through these data, network congestion characteristics that may occur in the polling scenario, such as communication delay and packet loss rate, can be extracted. These characteristics are important bases for evaluating the effectiveness and optimization of the polling mechanism. According to the network congestion data of the polling scenario, control optimization processing is performed on the node interactive polling mechanism. The optimized polling mechanism can more effectively cope with network congestion problems and ensure the stability and performance of the system under different load conditions. Combining the node interactive polling optimization mechanism and the node resource allocation priority data, multi-agent resource competition strategies are formulated. These strategies include resource dynamic adjustment, load balancing measures, etc., aiming to maximize the utilization efficiency of system resources and the overall performance, while ensuring the fair allocation of resources among nodes.

[0041] Preferably, the design of the node polling mechanism according to the node resource allocation priority data and the limited range of node resource release includes the following steps:

[0042] Calculate the limited resource release difference between nodes with different priorities for the limited range of node resource release according to the node resource allocation priority data to obtain the limited resource release difference data;

[0043] Based on the limited resource release difference data, evaluate the stability of resource release between nodes with different priorities to obtain the resource release stability evaluation data;

[0044] Compensate the resource release related factors for the node resource allocation priority data according to the limited resource release difference data and the resource release stability evaluation data to obtain the node resource allocation priority compensation data;

[0045] Design the node interactive polling mechanism according to the node resource allocation priority compensation data to obtain the node interactive polling mechanism.

[0046] The present invention calculates the limited resource release differences between nodes with different priorities based on the node resource allocation priority data. This step aims to quantify the differences in resource release among different nodes and determine which nodes may require more frequent or less resource release to optimize resource utilization efficiency. Based on the limited resource release difference data, a resource release stability assessment is carried out. This assessment can help determine whether the resource release patterns of nodes with different priorities are stable, whether there are frequent fluctuations or instability. The stability assessment data provides a basis for the subsequent design of the polling mechanism. Combining the limited resource release difference data and the stability assessment data, the node resource allocation priority data is compensated. This process can adjust the resource allocation priority according to the differences and stability between nodes to ensure more fair resource allocation during polling and avoid uneven resource release. Based on the node resource allocation priority compensation data, an interactive polling mechanism between nodes is designed. This design takes into account the priority and stability characteristics of different nodes to ensure efficient utilization and fair allocation of resources in a multi-node system. The design of the polling mechanism should be able to maximize the overall performance of the system and reduce the risk of potential system bottlenecks.

[0047] Preferably, controlling and optimizing the node interactive polling mechanism according to the network congestion data of the polling scenario includes the following steps:

[0048] Calculate the network congestion connection volume for the network congestion data of the polling scenario to obtain network congestion connection volume data;

[0049] Capture the congestion traffic surge for the network congestion data of the polling scenario according to the network congestion connection volume data to obtain congestion traffic surge capture data;

[0050] Conduct a weighted linear growth regression analysis on the congestion traffic surge capture data to obtain traffic surge weighted linear regression data;

[0051] Perform a segmented processing of the surge trend on the congestion traffic surge capture data according to the traffic surge weighted linear regression data to obtain congestion traffic surge segmented data;

[0052] Use the Binary Increase Congestion algorithm to perform a burst segmented response control on the congestion traffic surge segmented data to obtain congestion traffic segmented response control data;

[0053] Control and optimize the node interactive polling mechanism according to the congestion traffic segmented response control data to obtain a node interactive polling optimization mechanism.

[0054] The present invention calculates the connection volume of the network congestion data in the polling scenario to obtain the network congestion connection volume data. This step can accurately reflect the current network load situation and provide basic data for subsequent optimization control. According to the network congestion connection volume data, the situation of a sudden increase in congestion traffic is captured. This includes identifying and recording the sudden increase in traffic in the network, that is, the sudden increase in congestion traffic. These data can help understand the causes and patterns of congestion. A weighted linear growth regression analysis is performed on the captured data of the sudden increase in congestion traffic to obtain the weighted linear regression data of the sudden increase in traffic. This analysis helps predict and simulate the trend of congestion traffic, so as to better master the change law of network load. According to the weighted linear regression data, the data of the sudden increase in congestion traffic is segmented. This step decomposes the congestion situation into different trend segments to more precisely respond to and control the change of congestion traffic. The Binary Increase Congestion (BIC) algorithm is used to perform a sudden segment response control on the segmented data of the sudden increase in congestion traffic. The BIC algorithm dynamically adjusts the traffic control strategy to quickly and effectively respond to network congestion and avoid the occurrence of long-term unpredictable congestion states. Based on the segmented response control data of congestion traffic, the node interaction polling mechanism is optimized. This includes adjusting the communication frequency, data transmission strategy, etc. between nodes according to the real-time network congestion situation to ensure that the system can still maintain stability and performance under high load and congestion conditions.

[0055] Preferably, step S3 includes the following steps:

[0056] Step S31: Normalize the multi-agent resource competition strategy to obtain the multi-agent resource competition normalization strategy;

[0057] Step S32: Perform hierarchical deep learning processing according to the multi-agent resource competition normalization strategy to obtain the hierarchical deep learning data of resource competition;

[0058] Step S33: Perform multi-agent hierarchical convergence behavior coupling based on the hierarchical deep learning data of resource competition to obtain the hierarchical convergence behavior coupling data.

[0059] By normalizing the resource competition strategies among multiple agents, the present invention can eliminate the unfairness or resource waste caused by different resource demands among different agents. This step ensures that under competitive conditions, each agent can reasonably obtain and utilize resources according to its needs, thereby improving the efficiency and fairness of the overall system. Using the normalized competition strategy, hierarchical deep learning processing is carried out. The main effect of this step is to analyze and predict the behavior and decision-making patterns of multiple agents during the resource competition process through deep learning methods. This kind of data can provide a strong basis for subsequent agent collaboration and decision-making, enabling the system to more intelligently adjust the resource allocation strategy to cope with the changing competition environment. Based on the hierarchical deep learning data of resource competition, the convergence behaviors of multiple agents are analyzed and coupled. The main effect of this step is to promote collaboration and convergence among agents in the system. By understanding the competitive impacts at different levels, the collaborative behaviors among agents are optimized. The coupled data reflects the mutual influence and adjustment process of agents in the resource competition in the system, thereby improving the stability and performance of the overall system.

[0060] Preferably, step S33 includes the following steps:

[0061] Step S331: Extract the hierarchical instability state among multiple agents from the hierarchical deep learning data of resource competition to obtain the hierarchical instability state data of learning behavior;

[0062] Step S332: Calculate the instability coefficient between multiple levels for the hierarchical instability state data of learning behavior to obtain the hierarchical instability coefficient;

[0063] Step S333: Screen the initial instability values at multiple points for the hierarchical instability state data of learning behavior according to the hierarchical instability coefficient to obtain the hierarchical instability multi-point screening data;

[0064] Step S334: Perform iterative calculation on the hierarchical instability state data of learning behavior according to the hierarchical instability multi-point screening data to obtain the hierarchical instability multi-point iterative data;

[0065] Step S335: Couple the hierarchical convergence behaviors of multiple agents according to the hierarchical instability multi-point iterative data to obtain the hierarchical convergence behavior coupling data.

[0066] The present invention identifies hierarchical instability states existing in a multi-agent system. These states may indicate unstable behaviors of certain agents or groups of agents in the system during resource competition or learning processes. By extracting these instability states, the system can identify and further analyze the causes leading to instability, and thus take appropriate measures to stabilize the system. Calculate the instability coefficients between different levels (which may be groups of agents, individuals, etc.). The instability coefficients reflect the degree of instability within each level and between different levels. Through this step, the system can quantify the degree of the instability phenomenon, helping to optimize the cooperation and resource allocation strategies among agents. According to the hierarchical instability coefficients, select appropriate initial instability points, which may be key points or starting conditions causing instability. This screening can help the system more accurately locate and analyze the root causes leading to the instability states, providing an important basis for subsequent optimization and control. Based on the selected initial instability points, perform iterative calculations to analyze the continuous development and evolution process of the instability states. This process not only helps to understand the dynamic characteristics of the instability phenomenon but also enables the provision of more refined and effective control strategies for system design. The final step is to analyze the hierarchical convergence behavior of the multi-agent system based on the iterative data. These data reveal the coupling relationships and influences among agents at different levels during the transition from the instability state to the stable state. By analyzing the coupling data, the system can optimize the cooperation and regulation mechanisms among agents, achieving more effective system convergence and resource utilization.

[0067] Preferably, the present invention further provides a multi-agent deep reinforcement learning system for executing the multi-agent deep reinforcement learning method as described above. The multi-agent deep reinforcement learning system includes:

[0068] A multi-agent interaction graph construction module, which is used to extract the operation logs of multi-agents to obtain multi-agent operation log data; perform multi-agent interaction structure analysis on the multi-agent operation log data to obtain multi-agent interaction structure data; construct a multi-agent interaction graph according to the multi-agent interaction structure data to obtain a multi-agent interaction graph;

[0069] A multi-agent resource competition strategy formulation module, which is used to label the load nodes of multi-agents on the multi-agent interaction graph to obtain multi-agent load node labeling data; calculate the resource release amounts between different nodes for the multi-agent load node labeling data to obtain load node resource release amount data; calculate the limited resource release intervals between different load nodes for the load node resource release amount data to obtain node resource release limited intervals; formulate a multi-agent resource competition strategy according to the node resource release limited intervals to obtain a multi-agent resource competition strategy;

[0070] A multi-agent hierarchical convergence coupling module is used to perform hierarchical deep learning processing according to the multi-agent resource competition strategy to obtain resource competition hierarchical deep learning data; perform multi-agent hierarchical convergence behavior coupling based on the resource competition hierarchical deep learning data to obtain hierarchical convergence behavior coupling data;

[0071] A competition balance model construction module is used to construct a multi-agent resource competition balance model according to the multi-agent resource competition strategy and the hierarchical convergence behavior coupling data to obtain a multi-agent resource competition balance model, and send the multi-agent resource competition balance model to the cloud platform to execute multi-agent deep reinforcement learning.

[0072] The beneficial effects of the present invention are as follows. By analyzing the operation log data of multi-agents, the interaction structure and pattern between each agent can be deeply understood. This helps to identify key interaction nodes, frequently occurring interaction events, and the time and space distribution rules of interactions. The interaction graph constructed based on the interaction structure data can help to discover bottlenecks and optimization points in the system. By identifying and optimizing the key interactions between multi-agents, the overall efficiency and response speed of the system can be improved. Analyzing the interaction graph helps to identify potential system fault points or unstable factors. By better managing and regulating the interactions between multi-agents, the possibility of system crashes or failures can be reduced, and the stability and reliability of the system can be enhanced. According to the load node annotation data and resource release amount data, the resource usage between different nodes can be accurately calculated. This helps to optimize the allocation and utilization of resources, avoid waste and overconsumption of resources, and improve the overall resource utilization efficiency of the system. By calculating the limited interval of resource release and formulating the multi-agent resource competition strategy, the resource competition between different nodes can be effectively managed and mediated. This helps to reduce conflicts and efficiency losses caused by resource competition, and maintain the stability and predictability of system operation. Formulating the resource competition strategy according to the limited interval of node resource release can make the system more flexible and adaptable to different operating environments and workloads. This helps the system to quickly adjust and respond when facing changes and challenges, and maintain efficient operation. Therefore, the present invention provides an optimization process for a traditional multi-agent deep reinforcement learning method, solves the problems that it is difficult to balance resource competition between multi-agents and multi-agents are prone to fall into an unstable state in the traditional multi-agent deep reinforcement learning method, improves the ability of multi-agents to balance resource competition, and reduces the problem that multi-agents are prone to fall into an unstable state. Description of the Drawings

[0073] Figure 1 It is a schematic flow chart of the steps of a multi-agent deep reinforcement learning method;

[0074] Figure 2 ForFigure 1 Schematic diagram of the detailed implementation steps of step S2 in

[0075] Figure 3 is Figure 1 Schematic diagram of the detailed implementation steps of step S3 in

[0076] The realization, functional characteristics, and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. Specific implementation manners

[0077] The technical method of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0078] In addition, the accompanying drawings are only schematic diagrams of the present invention and are not necessarily drawn to scale. The same reference numerals in the drawings represent the same or similar parts, and thus the repeated description thereof will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. The functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor methods and / or microcontroller methods.

[0079] It should be understood that although the terms "first", "second", etc. may be used here to describe various units, these units should not be limited by these terms. These terms are only used to distinguish one unit from another. For example, without departing from the scope of the exemplary embodiments, the first unit can be called the second unit, and similarly the second unit can be called the first unit. The term "and / or" used here includes any and all combinations of one or more of the listed related items.

[0080] To achieve the above object, please refer to Figures 1 to 3 , a multi-agent deep reinforcement learning method, the method includes the following steps:

[0081] Step S1: Extract the operation logs of the multi-agent to obtain multi-agent operation log data; analyze the multi-agent interaction structure of the multi-agent operation log data to obtain multi-agent interaction structure data; construct a multi-agent interaction graph based on the multi-agent interaction structure data to obtain a multi-agent interaction graph;

[0082] Step S2: Perform multi-agent load node annotation on the multi-agent interaction graph to obtain multi-agent load node annotation data; calculate the resource release amounts between different nodes for the multi-agent load node annotation data to obtain load node resource release amount data; calculate the limited resource release intervals between different load nodes for the load node resource release amount data to obtain node resource release limited intervals; formulate a multi-agent resource competition strategy based on the node resource release limited intervals to obtain a multi-agent resource competition strategy;

[0083] Step S3: Perform hierarchical deep learning processing based on the multi-agent resource competition strategy to obtain resource competition hierarchical deep learning data; perform multi-agent hierarchical convergence behavior coupling based on the resource competition hierarchical deep learning data to obtain hierarchical convergence behavior coupling data;

[0084] Step S4: Construct a multi-agent resource competition balance model according to the multi-agent resource competition strategy and the hierarchical convergence behavior coupling data to obtain a multi-agent resource competition balance model, and send the multi-agent resource competition balance model to the cloud platform to perform multi-agent deep reinforcement learning.

[0085] In the embodiment of the present invention, refer to Figure 1 As described, it is a schematic diagram of the step process of a multi-agent deep reinforcement learning method of the present invention. In this example, the multi-agent deep reinforcement learning method includes the following steps:

[0086] Step S1: Extract the operation logs of multi-agents to obtain multi-agent operation log data; analyze the multi-agent interaction structure for the multi-agent operation log data to obtain multi-agent interaction structure data; construct a multi-agent interaction graph according to the multi-agent interaction structure data to obtain a multi-agent interaction graph;

[0087] In the embodiments of the present invention, operation data is extracted from the log systems of each agent. This data may include event logs, message logs, error logs. Ensure that the log data formats are unified, and common formats include JSON, XML, CSV. Extract key events and interaction information from the logs, such as messages sent by agents, received messages, called services, etc. Ensure that all events are arranged in chronological order for subsequent analysis. Identify the interaction relationships between agents, construct a message passing graph between agents, count the interaction frequencies between each agent, identify high-frequency interactions and low-frequency interactions, organize the extracted interaction relationships and frequency information into structured data, usually stored in the form of a graph structure. Each agent is used as a node in the graph, and the interactions between agents are used as edges in the graph. The weight of the edge can be determined according to the interaction frequency. Use graph algorithms to construct an agent interaction map, such as Dijkstra algorithm, Floyd-Warshall algorithm, etc., for path finding and optimization. Use Graphviz, Gephi or other graph visualization tools to visualize the interaction map, analyze the centrality of each node in the graph, identify key agents, identify the community structure in the agent group, find the functional clusters of agents, and analyze the shortest path, longest path, etc. between agents.

[0088] Step S2: Perform multi-agent load node annotation on the multi-agent interaction map to obtain multi-agent load node annotation data; calculate the resource release amounts between different nodes for the multi-agent load node annotation data to obtain load node resource release amount data; calculate the limited resource release intervals between different load nodes for the load node resource release amount data to obtain node resource release limited intervals; formulate a multi-agent resource competition strategy based on the node resource release limited intervals to obtain a multi-agent resource competition strategy;

[0089] In the embodiments of the present invention, load indicators are defined, such as CPU usage rate, memory usage rate, bandwidth usage rate, etc. Load data of each agent is collected from the operation logs, and its CPU, memory, and bandwidth usage conditions are recorded. The load value of each agent is calculated. For example, the average CPU usage rate and the maximum memory usage rate are calculated. According to the load calculation results, the nodes with high load values are labeled as high-load nodes, and the nodes with low load values are labeled as low-load nodes. A resource usage model is established to describe the resource usage conditions of different load nodes. The resource release amount is defined, such as the amount of CPU, memory, and other resources that can be released by a node within a certain period of time. The labeled data of the load nodes is analyzed, and the resource release amount between different nodes is calculated. For example, the amount of resources released by a high-load node when the load decreases is calculated. A limited interval of resource release is defined, such as the upper and lower limits of the resource release amount. The limited interval of resource release between different load nodes is calculated. For example, the interval of the resource release amount of a high-load node within a certain load decrease range is calculated. The calculation results are organized into structured data, and the limited interval of resource release for each node is labeled. A resource competition strategy model is established to describe the competition mechanism of resources between different load nodes. According to the actual operation conditions, the resource competition strategy is optimized to ensure the efficiency and fairness of resource allocation. According to the actual operation conditions, the resource competition strategy is optimized to ensure the efficiency and fairness of resource allocation. The resource competition strategy is deployed to the multi-agent system, and its operation effect is monitored and adjusted as needed. (For example, there is a system composed of three agents (A, B, and C). It is necessary to label the load nodes, calculate the resource release amount, calculate the limited interval of resource release, and formulate the resource competition strategy for it. The CPU usage rate and memory usage rate data of A, B, and C are collected, and their average load values are calculated. A is labeled as a high-load node, with a CPU usage rate of 90% and a memory usage rate of 80%. B and C are low-load nodes, with CPU usage rates of 30% and 40% respectively, and memory usage rates of 20% and 30% respectively. The load data of A is analyzed, and the amount of CPU resources (20%) and memory resources (10%) that can be released when its load drops to 70% are calculated. The limited interval of resource release for A is defined as CPU release amount [20% - 30%], and memory release amount [10% - 20%]. The limited interval of resource release for B and C is defined as CPU release amount [10% - 20%], and memory release amount [5% - 15%]. A resource competition strategy model is established to describe the competition mechanism of resources among A, B, and C. Specific resource competition strategy rules are formulated, such as preferentially allocating the resources released by A to B and C, or allocating them according to the priority of load decrease. The resource competition strategy is deployed to the multi-agent system, and its operation effect is monitored and adjusted as needed).

[0090] Step S3: Perform hierarchical deep learning processing according to the multi-agent resource competition strategy to obtain resource competition hierarchical deep learning data; perform multi-agent hierarchical convergence behavior coupling based on the resource competition hierarchical deep learning data to obtain hierarchical convergence behavior coupling data;

[0091] In the embodiment of the present invention, features are extracted from the resource competition strategy data as the input of the deep learning model. A hierarchical deep learning model is designed. Each layer of the model corresponds to different levels of resource competition strategies. According to the complexity of the resource competition strategies, the model is divided into multiple layers, such as an input layer, hidden layers (multiple layers), and an output layer. The resource competition strategy data is divided into a training set, a validation set, and a test set. The training set is used to train the model, adjust the model parameters, and optimize the loss function. The validation set is used to adjust the model hyperparameters to avoid overfitting. The test set is used to evaluate the model performance to ensure the generalization ability of the model. Optimization algorithms (such as Adam, SGD, etc.) are used to optimize the model parameters to improve the accuracy and efficiency of the model. The model structure is adjusted, the number of layers is increased or decreased, and the number of neurons in each layer is adjusted to improve the model performance. The output data of the hierarchical deep learning model is collected and sorted as the input data for the hierarchical convergence behavior. Specific indicators for hierarchical convergence behavior coupling are defined, such as resource allocation efficiency, competition strategy consistency, etc. A hierarchical coupling model is designed to describe the behavior coupling relationship between different levels. The hierarchical deep learning data is analyzed to identify the behavior patterns of each layer. The behavior coupling degree between each layer is calculated, such as through similarity calculation, correlation analysis, etc. The convergence behavior between each layer is analyzed to evaluate its stability and consistency. The results of the behavior coupling analysis are organized into structured data to record the coupling relationship and convergence behavior between each layer. Optimization strategies are formulated to improve the behavior coupling degree between each layer and enhance the stability and consistency of the system. The optimization strategies are applied to the multi-agent system, and its running effect is monitored and adjusted as needed.

[0092] Step S4: Construct a multi-agent resource competition balance model according to the multi-agent resource competition strategy and the hierarchical convergence behavior coupling data to obtain a multi-agent resource competition balance model, and send the multi-agent resource competition balance model to the cloud platform to perform multi-agent deep reinforcement learning.

[0093] In the embodiments of the present invention, multi-agent resource competition strategy data is collected, and the results of hierarchical convergence behavior coupling analysis are collected. A multi-agent resource competition balance model is defined, considering indicators such as resource utilization efficiency, load balance, and response time. The structure of the balance model is designed, including an input layer, a hidden layer, and an output layer. The input layer contains resource competition strategy data and hierarchical convergence behavior coupling data, and the output layer is the resource allocation decision. Features are extracted from the resource competition strategy data and hierarchical convergence behavior coupling data as the input of the model, and the features that have the greatest impact on resource competition balance are selected, such as CPU usage rate, memory usage rate, bandwidth usage rate, node coupling degree, etc. The data is divided into a training set, a validation set, and a test set. The training set is used to train the balance model, and the loss function is optimized to enable the model to effectively balance resource competition. The validation set is used to adjust the model hyperparameters to avoid overfitting. The test set is used to evaluate the model performance to ensure the generalization ability of the model. Optimization algorithms (such as Adam, SGD, etc.) are used to optimize the model parameters to improve the accuracy and efficiency of the model. The model structure is adjusted, the number of layers is increased or decreased, and the number of neurons in each layer is adjusted to improve the model performance. The required computing resources and environment are configured on the cloud platform, including GPU / CPU, storage, operating system, etc. The trained multi-agent resource competition balance model is packaged to ensure that all dependent libraries and configuration files are complete. The model is uploaded to the cloud platform and necessary configuration and deployment are carried out. The reinforcement learning library (such as TensorFlow, PyTorch, etc.) is installed and configured to ensure that the cloud platform has the ability to execute deep reinforcement learning. The training environment is configured, including the simulation environment of the multi-agent system, the definition of the reward function, etc. (for example, a suitable deep reinforcement learning algorithm (such as DQN, DDPG, PPO, etc.) is selected for the training of the multi-agent resource competition balance model. The parameters of the reinforcement learning algorithm are configured, including the learning rate, discount factor, exploration strategy, etc. The strategies of the multi-agents are initialized, and the initial state and action space are defined. The multi-agent interaction is executed in the simulation environment, and the strategies are continuously optimized through exploration and exploitation. The reward function is defined to encourage balancing resource competition, improving resource utilization efficiency, and system stability. Through multiple iterative trainings, the strategies of the multi-agents are continuously optimized to achieve the goal of resource competition balance. The training process is monitored in real time, and the training logs and metrics are recorded. The trained strategies are evaluated using the validation set and the test set to ensure their effectiveness and robustness in the actual environment. The trained strategies are deployed to the actual multi-agent system, and their running effects are monitored).

[0094] Preferably, step S1 includes the following steps:

[0095] Step S11: Extract the operation logs of the multi-agents to obtain the multi-agent operation log data;

[0096] Step S12: Perform task processing type marking among different agents on the multi-agent operation log data to obtain task processing type marked data;

[0097] Step S13: Analyze the multi-agent interaction structure of the task processing type marked data based on the multi-agent operation log data to obtain multi-agent interaction structure data;

[0098] Step S14: Analyze the multi-scale characteristics of the multi-agent interaction structure data to obtain interaction structure multi-scale characteristic data;

[0099] Step S15: Construct a multi-agent interaction graph based on the multi-agent interaction structure data and the interaction structure multi-scale characteristic data to obtain a multi-agent interaction graph.

[0100] In the embodiments of the present invention, operation data is extracted from the log systems of each agent, including event logs, message logs, and error logs, to ensure the uniformity of the log data format. Common formats include JSON, XML, and CSV. Different types of task processing types are defined, such as computing tasks, data transmission tasks, storage tasks, etc. According to the information such as keywords and event types in the operation logs, marking rules for task processing types are formulated, and scripts or tools are used to automatically identify and mark the task processing types of each agent. The task processing type marking data is stored in a database and associated with the original log data. Key events and interaction information are extracted from the logs, such as messages sent by agents, received messages, and called services, and all events are ensured to be arranged in chronological order for subsequent analysis. The interaction relationships between agents are identified to construct a message passing graph between agents, and the interaction frequencies between each agent are statistically analyzed to identify high-frequency interactions and low-frequency interactions; the interaction data is analyzed in different time windows (such as seconds, minutes, hours) to identify short-term and long-term interaction patterns, and the interaction characteristics of agents under different network topologies are analyzed to identify local and global interaction patterns. The local interaction characteristics are analyzed, such as the interaction between subgroups in the agent population, and the global interaction characteristics are analyzed, such as the interaction network topology of the entire system. The interaction characteristic data at different scales is organized into structured data to record the interaction characteristics at different time and space scales; each agent is used as a node in the graph, and the interaction between agents is used as an edge in the graph. The weight of the edge can be determined according to the interaction frequency. Graph algorithms are used to construct an agent interaction graph, such as the Dijkstra algorithm, the Floyd-Warshall algorithm, etc., for path finding and optimization. Graphviz, Gephi, or other graph visualization tools are used to visualize the interaction graph, and the centrality of each node in the graph is analyzed to identify key agents, and the community structure in the agent population is identified to find the functional clusters of agents, so as to construct a multi-agent interaction graph and obtain a multi-agent interaction graph (for example, there is a system composed of three agents A, B, and C. Operation data is extracted from the log systems of A, B, and C, redundant information is deleted, the log data is uniformly converted into the JSON format, computing tasks, data transmission tasks, and storage tasks are defined, and scripts are used to automatically identify and mark the task types of A, B, and C. The marking results are manually verified. Events such as A sending messages, B receiving messages, and C calling services are extracted from the logs, the interaction relationships between A and B, B and C, and C and A are identified, the interaction frequencies are statistically analyzed, the interaction relationships and frequency information are organized into structured data, the interaction data is analyzed in minutes and hours to identify short-term and long-term interaction patterns, and the local characteristics, such as the interaction pattern between A and B, are analyzed;Analyze the global characteristics, such as the interaction topology of the entire system, organize the multi-scale characteristic data into structured data, define A, B, and C as nodes, establish the edges of A-B, B-C, and C-A, use Graphviz to visualize the graph spectrum, display the interaction relationships between agents, analyze the centrality of A, and identify it as the key agent of the system; identify the community structure in the agent group).;

[0101] Preferably, step S2 includes the following steps:

[0102] Step S21: Perform multi-agent load node annotation on the multi-agent interaction graph spectrum according to the multi-agent operation log data to obtain multi-agent load node annotation data;

[0103] Step S22: Sort the resource shortage between different load nodes for the multi-agent load node annotation data to obtain load node resource shortage sorting data;

[0104] Step S23: Calculate the resource release amount between different nodes for the multi-agent load node annotation data according to the load node resource shortage sorting data and the multi-agent operation log data to obtain load node resource release amount data;

[0105] Step S24: Calculate the limited resource release interval between different load nodes for the load node resource release amount data to obtain the node resource release limited interval;

[0106] Step S25: Evaluate the performance difference between different load nodes according to the load node resource release amount data and the node resource release limited interval to obtain node performance difference evaluation data;

[0107] Step S26: Formulate a multi-agent resource competition strategy according to the node performance difference evaluation data and the node resource release limited interval to obtain a multi-agent resource competition strategy.

[0108] As an example of the present invention, refer to Figure 2 As shown, in this example, step S2 includes:

[0109] Step S21: Perform multi-agent load node annotation on the multi-agent interaction graph spectrum according to the multi-agent operation log data to obtain multi-agent load node annotation data;

[0110] In the embodiments of the present invention, the operation log data of multiple agents is collected, including resource usage conditions such as CPU usage rate, memory usage rate, and bandwidth usage rate. The multi-agent interaction graph obtained in step S1 is used as the basis for labeling load nodes, and the criteria for load nodes are defined, such as high CPU usage rate and high memory usage rate. According to the operation log data, the resource usage of each agent is calculated, and the load nodes are identified. According to the load calculation results, the agents are divided into high-load nodes, medium-load nodes, and low-load nodes, and load labels such as "high load", "medium load", and "low load" are labeled for each node. The load node labeling data is stored in the database and associated with the multi-agent interaction graph.

[0111] Step S22: Sort the resource scarcities among different load nodes for the multi-agent load node labeling data to obtain the load node resource scarcity sorting data;

[0112] In the embodiments of the present invention, the resource usage of each load node is analyzed, such as CPU usage rate, memory usage rate, and bandwidth usage rate, and the indicators of resource scarcity are defined, such as resource usage rate and resource remaining amount. A sorting algorithm (such as quicksort, heapsort, etc.) is used to sort the load nodes according to resource scarcity to obtain the load node resource scarcity sorting data, which is arranged from the most scarce to the least scarce. The load node resource scarcity sorting data is stored in the database for subsequent processing.

[0113] Step S23: Calculate the resource release amounts among different nodes for the multi-agent load node labeling data according to the load node resource scarcity sorting data and the multi-agent operation log data to obtain the load node resource release amount data;

[0114] In the embodiments of the present invention, the load node resource scarcity sorting data is extracted from the database, and the multi-agent operation log data is extracted for resource release amount calculation. The indicators of resource release amount are defined, such as CPU release amount, memory release amount, bandwidth release amount, etc. According to the resource scarcity sorting results, the conditions and thresholds for resource release are determined. According to the scarcity sorting data and the operation log data, the resource release amount of each load node is calculated. For each pair of load nodes, the resource release amount between them is calculated. (For example, the load node resource scarcity sorting data and the operation log data are extracted from the database, the CPU release amount is defined as the current CPU usage rate minus the target usage rate such as 70%, and the release condition is set: the high-load node needs to release resources to reduce the scarcity degree. Calculate the CPU release amount of A, such as the current usage rate of 90% and the target usage rate of 70%, and the release amount is 20%). The calculated load node resource release amount data is stored in the database.

[0115] Step S24: Calculate the limited resource release intervals among different load nodes for the load node resource release amount data to obtain the node resource release limited intervals;

[0116] In the embodiment of the present invention, the load node resource release amount data is extracted from the database, the multi-agent load node annotation data is extracted from the database, a calculation model for the limited resource release intervals is established, the resource usage of each node and the total system resources are considered, interval parameters are defined, such as the minimum release amount, the maximum release amount, and the average release amount. According to the release amount data, the resource release amounts of each node are divided into different intervals (such as low, medium, high). According to actual requirements, the interval parameters are adjusted to ensure the rationality and operability of each interval. The node resource release limited interval data is stored in the database and associated with the load node resource release amount data.

[0117] Step S25: Evaluate the performance differences among different load nodes according to the load node resource release amount data and the node resource release limited intervals to obtain the node performance difference evaluation data;

[0118] In the embodiment of the present invention, the load node resource release amount data is extracted from the database. The node resource release limited interval data is extracted from the database. Indexes for evaluating performance differences are defined, such as response time, throughput, resource utilization rate, etc. Weights are set for each performance index to reflect its relative importance. A performance difference evaluation model (such as the weighted scoring method, the analytic hierarchy process, etc.) is used to evaluate each node. According to the evaluation indexes and weights, each node is scored to calculate its performance difference score. The node performance difference evaluation data is stored in the database and associated with the node resource release amount data and the limited interval data.

[0119] Step S26: Develop a multi-agent resource competition strategy according to the node performance difference evaluation data and the node resource release limited intervals to obtain the multi-agent resource competition strategy.

[0120] In the embodiment of the present invention, the node performance difference evaluation data is extracted from the database, and the node resource release limited interval data is extracted from the database. The goals of the resource competition strategy are defined, such as maximizing system performance, minimizing resource conflicts, etc. A decision-making model for the resource competition strategy is established, considering the performance differences of each node and the node resource release limited intervals. Specific resource competition strategy rules are developed, such as resource priority allocation, dynamic adjustment mechanisms, etc. An optimization algorithm (such as the genetic algorithm, the particle swarm algorithm, etc.) is used to optimize the resource competition strategy to ensure the effectiveness and stability of the strategy. The multi-agent resource competition strategy is stored in the database for subsequent execution and adjustment.

[0121] Preferably, step S24 includes the following steps:

[0122] Step S241: Conduct a resource release volume fluctuation analysis among different nodes based on the load node resource release volume data to obtain node resource release volume fluctuation data;

[0123] Step S242: Calculate the extreme difference of resource release fluctuations among different load nodes for the node resource release volume fluctuation data to obtain the extreme difference of resource release fluctuations;

[0124] Step S243: Conduct a local similar fluctuation structure analysis among different load nodes for the node resource release volume fluctuation data based on the extreme difference of resource release fluctuations to obtain local similar fluctuation structure data;

[0125] Step S244: Perform a non - linear iterative numerical approximation on the node resource release volume fluctuation data based on the local similar fluctuation structure data to obtain release volume iterative numerical approximation data;

[0126] Step S245: Conduct a local convex optimization process on the local similar fluctuation structure data based on the release volume iterative numerical approximation data to obtain local fluctuation convex optimization data;

[0127] Step S246: Calculate the limited interval of resource release among different load nodes for the node resource release volume fluctuation data based on the release volume iterative numerical approximation data and the local fluctuation convex optimization data to obtain the limited interval of node resource release.

[0128] In the embodiments of the present invention, data on the resource release amount of load nodes is extracted from a database, and an analysis time window (such as minutes, hours, days) is defined for fluctuation analysis. The resource release amount data of each node is converted into time series data, and the fluctuation of the resource release amount within each time window is calculated, including indicators such as mean, variance, and standard deviation. The calculated fluctuation data of the node resource release amount is stored in the database and associated with the original resource release amount data. The maximum and minimum values of the resource release amount of each node are identified, and the extreme difference of the resource release fluctuation of each node is calculated, that is, the difference between the maximum value and the minimum value. The extreme differences of different nodes are compared to identify the node with the largest and smallest resource release fluctuations. The calculated extreme difference data of the resource release fluctuation is stored in the database and associated with the fluctuation data. A similarity measurement method for the resource release amount fluctuation between nodes is defined, such as dynamic time warping (DTW), Euclidean distance. Using the similarity measurement method, the local similar fluctuation structures between different nodes are identified, and node pairs with similar fluctuation patterns are identified, and their local fluctuation characteristics are analyzed. The local similar fluctuation structure data is organized into structured data for subsequent analysis. (For example, there is a system composed of three agents A, B, and C. The resource release amount data of A, B, and C is extracted from the database. The analysis time window is defined as hours, and the resource release amount data of A, B, and C is converted into hourly time series data. The mean, variance, and standard deviation of the resource release amount within each hour are calculated. The resource release amount fluctuation data of A, B, and C is extracted from the database. The maximum value of the resource release amount of A is calculated as 50%, the minimum value is 10%, and the extreme difference is 40%. The extreme differences of the resource release amounts of B and C are calculated, and the extreme differences of A, B, and C are compared to identify the node with the largest and smallest resource release fluctuations. Using the dynamic time warping (DTW) method, the similarity of the resource release amount fluctuations between A, B, and C is calculated, and it is identified that A and B have similar fluctuation patterns. The local fluctuation characteristics of A and B are analyzed, and their local similar fluctuation structures are identified. The local similar fluctuation structure data is organized into structured data); An appropriate non-linear iterative algorithm is selected, such as Newton's iterative method, conjugate gradient method, and the initial conditions and parameters are set, including the initial iteration value, convergence condition. The node resource release amount fluctuation data is iteratively calculated to gradually approach the true value. After each iterative step, the accuracy of the result is verified to ensure that the iterative result converges step by step. The calculated iterative numerical approximation data of the release amount is stored in the database and associated with the original fluctuation data. The objective function of local fluctuation convex optimization is set, such as minimizing the resource release fluctuation, and the constraint conditions in the optimization process are defined, such as total resource constraint, node performance constraint. An appropriate convex optimization algorithm is selected, such as interior point method, gradient descent method. According to the data characteristics, the parameters of the optimization algorithm are tuned to ensure the stability and effectiveness of the optimization process. The iterative approximation data is optimized to obtain the convex optimization result of the local fluctuation.During the optimization process, verify the results of each step to ensure that the optimization process converges to a local optimal solution. Define a finite interval for resource release based on the iterative approximation data and the optimization data. Establish a calculation model for the finite interval of resource release, considering the resource usage among nodes and the total system resources. According to the calculation model, divide the node resource release amount into different finite intervals (such as low, medium, high). Adjust the interval parameters according to the actual requirements and optimization results to ensure the rationality and operability of each interval. Store the calculated finite interval data of node resource release in the database and associate it with the iterative approximation data and the optimization data (for example, extract the local similarity fluctuation structure data and resource release amount fluctuation data of A, B, and C from the database, select the Newton iteration method as the non-linear iterative algorithm, set the initial iteration value as the current resource release amount, and the convergence condition as the error being less than 1%. Perform iterative calculations on the resource release amount fluctuation data of A, B, and C to gradually approach the true value. After each iterative step, verify the accuracy of the result to ensure that the iterative result converges step by step. Extract the iterative approximation data of the release amount and the local similarity fluctuation structure data of A, B, and C from the database, set the optimization goal as minimizing the resource release fluctuation, define the total resource constraint and node performance constraint, select the gradient descent method as the convex optimization algorithm, tune the parameters of the optimization algorithm to ensure the stability and effectiveness of the optimization process, perform optimization calculations on the iterative approximation data of A, B, and C to obtain the convex optimization result of local fluctuation, verify the result of each step to ensure that the optimization process converges to a local optimal solution, define a finite interval for resource release based on the iterative approximation data and the optimization data, establish a calculation model for the finite interval of resource release, considering the resource usage among nodes and the total system resources, divide the resource release amounts of A, B, and C into different finite intervals such as low 10% - 20%, medium 20% - 40%, high 40% - 50%, and adjust the interval parameters according to the actual requirements and optimization results to ensure the rationality and operability of each interval).

[0129] Preferably, step S26 includes the following steps:

[0130] Step S261: Calculate the processing efficiency of the same task amount among different nodes based on the node performance difference evaluation data to obtain the node processing efficiency difference data;

[0131] Step S262: Determine the resource allocation priority among different nodes according to the node processing efficiency difference data and the finite interval of node resource release to obtain the node resource allocation priority data;

[0132] Step S263: Design a node interaction polling mechanism according to the node resource allocation priority data and the finite interval of node resource release to obtain the node interaction polling mechanism;

[0133] Step S264: Conduct a running scenario simulation on the node interaction polling mechanism to obtain polling running scenario simulation data; extract network congestion characteristics based on the polling running scenario simulation data to obtain polling scenario network congestion data;

[0134] Step S265: Perform control optimization processing on the node interaction polling mechanism according to the polling scenario network congestion data to obtain a node interaction polling optimization mechanism;

[0135] Step S266: Develop a multi-agent resource competition strategy based on the node interaction polling optimization mechanism and the node resource allocation priority data to obtain a multi-agent resource competition strategy.

[0136] In the embodiment of the present invention, a calculation formula for the processing efficiency of the same task volume is set, such as efficiency = task volume / (time multiplied by resource consumption), calculate the efficiency of each node when processing the same task volume to obtain node processing efficiency difference data, set an algorithm for resource allocation priority, such as setting the priority according to the high and low efficiency and the limited interval of resource release, calculate the resource allocation priority of each node according to the processing efficiency difference data and the limited interval of resource release, and the priority calculation formula is multiplied by the processing efficiency + multiplied by (1 - )), where and are weight parameters, and obtain the node resource allocation priority data, define the interaction polling rules between nodes, usually designed according to the resource allocation priority and the resource release situation, design polling strategies such as time slice polling, weighted polling, etc., to ensure that high-priority nodes obtain more resources, set polling parameters including time slice size, polling order, priority weight, and construct a node interaction polling mechanism model to ensure that the polling strategy can be dynamically adjusted (for example, there is a system composed of three agents A, B, and C, count the time required for A, B, and C to process the same task volume, for example, A is 10 seconds, B is 12 seconds, C is 8 seconds, calculate the processing efficiency, A is 0.1 task / second, B is 0.083 task / second, C is 0.125 task / second, compare the processing efficiency to obtain the processing efficiency difference data of A, B, and C, define the priority rule, and set the weight parameter = 0.6, = 0.4, calculate the priority score, sort A, B, and C according to the score, define the polling rule, allocate time slices according to the priority level, with C being 50%, A being 30%, and B being 20%. Design a time slice polling strategy to ensure that C obtains more resources. Set polling parameters such as the time slice size being 100 milliseconds, construct a node interaction polling mechanism model to ensure that C, A, and B are polled in the order of priority. Implement the polling mechanism algorithm to ensure that high-priority nodes obtain resources first). Select and build a suitable simulation platform such as MATLAB, Simulink, OMNeT++. Select and build a suitable simulation platform such as MATLAB, Simulink, OMNeT++. Execute the node interaction polling mechanism on the simulation platform to simulate the actual operation of the multi-agent system. During the simulation process, collect various data of the polling operation scenario in real time, including task processing time, resource usage, and network transmission. Analyze the simulation data to extract characteristic data of network congestion such as transmission delay, packet loss rate, and bandwidth utilization. Define the goals of control optimization such as minimizing transmission delay, reducing packet loss rate, and increasing bandwidth utilization, etc. Set the constraints during the optimization process such as the total amount of resources and node priorities. Control and optimize the node interaction polling mechanism, adjust the polling parameters and strategies to achieve the optimization goal. During the optimization process, verify the results of each step to ensure that the optimization process converges to the optimal solution. According to the optimized interaction polling mechanism and resource allocation priority, construct a multi-agent resource competition strategy model. Define resource competition strategy rules such as priority allocation rules and resource scheduling rules.

[0137] Preferably, the design of the node polling mechanism according to the node resource allocation priority data and the limited interval of node resource release includes the following steps:

[0138] Calculate the limited difference in resource release between nodes with different priorities for the limited interval of node resource release according to the node resource allocation priority data to obtain limited difference data in resource release;

[0139] Based on the limited difference data in resource release, conduct an assessment of the resource release stability between nodes with different priorities to obtain resource release stability assessment data;

[0140] Compensate the node resource allocation priority data for the associated factors of resource release according to the limited difference data in resource release and the resource release stability assessment data to obtain node resource allocation priority compensation data;

[0141] Design the node interaction polling mechanism according to the node resource allocation priority compensation data to obtain the node interaction polling mechanism.

[0142] In the embodiments of the present invention, nodes are divided into different priority groups according to resource allocation priority data. Calculate the limited resource release differences between different priority groups. For example: Limited resource release difference = |Average resource release interval of high-priority group - Average resource release interval of low-priority group| to obtain resource release stability evaluation data. Set a compensation factor based on the limited resource release difference and stability. Compensate the node resource allocation priority according to the compensation factor to obtain node resource allocation priority compensation data. Define node interaction polling rules based on the compensated resource allocation priority and design specific polling strategies such as time slice polling, weighted polling, etc. to ensure that high-priority nodes obtain more resources. Set polling parameters including time slice size, polling order, and priority weight, and construct a node interaction polling mechanism model to ensure that the polling strategy can be dynamically adjusted. Implement the polling mechanism algorithm to ensure the executability and effectiveness of the algorithm.

[0143] Preferably, controlling and optimizing the node interaction polling mechanism according to the network congestion data of the polling scenario includes the following steps:

[0144] Calculate the network congestion connection volume for the network congestion data of the polling scenario to obtain network congestion connection volume data;

[0145] Capture the congestion traffic surge for the network congestion data of the polling scenario according to the network congestion connection volume data to obtain congestion traffic surge capture data;

[0146] Conduct a weighted linear growth regression analysis on the congestion traffic surge capture data to obtain traffic surge weighted linear regression data;

[0147] Perform a segmented processing of the surge trend on the congestion traffic surge capture data according to the traffic surge weighted linear regression data to obtain congestion traffic surge segmented data;

[0148] Use the Binary Increase Congestion algorithm to perform a burst segmented response control on the congestion traffic surge segmented data to obtain congestion traffic segmented response control data;

[0149] Conduct a control optimization process on the node interaction polling mechanism according to the congestion traffic segmented response control data to obtain a node interaction polling optimization mechanism.

[0150] In the embodiments of the present invention, first, network congestion data in a polling scenario is obtained and analyzed. Calculation metrics for the network congestion connection volume are defined, such as the number of connections, bandwidth occupancy rate, etc. By summarizing and statistically analyzing the number of network connections and bandwidth usage at each time point, the network congestion connection volume of the entire system is calculated. The specific method is to extract the connection number and bandwidth occupancy data of each node, and then summarize these data to calculate the total network connection volume. Finally, the calculated network congestion connection volume data is stored in a database to provide basic data for subsequent analysis and optimization. According to the network congestion connection volume data obtained in the first step, it is further analyzed to capture the phenomenon of sudden increase in network traffic. Capture metrics for the sudden increase in traffic are defined, such as the instantaneous connection number growth rate or the magnitude of sudden traffic increase, etc. During the data analysis process, a threshold is set. When the network traffic or the number of connections grows beyond the set threshold, it is marked as a sudden traffic increase event. The specific method is to perform a difference calculation on the time series data to identify significant growth events that occur within a short period of time. The data of these sudden increase events is stored in a database for subsequent weighted linear growth regression analysis. For the data of the sudden traffic increase events captured in the second step, weighted linear growth regression analysis is performed. First, weights are assigned to each sudden increase event, usually weighted according to the importance or occurrence frequency of the event. Then, a weighted linear regression model is used to fit the data, calculate the regression coefficients, and analyze the trend and characteristics of traffic growth. The specific method is to describe the variation law of the sudden traffic increase through a linear regression equation. The result data of the regression analysis is stored in a database to provide a basis for subsequent trend segmentation processing. According to the results of the weighted linear regression analysis, trend segmentation processing is performed on the captured data of the congested traffic sudden increase. First, rules for trend segmentation are defined, such as segmenting according to a time window or traffic growth rate. Then, the sudden traffic increase data is segmented according to the set rules to identify the traffic growth trends in different time periods. The specific method is to use the sliding window technique to divide the time series data into multiple consecutive time periods, and analyze and describe the traffic changes within each time period. Finally, the segmented data of the congested traffic sudden increase is obtained and stored in a database for subsequent bursty segmentation response control. The BinaryIncrease Congestion (BIC) algorithm is used to perform bursty segmentation response control on the segmented data of the congested traffic sudden increase obtained in the fourth step. The BIC algorithm is an algorithm for network traffic control that can effectively respond to the sudden changes in network congestion. The specific method is to perform binary search and optimization adjustment on the segmented data of the sudden traffic increase to identify the key time points of sudden traffic increase and take corresponding control measures. The calculated congested traffic segmented response control data can provide precise control parameters and strategies for the optimization of the node interaction polling mechanism. According to the congested traffic segmented response control data obtained in the fifth step, the node interaction polling mechanism is optimized.First, analyze and understand the congestion response control data to identify potential bottlenecks and optimization spaces in the system. Then, adjust the parameters and strategies of the node interaction polling mechanism based on this data, such as adjusting the allocation of time slices, priority weights, etc., to improve the overall resource utilization efficiency and network performance of the system. Finally, apply the optimized node interaction polling mechanism to the actual system for testing and verification to ensure the effectiveness and stability of the optimization measures. Through the above steps, the control optimization of the node interaction polling mechanism is completed, and the operation efficiency and stability of the multi-agent system in high-load and complex network environments are improved.

[0151] Preferably, step S3 includes the following steps:

[0152] Step S31: Normalize the multi-agent resource competition strategy to obtain the multi-agent resource competition normalization strategy;

[0153] Step S32: Perform hierarchical deep learning processing according to the multi-agent resource competition normalization strategy to obtain the resource competition hierarchical deep learning data;

[0154] Step S33: Based on the resource competition hierarchical deep learning data, perform multi-agent hierarchical convergence behavior coupling to obtain the hierarchical convergence behavior coupling data.

[0155] As an example of the present invention, refer to Figure 3 As shown, in this example, step S3 includes:

[0156] Step S31: Normalize the multi-agent resource competition strategy to obtain the multi-agent resource competition normalization strategy;

[0157] In the embodiment of the present invention, obtain and sort out the multi-agent resource competition strategy data, which usually includes information on resource allocation, usage efficiency, competition priorities, etc. among different agents. Standardize and normalize this data to eliminate the dimensional differences between different data dimensions and make the data more comparable. Specific methods include min-max normalization and Z-score standardization. Normalize each strategy parameter and scale it to the interval [0, 1]. Through normalization, the multi-agent resource competition normalization strategy is obtained, providing a consistent and standard data input for subsequent hierarchical deep learning processing. The key point of this step is to ensure the integrity and consistency of the data while trying to retain the characteristics and distribution of the original data to more accurately reflect the actual resource competition situation.

[0158] Step S32: Perform hierarchical deep learning processing according to the multi-agent resource competition normalization strategy to obtain the resource competition hierarchical deep learning data;

[0159] In the embodiments of the present invention, a hierarchical deep learning model is designed and trained based on the normalized multi-agent resource competition strategy. First, determine the structure of the model, including the number of input layers, hidden layers, and output layers and the configuration of their neurons. For the multi-agent resource competition problem, it is usually necessary to construct a multi-layer neural network to capture complex non-linear relationships. Then, prepare the training data, and divide the normalized strategy data into a training set and a validation set. The model is trained through the backpropagation algorithm and the gradient descent optimizer, continuously adjusting the network weights and biases to minimize the loss function. During the specific training process, a cross-validation method can be adopted to prevent overfitting. After training, the validation set data is used to evaluate the model to test its accuracy and robustness in predicting resource competition strategies. Through this process, hierarchical deep learning data for resource competition is obtained, providing basic data and model support for the next multi-agent hierarchical convergence behavior coupling.

[0160] Step S33: Perform multi-agent hierarchical convergence behavior coupling based on the hierarchical deep learning data for resource competition to obtain hierarchical convergence behavior coupling data.

[0161] In the embodiments of the present invention, the hierarchical deep learning data for resource competition is used to perform hierarchical convergence behavior coupling analysis on multi-agents. First, define the hierarchical structure and behavior model of the agents, including the top-level decision-making layer, the middle-level coordination layer, and the bottom-level execution layer, etc. Based on the data output by the deep learning model, analyze the behavior coupling relationship between different levels. For example, by analyzing the changes in the top-level decision-making, predict its impact on the middle-level coordination and the bottom-level execution. Use the multi-agent reinforcement learning algorithm to jointly optimize the behaviors of different levels to ensure the collaborative work of the entire system in a resource competition environment. Specific methods include policy gradient methods, Q-learning, etc. Through these algorithms, optimize the cooperation strategies between levels to improve the resource utilization efficiency and stability of the overall system. In actual operation, simulation and experimental methods can be adopted to verify the effectiveness and performance of the coupling model. Finally, hierarchical convergence behavior coupling data is obtained, providing support and guarantee for the efficient operation of the multi-agent system in a complex resource competition environment.

[0162] Preferably, step S33 includes the following steps:

[0163] Step S331: Extract the hierarchical instability state among multi-agents from the hierarchical deep learning data for resource competition to obtain learning behavior hierarchical instability state data;

[0164] Step S332: Calculate the instability coefficient between multiple levels for the learning behavior hierarchical instability state data to obtain the hierarchical instability coefficient;

[0165] Step S333: Screen the initial instability values at multiple points for the hierarchical instability state data of the learning behavior according to the hierarchical instability coefficient, and obtain the hierarchical instability multi-point screening data;

[0166] Step S334: Perform iterative calculations on the hierarchical instability state data of the learning behavior according to the hierarchical instability multi-point screening data, and obtain the hierarchical instability multi-point iterative data;

[0167] Step S335: Perform multi-agent hierarchical convergence behavior coupling according to the hierarchical instability multi-point iterative data, and obtain the hierarchical convergence behavior coupling data.

[0168] In the embodiments of the present invention, hierarchical deep learning data is utilized to extract the hierarchical instability states among multiple agents. The goal of this step is to identify and analyze the possible hierarchical instability phenomena in the system, that is, the coordination and decision-making between different levels may lead to a decrease in the overall efficiency of the system or an unstable state due to resource competition. Through the data output by the deep learning model, the resource competition situation between different levels can be observed. For example, how the changes in the top decision-making layer affect the behaviors of the middle-level coordination and the bottom-level execution. The specific operations include performing time series analysis and feature extraction on the deep learning data to identify the hierarchical instability events in the system. This may involve anomaly detection algorithms or pattern recognition techniques to capture and mark the system instability signals that appear during the resource competition process. These signals can be the change patterns of the behaviors at each level, such as frequent changes in the decision-making layer and a decrease in the execution efficiency of the coordination layer. Through effective data processing and analysis, the hierarchical instability state data of the learning behavior is obtained, laying a foundation for the subsequent calculation of the instability coefficient. Using the information extracted from the hierarchical instability state data, the instability coefficient between multiple levels is calculated. The instability coefficient is used to quantify the intensity and influence degree of the instability states between different levels and is one of the important indicators for evaluating the system stability and performance. During the calculation process, statistical analysis methods and mathematical modeling techniques can be adopted, comprehensively considering the time series data of the behaviors at each level and the frequency and duration of the instability events. Specifically, first, the instability events at each level are statistically analyzed and classified to determine their occurrence frequency and intensity. Then, through an appropriate mathematical model or algorithm, these data are converted into the instability coefficient, reflecting the influence degree of the resource competition between different levels. The calculation of the instability coefficient not only depends on the quantity and quality of the data but also takes into account the complexity and non-linear characteristics of the behavior coupling between different levels. Through this step, the hierarchical instability coefficient is obtained, providing a quantitative basis for the subsequent screening and iterative calculation. According to the hierarchical instability coefficient, multi-point instability initial value screening is performed on the hierarchical instability state data of the learning behavior. The purpose of this step is to screen out the most representative and significant instability initial value points from all the instability states for subsequent iterative calculation and in-depth analysis. The specific operations include combining the instability coefficient and the characteristics of the time series data to identify those instability initial value points that have an important influence during the resource competition process. This may involve pattern matching and clustering analysis techniques to ensure that the selected instability initial values can effectively represent the system's response under different resource competition conditions. The diversity and representativeness of the instability initial values need to be considered during the screening process to ensure the comprehensiveness and accuracy of the subsequent analysis. Using the screened hierarchical instability initial value point data, iterative calculation is performed on the hierarchical instability state data of the learning behavior. In this step, by applying numerical simulation and computational experiment techniques, the evolution process and influence effects of different instability initial value points under resource competition conditions are analyzed. The specific operations include establishing a mathematical model or a simulation platform to simulate the system response and subsequent evolution triggered by the instability initial value points.By continuously performing iterative calculations, observe the dynamic changes of the system under various competitive conditions, and analyze its stability and behavioral characteristics. This process can not only deeply understand the resource competition mechanism between different levels in the system, but also reveal potential problems and optimization spaces that may exist. Finally, obtain hierarchical instability multi-point iterative data, providing a scientific basis for the in-depth understanding and optimization of the coupling of multi-agent hierarchical convergence behaviors.

[0169] Preferably, the present invention also provides a multi-agent deep reinforcement learning system for executing the multi-agent deep reinforcement learning method as described above. The multi-agent deep reinforcement learning system includes:

[0170] A multi-agent interaction graph construction module, which is used to extract the operation logs of multi-agents to obtain multi-agent operation log data; perform multi-agent interaction structure analysis on the multi-agent operation log data to obtain multi-agent interaction structure data; construct a multi-agent interaction graph according to the multi-agent interaction structure data to obtain a multi-agent interaction graph;

[0171] A multi-agent resource competition strategy formulation module, which is used to label the load nodes of multi-agents on the multi-agent interaction graph to obtain multi-agent load node labeling data; calculate the resource release amounts between different nodes for the multi-agent load node labeling data to obtain load node resource release amount data; calculate the limited resource release intervals between different load nodes for the load node resource release amount data to obtain node resource release limited intervals; formulate a multi-agent resource competition strategy according to the node resource release limited intervals to obtain a multi-agent resource competition strategy;

[0172] A multi-agent hierarchical convergence coupling module, which is used to perform hierarchical deep learning processing according to the multi-agent resource competition strategy to obtain resource competition hierarchical deep learning data; perform multi-agent hierarchical convergence behavior coupling based on the resource competition hierarchical deep learning data to obtain hierarchical convergence behavior coupling data;

[0173] A competition balance model construction module, which is used to construct a multi-agent resource competition balance model according to the multi-agent resource competition strategy and the hierarchical convergence behavior coupling data to obtain a multi-agent resource competition balance model, and send the multi-agent resource competition balance model to the cloud platform to execute multi-agent deep reinforcement learning.

[0174] Therefore, from any perspective, the embodiments should be regarded as exemplary and non-restrictive. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the application documents are intended to be included in the present invention.

[0175] The above are only specific embodiments of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather to the widest scope consistent with the principles and novel features invented herein.

Claims

1. A multi-agent deep reinforcement learning method, characterized in that, Including the following steps: Step S1: Extract the running logs of multi - agents to obtain multi - agent running log data; Perform multi - agent interaction structure analysis on the multi - agent running log data to obtain multi - agent interaction structure data; Construct a multi - agent interaction graph based on the multi - agent interaction structure data to obtain a multi - agent interaction graph; Step S2: Perform multi - agent load node annotation on the multi - agent interaction graph to obtain multi - agent load node annotation data; Calculate the resource release amount between different nodes for the multi - agent load node annotation data to obtain load node resource release amount data; Calculate the limited resource release interval between different load nodes for the load node resource release amount data to obtain the node resource release limited interval; Develop a multi - agent resource competition strategy based on the node resource release limited interval to obtain a multi - agent resource competition strategy; Step S3: Perform hierarchical deep learning processing according to the multi - agent resource competition strategy to obtain resource competition hierarchical deep learning data; Perform multi - agent hierarchical convergence behavior coupling based on the resource competition hierarchical deep learning data. Through continuous iterative calculation, observe the dynamic changes of the system under various competition conditions, analyze its stability and behavior characteristics, and obtain hierarchical convergence behavior coupling data; Step S4: Construct a multi - agent resource competition balance model according to the multi - agent resource competition strategy and the hierarchical convergence behavior coupling data to obtain a multi - agent resource competition balance model, and send the multi - agent resource competition balance model to the cloud platform to perform multi - agent deep reinforcement learning.

2. The multi-agent deep reinforcement learning method according to claim 1, wherein Step S1 includes the following steps: Step S11: Extract the running logs of multi - agents to obtain multi - agent running log data; Step S12: Mark the task processing types between different agents for the multi - agent running log data to obtain task processing type marking data; Step S13: Perform multi - agent interaction structure analysis on the task processing type marking data according to the multi - agent running log data to obtain multi - agent interaction structure data; Step S14: Perform multi - scale characteristic analysis on the multi - agent interaction structure data to obtain interaction structure multi - scale characteristic data; Step S15: Construct a multi - agent interaction graph according to the multi - agent interaction structure data and the interaction structure multi - scale characteristic data to obtain a multi - agent interaction graph.

3. The multi-agent deep reinforcement learning method according to claim 1, characterized in that Step S2 includes the following steps: Step S21: Perform multi - agent load node annotation on the multi - agent interaction graph according to the multi - agent running log data to obtain multi - agent load node annotation data; Step S22: Sort the resource scarcity of different load nodes for the multi - agent load node annotation data to obtain load node resource scarcity sorting data; Step S23: Calculate the resource release amount between different nodes for the multi - agent load node annotation data according to the load node resource scarcity sorting data and the multi - agent running log data to obtain load node resource release amount data; Step S24: Calculate the limited resource release interval between different load nodes for the load node resource release amount data to obtain the node resource release limited interval; Step S25: Evaluate the performance differences between different load nodes based on the load node resource release amount data and the limited range of node resource release, and obtain the node performance difference evaluation data; Step S26: Develop a multi-agent resource competition strategy based on the node performance difference evaluation data and the limited range of node resource release, and obtain the multi-agent resource competition strategy.

4. The multi-agent deep reinforcement learning method according to claim 3, characterized in that Step S24 includes the following steps: Step S241: Analyze the resource release amount fluctuations between different nodes based on the load node resource release amount data, and obtain the node resource release amount fluctuation data; Step S242: Calculate the extreme difference of resource release fluctuations between different load nodes for the node resource release amount fluctuation data, and obtain the extreme difference of resource release fluctuations; Step S243: Analyze the local similar fluctuation structure between different load nodes for the node resource release amount fluctuation data based on the extreme difference of resource release fluctuations, and obtain the local similar fluctuation structure data; Step S244: Perform non-linear iterative numerical approximation on the node resource release amount fluctuation data based on the local similar fluctuation structure data, and obtain the iterative numerical approximation data of the release amount; Step S245: Perform local convex optimization on the local similar fluctuation structure data based on the iterative numerical approximation data of the release amount, and obtain the local fluctuation convex optimization data; Step S246: Calculate the limited range of node resource release between different load nodes for the node resource release amount fluctuation data based on the iterative numerical approximation data of the release amount and the local fluctuation convex optimization data, and obtain the limited range of node resource release.

5. The multi-agent deep reinforcement learning method according to claim 3, wherein Step S26 includes the following steps: Step S261: Calculate the processing efficiency of the same task volume between different nodes based on the node performance difference evaluation data, and obtain the node processing efficiency difference data; Step S262: Determine the resource allocation priority between different nodes based on the node processing efficiency difference data and the limited range of node resource release, and obtain the node resource allocation priority data; Step S263: Design a node interaction polling mechanism based on the node resource allocation priority data and the limited range of node resource release, and obtain the node interaction polling mechanism; Step S264: Simulate the operation scenario of the node interaction polling mechanism, and obtain the polling operation scenario simulation data; Extract the network congestion characteristics based on the polling operation scenario simulation data, and obtain the network congestion data of the polling scenario; Step S265: Perform control optimization on the node interaction polling mechanism based on the network congestion data of the polling scenario, and obtain the optimized node interaction polling mechanism; Step S266: Develop a multi-agent resource competition strategy based on the optimized node interaction polling mechanism and the node resource allocation priority data, and obtain the multi-agent resource competition strategy.

6. The multi-agent deep reinforcement learning method according to claim 5, characterized in that, Designing a node polling mechanism based on the node resource allocation priority data and the limited range of node resource release includes the following steps: Calculate the limited difference in resource release between nodes with different priorities for the limited range of node resource release based on the node resource allocation priority data, and obtain the limited difference in resource release data; Based on the resource release limited difference data, the resource release stability evaluation between nodes with different priorities is carried out to obtain the resource release stability evaluation data; According to the resource release limited difference data and the resource release stability evaluation data, the resource release correlation factor compensation is carried out on the node resource allocation priority data to obtain the node resource allocation priority compensation data; According to the node resource allocation priority compensation data, the node interaction polling mechanism is designed to obtain the node interaction polling mechanism.

7. The multi-agent deep reinforcement learning method according to claim 5, wherein The control and optimization of the node interaction polling mechanism according to the polling scenario network congestion data includes the following steps: Calculate the network congestion connection volume for the polling scenario network congestion data to obtain the network congestion connection volume data; According to the network congestion connection volume data, capture the congestion traffic surge for the polling scenario network congestion data to obtain the congestion traffic surge capture data; Conduct weighted linear growth regression analysis on the congestion traffic surge capture data to obtain the traffic surge weighted linear regression data; According to the traffic surge weighted linear regression data, perform segmented processing on the congestion traffic surge capture data to obtain the congestion traffic surge segmented data; Use the Binary Increase Congestion algorithm to perform burst segmented response control on the congestion traffic surge segmented data to obtain the congestion traffic segmented response control data; According to the congestion traffic segmented response control data, perform control and optimization on the node interaction polling mechanism to obtain the node interaction polling optimization mechanism.

8. The multi-agent deep reinforcement learning method according to claim 1, characterized in that Step S3 includes the following steps: Step S31: Normalize the multi-agent resource competition strategy to obtain the multi-agent resource competition normalization strategy; Step S32: Perform hierarchical deep learning processing according to the multi-agent resource competition normalization strategy to obtain the resource competition hierarchical deep learning data; Step S33: Based on the resource competition hierarchical deep learning data, perform multi-agent hierarchical convergence behavior coupling to obtain the hierarchical convergence behavior coupling data.

9. The multi-agent deep reinforcement learning method according to claim 8, characterized in that Step S33 includes the following steps: Step S331: Extract the hierarchical instability state among multi-agents from the resource competition hierarchical deep learning data to obtain the learning behavior hierarchical instability state data; Step S332: Calculate the instability coefficient between multiple levels for the learning behavior hierarchical instability state data to obtain the level instability coefficient; Step S333: According to the level instability coefficient, screen the initial instability values at multiple points for the learning behavior hierarchical instability state data to obtain the hierarchical instability multi-point screening data; Step S334: Perform iterative calculation on the learning behavior hierarchical instability state data according to the hierarchical instability multi-point screening data to obtain the hierarchical instability multi-point iterative data; Step S335: According to the hierarchical instability multi-point iterative data, perform multi-agent hierarchical convergence behavior coupling to obtain the hierarchical convergence behavior coupling data.

10. A multi-agent deep reinforcement learning system, characterized in that, For implementing the multi-agent deep reinforcement learning method as described in claim 1, the multi-agent deep reinforcement learning system includes: The multi-agent interaction graph construction module is used to extract the operation logs of multi-agents to obtain multi-agent operation log data; analyze the multi-agent interaction structure of the multi-agent operation log data to obtain multi-agent interaction structure data; construct a multi-agent interaction graph based on the multi-agent interaction structure data to obtain a multi-agent interaction graph; The multi-agent resource competition strategy formulation module is used to label the multi-agent load nodes of the multi-agent interaction graph to obtain multi-agent load node labeling data; calculate the resource release amount between different nodes of the multi-agent load node labeling data to obtain load node resource release amount data; calculate the limited resource release interval between different load nodes of the load node resource release amount data to obtain the node resource release limited interval; formulate a multi-agent resource competition strategy based on the node resource release limited interval to obtain a multi-agent resource competition strategy; The multi-agent hierarchical convergence coupling module is used to perform hierarchical deep learning processing according to the multi-agent resource competition strategy to obtain resource competition hierarchical deep learning data; perform multi-agent hierarchical convergence behavior coupling based on the resource competition hierarchical deep learning data to obtain hierarchical convergence behavior coupling data; The competition balance model construction module is used to construct a multi-agent resource competition balance model according to the multi-agent resource competition strategy and the hierarchical convergence behavior coupling data to obtain a multi-agent resource competition balance model, and send the multi-agent resource competition balance model to the cloud platform to perform multi-agent deep reinforcement learning.

Citation Information

Patent Citations

  • Virtual network function migration method and device based on hierarchical reinforcement learning

    CN114785693A

  • Null strategy method for improving overall benefit of multi-agent complex game system

    CN115688922A