Multi-dimensional resource dynamic game optimization control method and system based on multi-agent system
By constructing a multi-dimensional heterogeneous resource coupling map and collecting intelligent proxy strategy trajectories, calculating dual-index indexes and inputting game optimization prediction model, the problem of policy divergence and resource conflict in multi-agent systems is solved, and adaptive convergence of strategy evolution and global optimization of resource allocation is realized.
Patent Information
- Application Number
- CN202510676913.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-26
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-05-26
AI Technical Summary
In multi-agent systems, with the expansion of system scale and the complexity of resource dimensions, the policy behavior and resource competition between multiple agents show a highly nonlinear and dynamically coupled game relationship, resulting in divergence of strategy, frequent resource conflicts and degradation of overall system operation efficiency.
By constructing a multi-dimensional heterogeneous resource coupling map, model the coordination and suppression relationship between resource types to form a tensor form map; collect the strategy trajectory of the intelligent agent, calculate the group strategy divergence index and resource coupling tension index; input the double-index index index into the game optimization prediction model, output the strategy behavior convergence risk score, and adjust the strategy or update it uniformly according to the score to form an adaptive convergence feedback closed loop.
It realizes dynamic evaluation and prediction of the overall behavior deviation degree and resource conflict state of multi-agent systems, improves the adaptability and convergence of strategy evolution, and ensures the coordination and resource scheduling efficiency of multi-agent systems.
Smart Images

Figure CN120197716A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of multi-agent system decision control, and more specifically, to a multi-dimensional resource dynamic game optimization control method and system based on a multi-agent system. Background Art
[0002] In recent years, in distributed autonomous decision-making systems with agent-driven as the core, multi-agent systems have been widely applied to scenarios such as intelligent manufacturing, vehicle networking, data center resource scheduling, and robot swarm control due to their distribution, scalability, and environmental adaptability. In such systems, each intelligent agent usually has the ability of autonomous strategy evolution and makes parallel decisions around multiple heterogeneous resources. However, with the expansion of the system scale and the complexity of resource dimensions, the strategic behaviors and resource competitions among multiple agents often present highly non-linear and dynamically coupled game relationships, which easily lead to problems such as strategy divergence, frequent resource conflicts, and a decline in the overall system operation efficiency.
[0003] Traditional resource scheduling methods or strategy optimization mechanisms often only consider static resource allocation rules or optimal response strategies based on a single agent, and it is difficult to effectively capture the evolution trend of group behaviors in a multi-agent system, nor can they model the dynamic cooperation and conflict relationships among resources over time, lacking the guarantee of overall behavior predictability and system stability.
[0004] Especially in a multi-dimensional heterogeneous resource environment such as computing resources, communication resources, energy consumption resources, and storage resources, there are cooperative relationships (such as computing and communication jointly acting on efficient task completion) and inhibitory relationships (such as energy consumption limitations inhibiting high-frequency computing operations) among different resource types, while traditional models cannot dynamically identify the evolution paths of such relationships. In addition, the mutual influence among multi-agent strategic behaviors often shows non-explicit deviations and confrontations, and traditional methods often cannot accurately identify whether the group behavior is in an unbalanced state, lacking an efficient risk discrimination mechanism. Therefore, a multi-dimensional resource dynamic game optimization control method and system based on a multi-agent system are proposed herein to solve the above problems. Summary of the Invention
[0005] To achieve the above object, the present invention provides the following technical solutions:
[0006] A multi-dimensional resource dynamic game optimization control method based on a multi-agent system, comprising the following steps:
[0007] Step 1, construct a multi-dimensional heterogeneous resource coupling map, and form a graph structure with resource types as nodes and dynamic weights as edges by modeling the cooperative and inhibitory relationships among resource types within the same control cycle, thereby constituting a tensor-form map for resource state evolution;
[0008] Step 2: Collect the policy trajectories of all intelligent agents over multiple consecutive control cycles, calculate the group policy divergence index and the group resource coupling tension index based on the dynamic change of edge weights in the resource coupling map, which are used as double-index indicators to characterize the deviation degree of the overall behavior of the multi-agent system and the resource conflict state;
[0009] Step 3: Input the double-index indicators into the trained game optimization prediction model. The model outputs a policy behavior convergence risk score. If the score is lower than the first threshold, the policy remains unchanged; if the score is higher than the first threshold, enter the refined analysis process in Step 4;
[0010] Step 4: Individually evaluate the policy trajectories and resource behaviors of all intelligent agents, calculate the individual policy deviation index of each agent, and compare this index with the second threshold. If the proportion of agents exceeding the second threshold in all agents is lower than the set ratio, only adjust the policy evolution direction and resource selection behavior of the agents exceeding the standard, that is, the agents exceeding the second threshold; if the proportion is not lower than the set ratio, perform a unified policy update operation for all agents;
[0011] Step 5: After the policy adjustment or unified update is completed, update the edge weight parameters and agent behavior trajectory data in the resource coupling map, which are used as the basis for the next round of evaluation and optimization, forming a continuously iterative dynamic feedback loop to achieve the adaptive convergence of policy evolution and the global optimization of resource allocation behavior.
[0012] In a preferred implementation, during the process of constructing a multi-dimensional heterogeneous resource coupling map, the resource types are limited to at least include computing resources, storage resources, communication resources, and energy consumption resources. In each control cycle, collect the usage intensity data of various resources as the original observation values, and use the original observation values to calculate the synergy function and inhibition function between resource types;
[0013] The synergy function is used to reflect the positive mutual influence degree between two resource types in actual tasks, and the inhibition function is used to reflect the negative restrictive relationship generated between two resource types in the allocation behavior;
[0014] Combine the calculation results of the synergy function and the inhibition function into a composite weight value, and use this composite weight value as the dynamic weight of the edge to connect the corresponding resource type nodes, thereby constructing a graph structure with resource types as nodes and dynamic weights as edges, and applying a time series identifier to this graph structure to record the evolution changes of this structure in multiple consecutive control cycles.
[0015] In a preferred embodiment, during the process of constructing a multi-dimensional heterogeneous resource coupling map, the calculation methods of the synergy function and the inhibition function between resource types are based on the resource synchronization usage rate and the conflict occupancy rate within a control period, specifically as follows:
[0016] Within each control period, count the number of resource allocation events participated by each resource type in each intelligent agent behavior, and calculate the joint usage frequency of different resource types being called simultaneously during the same time period. Divide this joint usage frequency by the geometric mean of the individual usage frequencies of each resource to obtain the initial value of the synergy function;
[0017] By comparing the number of allocation conflict events between different resource types, that is, the number of times two resource types cannot be allocated simultaneously due to quota restrictions in the resource allocation decision, and combining the weight of the task failure events caused by the agent behavior, calculate the initial value of the inhibition function of one resource on another resource;
[0018] Synthesize the synergy function and the inhibition function into the dynamic weight of the edge through normalization mapping, complete the update of the edge weight between resource types in the graph structure, and use it as the edge weight parameter for constructing a tensor-form map reflecting the evolution of the resource state.
[0019] In a preferred embodiment, during the process of calculating the population strategy divergence index, first perform time series encoding on the strategy trajectories of all intelligent agents within multiple consecutive control periods, represent each strategy trajectory as a set of vectors with a fixed length. Within the same control period, calculate the mean Euclidean distance between the strategy vector of each agent and the strategy vectors of all other agents, and use the ratio of this mean to the standard deviation of the change of the agent's own strategy vector to obtain the divergence score of this agent; then normalize the divergence scores of all agents and perform mean aggregation to obtain the population strategy divergence index, which is used to measure the deviation amplitude, behavior divergence degree, and the consistency strength of the evolution trend among all intelligent agents in terms of their strategy behaviors.
[0020] In a preferred embodiment, during the process of calculating the population resource coupling tension index based on the dynamic change of the edge weight in the resource coupling map, first extract the weight change values of all edges within two adjacent control periods, and calculate the tension strength of each edge during this change process, and its calculation method is the square of the difference between the weight of the current period and the previous period of this edge multiplied by the initial weight factor; then, with each resource type as the center, summarize the tension strength values of all its connected edges, and calculate the local tension response value of this resource type; then take the weighted average of the local tension response values of all resource types to generate the population resource coupling tension index, which is used to measure the degree of structural tension and the resource coordination stability at the system level that occur during the evolution of the resource state in the current control period.
[0021] In a preferred embodiment, during the process of inputting the double-exponential index into the trained game optimization prediction model and outputting the policy behavior convergence risk score by the model, the group strategy divergence index and the group resource coupling tension index are first subjected to normalization mapping, and a double-index input vector is formed according to a preset dimension. This input vector, as the expression form of the current multi-agent system behavior state, is input into the game optimization prediction model;
[0022] The game optimization prediction model is trained based on the historical evolution data within the control period and adopts a non-linear prediction structure integrated by a feedforward neural network and a graph correlation reasoning structure. By fitting and predicting the position relationship of the input vector in the historical behavior mapping space, it outputs the convergence risk score of the current behavior state in the process of policy evolution. This score is used to reflect the behavior consistency, evolution stability, and future convergence trend of the current multi-agent system in the policy dimension, and is used to determine whether to trigger the next refined analysis process.
[0023] In a preferred embodiment, during the process of separately evaluating the policy trajectories and resource behaviors of all intelligent agents and calculating the individual policy deviation index of each agent, first, the policy trajectories of each agent within multiple consecutive control periods are selected, and a sequence of policy behavior vectors of each agent in each control period is constructed;
[0024] After standardizing the vector sequence, the change amplitude of the policy behavior vectors between adjacent control periods is obtained, and the average change amplitude of the agent is used as the internal variability of its policy trajectory;
[0025] The sequence of policy behavior vectors of other agents among all intelligent agents is selected, and the average vector difference of the policy of this agent relative to all the other agents is constructed as the external difference of its behavior deviation degree;
[0026] The internal variability of the policy trajectory and the external difference of the behavior deviation degree are synthesized into a single index in the form of a weighted linear combination, and this index is smoothed within a fixed window. Finally, the individual policy deviation index of this agent at the current stage is generated, which is used to measure the instability of the agent's policy behavior, the difference from the group behavior, and the potential abnormal behavior trend.
[0027] In a preferred embodiment, when only the over-standard agents, i.e., the agents exceeding the second threshold, are adjusted in terms of the policy evolution direction and resource selection behavior, a behavior direction correction factor is set based on the difference direction between the policy behavior vector of the over-standard agent in the current control period and the average policy behavior vector of all intelligent agents. This correction factor is the negative of the cosine value of the angle between the difference vector and the unit convergence vector multiplied by the non-zero proportionality coefficient γ. This correction factor is used in the policy generation vector superposition term of the over-standard agent in the next control period to adjust the direction of the policy behavior evolution path.
[0028] Collect the resource set used by the over-standard agent in the current control period, and form a deviation resource set with the resource having the largest difference in the current resource usage frequency ranking of all agents. Set the resource allocation priority compression factor η of the over-standard agent in the next period as the initial resource selection probability multiplied by [1 - fd / (ft + C)], where fd is the resource usage frequency of the agent, ft is the average resource usage frequency of the group, and C is a preset non-zero constant. And set that this compression operation only acts on the resource types in the deviation resource set, so as to achieve dynamic constraints on resource selection behavior.
[0029] In a preferred embodiment, if the proportion of agents exceeding the second threshold in all agents is not less than the set proportion, then perform a unified policy update operation for all agents, including extracting the average policy trajectory vector of all agents in the past N control periods as the reference behavior model, obtaining the average deviation degree of all agent behavior vectors in this model, and using the policy behavior vector with the smallest deviation degree as the convergence direction vector. Reset the initial value of the policy generation parameter of all agents as this convergence direction vector multiplied by the initialization scale factor, and sort all resource behavior parameters according to the resource selection frequency, and map them to the preferred resource interval constructed based on the group resource co-occurrence probability. Apply this mapping structure uniformly in the new control period to achieve the synchronous convergence of agent behavior and resource selection behavior, and improve the overall behavior consistency and resource scheduling coordination.
[0030] In a preferred embodiment, a multi-agent system-based multi-dimensional resource dynamic game optimization control system includes:
[0031] A resource map builder, used to build a multi-dimensional heterogeneous resource coupling map. By modeling the cooperation relationship and inhibition relationship between resource types in the same control period, generate a graph structure with resource types as nodes and dynamic weights as edges, and output a tensor-form map to depict the resource state evolution.
[0032] A behavior index calculator is used to collect the policy trajectories of all intelligent agents in multiple consecutive control cycles, calculate the group policy divergence index and the group resource coupling tension index obtained from the dynamic changes of the edge weights in the resource coupling map, and generate a double-index indicator to characterize the overall behavior deviation degree and resource conflict state of the multi-agent system;
[0033] A game risk discriminator is used to receive the double-index indicator and input it into a trained game optimization prediction model, output a policy behavior convergence risk score, and based on the comparison result between this score and the first threshold, determine whether to trigger the subsequent analysis and adjustment process;
[0034] A local evolution regulator is used to separately evaluate the policy trajectories and resource behaviors of all intelligent agents, calculate the individual policy deviation index, and based on the comparison result between this index and the second threshold, determine whether the proportion of over-standard agents is lower than the set proportion. If it is lower, only adjust the policy evolution direction and resource selection behavior of the over-standard agents. If it is not lower, perform a unified policy update operation for all agents;
[0035] An evolution feedback updater is used to update the edge weight parameters and agent behavior trajectory data in the resource coupling map after the policy adjustment or unified update is completed, provide a new input basis for the next control cycle, and form an adaptive convergence feedback loop for policy evolution.
[0036] The technical effects and advantages of the present invention:
[0037] By constructing a graph structure with resource types as nodes and dynamic weights as edges, the present invention realizes the modeling of the cooperation relationship and inhibition relationship between multi-dimensional heterogeneous resources, and further transforms this graph structure into a tensor form map as the basic expression of resource state evolution. This processing method enables the explicit expression and dynamic tracking of the complex dependence relationship between resources, not only improves the resolvability of resource modeling, but also enhances the system's perception ability of structural changes at the resource level, thus contributing to more accurate evaluation and adjustment at the resource level in the subsequent policy optimization process.
[0038] By collecting the policy trajectories of all intelligent agents in multiple consecutive control cycles, calculating the group policy divergence index and the group resource coupling tension index based on the dynamic changes of the edge weights in the resource coupling map, and jointly inputting the two as double-index indicators into a trained game optimization prediction model, the present invention realizes the synchronous evaluation and prediction of the overall behavior deviation degree and resource conflict state of the multi-agent system. This double-index mechanism breaks through the limitations of single behavior judgment or resource monitoring, enabling the system to dynamically perceive whether there is a non-convergence risk in the current game situation, providing a scientific basis and quantifiable trigger conditions for subsequent detailed analysis and behavior intervention.
[0039] During the strategy evolution process of the present invention, a mechanism is designed to trigger the refinement analysis process based on the convergence risk score of the strategy behavior, and local behavior adjustment or overall strategy update operations are respectively executed according to whether the individual strategy deviation index exceeds the threshold and the proportion of over-standard agents. After completing the behavior adjustment, the system synchronously updates the edge weight parameters in the resource coupling graph and the agent behavior trajectory data, thus constituting a complete behavior-resource-feedback closed-loop optimization path. This mechanism significantly improves the self-adaptability and convergence of the strategy evolution, and can continuously maintain the coordination and resource scheduling efficiency of the multi-agent system operation in a dynamic environment. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] For the convenience of those skilled in the art to understand, the present invention will be further described below with reference to the accompanying drawings;
[0041] Figure 1 It is the schematic diagram of the multi-dimensional resource dynamic game optimization control method based on a multi-agent system in the present invention.
[0042] Figure 2 It is the schematic diagram of the multi-dimensional resource dynamic game optimization control system based on a multi-agent system in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0043] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0044] Refer to Figure 1 - Figure 2 The following embodiments are obtained:
[0045] Embodiment 1: The multi-dimensional resource dynamic game optimization control method based on a multi-agent system includes the following steps:
[0046] Step 1: Construct a multi-dimensional heterogeneous resource coupling graph. By modeling the cooperation and inhibition relationships among resource types within the same control cycle, a graph structure with resource types as nodes and dynamic weights as edges is formed, constituting a tensor-form graph for resource state evolution; constructing a multi-dimensional heterogeneous resource coupling graph means establishing a structured expression model that can describe the interaction relationships among resources during the operation of a multi-agent system. In this step, each resource type is abstracted as a node in the graph structure, and based on the degree of cooperation and inhibition between resource types, the edges between them are constructed. The weight of each edge is dynamically updated according to the actually observed resource usage data, so as to realize the evolution of the edge weight over time. This graph structure is adjusted according to real-time data in each control cycle, so as to accurately reflect the coupling strength and directionality between current resource states. Finally, this graph structure is encoded into a tensor form for subsequent processing and model input, serving as the input basis and behavioral background representation of resource state evolution behavior.
[0047] Step 2: Collect the policy trajectories of all intelligent agents in multiple consecutive control cycles, calculate the group policy divergence index, and the group resource coupling tension index based on the dynamic change of edge weights in the resource coupling graph, as a double-index indicator to characterize the deviation degree of the overall behavior of the multi-agent system and the resource conflict state; collecting the policy trajectories of all intelligent agents in multiple consecutive control cycles means collecting the policy behavior data of each agent in a continuous time period from the system level and transforming it into a set of policy trajectories expressed in time series. After obtaining all the policy trajectories, first conduct statistical analysis at the group behavior level. By comparing the policy paths among agents, an overall index reflecting the consistency and divergence degree of behaviors among agents is calculated, called the group policy divergence index; at the same time, based on the change of edge weights over time in the resource coupling graph constructed in the first step, the degree of tension change of resources between different control cycles is calculated, so as to obtain the group resource coupling tension index. The two together serve as a double-index indicator to characterize the current system operation state, reflecting the behavior deviation trend on the one hand and revealing the evolution level of resource conflicts on the other hand, with the significance of behavior-resource cooperation evaluation.
[0048] Step 3: Input the double-exponential index into the trained game optimization prediction model. The model outputs a policy behavior convergence risk score. If the score is lower than the first threshold, the policy remains unchanged. If the score is higher than the first threshold, enter the refined analysis process in Step 4. Input the double-exponential index obtained in Step 2 into the trained game optimization prediction model, aiming to predict whether there is a potential non-convergence risk in the system based on the current policy behavior state and resource coupling characteristics. This prediction model is a non-linear structure model pre-trained with historical multi-cycle evolution data, which can identify whether the complex game situation represented by the double index will cause problems such as violent fluctuations in policy behavior, local optimal traps, or imbalances in resource competition. The model outputs a policy behavior convergence risk score in numerical form, representing the stability of the evolution path of the multi-agent system at the current stage. If the score result is lower than the first threshold, it indicates that the current state of the system is acceptable and no adjustment is needed. If the score exceeds the first threshold, it means that the system may face the risk of policy divergence or increased resource conflicts and needs to further enter the subsequent refined evaluation stage.
[0049] Step 4: Individually evaluate the policy trajectories and resource behaviors of all intelligent agents, calculate the individual policy deviation index of each agent, and compare this index with the second threshold. If the proportion of agents exceeding the second threshold in all agents is lower than the set ratio, only adjust the policy evolution direction and resource selection behavior of the agents exceeding the standard, that is, the agents exceeding the second threshold. If this proportion is not lower than the set ratio, perform a unified policy update operation for all agents. After entering the refined analysis process, more detailed analysis of the policy trajectories and resource behaviors of all intelligent agents is required. The system will analyze the behavior trajectories of each agent one by one, calculate its behavior fluctuations and deviation degrees within multiple control cycles, so as to form an individual policy deviation index. This index represents the consistency level between a single agent and the overall behavior pattern. The system compares the individual policy deviation index of each agent with the second threshold, identifies the agents exceeding the standard, and counts the proportion of these agents in all agents. If this proportion is lower than the preset set ratio threshold, it indicates that the problem mainly focuses on the behavior deviations of a small number of individuals, and the system will only perform targeted adjustments on these agents exceeding the standard, including methods such as correcting the policy direction and compressing the resource selection priority. If the proportion of agents exceeding the standard is not lower than the set ratio, it indicates that the group behavior deviates significantly, and the system will trigger a unified policy reconstruction process for all agents to achieve an overall recalibration of the system behavior.
[0050] Step 5. After the strategy adjustment or unified update is completed, update the edge weight parameters and agent behavior trajectory data in the resource coupling graph, which serves as the basis for the next round of evaluation and optimization, forming a continuously iterative dynamic feedback loop to achieve the adaptive convergence of strategy evolution and the global optimization of resource allocation behavior. After completing the local adjustment for individual agents or the unified strategy update for all agents, it is necessary to synchronously update the aforementioned resource coupling graph and agent behavior trajectory data. The significance of this step is to incorporate the adjustment results in the current control cycle into the state initialization and evaluation calculation of the next cycle. Specifically, it includes the adjustment of the edge weights in the resource coupling graph to reflect the current collaborative or conflicting relationships between resources after adjustment; it also includes the update of the agent behavior trajectory to record the behavior output of each agent in the latest control cycle. Through this update mechanism, a complete closed-loop can be formed, enabling the system to make accurate judgments based on the latest state in each control cycle, achieving adaptive convergence during the strategy evolution process, and at the same time promoting the evolution of resource allocation behavior towards a globally optimal stable structure, ensuring the efficiency and stability of the entire multi-agent system during long-term operation.
[0051] In the process of constructing a multi-dimensional heterogeneous resource coupling graph, first classify and divide the resource types involved in the multi-agent system, and clearly define them to include at least computing resources, storage resources, communication resources, and energy consumption resources. Computing resources refer to the processing capabilities used to execute agent strategy logic, such as the time slices of a central processing unit; storage resources refer to the storage capacity used for data buffering or intermediate state preservation; communication resources refer to the data transmission bandwidth used for interactions between agents or between agents and the outside world; energy consumption resources refer to the actual electricity, energy, etc. consumed during the execution of behaviors. The above four types of resources constitute the core resource set of the system and have heterogeneity and dynamics.
[0052] In each control cycle, the system collects the usage conditions of the above four types of resources to obtain their usage intensity data. The usage intensity data refers to the frequency or occupancy ratio of each type of resource being called by each intelligent agent per unit time. This data is input as the original observation value into the subsequent relationship calculation process. Using this original observation value, calculate the synergy function and inhibition function between resource types in sequence. The synergy function is used to reflect the positive mutual influence degree generated by two resource types in actual agent tasks. Specifically, it means that the higher the degree of simultaneous invocation of two resource types within the same time period, the more collaborative they are in task execution, and the higher the value of the synergy function; for example, when a certain intelligent agent executes a video processing task, computing resources and communication resources are often simultaneously invoked, and at this time the value of the synergy function between the two is relatively high. The synergy function can be calculated through the proportional relationship between the joint frequency of resource usage and their respective independent usage frequencies.
[0053] The inhibition function is used to reflect the negative constraint relationship existing between two resource types in the resource allocation behavior. It refers to the situation where when there is a quota conflict between two resources, they cannot be allocated simultaneously, or there is an exclusion phenomenon due to restrictive conditions, there is a resource competition relationship between them, and the value of the inhibition function increases. The inhibition function can be estimated by combining indicators such as the number of resource conflict events and the failure rate of agent tasks. For example, if the computing resources are highly invoked in a certain control cycle, resulting in the restricted invocation of energy consumption resources, it indicates that there is a certain degree of inhibition between the two.
[0054] The calculation results of the synergy function and the inhibition function are linearly weighted and synthesized to form a composite weight value. The linear weighted synthesis is not limited to a specific form, and those skilled in the art in the prior art can choose to use it according to actual needs. This weight value is used to express the comprehensive interaction relationship between two resource types in the current control cycle. The composite weight value is directly used as the dynamic weight of the edge between the resource type nodes in the graph structure, and this weight can be updated as the control cycle changes.
[0055] Specifically: In each control cycle, count the number of resource allocation events participated by each resource type in each intelligent agent behavior, and calculate the joint usage frequency of different resource types being invoked simultaneously in the same time period. Divide this joint usage frequency by the geometric mean of the individual usage frequencies of each resource to obtain the initial value of the synergy function; by comparing the number of allocation conflict events between different resource types, that is, the number of times when two resource types cannot be allocated simultaneously due to quota restrictions in the resource allocation decision, and combining the weight of the task failure event caused by the agent behavior, calculate the initial value of the inhibition function of one resource to another resource; synthesize the synergy function and the inhibition function through a normalization mapping into the dynamic weight of the edge, complete the update of the edge weight between the resource types in the graph structure, and use it as the edge weight parameter for constructing a tensor form graph spectrum reflecting the evolution of the resource state.
[0056] The purpose of constructing a graph structure with resource types as nodes and composite weights as edges is to establish a visual expression form that can represent the multi-dimensional interaction logic between resources. Through this structure, the system can track the evolution trend of resource relationships in the form of a graph. To achieve the tracking of the resource state over time, a time sequence identifier is further imposed on this graph structure. The time sequence identifier refers to annotating the corresponding control cycle number or timestamp information for each change in the edge weight, so as to record the evolution history of this graph structure in multiple consecutive control cycles. Finally, based on the above structured process, the formed resource coupling graph spectrum will be input into the subsequent evaluation and decision-making modules as a tensor form graph spectrum, constituting an important basis for resource state modeling in the game optimization control process of the multi-agent system.
[0057] In the process of calculating the group strategy divergence index, it is first necessary to obtain the behavioral data of all intelligent agents executing strategies within multiple consecutive control cycles to form the strategy trajectory of each agent. The strategy trajectory refers to the continuous decision-making records made by a certain intelligent agent according to the environmental state within multiple control cycles, which can be expressed as a time series of behavioral decisions.
[0058] For the convenience of unified processing and analysis, each strategy trajectory is encoded into a time series, that is, the original behavior sequence is transformed into a set of fixed-length numerical vectors in chronological order. Each vector represents the strategy behavior of the agent within one control cycle. For example, the vector can include dimensions such as strategy type encoding, resource usage preference, and current state feedback. After all strategy trajectories are encoded, a structured vector set of agent behavior time series is formed.
[0059] Within any one control cycle, for the strategy vector of each agent, calculate the Euclidean distance between it and the strategy vectors of all other agents in the system during this cycle. The Euclidean distance refers to the straight-line distance between two vectors in a multi-dimensional space, which is used to measure the difference in behavioral characteristics between two agents. Average the numerical values of the Euclidean distances between this agent and all other agents to obtain the mean behavioral distance of this agent relative to the group during this cycle. The larger this value is, the higher the "dispersion degree" of this agent's behavior in the group.
[0060] Calculate the standard deviation of the changes between the strategy vectors of this agent within the entire control cycle sequence. This standard deviation is used to represent the degree of internal fluctuation of its strategy behavior. Use the mean behavioral distance of this agent within a single cycle as the numerator and its cross-cycle standard deviation as the denominator to calculate the divergence score of this agent. This score reflects the proportional relationship between the degree of deviation of its behavior from the external group and the stability of its own behavior.
[0061] For example, if an agent has small behavioral fluctuations in each cycle but has a large behavioral difference from other agents, its divergence score will be high, indicating that its behavior deviates from the group trend. On the contrary, if this agent has both stable behavior and is consistent with the group strategy, its divergence score will be low, indicating that this agent's behavior is in the core area of group convergence.
[0062] Collect the divergence scores of all intelligent agents in the system and perform normalization, that is, through a unified proportional transformation, make the scores of all agents fall within the same dimension range (such as between zero and one), eliminating the influence brought by different scales. Then, take the arithmetic mean of all the normalized divergence scores, that is, calculate the equal weights of the contributions of each agent, and obtain the group strategy divergence index under the current control period. This group strategy divergence index is used to measure the overall consistency level of all intelligent agents in the policy behavior during the current stage. The higher the index value, the greater the degree of divergence between agent behaviors, and the system policy evolution shows a discrete trend; the lower the index value, the more convergent the agent strategies are, and the system policy evolution direction is unified and stable, which can be used as one of the important bases for the subsequent game optimization model to judge whether adjustment is needed.
[0063] In the process of calculating the group resource coupling tension index based on the dynamic change of the edge weights in the resource coupling map, it is first necessary to clarify that the structure of the resource coupling map has been constructed in the control period. This map consists of several resource type nodes and the edges between them, and each edge has an edge weight value representing the relationship strength between resources. The edge weight value is a dynamic parameter and is updated continuously with the control period. To obtain the degree of change in the relationship strength between resources in different periods, extract the edge weight values of all edges in two adjacent control periods, and calculate the weight change value of each edge between these two periods. The weight change value refers to the edge weight value of the current control period minus the edge weight value of the previous control period, indicating the fluctuation range of the resource interaction intensity.
[0064] Calculate the tension intensity of each edge in this change process. The calculation method of the tension intensity is the square of the weight difference between the current period and the previous period of this edge, which is used to highlight the edges with larger change amplitudes; and multiply it by the weight value of this edge in the initial state as an importance adjustment factor for its change amplitude, and finally obtain the tension intensity of this edge. This process can be interpreted as that the edges of "high change + high initial coupling relationship" will generate stronger tension responses.
[0065] Taking each resource type node as the center, traverse all the edges connected to this node in the graph, and summarize the corresponding tension intensity values. This summarization process is to calculate the total tension response of each resource type to other resource types in this period, and obtain the local tension response value of this resource type. The larger the local tension response value, the higher the structural pressure faced by this resource type in the current control period.
[0066] For example, in a certain period, the edge weight between communication resources and energy consumption resources fluctuates violently, and the initial coupling weight is large. Then the tension intensity corresponding to this edge is large, and the local tension response values on the communication resource and energy consumption resource nodes increase accordingly, indicating that there are serious collaborative fluctuations or competitive imbalances between these two types of resources in this period. Finally, the local tension response values of all resource types are weighted and averaged to generate the final group resource coupling tension index. The weighting factor can be set according to the importance level of each resource type in the system or its proportion of basic resources. The generated index is used to measure the degree of system tension presented in the process of the overall resource state evolution in the current control period. The higher the index value, the more violent collaborative or conflict changes exist between resources in this period; the lower the value, the more stable the collaborative relationship between resources and the system is in a state of resource coordination and balance. As a resource structural indicator in the double-index evaluation system, the group resource coupling tension index, combined with the group strategy divergence index, is used to jointly determine whether subsequent strategy intervention and resource optimization behaviors need to be triggered.
[0067] In the process of inputting the double-index indicators into the trained game optimization prediction model and obtaining the strategy behavior convergence risk score output by the model, first, the group strategy divergence index and the group resource coupling tension index are normalized and mapped. Normalization mapping means converting the original indicator values to a unified numerical interval (such as 0 to 1) to eliminate the dimensional differences and the influence of measurement units, so that the two indexes have equivalent expression capabilities in the model. This normalization method can adopt the maximum-minimum normalization method or the Z-score normalization method to ensure the comparability of index values in different stages.
[0068] According to the preset dimension requirements of the model input structure, the two normalized indexes are combined into a double-index input vector in a fixed order. This input vector is used as the expression form of the behavior state of the current multi-agent system in the current control period, that is, the joint characterization vector of the system behavior structure and resource coupling characteristics. This input vector is fed into the trained game optimization prediction model, which is a prediction network structure constructed based on data from multiple behavior stages in the historical control period and has the ability of nonlinear modeling. Among them, "completed training" means that the model has been trained and optimized through the historical behavior data sample set to form a mapping relationship capable of predicting risks for the input state.
[0069] The game optimization prediction model adopts a non - linear prediction structure constructed by integrating a feed - forward neural network and a graph - associated reasoning structure. The feed - forward neural network is a basic artificial neural network architecture suitable for learning the mapping from input vectors to output values and can extract complex non - linear features from static data. The graph - associated reasoning structure is used to process the part of the input state involving the evolution of the resource graph. It can model the evolution law of resource tension in the graph space based on the edge - weight change trend between historical graph structures, enhancing the model's ability to express the mutual influence between resources and behaviors.
[0070] The model searches for similar behavior patterns in the historical behavior mapping space constructed by the double - index input vector in its training memory. According to the position relationship of the input in this space, it performs non - linear fitting and inference prediction. Finally, the model outputs a continuous numerical value as the convergence risk score of the current input state in the process of policy evolution. This score represents the possibility of unstable evolution or potential divergence trend of the entire multi - agent system at the behavioral policy level during the current control period. For example, when both the group policy divergence index and the resource coupling tension index are in the middle - to - high range, indicating that the two - dimensional state of behavior and resources is relatively chaotic at the current stage, the model may output a high - risk score value; conversely, if both indices remain low and fluctuate smoothly, the model outputs a low - risk score. This policy behavior convergence risk score is used as the judgment basis in the subsequent control process. If this score is lower than the preset first threshold, the system maintains the current policy unchanged; if it is higher than the first threshold, the system enters the next step of detailed analysis process to separately evaluate the behaviors of each agent and trigger the policy intervention mechanism.
[0071] In the process of separately evaluating the policy trajectories and resource behaviors of all intelligent agents and calculating the individual policy deviation index of each agent, first, the policy trajectories of each agent in multiple consecutive control cycles are selected. The policy trajectory refers to the sequence of behavioral policies sequentially executed by the agent in a continuous time series, which can reflect its behavior evolution path.
[0072] The policy trajectory is transformed into a structured representation. Specifically, in each control cycle, the policy behavior vector of the agent is constructed. This vector contains key behavior features such as its policy decision, resource selection preference, and environmental state perception in this cycle. The policy behavior vectors in all control cycles are arranged in chronological order to form the policy behavior vector sequence of the agent.
[0073] Among them, policy decision means that the agent selects the final execution action from a set of optional actions according to its policy function within the current control cycle. For example, when faced with multiple task choices, it decides whether to execute a "high-energy consumption and high-return task" or a "low-resource consumption and conservative strategy". This feature can be written into a vector by numbering different policy actions and mapping them into discrete numerical forms. For example, for "Task A", the value of the first dimension in the vector is 1, and for "Task B", it is 2, etc.
[0074] Resource selection preference means which resource types the agent tends to preferentially use for behavior execution within the current control cycle. This preference can be expressed by statistical resource request order or weights. For example, if the agent's resource request ranking in this cycle is "communication resource > storage resource > computing resource", then a set of resource preference weight values can be assigned to it in the vector, such as [0.5, 0.3, 0.2].
[0075] Environmental state perception means the subjective reception and internal coding result of the agent for the environmental state before making a policy behavior. For example, state indicators such as the current task complexity, the density of neighboring agents, and the network communication delay. Usually, it can be composed of the state feature codes output by its perception module. For example, indicators such as high task complexity, 20% remaining battery, and weak communication signal can be standardized and used as vector dimensions for input, such as [0.8, 0.2, 0.1].
[0076] Concatenating the above three types of features - policy decision, resource selection preference, and environmental state perception - according to the preset dimensions constitutes the policy behavior vector of the agent within the current control cycle. For example, a vector containing 3-dimensional policy decision coding, 3-dimensional resource preference values, and 3-dimensional environmental state perception can be represented as a 9-dimensional vector: [1, 0, 0, 0.6, 0.3, 0.1, 0.7, 0.5, 0.2]. Among them, the first dimension is the selected policy coding, the second to fourth dimensions are the resource preference ratios, and the fifth to seventh dimensions are the state perception parameters. Arranging the policy behavior vectors within all control cycles in chronological order forms the policy behavior vector sequence of the agent, that is, a two-dimensional structure of T×D, where T is the number of control cycles and D is the dimension of each behavior vector. This sequence is used to capture the time dynamic characteristics of the agent's behavior evolution and is used in subsequent analyses for calculations such as behavior volatility and group consistency. For example, if an agent generates 5 policy behavior vectors within 5 control cycles, and each vector has the above 9-dimensional structure, then its policy behavior vector sequence is a 5-row and 9-column matrix.
[0077] Normalizing the vector sequence means performing unified numerical adjustment on each dimension of each vector so that it is expressed on a comparable scale, eliminating the interference caused by inconsistent original data dimensions or different numerical distributions. The normalization method can adopt the min-max scaling or zero-mean normalization method.
[0078] Calculate the magnitude of the change in the policy behavior vector of the agent between adjacent control cycles, that is, the degree of change between every two consecutive vectors. Average these magnitudes of change to obtain the internal variability of the policy trajectory of the agent in this sequence. This metric is used to represent the fluctuations in the agent's own behavior over time. The larger the value, the more unstable its behavior and the more frequent the decision changes.
[0079] Select the policy behavior vector sequences of other agents among all intelligent agents, align them uniformly to the same time axis and dimension as the current agent, and calculate the average vector difference between the current policy behavior vector of the agent and the policy behavior vectors of all other agents in the same control cycle as the degree of behavioral deviation between it and the group. This value represents the gap between the current decision of the agent and the group policy center.
[0080] Perform a weighted linear combination of the above internal variability of the policy trajectory and the degree of behavioral deviation, that is, according to two preset weight factors, add them after weighting to form a single behavioral metric for the agent in the current control cycle. This metric combines the two dimensions of behavioral stability and group consistency and is used to comprehensively evaluate the abnormality of the agent's current behavior. To further eliminate the interference of occasional fluctuations, smooth this behavioral metric within a fixed sliding window. The sliding window is an interval of a set number of consecutive control cycles, and a weighted moving average method can be used to smooth the metric value to highlight the behavioral trend and weaken the temporary interference values. Finally, generate the individual policy deviation index of the agent in the current stage. The larger the value of this index, the greater the behavioral fluctuations of the agent and the more serious its deviation from the group, indicating a possible abnormal behavioral trend; the smaller the value, the more stable its behavior and the more consistent with the group policy, and it is a convergent agent in the system. For example, if an agent's policy changes continuously in the last 5 cycles and there is a significant gap between its policy and the average behavior of other agents in most cycles, its individual policy deviation index will be significantly higher than that of other agents, and the system will identify it as a potential source of instability. This individual policy deviation index will be used in subsequent behavior judgments as the decision basis for whether to correct the policy evolution direction and adjust the resource selection of the agent.
[0081] When only adjusting the policy evolution direction and resource selection behavior for over-standard agents, i.e., agents exceeding the second threshold, first obtain the policy behavior vector of the over-standard agent in the current control cycle. This vector is a structured encoded behavior expression that already includes multiple dimensions such as its policy decision encoding, resource preference value, and environmental state perception parameters. At the same time, the system obtains the mean vector of the policy behaviors of all intelligent agents in this control cycle, that is, the policy behavior vectors of all agents are averaged by dimension to serve as the group behavior center expression for the current cycle. By calculating the difference direction between the vector of this over-standard agent and this mean vector, a directional indicator of its current behavior deviating from the group evolution trend can be obtained. Based on this difference direction, a behavior direction correction factor is set. This factor is used to measure the degree of consistency between this deviation direction and a preset unit convergence vector. The unit convergence vector represents the standard evolution direction obtained by modeling the policy trajectory convergence trend during the training phase of the system and is a normalized reference direction vector.
[0082] The calculation method of the behavior direction correction factor is as follows: calculate the cosine value of the angle between the difference direction vector and the unit convergence vector (used to represent the direction similarity between the two), then take its negative number (representing that the correction direction should be adjusted reversely) and multiply it by a non-zero proportionality coefficient γ (gamma). γ is an adjustable control parameter that determines the adjustment amplitude. For example, if the angle is 60°, the cosine value is 0.5, and γ is 0.8, then the correction factor is -0.4. This value indicates that the agent's behavior deviates from the group direction and should callback a certain weight in the opposite direction in the next cycle. Apply this correction factor to the superposition term of the policy generation vector of the over-standard agent in the next control cycle, that is, add the vector component of this correction direction in the behavior output stage of the policy generation model to achieve fine-tuning of the behavior path and make it gradually approach the group convergence center.
[0083] Collect the resource set used by the over-standard agent in the current control cycle, that is, the resource types actually applied for or occupied during the execution of the policy, and compare it with the resource usage frequency ranking of all agents in the current cycle to find the resource types with relatively high usage frequency by this agent but relatively low group average usage frequency, which are defined as the deviation resource set. These resource types represent that the agent has inconsistent resource preferences.
[0084] Set the resource allocation priority compression factor η (eta) for the next control cycle of the out-of-standard agent. This factor is used to suppress the probability of the agent using deviated resources in the resource selection strategy. The calculation formula for η is: the initial resource selection probability multiplied by [1 - fd / (ft + C)], where fd is the usage frequency of a certain resource currently used by the agent (such as the proportion of the number of calls to the total available number of resources); ft is the average usage frequency of the group for this resource; C is a preset non-zero constant used to avoid a zero denominator and introduce suppression smoothness. This compression operation is only applied to the resource types in the deviated resource set and does not affect the agent's degree of freedom of behavior on the mainstream resources of the group. What is ultimately achieved is: dynamically adjusting the resource policy function of the agent in the next cycle, making it assign a lower adoption probability to deviated resources, guiding its resource usage behavior to approach the group consensus preference, so as to achieve the goal of behavior collaborative control at the resource level.
[0085] Suppose an agent A frequently uses "communication resources" and "energy consumption resources" in the current cycle. However, according to the group statistics, the usage frequency of "energy consumption resources" is extremely low. Then, the deviation frequency of agent A from the group average in "energy consumption resources" is very large. At this time: if the usage frequency fd = 0.6, the group average ft = 0.2, and the constant C = 0.1, then the compression factor η = the original adoption probability × [1 - 0.6 / (0.2 + 0.1)] = the original adoption probability × [1 - 2.0] = the original adoption probability × (-1.0). Since η cannot be negative, the system sets the lower limit of η to 0, that is, this resource will be completely suppressed in the next cycle.
[0086] If the proportion of agents exceeding the second threshold in all agents is not less than the set ratio, it indicates that there is a widespread deviation phenomenon in the process of policy behavior evolution of the current multi-agent system, which belongs to the group-level convergence risk state. At this time, the system will no longer only adjust individual agents, but trigger a unified policy update operation for all agents to restore behavior consistency and improve resource allocation coordination. This update process first includes: extracting the mean vector of the policy trajectories of all agents in the past N control cycles. The mean vector of the policy trajectories refers to the eigenvector representing its long-term policy preference obtained by weighted averaging the policy behavior vectors of each agent within N control cycles. After calculating the trajectory mean vectors for all agents respectively, their set is input into the unified analysis model. In this model, further obtain the average deviation degree of all agent behavior vectors relative to the trajectory mean vector. The average deviation degree is a measure of the distance scale between an agent's behavior and the group consensus behavior, usually calculated by the vector Euclidean distance. After calculating the average deviation degrees of all agents, select the policy behavior vector with the smallest deviation degree among them as the most representative behavior in the current system state. This vector is defined as the convergence direction vector, that is, the behavior target direction that all agents should jointly approach in the future policy evolution process.
[0087] Based on this convergence direction vector, the system will reset the initial values of the policy generation parameters of all agents. Specifically, multiply the convergence direction vector by an initialization scale factor, which is a preset positive number used to control the amplitude and convergence speed of the newly generated policy. A larger value indicates a more significant adjustment amplitude, while a smaller value indicates a smoother behavior transition. This operation is equivalent to synchronously setting the initial conditions of each agent's policy function to the weighted starting point in the direction of consensus evolution, ensuring consistent behavior orientation.
[0088] The system also collaboratively reconstructs the resource preferences of the agents. First, sort all resource behavior parameters according to the resource selection frequency, which represents the cumulative proportion of times each resource type is used by all agents within N control cycles. Identify the main resource set and the marginal resource set based on this sorting result. Then, based on the historical behavior data of the group, calculate the co-occurrence probability of each resource type during the same policy evolution stage, that is, the joint probability value of a certain resource co-occurring with other resources in the agent's behavior. According to the co-occurrence probability matrix, construct the preferred resource interval between resource types, that is, recommend a set of resource combination areas with high co-occurrence and low conflict for each policy type. Use the above resource interval as the mapping reference structure to uniformly map and reset the resource behavior parameters of the agents, making their resource selection logic more inclined to enter the preferred resource set and avoiding the intensification of system resource competition caused by individual preferences. In the new control cycle, all agents uniformly apply this behavior-resource synchronization mapping structure, that is, using the convergence direction vector as the policy direction constraint and the preferred resource interval as the resource selection boundary, so as to achieve the synchronous convergence of policy evolution and resource behavior.
[0089] Embodiment 2: A multi-agent system-based multi-dimensional resource dynamic game optimization control system, including:
[0090] A resource graph builder for constructing a multi-dimensional heterogeneous resource coupling graph. By modeling the collaborative and inhibitory relationships between resource types within the same control cycle, a graph structure with resource types as nodes and dynamic weights as edges is generated, and a tensor-form graph is output to depict the evolution of resource states;
[0091] A behavior index calculator for collecting the policy trajectories of all intelligent agents in multiple consecutive control cycles, calculating the group policy divergence index and the group resource coupling tension index obtained from the dynamic changes of the edge weights in the resource coupling graph, and generating a double-index indicator to characterize the overall behavior deviation degree and resource conflict state of the multi-agent system;
[0092] A game risk discriminator for receiving the double-index indicator and inputting it into a trained game optimization prediction model, outputting a policy behavior convergence risk score, and based on the comparison result between this score and the first threshold, determining whether to trigger subsequent analysis and adjustment processes;
[0093] A local evolution regulator is used to separately evaluate the policy trajectories and resource behaviors of all intelligent agents, calculate the individual policy deviation index, and based on the comparison result between this index and a second threshold, determine whether the proportion of over-standard agents is lower than a set proportion. If it is lower, only adjust the policy evolution direction and resource selection behavior of the over-standard agents. If it is not lower, perform a unified policy update operation for all agents;
[0094] An evolution feedback updater is used to update the edge weight parameters in the resource coupling map and the agent behavior trajectory data after the policy adjustment or unified update is completed, provide a new input basis for the next control cycle, and form an adaptive convergence feedback loop for policy evolution.
[0095] All the above formulas are dimensionless and take their numerical calculations. The formula is obtained by collecting a large amount of data for software simulation to get a formula closest to the real situation. The preset parameters in the formula are set by those skilled in the art according to the actual situation.
[0096] It should be understood that in various embodiments of the present application, the magnitudes of the serial numbers of the above processes do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0097] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0098] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0099] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in the present application, and all should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A multi-dimensional resource dynamic game optimization control method based on a multi-agent system, characterized in that Including the following steps: Step 1: Construct a multi-dimensional heterogeneous resource coupling graph. By modeling the cooperation and inhibition relationships between resource types within the same control period, a graph structure with resource types as nodes and dynamic weights as edges is formed, constituting a tensor-form graph for resource state evolution; Step 2: Collect the policy trajectories of all intelligent agents in multiple consecutive control periods, calculate the group policy divergence index, and the group resource coupling tension index based on the dynamic change of edge weights in the resource coupling graph, as a double-index indicator to characterize the deviation degree of the overall behavior of the multi-agent system and the resource conflict state; Step 3: Input the double-index indicator into the trained game optimization prediction model. The model outputs a policy behavior convergence risk score. If the score is lower than the first threshold, the policy remains unchanged; if the score is higher than the first threshold, enter the refinement analysis process in Step 4; Step 4: Individually evaluate the policy trajectories and resource behaviors of all intelligent agents, calculate the individual policy deviation index of each agent, and compare this index with the second threshold. If the proportion of agents exceeding the second threshold among all agents is lower than the set ratio, only adjust the policy evolution direction and resource selection behavior of the agents exceeding the standard, that is, the agents exceeding the second threshold; if the proportion is not lower than the set ratio, perform a unified policy update operation for all agents; Step 5: After the policy adjustment or unified update is completed, update the edge weight parameters and agent behavior trajectory data in the resource coupling graph, and use this as the basis for the next round of evaluation and optimization, forming a continuously iterative dynamic feedback loop to achieve the adaptive convergence of policy evolution and the global optimization of resource allocation behavior.
2. The multi-dimensional resource dynamic game optimization control method based on a multi-agent system according to claim 1, wherein, During the process of constructing the multi-dimensional heterogeneous resource coupling graph, the resource types are limited to at least include computing resources, storage resources, communication resources, and energy consumption resources. In each control period, collect the usage intensity data of various resources as the original observation values, and use the original observation values to calculate the cooperation degree function and inhibition degree function between resource types; The cooperation degree function is used to reflect the positive mutual influence degree between two resource types in actual tasks, and the inhibition degree function is used to reflect the negative restriction relationship generated between two resource types in the allocation behavior; Combine the calculation results of the cooperation degree function and the inhibition degree function into a composite weight value, and use this composite weight value as the dynamic weight of the edge to connect the corresponding resource type nodes, thereby constructing a graph structure with resource types as nodes and dynamic weights as edges, and applying a time series identifier to this graph structure to record the evolution changes of this structure in multiple consecutive control periods.
3. The multi-dimensional resource dynamic game optimization control method based on a multi-agent system according to claim 2, characterized in that During the process of constructing the multi-dimensional heterogeneous resource coupling graph, the calculation methods of the cooperation degree function and the inhibition degree function between resource types are based on the resource synchronous usage rate and conflict occupancy rate within the control period, specifically: In each control period, count the number of resource allocation events participated by each resource type in the behaviors of each intelligent agent, and calculate the joint usage frequency of different resource types being called simultaneously during the same time period. Divide this joint usage frequency by the geometric mean of the individual usage frequencies of each resource to obtain the initial value of the cooperation degree function; By comparing the number of allocation conflict events between different resource types, i.e., the number of times when two resource types cannot be allocated simultaneously due to quota restrictions in resource allocation decisions, and combining the weights of task failure events caused by agent behavior, the initial value of the inhibition function of one resource on another resource is calculated; The synergy function and the inhibition function are synthesized into the dynamic weight of the edge through normalization mapping to complete the update of the edge weights between resource types in the graph structure, which is used as the edge weight parameter for constructing the tensor-form graph reflecting the evolution of resource states.
4. The multi-dimensional resource dynamic game optimization control method based on a multi-agent system according to claim 3, characterized in that, In the process of calculating the group strategy divergence index, first, the policy trajectories of all intelligent agents in multiple consecutive control cycles are encoded in a time series, and each policy trajectory is represented as a set of vectors with a fixed length. In the same control cycle, the mean Euclidean distance between the policy vector of each agent and the policy vectors of all other agents is calculated, and the ratio of this mean to the standard deviation of the change of the agent's own policy vector is used to obtain the divergence score of the agent; then, the divergence scores of all agents are normalized and then aggregated by mean to obtain the group strategy divergence index, which is used to measure the deviation amplitude, behavioral disagreement degree, and the consistency strength of the evolution trend among all intelligent agents in terms of policy behavior.
5. The multi-dimensional resource dynamic game optimization control method based on a multi-agent system according to claim 4, characterized in that In the process of calculating the group resource coupling tension index based on the dynamic change of edge weights in the resource coupling graph, first, the weight change values of all edges are extracted in two adjacent control cycles, and the tension strength of each edge in this change process is calculated, and its calculation method is the square of the weight difference between the current cycle and the previous cycle of this edge multiplied by the initial weight factor; then, with each resource type as the center, the tension strength values of all its connected edges are summarized to calculate the local tension response value of this resource type; then, the weighted average of the local tension response values of all resource types is taken to generate the group resource coupling tension index, which is used to measure the degree of structural tension and the resource coordination stability at the system level during the evolution of resource states in the current control cycle.
6. The multi-dimensional resource dynamic game optimization control method based on a multi-agent system according to claim 5, wherein, In the process of inputting the double-index indicator into the trained game optimization prediction model and outputting the policy behavior convergence risk score by this model, first, the group strategy divergence index and the group resource coupling tension index are normalized and mapped, and a double-index input vector is formed according to the preset dimension. This input vector is used as the expression form of the current multi-agent system behavior state and is input into the game optimization prediction model; This game optimization prediction model is trained based on the historical evolution data in the control cycle and adopts a nonlinear prediction structure integrated by a feedforward neural network and a graph correlation reasoning structure. By fitting and predicting the position relationship of the input vector in the historical behavior mapping space, the convergence risk score of the current behavior state in the process of policy evolution is output.
7. The multi-dimensional resource dynamic game optimization control method based on a multi-agent system according to claim 6, characterized in that, In the process of separately evaluating the policy trajectories and resource behaviors of all intelligent agents and calculating the individual policy deviation index of each agent, first, the policy trajectories of each agent in multiple consecutive control cycles are selected, and the sequence of policy behavior vectors of each agent in each control cycle is constructed; After normalizing the vector sequence, obtain the change amplitude of the policy behavior vectors between adjacent control cycles, and use the average change amplitude of the agent as its internal variability of the policy trajectory; Select the policy behavior vector sequences of other agents among all intelligent agents, and construct the average vector difference of the policy of this agent relative to all the other agents as the external difference of its behavior deviation degree; Synthesize the internal variability of the policy trajectory and the external difference of the behavior deviation degree into a single index in the form of a weighted linear combination, and smooth this index within a fixed window, and finally generate the individual policy deviation index of this agent in the current stage.
8. The multi-dimensional resource dynamic game optimization control method based on a multi-agent system according to claim 7, characterized in that When only adjusting the policy evolution direction and resource selection behavior of the over-standard agents, that is, the agents exceeding the second threshold, based on the difference direction between the policy behavior vector of the over-standard agent in the current control cycle and the mean vector of the policy behavior of all intelligent agents, set a behavior direction correction factor, which is the negative of the cosine value of the angle between the difference vector and the unit convergence vector multiplied by the non-zero proportionality coefficient γ, and use this correction factor in the policy generation vector superposition term of the next control cycle of the over-standard agent to adjust the direction of the policy behavior evolution path; Collect the resource set used by the over-standard agent in the current control cycle, and form a deviation resource set with the resource with the largest difference in the resource usage frequency ranking of all current agents. Set the resource allocation priority compression factor η of the over-standard agent in the next cycle as the initial resource selection probability multiplied by [1 - fd / (ft + C)], where fd is the resource frequency used by the over-standard agent, ft is the average resource usage frequency of the group, and C is a preset non-zero constant.
9. The multi-dimensional resource dynamic game optimization control method based on a multi-agent system according to claim 8, wherein, If the proportion of agents exceeding the second threshold in all agents is not less than the set proportion, then perform a unified policy update operation for all agents, including extracting the mean vector of the policy trajectories of all agents in the past N control cycles as the reference behavior model, obtaining the average deviation degree of all agent behavior vectors in this model, and using the policy behavior vector with the smallest deviation degree as the convergence direction vector, resetting the initial value of the policy generation parameter of all agents as the convergence direction vector multiplied by the initialization scale factor, and sorting all resource behavior parameters according to the resource selection frequency, mapping them to the preferred resource interval constructed based on the group resource co-occurrence probability and uniformly applying them in the new control cycle.
10. A multi-dimensional resource dynamic game optimization control system based on a multi-agent system, based on the multi-dimensional resource dynamic game optimization control method according to any one of claims 1-9, characterized in that, Including: A resource map builder, which is used to build a multi-dimensional heterogeneous resource coupling map. By modeling the cooperation relationship and inhibition relationship between resource types in the same control cycle, generate a graph structure with resource types as nodes and dynamic weights as edges, and output a tensor-form map to depict the evolution of resource states; A behavior index calculator, which is used to collect the policy trajectories of all intelligent agents in multiple consecutive control cycles, calculate the group policy divergence index and the group resource coupling tension index obtained from the dynamic change of the edge weights in the resource coupling map, and generate a double-index indicator to characterize the overall behavior deviation degree and resource conflict state of the multi-agent system; A game risk discriminator, which is used to receive double exponential indicators and input them into a trained game optimization prediction model, output a policy behavior convergence risk score, and judge whether to trigger subsequent analysis and adjustment processes based on the comparison result between the score and the first threshold; A local evolution regulator, which is used to separately evaluate the policy trajectories and resource behaviors of all intelligent agents, calculate the individual policy deviation index, and judge whether the proportion of over-standard agents is lower than the set proportion based on the comparison result between the index and the second threshold. If it is lower, only adjust the policy evolution direction and resource selection behavior of over-standard agents. If it is not lower, perform a unified policy update operation for all agents; An evolution feedback updater, which is used to update the edge weight parameters in the resource coupling map and the agent behavior trajectory data after the policy adjustment or unified update, provide a new input basis for the next control cycle, and form an adaptive convergence feedback loop for policy evolution.
Citation Information
Patent Citations
Network risk analysis method and system based on multilevel game model
CN119254483A
Game behavior dynamic evolution and strategy deduction optimization method and system in space field
CN119539090A
Dynamic optimization method and system based on calculation engine model driving operator chain
CN119739534A
Distributed industrial energy operation optimization platform automatically constructing intelligent models and algorithms
US11487273B1
Cited By
Marine mobile distributed detection network construction method
CN120528949A
Efficient agent pool management method based on dynamic selection and intelligent optimization
CN120750935A
Multi-task collaborative incremental webpage data acquisition method and system
CN120780428A