An Operational Optimization Algorithm Development System under Intelligent Games
By developing an operation optimization algorithm development system under intelligent game in an intelligent game environment, the problems of low collaboration efficiency, poor decision quality and insufficient adaptability of multi-agent systems in large-scale, high-dimensional, sparse rewards and dynamic changes are solved, and the effects of efficient exploration, low overhead collaboration and strategic robustness are achieved.
Patent Information
- Application Number
- CN202510415257.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-04-03
AI Technical Summary
In the intelligent game environment, multi-agent systems face the problems of low collaboration efficiency, poor decision-making quality and insufficient adaptability, especially in a large-scale, high-dimensional, sparse rewards and dynamic changes.
Develop a system for operation optimization algorithm development under intelligent game, including multi-level reward shaping module, pheromone indirect communication network module, agent role adaptive differentiation module, curiosity-driven exploration processing module, and Nash equilibrium and multi-agent value decomposition module to improve the collaboration efficiency and decision-making quality of multi-agents.
Through this system, multiple agents have efficient exploration capabilities in sparse reward environments, realize low-overhead collaborative optimization, improve the robustness of the strategy and the efficiency of computing resource utilization, and significantly optimize the generalization ability and performance performance in different game environments.
Smart Images

Figure CN119918572B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of artificial intelligence and operational research optimization, and more specifically, it relates to a system for developing an operational research optimization algorithm under intelligent games. Background Art
[0002] With the rapid development of artificial intelligence technology, the problem of collaborative decision-making and optimization of multi-agent systems in complex environments has attracted increasing attention. In an intelligent game environment, there are both cooperation and competition relationships among multi-agents, and their operational research optimization faces many challenges: First, traditional operational research optimization methods are difficult to handle scenarios with a large number of agents and a high-dimensional state-action space; Second, in a sparse reward signal environment, the exploration efficiency of agents is low, and it is difficult to discover valuable strategies; Third, the communication cost between agents is high, and it is difficult to achieve efficient cooperation; Finally, the dynamic changes in the game environment require the optimization strategy to have strong adaptability and robustness.
[0003] Currently, the methods for solving the above problems mainly include: multi-agent cooperation methods based on reinforcement learning, equilibrium solving methods based on game theory, and collaborative optimization methods based on swarm intelligence. However, these methods still have serious deficiencies in dealing with large-scale, high-dimensional, sparse reward, and dynamically changing intelligent game environments: Exploration efficiency problem: In a sparse reward environment, existing methods often rely on random exploration, resulting in a slow learning process and difficulty in discovering high-value strategies; Communication overhead problem: Most multi-agent cooperation methods rely on explicit communication, and as the number of agents increases, the communication overhead grows exponentially; Curse of dimensionality problem: As the number of agents and the dimensions of the state-action space increase, the policy search space expands exponentially, and existing methods are difficult to effectively handle; Cooperation-competition balance problem: In a mixed-motivation environment, it is difficult to balance the relationship between the individual goals and collective goals of agents; Generalization ability problem: The learned policies are usually over-specialized to a specific environment and are difficult to transfer and apply to new environments.
[0004] Therefore, there is an urgent need to develop an operational research optimization system that can effectively solve the above problems and improve the cooperation efficiency, decision-making quality, and adaptability of multi-agents in an intelligent game environment. Summary of the Invention
[0005] The present invention provides a system for developing an operational research optimization algorithm under intelligent games, which solves the technical problems of insufficient cooperation efficiency, decision-making quality, and adaptability of multi-agents in an intelligent game environment in related technologies.
[0006] The present invention provides a system for developing an operational research optimization algorithm under intelligent games, including:
[0007] Multi-level reward shaping module, which integrates intrinsic motivation drive and external environmental reward signals to generate a composite reward function, guiding multi-agent to efficiently explore and learn in a sparse reward environment;
[0008] Pheromone indirect communication network module, allowing multi-agent to share experience and knowledge without explicit communication, realizing low-overhead collaborative optimization;
[0009] Agent role adaptive differentiation module, enabling individuals in a multi-agent group to dynamically select specialized roles according to their own characteristics and environmental needs, forming a synergistic and complementary self-organization structure;
[0010] Curiosity-driven exploration processing module, guiding multi-agent to efficiently explore and exploit in a complex game space by analogy with the potential energy and entropy increase principles in physical systems;
[0011] Nash equilibrium and multi-agent value decomposition module, constructing an operational optimization framework that can coordinate the cooperation and competition relationships of multi-agent in an intelligent game environment.
[0012] Furthermore, the multi-level reward shaping module specifically includes:
[0013] Three-source reward model unit, classifying rewards into three types: external environmental rewards, intrinsic curiosity rewards, and prediction error rewards;
[0014] State prediction model unit, for each agent, predicting the next state and calculating the prediction error with the true next state;
[0015] State rarity evaluation unit, modeling the access frequency of states, calculating state rarity, and forming prediction error rewards;
[0016] Adaptive weight adjustment unit, generating a dynamically adjusted composite reward function, where the weight coefficients are dynamically adjusted with the training process.
[0017] Furthermore, the pheromone indirect communication network module specifically includes:
[0018] Game state action space graph structure representation unit, where the vertex set represents game states and the edge set represents possible state transition actions;
[0019] Pheromone concentration correlation unit, associating pheromone concentration with each edge, and initializing the pheromone concentration of all edges to a unified small positive number;
[0020] Pheromone update rule unit, after each time step, updating the pheromone concentration according to the pheromone evaporation coefficient and the pheromone increment released by the agent;
[0021] The pheromone increment calculation unit associates the pheromone increment with the reward obtained by the agent and the action quality;
[0022] The pheromone-guided action selection unit constructs a pheromone-based action selection probability distribution to guide the agent in making decisions.
[0023] Furthermore, the agent role adaptive differentiation module specifically includes:
[0024] The functional role type definition unit includes an explorer role, a developer role, a coordinator role, and a defender role;
[0025] The agent ability index vector construction unit characterizes the potential of the agent in various roles;
[0026] The role fitness evaluation unit evaluates the role fitness based on the current game environment state, the agent ability index vector, and the historical performance record;
[0027] The role assignment vector calculation unit calculates the role assignment vector based on the multi-objective optimization principle;
[0028] The role conversion dynamic adjustment unit includes a role switching inertia factor, a group diversity evaluation function, and an agent role complementarity evaluation;
[0029] The role behavior strategy adaptation unit adjusts the agent's behavior strategy according to the assigned role.
[0030] Furthermore, the curiosity-driven exploration processing module specifically includes:
[0031] The knowledge potential field design unit characterizes the knowledge distribution in the game state space;
[0032] The entropy-driven exploration mechanism unit defines the state entropy to measure the uncertainty of the state and introduces an entropy regularization term to encourage strategy exploration;
[0033] The stochastic resonance exploration enhancement unit introduces an appropriate amount of noise in the agent's decision-making process;
[0034] The state curiosity degree evaluation unit combines the knowledge potential and the access frequency;
[0035] The exploration-exploitation dynamic balance unit realizes the dynamic balance between the exploration strategy and the exploitation strategy;
[0036] The memory-guided exploration unit uses historical experience to guide future exploration.
[0037] Furthermore, the Nash equilibrium and multi-agent value decomposition module specifically includes:
[0038] The state-action value function construction unit characterizes the individual value and the collective value respectively;
[0039] The multi-agent value decomposition network unit includes a state representation network, an individual advantage network, and a hybrid value decomposition network;
[0040] The Nash equilibrium action selection unit solves the game problem in each state;
[0041] The contribution degree decomposition unit evaluates the contribution of each agent to the overall goal;
[0042] The cooperation-competition balance adjustment unit realizes dynamic cooperation-competition balance adjustment;
[0043] The multi-agent collaborative learning unit integrates the above components to achieve collaborative learning.
[0044] Furthermore, the composite reward function in the multi-level reward shaping module is defined as:
[0045] ;
[0046] where is the total reward of agent , represents the sparse reward signal directly obtained by agent from the environment, represents the intrinsic exploration motivation of the -th agent for the unexplored area, represents the error reward of the -th agent for the dynamic prediction of the environment, and respectively represent the adaptive weight coefficients of the curiosity reward and the prediction error reward.
[0047] Furthermore, the pheromone update rule in the pheromone indirect communication network module is defined as:
[0048] ;
[0049] where represents the pheromone concentration of edge at time step ; represents the pheromone concentration of edge at time step , is the pheromone evaporation coefficient, , controlling the forgetting rate of historical information; is the total number of agents; is the -th agent at time step releasing the pheromone increment on edge ; k represents the agent number, and its value range is from 1 to 。
[0050] Furthermore, the state curiosity evaluation model in the curiosity-driven exploration processing module is defined as:
[0051] ;
[0052] where represents the curiosity degree of state , represents the knowledge potential energy of state , represents the novelty degree of state , calculated by measuring the distance from the state in the memory bank; represents the novelty weight.
[0053] A computer-readable storage medium is used to store computer-readable instructions, which can run an operation research optimization algorithm development system for intelligent games as described above when read by a computer.
[0054] The beneficial effects of the present invention are as follows: The system has efficient exploration ability in a sparse reward environment through a multi-level reward shaping system and a curiosity-driven exploration process; the cooperation ability of the multi-agent system is greatly improved through a pheromone indirect communication network and an agent role adaptive differentiation mechanism; the robustness of the strategy is significantly improved based on the Nash equilibrium and multi-agent value decomposition optimization framework; the utilization of computing resources is significantly optimized through agent role adaptive differentiation and value decomposition technology;
[0055] The system demonstrates strong generalization ability among different game environments; after zero-shot transfer in an unseen game environment, the performance retention rate reaches 76% of the original environment, exceeding the existing methods by about 30%; in cross-domain generalization tests, only about 25% of the fine-tuning data is required to achieve performance similar to that of a dedicated algorithm; the strategy abstraction ability is improved by about 63%, enabling the learned strategy to be directly applied in a variety of similar environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] Figure 1 It is a module diagram of an operation research optimization algorithm development system for intelligent games of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0057] The subject matter described herein will now be discussed with reference to example embodiments. It should be understood that discussing these embodiments is only to enable those skilled in the art to better understand and thus implement the subject matter described herein, and that changes can be made to the functions and arrangements of the elements discussed without departing from the scope of protection of the content of this specification. Each example can omit, substitute, or add various processes or components as needed. Additionally, the features described in some examples can also be combined in other examples.
[0058] In at least one embodiment of the present invention, an operation research optimization algorithm development system under intelligent gaming is disclosed, as Figure 1 shown, including:
[0059] A multi-level reward shaping module that integrates intrinsic motivation drive and external environment reward signals to generate a composite reward function, guiding multi-agents to efficiently explore and learn in a sparse reward environment:
[0060] Specifically as follows:
[0061] Step 1.1, establish a three-source reward model, and divide the rewards into external environment rewards , intrinsic curiosity rewards , and prediction error rewards into three types, where:
[0062] represents the sparse reward signal directly obtained from the environment;
[0063] represents the intrinsic exploration motivation of the agent for the unexplored area;
[0064] represents the error reward of the agent's dynamic prediction of the environment.
[0065] Step 1.2, for each agent , construct a state prediction model , which receives the current state and the action as inputs, predicts the next state , and calculates the prediction error with the true next state :
[0066] ;
[0067] where represents the prediction error of the th agent at time step , represents the square of the Euclidean distance, used to measure the difference between the predicted state and the true state; represents the current time step, represents the agent number;
[0068] Construct a curiosity reward based on this prediction error:
[0069] ;
[0070] where represents the intrinsic exploration motivation of the -th agent for the unexplored area, is the curiosity intensity parameter, which controls the intensity of the curiosity reward.
[0071] Step 1.3, design a state rarity evaluator to model the access frequency of the state and calculate the state rarity to form a prediction error reward:
[0072] ;
[0073] where represents the error reward of the -th agent's prediction of the environmental dynamics, is the prediction reward intensity parameter, which controls the intensity of the prediction error reward, represents the rarity of the state , and the higher the rarity, the larger the value.
[0074] Step 1.4, construct an adaptive weight adjustment module to generate a dynamically adjusted composite reward function:
[0075] ;
[0076] where is the total reward of agent , represents the sparse reward signal directly obtained by agent from the environment, represents the current time step, represents the agent number;
[0077] is the adaptive weight coefficient of the curiosity reward, which is dynamically adjusted with the training process, indicating that as the training progresses, the curiosity reward gradually decreases:
[0078] ;
[0079] where is the initial curiosity weight, is the curiosity weight decay rate, represents the exponential function, represents the current training step number;
[0080] The adaptive weight coefficient for the prediction error reward is dynamically adjusted with the training process, indicating that as the model's prediction ability improves, the prediction error reward gradually increases:
[0081] ;
[0082] Among them, is the maximum weight of the prediction error reward, is the growth rate of the prediction error weight.
[0083] Through this multi-level reward shaping system, the problem of insufficient exploration motivation of agents in sparse reward environments is solved, and the balanced exploration of agents in the game space is promoted, avoiding premature convergence to local optimal solutions.
[0084] The pheromone indirect communication network module allows multiple agents to share experiences and knowledge without explicit communication, achieving low-overhead collaborative optimization:
[0085] Specifically as follows:
[0086] Step 2.1, construct a graph structure representation of the game state-action space:
[0087] ;
[0088] Among them represents the graph structure of the game state-action space, represents the game state, represents the possible state transition actions, and each edge corresponds to the transition from state to state of.
[0089] Step 2.2, associate the pheromone concentration with each edge , initialize the pheromone concentration of all edges to a unified small positive number , and establish a pheromone storage matrix , whose element represents the pheromone concentration value on edge .
[0090] Step 2.3, define the pheromone update rule. After each time step , update the pheromone concentration according to the following formula:
[0091] ;
[0092] Among them represents edge at time step The pheromone concentration; Indicates an edge At time step The pheromone concentration, Is the pheromone evaporation coefficient, , controlling the forgetting rate of historical information; Is the total number of agents; Is the th agent at time step On the edge The increment of pheromone released; k represents the agent number, and the value range is from 1 to .
[0093] Step 2.4, design a method for calculating the pheromone increment and associate it with the reward and action quality obtained by the agent:
[0094] ;
[0095] Where Represents the th agent at time step On the edge The increment of pheromone released, Is the pheromone release intensity coefficient, controlling the magnitude of pheromone release; Is the agent At time step The total reward obtained.
[0096] Step 2.5, construct a probability distribution of pheromone-guided action selection;
[0097] For the agent In the state When choosing to transfer to the state The probability Is calculated as follows:
[0098] ;
[0099] Where Is the set of all possible successor states of the state ; Is the heuristic evaluation value of the agent For the transfer ; Is the pheromone importance parameter, controlling the influence weight of pheromone concentration in decision-making; Is the heuristic information importance parameter, controlling the influence weight of heuristic information in decision-making; Represents a state in the set of all possible successor states of the state , traversing all possible successor states; Represents the state The set of all possible successor states; Represents the agent For the state The heuristic evaluation value, traversing all possible successor states; Represents the state At time step The pheromone concentration value, traversing all possible successor states.
[0100] Through this pheromone indirect communication network, the multi-agent system can achieve experience sharing under the condition of low communication overhead, accelerate the collective learning process, and improve the exploration efficiency of valuable areas in the game environment.
[0101] The agent role adaptive differentiation module enables individuals in the multi-agent group to dynamically select specialized roles according to their own characteristics and environmental needs, forming a self-organizing structure of coordination and complementarity:
[0102] Specifically as follows:
[0103] Step 3.1, define multiple functional role types, including but not limited to:
[0104] Explorer role: responsible for discovering new areas in the game space;
[0105] Exploiter role: responsible for deeply exploring known high-value areas;
[0106] Coordinator role: responsible for balancing exploration and exploitation in the group;
[0107] Defender role: responsible for coping with the adversarial strategies of opponents.
[0108] Step 3.2, construct the agent ability index vector , representing the agent The potential in various roles, including:
[0109] Exploration ability: , measuring the efficiency of the th agent to discover new states;
[0110] Development ability: , measuring the ability of the th agent to optimize known strategies;
[0111] Coordination ability: , measuring the ability of the th agent to balance group behavior;
[0112] Defense ability: , which measures the ability of the th agent to respond to adversarial strategies;
[0113] Step 3.3, design the role fitness evaluation function:
[0114] ;
[0115] Among them represents the fitness evaluation function of the th agent in the current game environment state , the agent ability index vector , and the historical performance record . represents the current game environment state of the th agent, and i represents the agent number; represents the agent ability index vector of the th agent; represents the historical performance record of the th agent.
[0116] Step 3.4, calculate the role assignment vector based on the multi-objective optimization principle:
[0117] ;
[0118] ;
[0119] Among them is the role assignment vector of the th agent; The function realizes the normalization of role probabilities; represents the probability of role ; represents the exponential function, traverses all possible roles.
[0120] Step 3.5, construct a dynamic adjustment mechanism for role conversion, including the following key components:
[0121] Role switching inertia factor , which prevents instability caused by frequent role switching;
[0122] Group diversity evaluation function to ensure the balance of role distribution in the group:
[0123] ;
[0124] Among them represents the group diversity evaluation function, , , represent the role assignment vectors of the 1st, 2nd, th agents respectively, where represents the total number of agents;
[0125] Agent Role Complementary Evaluation , which encourages the formation of complementary roles, where represents the role assignment vector of the th agent, represents the role assignment vector of the th agent, and both represent the agent numbers.
[0126] Step 3.6, implement the role behavior strategy adapter to adjust the behavior strategy of the agent according to the assigned role:
[0127] For the explorer role, increase the curiosity reward weight:
[0128] ;
[0129] where represents the curiosity reward weight of the explorer role at the th time step, represents the curiosity reward weight of the explorer role at the th time step, represents the adjustment factor of the curiosity reward weight of the explorer role, represents the probability of the explorer role of the th agent;
[0130] For the developer role, reduce the random exploration probability and increase the strategy certainty;
[0131] For the coordinator role, increase the sensitivity to the pheromone network:
[0132] ;
[0133] where represents the pheromone reward weight of the coordinator role at the th time step, represents the pheromone reward weight at the th time step;
[0134] For the defender role, increase the modeling accuracy and response speed of the opponent's behavior.
[0135] Through the agent role adaptive differentiation system, the multi-agent group forms a collaborative and complementary functional division of labor, improving the overall decision-making efficiency and robustness, and adapting to the diverse challenges in the complex game environment.
[0136] The curiosity-driven exploration processing module guides multi-agent to conduct efficient exploration and exploitation in the complex game space by analogy with the principles of potential energy and entropy increase in physical systems:
[0137] Specifically as follows:
[0138] Step 4.1, design a knowledge potential field to represent the knowledge distribution in the game state space:
[0139] ;
[0140] Where represents the knowledge potential of state , represents the posterior probability of state under the condition of the given current data set , represents the logarithmic function;
[0141] The knowledge potential field has the following characteristics: the potential energy in the fully explored area is low, and the potential energy in the unexplored area is high.
[0142] Step 4.2, construct an entropy-driven exploration mechanism to define the state entropy to measure the uncertainty of the state:
[0143] ;
[0144] Where represents the entropy of state , is the strategy of choosing action in state , is the set of all possible actions;
[0145] Introduce an entropy regularization term to encourage strategy exploration:
[0146] ;
[0147] Where represents the expected return of the strategy, represents the expected return of the strategy, represents the weight of the entropy regularization term, represents the entropy of the strategy, represents the th time step discount factor, represents the th time step expected total return.
[0148] Step 4.3, implement an exploration enhancement mechanism based on stochastic resonance, and introduce appropriate noise in the agent decision-making process:
[0149] The action selection probability is updated to:
[0150] ;
[0151] where represents the action selection probability at the time step, represents the action selection probability at the time step, represents the noise, represents the normal distribution, represents the state noise variance:
[0152] ;
[0153] where represents the initial value of the noise variance, represents the knowledge potential of the state ;
[0154] By introducing an appropriate amount of noise in the agent's decision-making process, the sensitivity of the system to key decision points is enhanced; this mechanism enables the agent to more easily discover and respond to important state transition points during exploration, avoiding being trapped in local optimal solutions; when the system approaches a key decision point, an appropriate amount of noise can instead improve the signal detection ability, similar to the stochastic resonance phenomenon in physical systems, thereby enhancing the exploration efficiency and decision-making quality of the agent in complex game environments.
[0155] Step 4.4, construct a state curiosity evaluation model , combining knowledge potential and access frequency:
[0156] ;
[0157] where represents the curiosity of the state , represents the knowledge potential of the state , represents the novelty of the state , calculated by measuring the distance from the states in the memory bank; represents the novelty weight.
[0158] Step 4.5, implement a dynamic balance mechanism between exploration and exploitation strategies:
[0159] Introduce a temperature parameter to adjust the exploration-exploitation balance:
[0160] ;
[0161] where Denotes the action selection probability at the time step, denotes the state under which the action is taken, denotes the state under which the action is taken, denotes the temperature parameter, denotes the set of all possible actions, and both belong to , denotes the exponential function;
[0162] The temperature parameter is dynamically adjusted with the training process:
[0163] ;
[0164] where denotes the temperature parameter at the time step, denotes the initial value of the temperature parameter, denotes the decay factor of the temperature parameter, denotes the time step, denotes the minimum value of the temperature parameter;
[0165] Design a trigger-based temperature reset mechanism to increase the temperature when exploration stagnates:
[0166] ;
[0167] where denotes the temperature parameter at the time step, denotes the reset value of the temperature parameter, denotes the change in the value function, is the stagnation determination threshold, denotes the step threshold for continuous stagnation.
[0168] Step 4.6, construct a memory-guided exploration mechanism to use historical experience to guide future exploration:
[0169] Maintain a memory bank representing the historical exploration trajectory:
[0170] ;
[0171] where is the memory bank capacity, , , , respectively denote the The state, action, reward, and next state of a record belong to ;
[0172] Calculate state similarity and novelty based on the memory bank:
[0173] ;
[0174] where represents the similarity between state and the states in the memory bank ; represents the distance between state and the next state ; represents the minimum value;
[0175] Construct a memory-guided exploration reward:
[0176] ;
[0177] where represents the memory-guided exploration reward of state ; represents the memory-guided exploration reward weight.
[0178] Through the curiosity-driven exploration process, the multi-agent system can efficiently explore in the complex game space, balance the relationship between local exploitation and global search, and improve the search efficiency and solution quality of the optimization algorithm.
[0179] Nash equilibrium and multi-agent value decomposition module to construct an operational optimization framework that can coordinate the cooperation and competition relationships of multi-agents in the intelligent game environment:
[0180] Specifically as follows:
[0181] Step 5.1, construct a state-action value function from the perspective of game theory to represent individual value and collective value respectively:
[0182] Individual value function : Evaluate the expected return of agent taking action in state ; belong to ; represents the action space of agent ;
[0183] Joint value function : Evaluate the joint action taken by all agents in state :
[0184] ;
[0185] wherein represents the joint action of all agents, , , respectively represent the actions of the 1st, 2nd, th agent, represents the total number of agents, belongs to , that is:
[0186] ;
[0187] wherein represents the joint action space of all agents, , , respectively represent the individual action spaces of the 1st, 2nd, th agent, represents the total number of agents;
[0188] Design a mixed value function, combining individual value and joint action influence:
[0189] ;
[0190] wherein represents the mixed value function of agent taking action in state , represents the joint action of all other agents except agent .
[0191] Step 5.2, implement a multi-agent value decomposition network based on Nash equilibrium, including the following key components:
[0192] State representation network : Extract high-level feature representations from the original game state , wherein represents the network parameters;
[0193] Individual advantage network : Evaluate the advantage of the action of agent relative to the average level;
[0194] Mixed value decomposition network:
[0195] ;
[0196] wherein , , respectively represent the individual value functions of the 1st, 2nd, and th agents, represents the total number of agents, represents the mixed value decomposition network, represents the joint value function of all agents, , , respectively represent the 1st, 2nd, and th actions of the agents.
[0197] Step 5.3, design an action selection mechanism based on Nash equilibrium, and solve the following game problem at each state :
[0198] For each agent , solve the optimal response strategy:
[0199] ;
[0200] where represents the optimal action of agent , represents the mixed value function of agent when taking action at state , represents the joint action of all other agents except agent , represents the maximization symbol;
[0201] Iteratively solve the process until it converges to the Nash equilibrium point:
[0202] ;
[0203] where represents the joint optimal action of all agents, , , respectively represent the 1st, 2nd, and th optimal actions of the agents, represents the total number of agents;
[0204] Introduce computational efficiency optimization techniques and use approximate Nash equilibrium to solve.
[0205] Step 5.4, construct a contribution decomposition mechanism to evaluate the contribution of each agent to the overall goal:
[0206] Design a contribution evaluation function inspired by the Shapley value:
[0207] ;
[0208] where represents the contribution degree of agent taking a joint action in state . represents a subset of agents, and the summation is performed over all subsets that do not contain agent . represents the set of agents, represents the total number of agents, represents the subset 's joint value function;
[0209] Adjust the reward distribution based on the contribution degree evaluation:
[0210] ;
[0211] where represents the adjusted reward of agent , represents the original reward of agent , represents the contribution degree reward weight, represents the contribution degree of agent .
[0212] Step 5.5, implement a dynamic cooperation - competition balance adjustment mechanism:
[0213] Introduce a cooperation - competition coefficient to dynamically adjust the weights of individual goals and collective goals;
[0214] Hybrid objective function:
[0215] ;
[0216] where represents the hybrid objective function of agent , represents the individual objective function of agent , represents the collective objective function, represents the cooperation - competition coefficient;
[0217] Adaptive adjustment based on environmental feedback value:
[0218] ;
[0219] where represents the cooperation - competition coefficient at the th time step, represents the cooperation-competition coefficient at the th time step, represents the cooperation-competition coefficient adjustment factor, represents the change in the collective value function at the th time step.
[0220] Step 5.6, design a multi-agent collaborative learning algorithm to integrate the above components:
[0221] Construct a centralized training and distributed execution architecture;
[0222] Design agent interaction modeling based on the attention mechanism:
[0223] ;
[0224] where represents the interaction modeling result of agent , is the attention weight of agent to agent , represents the state of agent , represents the action of agent , is the interaction modeling function with parameter ;
[0225] Implement an experience replay pool isolation and sharing mechanism to balance individual experience and group experience;
[0226] By integrating the Nash equilibrium and multi-agent value decomposition techniques, this algorithm can coordinate the cooperation and competition relationships among multiple agents in a complex game environment, find the game equilibrium solution, optimize the overall operation objective, and effectively address the decision optimization challenges in the high-dimensional game space.
[0227] A computer-readable storage medium for storing computer-readable instructions that, when read by a computer, can run an operation optimization algorithm development system for intelligent gaming as described above.
[0228] Here, the present invention provides an implementation example: complex supply chain network optimization;
[0229] In this example, the above method is applied to an optimization problem of a complex supply chain network with multiple participants.
[0230] This scenario has the following characteristics:
[0231] The network contains 50 nodes, representing raw material suppliers, manufacturers, distributors, retailers, and consumers respectively;
[0232] Each node is controlled by an agent, with its own objective function and decision-making power;
[0233] There are complex cooperation and competition relationships among the nodes, forming a multi-level game structure;
[0234] The system faces uncertain factors such as random demand fluctuations, supply interruption risks, and market competition;
[0235] The overall goal is to optimize the global supply chain efficiency and robustness while ensuring the interests of all parties involved.
[0236] First, construct a graph structure representation of the supply chain network, where the nodes are the parties involved and the edges are the supply-demand relationships;
[0237] Define the state space , which includes key indicators such as inventory levels, order status, and production capacity;
[0238] Define the action space , which includes decision variables such as pricing, production volume, order quantity, and inventory strategies;
[0239] Set external reward signals to reflect business metrics such as profitability, cost, and service level.
[0240] Subsequently, construct a state prediction model to predict supply-demand changes and price fluctuations;
[0241] Design a three-source reward model, where the external reward is the profit metric and the intrinsic reward promotes exploration of market changes;
[0242] Implement adaptive weight adjustment. In the initial stage , , and then gradually shift to paying more attention to actual business metrics.
[0243] Then, establish a pheromone propagation channel on the supply chain network, corresponding to the order flow, information flow, and capital flow;
[0244] Set the pheromone evaporation coefficient , enabling the system to gradually forget outdated information;
[0245] The pheromone increment is positively correlated with the transaction success rate and profit margin, guiding multi-agents to form an efficient cooperation mode.
[0246] After that, define specific roles in the supply chain: demand forecaster, risk manager, inventory optimizer, price strategist;
[0247] Dynamically allocate roles based on historical performance and current market conditions;
[0248] Increase the proportion of risk managers when market fluctuations intensify and increase the proportion of price strategists during stable periods;
[0249] Furthermore, construct a knowledge potential energy field and assign high potential energy values to emerging markets and innovative supply models;
[0250] Implement strategy exploration based on stochastic resonance and increase the exploration intensity at market turning points;
[0251] Set the temperature parameter , , and gradually reduce the randomness as the training progresses.
[0252] Finally, construct a hybrid value function to balance the maximization of individual profits and the overall efficiency of the supply chain;
[0253] Implement a contribution decomposition mechanism to reasonably allocate the additional benefits brought by collaboration;
[0254] Design a dynamic cooperation-competition balance mechanism to adjust the value under different market conditions.
[0255] After applying this embodiment, compared with traditional supply chain optimization methods, the system has achieved performance improvements:
[0256] The global inventory turnover rate has increased by 38%, and the inventory cost has been reduced by 26%; the order fulfillment rate has increased by 32%, and the average delivery time has been shortened by 41%; the overall profit of the supply chain has increased by 23%, and at the same time, the profit distribution among all participants is more balanced; the cost caused by demand forecasting deviation has been reduced by 46%.
[0257] The system robustness has been enhanced: for the simulated supply interruption event, the recovery time has been shortened by 58%; the adaptability to random demand fluctuations has been improved by 65%; when new participants are introduced, the network reorganization efficiency has been increased by 73%; the speed of the strategy adapting to market changes has been accelerated by 3.2 times.
[0258] The computing performance has been optimized: the decision-making time has been reduced by 69% to meet the real-time response requirements; when the system is expanded to 100 participants, the performance only drops by 12%; the communication overhead has been reduced by 78% to reduce the infrastructure requirements; the energy consumption has been reduced by 43% to support the concept of green computing.
[0259] This application example fully demonstrates the practical value of the bio-inspired self-organizing multi-agent operations research optimization algorithm in complex intelligent game environments. Its key technologies can solve complex optimization problems in the real world and provide innovative technical solutions for fields such as supply chain management, resource scheduling, and market games.
[0260] The embodiments of the present invention have been described above. However, these embodiments are not limited to the specific implementation manners described above. The above specific implementation manners are merely illustrative rather than restrictive. Under the inspiration of this embodiment, those of ordinary skill in the art can also make more equivalent embodiments in various forms, all of which fall within the protection scope of this embodiment.
Claims
1. A system for developing operations optimization algorithms under intelligent game, characterized in that: include: Multi-level reward shaping module, including: The three-source reward model unit divides rewards into three types: external environmental rewards, intrinsic curiosity rewards, and prediction error rewards; The state prediction model unit predicts the next state for each agent and calculates the prediction error with the actual next state; The state rarity evaluation unit models the access frequency of the state, calculates the state rarity, and forms a prediction error reward; An adaptive weight adjustment unit that generates a dynamically adjusted compound reward function, where the weight coefficients are dynamically adjusted as the training progresses; Pheromone indirect communication network module, including: The game state action space graph structure represents the unit, where the vertex set represents the game state and the edge set represents the possible state transition actions; The pheromone concentration association unit associates the pheromone concentration for each edge and initializes the pheromone concentration of all edges to a uniform small positive number; The pheromone update rule unit updates the pheromone concentration after each time step according to the pheromone evaporation coefficient and the pheromone increment released by the agent; The pheromone increment calculation unit associates the pheromone increment with the reward and action quality obtained by the agent; The pheromone-guided action selection unit constructs the action selection probability distribution based on pheromones and guides the agent to make decisions; The agent role adaptive differentiation module includes: Functional role type definition units, including explorer role, developer role, coordinator role and defender role; The agent capability index vector building unit characterizes the agent's potential in various roles; The role fitness evaluation unit evaluates the role fitness based on the current game environment state, agent capability index vector and historical performance records; A role allocation vector calculation unit, which calculates the role allocation vector based on a multi-objective optimization principle; The role switching dynamic adjustment unit includes the role switching inertia factor, the group diversity evaluation function and the agent role complementarity evaluation; The role behavior strategy adaptation unit adjusts the agent's behavior strategy according to the assigned role; Curiosity-driven exploration processing module, including: The knowledge potential field design unit characterizes the knowledge distribution in the game state space; Entropy-driven exploration mechanism unit, which defines state entropy to measure state uncertainty and introduces entropy regularization term to encourage strategic exploration; Stochastic resonance exploration enhancement unit, which introduces appropriate amount of noise into the agent's decision-making process; The state curiosity assessment unit combines knowledge potential and access frequency; Exploration and development dynamic balance unit to achieve a dynamic balance between exploration strategy and development strategy; Memory-guided Exploration Unit, using historical experience to guide future exploration; Nash equilibrium and multi-agent value decomposition module, including: The state-action-value function constructs units that represent individual and collective values respectively; Multi-agent value decomposition network unit, including state representation network, individual advantage network and hybrid value decomposition network; Nash equilibrium action selection unit, solving the game problem in each state; Contribution decomposition unit, which evaluates the contribution of each agent to the overall goal; Cooperation and competition balance adjustment unit, to achieve dynamic cooperation and competition balance adjustment; The multi-agent collaborative learning unit integrates the state-action value function construction unit, the multi-agent value decomposition network unit, the Nash equilibrium action selection unit, the contribution decomposition unit and the cooperative competition balance adjustment unit to achieve collaborative learning.
2. According to the intelligent game operation optimization algorithm development system of claim 1, it is characterized in that: The composite reward function in the multi-level reward shaping module is defined as: ; in For intelligent agents Total rewards, Representing an Agent A sparse reward signal obtained directly from the environment, Indicates The agent’s intrinsic curiosity reward for unexplored areas, Indicates The error reward of each agent's prediction of the environment dynamics, and Represent the adaptive weight coefficients of curiosity reward and prediction error reward respectively.
3. The system for developing an operational optimization algorithm under intelligent game according to claim 1, characterized in that: The pheromone update rule in the pheromone indirect communication network module is defined as: ; in Represents edge At time step pheromone concentration; Represents edge At time step The pheromone concentration, is the pheromone evaporation coefficient, , control the forgetting rate of historical information; is the total number of agents; It is Agents at time step On the side Increased pheromone release; Indicates the agent number, ranging from 1 to .
4. The system for developing an operational optimization algorithm under intelligent game according to claim 1, characterized in that: The state curiosity evaluation model in the curiosity-driven exploration processing module is defined as: ; in Indicates status Curiosity Indicates status The knowledge potential, Indicates status The novelty of is calculated by measuring the distance from the state in the memory bank; represents the novelty weight.
5. A computer-readable storage medium, characterized in that: It is used to store computer-readable instructions, and when the computer-readable instructions are read by a computer, it can run an operations optimization algorithm development system under intelligent gaming as described in any one of claims 1-4.
Citation Information
Patent Citations
Reinforcement learning macro-module layout method based on curiosity driving
CN118536458A
Game strategy optimization method based on reinforcement learning
CN118940819A
Cited By
Multi-agent game construction method and multi-agent system
CN120560133A
A method for constructing a multi-agent game and a multi-agent system
CN120560133B