Intelligent Reinforcement Learning Training Method Based on Self-Game

By combining action space splitting, dynamic mask filtering, priority fusion and coordinated attention mechanism in the heterogeneous agent collaborative decision-making system, the problems of inefficient training, sparse gradients and decision distortion in traditional methods are solved, and more efficient, stable and adaptable collaborative decision-making is achieved.

CN119783759BActive Publication Date: 2025-06-17NO 15 INST OF CHINA ELECTRONICS TECH GRP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510278940.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-06-17
Estimated Expiration
2045-03-11

AI Technical Summary

Technical Problem

In the heterogeneous agent collaborative decision-making system, traditional methods have low training efficiency, sparse gradients, and difficulty in capturing the tactical behavior characteristics of specific types of agents due to high-dimensional joint action space, and illegal action filtering has problems such as constraint lag and deviation in exploration direction.

Method used

Through the typed splitting and shared feature extraction of action space, the scale of network parameters is reduced; the application of dynamic binary masks realizes real-time illegal action filtering; the priority score matrix is ​​built to optimize action selection strategies based on hierarchical importance sampling; the cross-attention network and co-gain coefficient are introduced to capture the tactical correlation between agents.

Benefits of technology

It significantly improves the performance and efficiency of the collaborative decision-making system of heterogeneous agents, alleviates the problems of inefficient training efficiency and gradient sparseness, ensures the compliance and efficiency of the decision-making process, improves the stability and environmental adaptability of the decision-making, and enhances the overall efficiency and robustness of collaborative decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119783759B_ABST
    Figure CN119783759B_ABST
Patent Text Reader

Abstract

The present invention discloses an intelligent reinforcement learning training method based on self-play, specifically related to the field of heterogeneous agent collaborative decision-making, and is used to solve the problems of real-time performance and policy stability in a high-dimensional heterogeneous action space. By means of typed splitting of the action space and shared feature extraction, the scale of network parameters is reduced; the application of dynamic binary masks realizes real-time filtering of illegal actions, overcomes the lag and exploration bias of relying on the reward and punishment mechanism, and ensures the compliance and efficiency of the decision-making process; the priority score matrix constructed by integrating tactical stability and environmental complexity, combined with hierarchical importance sampling, optimizes the action selection strategy and improves the stability and environmental adaptability of decision-making; the introduction of cross-attention networks and collaborative gain coefficients effectively captures and utilizes the tactical associations between agents, enhancing the overall effectiveness and robustness of collaborative decision-making.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of heterogeneous agent collaborative decision-making, and more specifically, to an intelligent reinforcement learning training method based on self-play. Background Art

[0002] In a heterogeneous agent collaborative decision-making system, since different types of agents need to execute high-dimensional heterogeneous actions (such as movement control, resource allocation, tactical response), traditional methods directly use the joint action space as the overall input to the policy network, resulting in an exponential expansion of the network parameter scale with the number of agents and the action dimension, leading to low training efficiency and sparse gradients. When the prior art adopts an end-to-end joint decision-making framework, it does not decouple the action space according to the differences in agent types, resulting in the difficulty for the policy network to capture the tactical behavior characteristics of specific types of agents; at the same time, the filtering of illegal actions relies on the ex-post punishment mechanism of the reward function, suffering from constraint lag and exploration direction deviation problems. When optimizing the multi-agent action coordination, the linear weighted fusion method ignores the non-linear tactical dependencies between actions, resulting in a high distortion rate of collaborative decision-making. The above defects seriously restrict the real-time performance and policy stability of heterogeneous agent collaborative decision-making in complex adversarial scenarios.

[0003] To solve the above problems, a technical solution is provided now. Summary of the Invention

[0004] To overcome the above-mentioned defects of the prior art, an embodiment of the present invention provides an intelligent reinforcement learning training method based on self-play. By splitting the overall action space according to agent type labels and extracting shared features, the scale of network parameters is reduced; the application of dynamic binary masks realizes the real-time filtering of illegal actions, overcoming the lag and exploration deviation of relying on the reward and punishment mechanism, and ensuring the compliance and efficiency of the decision-making process; the priority score matrix constructed by integrating tactical stability and environmental complexity, combined with hierarchical importance sampling, optimizes the action selection strategy, improving the stability and environmental adaptability of the decision-making; the introduction of cross-attention networks and collaborative gain coefficients effectively captures and utilizes the tactical correlations between agents, enhancing the overall effectiveness and robustness of collaborative decision-making, so as to solve the problems raised in the above background art.

[0005] To achieve the above object, the present invention provides the following technical solutions:

[0006] Split the overall action space into several sub-spaces according to agent type labels, and generate the original action probability matrix for each sub-space;

[0007] In the real-time environmental state, filter out illegal actions to obtain the set of legal action sequences for each sub-space;

[0008] In a legal action sequence, actions that balance tactical stability and environmental adaptability are preferentially selected to generate an action vector;

[0009] The action vector is input into a cross-attention network to calculate the tactical correlation matrix, and a collaborative gain coefficient is generated through matrix multiplication to correct the original action value, and finally a set of collaborative decision instructions is output.

[0010] In a preferred embodiment, the acquisition of the original action probability matrix includes the following:

[0011] Under the guidance of the agent type label, the overall action space is split into multiple independent subspaces; after splitting, the actions in each subspace are no longer coupled to each other, and each subspace corresponds to an independent policy network; for the generation of the action probability matrix within each subspace, first the global environmental state is input into a shared convolutional layer for feature extraction to obtain an environmental feature representation; then, for each subspace, an independent linear mapping weight and bias term are applied to transform the extracted features, and finally the original action probability matrix corresponding to the subspace is generated.

[0012] In a preferred embodiment, the process of obtaining the set of legal action sequences for each subspace is as follows:

[0013] Based on the analysis result of the real-time environmental state and the tactical rule base, a dynamic binary mask matrix is generated; if an action meets the current environmental and tactical requirements, the mask value corresponding to the action is 1; otherwise, the mask value is 0; the original action probability matrix and the dynamic binary mask matrix are multiplied element by element through the Hadamard product, ensuring that only legal actions are retained while illegal actions are set to zero; then, the Softmax function is applied to renormalize the filtered matrix to ensure that the sum of the probabilities of all legal actions is 1; the action probability matrix after the Hadamard product and renormalization is the set of legal action sequences for each subspace; each subspace only retains the actions that are legal under the current environment and tactical rules.

[0014] In a preferred embodiment, the tactical stability analysis is performed by calculating the tactical stability index, and the environmental adaptability analysis is performed by calculating the complexity sensitivity index.

[0015] In a preferred embodiment, the acquisition logic of the tactical stability index is as follows:

[0016] The historical action sequence of the agent is divided into the first half and the second half in chronological order; for these two subsequences, the dynamic time warping algorithm is applied to find the optimal alignment path between the two sequences and calculate the cumulative distance along this path; the cumulative distance is obtained by accumulating point by point by comparing the differences between the corresponding points in the sequences; the action sequence is converted into the state transition form of consecutive action pairs, the occurrence times of each action pair are counted, and the occurrence frequency of each action pair is calculated; based on the frequency distribution, the entropy value of the state transition is calculated using the Shannon entropy formula; the reciprocal of the dynamic time warping distance is combined with the negative exponent of the Shannon entropy value to calculate the tactical stability index.

[0017] In a preferred embodiment, the calculation process of the complexity-sensitive index is as follows:

[0018] For each environmental complexity index, the logistic map formula is applied for multiple iterations to generate a chaotic sequence; the index after chaotic transformation is regarded as the node of the graph, and an adjacency matrix is constructed, where the matrix elements are the reciprocals of the distances between the nodes; based on the adjacency matrix, the Laplacian matrix of the graph is calculated, specifically the degree matrix minus the adjacency matrix; the eigenvalues of the Laplacian matrix are solved, and the second smallest eigenvalue is extracted, which is called the algebraic connectivity; the algebraic connectivity is divided by the sum of the indices after chaotic transformation to calculate the complexity-sensitive index.

[0019] In a preferred embodiment, the process of obtaining the action vector is as follows:

[0020] The tactical stability index and the complexity-sensitive index are combined to calculate the priority score of each agent type; the priority scores of all agent types are formed into a vector, and the Softmax function is applied for normalization, and the normalization result is used as the priority score matrix; the original probability is multiplied element by element by the score of the corresponding type in the priority score matrix to obtain the adjusted probability distribution; the adjusted distribution is normalized, a random number is generated and sampled for specific actions to form the action vector.

[0021] In a preferred embodiment, the action vector is input into the cross-attention network to calculate the tactical relevance between agents. The specific process is as follows: for two agents, their action vectors are respectively transformed into query vectors, key vectors, and value vectors; by calculating the dot product of the query vector and the key vector, the attention weights are obtained; after the attention weights are normalized, they are weighted and summed with the value vectors to generate the tactical association vector of a single agent; the tactical association vectors of all agents are integrated to form a tactical association degree matrix, and the elements in it represent the tactical association strength between agents.

[0022] In a preferred embodiment, based on the tactical correlation matrix, a collaborative gain coefficient is generated through matrix multiplication operations to correct the action value. The specific steps are as follows: Multiply the tactical correlation matrix by the action vector matrix to generate a collaborative correction vector, where the elements represent the action adjustment values of the agents after considering tactical correlations; Subsequently, the collaborative correction vector is weighted and fused with the original action value to generate a collaborative gain coefficient.

[0023] In a preferred embodiment, the collaborative gain coefficient is input into the decision-making network to generate a final collaborative decision instruction set; The decision-making network converts the collaborative gain coefficient into an action probability distribution, and its internal parameters include a weight matrix and a bias vector, which are used to adjust the characteristics of the probability distribution; According to the generated action probability distribution, by selecting the action with the highest probability, a collaborative decision instruction set is formed.

[0024] The technical effects and advantages of the intelligent reinforcement learning training method based on self-play in the present invention:

[0025] By combining action space decoupling, dynamic mask filtering, priority fusion, and collaborative attention mechanism, the present invention significantly improves the performance and efficiency of the heterogeneous agent collaborative decision-making system. First, through the typed splitting and shared feature extraction of the action space, the scale of network parameters is reduced, alleviating the problems of low training efficiency and sparse gradients caused by the high-dimensional joint action space in traditional methods. Secondly, the application of the dynamic binary mask realizes the real-time filtering of illegal actions, overcoming the lag and exploration bias of relying on the reward and punishment mechanism, and ensuring the compliance and efficiency of the decision-making process. Further, the priority score matrix constructed by fusing tactical stability and environmental complexity, combined with hierarchical importance sampling, optimizes the action selection strategy, improving the stability of decision-making and environmental adaptability. Finally, the introduction of the cross-attention network and the collaborative gain coefficient effectively captures and utilizes the tactical correlations between agents, enhancing the overall effectiveness and robustness of collaborative decision-making. Furthermore, in complex adversarial scenarios, the real-time decision-making ability, strategy stability, and collaborative optimization level of the heterogeneous agent system are significantly improved. Brief Description of the Drawings

[0026] Figure 1 It is a schematic flow chart of the intelligent reinforcement learning training method based on self-play in the present invention. Detailed Embodiments

[0027] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0028] Example 1: Figure 1 The intelligent reinforcement learning training method based on self - game of the present invention is given, including:

[0029] S1: Split the overall action space into several sub - spaces according to the agent type label, and generate the original action probability matrix of each sub - space.

[0030] S2: In the real - time environment state, filter out illegal actions to obtain the set of legal action sequences of each sub - space.

[0031] S3: Among the legal action sequences, preferentially select actions that take into account both tactical stability and environmental adaptability to generate an action vector.

[0032] S4: Input the action vector into the cross - attention network to calculate the tactical correlation matrix, generate a collaborative gain coefficient through matrix multiplication to correct the original action value, and output the final collaborative decision instruction set.

[0033] In a complex scenario of multi - agent collaborative decision - making, different types of agents usually have different task objectives, action spaces, and behavioral strategies. Due to the functional differences of these agents, directly merging all action spaces into a unified input will lead to parameter explosion, increasing the computational overhead and gradient sparsity problems during the training process. Step S1 splits the overall action space according to the agent type label, and combines with a shared convolutional layer to extract environmental information, independently assigns a policy network to each sub - space, while maintaining the consistency of global environmental features.

[0034] Multi - agent collaborative decision - making refers to the situation where multiple independent agents cooperate and coordinate in a complex environment to jointly make decisions to achieve an overall goal or optimize system performance. In such a scenario, each agent usually has different functions, tasks, and behavioral spaces. There is both information interaction and mutual influence among agents. According to the environmental state and the behaviors of other agents, through sharing information, collaborative planning, and division of labor execution, the comprehensive response to the environment is finally achieved. Multi - agent collaborative decision - making is widely used in fields such as robot swarms, autonomous driving, energy scheduling, and financial investment. Its core challenge lies in how to handle the interactions and conflicts among agents while maintaining the efficiency and stability of the decision - making process.

[0035] Step S1 includes the following content:

[0036] Agent types are classified according to their functions and the tasks they perform. Each type of agent has a different action range and decision space. A typical classification method is based on the functions, tasks, and required operation dimensions of the platform. For example, if multiple roles such as patrolling, reconnaissance, and command are involved in a complex task, each role serves as an agent type and has a unique action space related to the task. Each agent type label provides the necessary information for subsequent action space splitting and policy training, ensuring that the policy network for each subspace can focus on specific behavior patterns and avoid unnecessary cross-interference.

[0037] Under the guidance of the agent type label, the overall action space is split into multiple independent subspaces. Each subspace represents the action set of a specific type of agent, and all actions within the action set have similar influence ranges and decision logics. After splitting, the actions of each subspace are no longer coupled to each other, and each subspace corresponds to an independent policy network. This approach effectively avoids the problems of gradient sparsity and low training efficiency caused by the high-dimensional joint action space in traditional methods. At the same time, the independence of subspace splitting reduces the interference of the action selection of various agents on the strategies of other agents, improving the stability during the training process.

[0038] In order to enable the policy networks of each subspace to uniformly perceive the global environmental state but not be affected by the differences between different agent types, the environmental situation features are extracted through a shared convolutional layer. The shared convolutional layer can transmit consistent environmental information between different subspaces, ensuring that the policy training of all agents is based on the same understanding of the global environment. For the generation of the action probability matrix within each subspace, first, the global environmental state is input into the shared convolutional layer for feature extraction to obtain the environmental feature representation. Then, for each subspace, independent linear mapping weights and bias terms are applied to transform the extracted features, and finally, the original action probability matrix corresponding to the subspace is generated.

[0039] By splitting the overall action space into independent subspaces for different agent types and using a shared convolutional layer to extract environmental information, this step effectively reduces the complexity of the policy network during the training process and improves the training efficiency. The policy network for each subspace focuses on the behavioral characteristics of a specific type of agent while maintaining a unified perception of the global environment, thus providing a stable and clear input for subsequent tactical rule parsing and priority decision-making.

[0040] Step S2 includes the following content:

[0041] In the complex multi-agent collaborative decision-making process, the real-time environmental state and the preset tactical rules jointly influence the agents' decisions. The goal of step S2 is to generate a set of legal actions by combining the tactical rules with the environmental state. The original action probabilities are screened and normalized through a dynamic binary mask matrix to ensure that each agent executes actions that conform to the current environment and task requirements.

[0042] The real-time environmental state refers to the dynamic characteristics of the current environment, usually including the state information of factors such as position, speed, obstacles, friendly and enemy units, etc. This information is collected and updated in real time through sensors, vision systems, or other monitoring methods. The tactical rule library is a database containing various strategic rules, and these rules determine the allowed action trajectories and behavior restrictions according to different tactical goals, task requirements, and environmental constraints. Based on these rules, the tactical rule library provides effective behavior guidance for the current situation every time the environment is updated.

[0043] According to the parsing results of the real-time environmental state and the tactical rule library, a dynamic binary mask matrix is generated. Each element in this matrix indicates whether a certain action is legal in a specific environment. When the real-time environment changes, the mask matrix is updated according to the rules in the tactical rule library and the current environmental conditions (such as the positions of the agents, the distribution of the targets, etc.). If an action conforms to the current environment and tactical requirements, the mask value corresponding to this action is 1; otherwise, the mask value is 0. The generation method of the mask matrix can be expressed as:

[0044]

[0045] where, is the action is the environmental state, indicates whether this action is valid in this state.

[0046] The original action probability matrix and the dynamic binary mask matrix are multiplied element by element through the Hadamard product. The Hadamard product ensures that only legal actions are retained, while illegal actions are set to zero. Then, the Softmax function is applied to renormalize the filtered matrix to ensure that the sum of the probabilities of all legal actions is 1. This process is carried out through the following formula:

[0047]

[0048] where, is the action probability after mask screening and Softmax renormalization, is the mask matrix, ensuring that the probability of illegal actions is 0.

[0049] The action probability matrix after Hadamard product and renormalization is the set of legal action sequences for each subspace. Each subspace only retains the actions that are legal under the current environment and tactical rules. Through this set of legal actions, subsequent decision-making processes can effectively avoid performing invalid actions or actions that do not meet the task requirements.

[0050] In step S2, by combining the real-time environmental state and the tactical rule library, a set of legal action sequences is dynamically generated. The mask matrix is used to filter out the actions that meet the current environment and task requirements. The Hadamard product and Softmax renormalization operations ensure that only legal actions can be selected in the final output action probability matrix. This process effectively reduces the action space of the agent, improves the real-time performance and stability of decision-making, and provides high-quality input for subsequent decision-making steps.

[0051] Step S3 includes the following:

[0052] Tactical stability analysis is carried out by calculating the tactical stability index, and environmental adaptability analysis is carried out by calculating the complexity sensitivity index.

[0053] The acquisition logic of the tactical stability index is as follows:

[0054] The calculation of the tactical stability index aims to quantify the consistency and predictability of the action selection pattern of the agent in historical decisions, so as to evaluate the stability of tactical behavior. The design combines the morphological analysis of time series and the entropy measure of information theory to ensure that the evaluation reflects both the continuity of the action pattern and the certainty of the selection behavior. The calculation process is divided into two stages: analyzing the temporal dynamic stability of the action sequence and quantifying the structural disorder degree of action selection. Finally, the results of the two parts are fused to generate a priority score.

[0055] The historical action sequence of the agent is divided into the first half and the second half in chronological order. For these two subsequences, the dynamic time warping algorithm (DTW) is applied to find the optimal alignment path between the two sequences and calculate the cumulative distance along this path. The cumulative distance is obtained by accumulating the differences between the corresponding points in the sequences point by point. The smaller the distance value, the less the change in the action pattern before and after, indicating higher temporal stability. The specific processing process is as follows:

[0056] Given the agent type In the time window The action sequence within where , represents the total number of action types; represents the length of the action sequence within the time window . The sequence is divided into the first half and the second half Calculate the dynamic time warping (DTW) distance between two subsequences DTW measures the similarity of action patterns through the optimal alignment path. The smaller the cumulative distance, the more consistent the front and back patterns are

[0057] Convert the action sequence into the state transition form of consecutive action pairs, count the occurrence times of each action pair, and calculate the occurrence frequency of each action pair. Based on the frequency distribution, apply the Shannon entropy formula to calculate the entropy value of the state transition. The formula is the sum of the product of each frequency and the logarithm of the corresponding frequency, taking the negative value. The lower the entropy value, the more orderly the state transition is, indicating a higher stability of action selection. Combine the reciprocal of the dynamic time warping distance with the negative exponent of the Shannon entropy value to calculate the tactical stability index. The specific operation is as follows: Take the reciprocal of the dynamic time warping distance, multiply it by the exponential form of the negative value obtained by taking the natural logarithm of the Shannon entropy value, and the result is used as the final index. The design ensures that when the distance is smaller and the entropy value is lower, the index is larger, reflecting higher tactical stability. The specific processing process is as follows

[0058] Convert the action sequence into the state transition probability distribution, and the state is defined as consecutive action pairs Count the frequencies of all state pairs where represents the th state pair. Calculate the Shannon entropy: The tactical stability index is fused through the following formula: where and are the action indices of the first half and the second half of the action sequence

[0059] The calculation process of the complexity-sensitive index is as follows

[0060] The calculation of the complexity-sensitive index aims to evaluate the impact of environmental complexity on the agent's decision-making by analyzing the non-linear dynamic characteristics of environmental features and the coupling relationship between indicators. Use chaos theory to amplify the unpredictability of the environment, and quantify the structural complexity between indicators through graph theory to ensure that the evaluation takes into account both the dynamic nature of single indicators and the relevance of multiple indicators. The calculation process is divided into two stages: non-linearly transform the environmental indicators, and analyze the interaction of the transformed indicators, and finally fuse the results to generate an index

[0061] For each environmental complexity indicator, apply the logistic map formula for multiple iterations to generate a chaotic sequence. The logistic map calculates the next iteration value in turn by setting the initial value and parameters, and the number of iterations is fixed. The chaotic sequence amplifies the small fluctuations of the indicator, reflects the potential instability of the environment, and generates the result as the basis for subsequent analysis. The specific processing process is as follows

[0062] Assume that the environmental complexity is determined by the indicator It is shown that for each is a scalar value. Applying the logistic map to each index generates a chaotic sequence: , with the initial value (normalized to [0, 1]), represents the number of iteration steps, is the parameter of the logistic map for chaotic map transformation, usually taking the value of 3.9; iterate = 10 times, and take . This transformation amplifies the non - linear characteristics of the index and simulates the dynamic effects of complex environments.

[0063] Regarding the index after chaotic transformation as the nodes of a graph, an adjacency matrix is constructed, and the matrix elements are the reciprocals of the distances between nodes. Based on the adjacency matrix, the Laplacian matrix of the graph is calculated, specifically the degree matrix minus the adjacency matrix. Solving the eigenvalues of the Laplacian matrix, the second - smallest eigenvalue is extracted, which is called algebraic connectivity. The larger the eigenvalue, the stronger the coupling between the indices and the more uniform the complexity distribution. Performing a ratio operation on the algebraic connectivity and the sum of the indices after chaotic transformation, the complexity - sensitive index is calculated. The specific operation is: divide the algebraic connectivity by the sum of all chaotic sequence values, and the result is used as the final index. When the design ensures that the index coupling is tight and the total complexity is moderate, the index is higher, reflecting the sensitivity of the environment to decision - making. The specific processing process is:

[0064] Regarding the transformed index as the nodes of a graph, an adjacency matrix is constructed, where . Calculate the Laplacian matrix ( is the degree matrix), and find its second - smallest eigenvalue . The complexity - sensitive index is defined as: , where reflects the coupling strength between the indices, and the larger the value, the more uniform the complexity distribution; and represent the indices of the nodes in the graph, and each node corresponds to the environmental index after chaotic map transformation and ; represents the number of index indicators of environmental complexity; the denominator normalizes the total complexity. Capturing the interaction between indices through graph theory.

[0065] The construction of the priority score matrix aims to comprehensively consider the tactical stability index and the complexity - sensitive index, providing a quantitative basis for the selection of agent types. Introducing the interaction between the two indices through a non - linear fusion formula, ensuring that the score reaches a peak when the tactical stability is high and the complexity sensitivity is moderate, taking into account both stability and adaptability. The calculation process includes two steps: non - linear fusion and normalization.

[0066] Combine the tactical stability index with the complexity sensitivity index to calculate the priority score for each agent type. The specific operation is as follows: Take the result of dividing the tactical stability index by the input of the complexity sensitivity index to the sigmoid function, where the sigmoid function is 1 divided by (1 plus the negative exponential power of e). This calculation makes the score most sensitive to the tactical stability when the complexity sensitivity index approaches a specific threshold. Form a vector of the priority scores for all agent types and apply the Softmax function for normalization. The Softmax function calculates the exponential value of each score, divides it by the sum of the exponential values of all scores, and generates a probability distribution. The normalization result serves as the priority score matrix, ensuring that the sum of the scores is 1 for subsequent operations. The specific processing process is as follows:

[0067] Combine the tactical stability index and the complexity sensitivity index to calculate the priority score , using a non - linear formula: . When the complexity sensitivity index is higher, the denominator increases, reducing the priority score; when the tactical stability index is higher, the priority score increases. The sigmoid function introduces non - linear interaction, enhancing the dynamic balance between the two indices. For all agent types , construct a score vector , and normalize it to a probability distribution through Softmax: .

[0068] According to the probability distribution in the priority score matrix, use the roulette wheel method to select the agent type. The roulette wheel method generates a random number between 0 and 1, determines which interval it falls into based on the cumulative probability sum, and selects the corresponding type. In the set of legal actions of the selected type, obtain the original action probability distribution. Multiply the original probability element - by - element with the score of the corresponding type in the priority score matrix to get the adjusted probability distribution. Normalize the adjusted distribution to ensure that the sum of the probabilities is 1. Based on the normalized distribution, generate a random number and sample a specific action to form an action vector. The specific processing process is as follows:

[0069] Sample the agent type according to the probability distribution , and in the set of legal actions of , sample the action according to the adjusted probability distribution to construct an action vector .

[0070] ​Step S3 generates a tactical stability index by analyzing the behavior trajectories of legal action sequences, and combines it with the complexity-sensitive index parsed from real-time sensor data to fuse and construct a priority score matrix, providing quantitative guidance for action selection. At the same time, a hierarchical importance sampling technique is used to optimize the generation of action vectors, ensuring that the decision-making takes into account both stability and adaptability, thus effectively making up for the deficiencies of traditional methods in high-dimensional heterogeneous action spaces and improving the real-time performance and policy stability of the heterogeneous agent collaborative decision-making system in complex adversarial scenarios.

[0071] In the heterogeneous agent collaborative decision-making system, the generation of action vectors is an important part of the decision-making process. Step S3 constructs a priority score matrix by fusing the tactical stability index and the complexity-sensitive index, and generates action vectors using a hierarchical importance sampling technique. However, this action vector only reflects the individual decision-making tendencies of agents in the current environment and does not fully consider the collaborative effects and tactical correlations between agents. Step S4 aims to capture the tactical correlations between agents by introducing a cross-attention network, generate a collaborative gain coefficient to correct the original action value, and finally output a collaborative decision instruction set, thereby improving the collaboration and overall effectiveness of the decision-making and ensuring stability and real-time performance in complex adversarial scenarios.

[0072] Step S4 includes the following:

[0073] Input the action vectors generated in the previous step into the cross-attention network to calculate the tactical correlations between agents. The cross-attention network operates through a specific mechanism. The specific process is as follows: For two agents, their action vectors are respectively transformed into query vectors, key vectors, and value vectors. By calculating the dot product of the query vector and the key vector, the attention weights are obtained, which reflect the degree of tactical attention of one agent to another. After the attention weights are normalized, they are weighted and summed with the value vectors to generate the tactical correlation vector of a single agent. The tactical correlation vectors of all agents are integrated to form a tactical correlation degree matrix, where the elements represent the strength of the tactical correlations between agents. This process uses the attention mechanism to accurately capture the dynamic tactical dependencies between agents.

[0074] Based on the tactical correlation degree matrix, generate a collaborative gain coefficient through matrix multiplication operations to correct the action value. The specific steps are as follows: Multiply the tactical correlation degree matrix by the action vector matrix to generate a collaborative correction vector, where the elements represent the action adjustment values of agents after considering tactical correlations. Subsequently, the collaborative correction vector is weighted and fused with the original action value to generate a collaborative gain coefficient. During the fusion process, the influence degree of the collaborative effect is controlled by adjusting parameters to ensure that the corrected action value reflects both individual decision-making and group collaboration effects. The weighted fusion adopts a linear combination method to reflect the gain adjustment effect of tactical correlations on the action value.

[0075] Input the collaborative gain coefficient into the decision-making network to generate the final collaborative decision instruction set. The decision-making network converts the collaborative gain coefficient into an action probability distribution, and its internal parameters include a weight matrix and a bias vector, which are used to adjust the characteristics of the probability distribution. According to the generated action probability distribution, a collaborative decision instruction set is formed by sampling or selecting the action with the highest probability. This process ensures that the instruction set reflects both the optimal decision at the individual level and the result of group collaborative optimization.

[0076] The collaborative decision instruction set is generated through iterative optimization in a multi-agent system, which is used to coordinate the behavior of agents and improve the decision-making efficiency. Its usage method is to represent the finally obtained instruction set as a set containing multiple scalar values, and each scalar value is input into the decision-making module of the agent as a control parameter to adjust the action selection probability distribution. Iterative optimization is achieved through a weight function and normalization processing. The weight function calculates the next-generation value based on the previous iteration result and the normal distribution parameters, and normalization processing ensures that the scalar value range is between 0 and 1. After multiple iterations, the instruction set converges to a stable state, guiding the agent group to achieve a global optimal decision in a complex dynamic environment. Its role is to enhance the collaboration and adaptability of decision-making behaviors, reduce conflicts and resource waste. In a high-dimensional heterogeneous action space, the application of the instruction set significantly improves the overall efficiency and real-time performance of group decision-making, and is applicable to distributed decision-making systems. Especially in complex adversarial scenarios that require efficient collaboration, it ensures that the system responds quickly and operates stably, and finally realizes the optimization of decision-making efficiency and the global optimal state.

[0077] Step S4 calculates the tactical correlation matrix through a cross-attention network, captures the tactical correlation between agents, generates a collaborative gain coefficient using matrix multiplication to correct the original action value, and finally outputs the collaborative decision instruction set. This method effectively integrates individual decision-making and group collaborative effects, makes up for the deficiencies of traditional methods in action collaborative optimization, improves the decision-making stability and real-time performance of heterogeneous agent systems in complex environments, and provides technical support for collaborative decision-making.

[0078] All the above formulas are dimensionless and take their numerical values for calculation. The formulas are obtained by collecting a large amount of data for software simulation to get a formula that is closest to the actual situation. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.

[0079] It should be noted that the system of the present invention can be deployed on the device itself to achieve embedded applications, or can also run on a PC or other terminals with a user interface, so as to meet various hardware environments and usage requirements.

[0080] Only some exemplary embodiments of the present invention have been described by way of illustration. Undoubtedly, for those of ordinary skill in the art, the described embodiments can be modified in various different ways without departing from the spirit and scope of the present invention. Therefore, the above drawings and description are illustrative in nature and should not be construed as limiting the scope of protection of the claims of the present invention.

[0081] It should be noted that in this text, if there are relational terms such as first and second, they are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the element.

[0082] As described above, this is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed in the present application can easily think of changes or substitutions, which should all be covered by the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.

Claims

1. An intelligent reinforcement learning training method based on self-game, characterized in that: Includes steps: The overall action space is split into several subspaces according to the agent type label, and the original action probability matrix of each subspace is generated. The agent is a robot cluster; In the real-time environment state, illegal actions are filtered out to obtain the legal action sequence set of each subspace; the real-time environment state refers to the dynamic characteristics of the current environment, including the position, speed and obstacle status information, which are collected and updated in real time through sensors, visual systems or other monitoring methods; In the legal action sequence, actions that take into account both tactical stability and environmental adaptability are prioritized to generate action vectors; The action vector is input into the cross-attention network to calculate the tactical relevance matrix, and the collaborative gain coefficient is generated through matrix multiplication to correct the original action value, and the final collaborative decision instruction set is output; The logic for obtaining the tactical stability index is as follows: The tactical stability analysis is conducted by calculating the tactical stability index, and the environmental adaptability analysis is conducted by calculating the complexity sensitivity index; The historical action sequence of the agent is divided into the first half and the second half in chronological order. For these two subsequences, the dynamic time warping algorithm is applied to find the optimal alignment path between the two sequences and calculate the cumulative distance along this path. The action sequence is converted into the state transition form of continuous action pairs, the frequency of occurrence of each action pair is calculated, and the entropy value of the state transition is calculated based on the frequency distribution. The inverse of the dynamic time warping distance is combined with the negative exponent of the Shannon entropy value to calculate the tactical stability index. The calculation process of complexity sensitivity index is as follows: For each environmental complexity index, the logistic mapping formula is applied for multiple iterations to generate a chaotic sequence. The index after chaotic transformation is regarded as the node of the graph, the adjacency matrix is ​​constructed, the Laplace matrix of the graph is calculated, the eigenvalue of the Laplace matrix is ​​solved, and the second smallest eigenvalue is extracted, which is called algebraic connectivity. The algebraic connectivity is compared with the sum of the index after chaotic transformation to calculate the complexity sensitivity index. The process of obtaining the action vector is as follows: Combine the tactical stability index with the complexity sensitivity index to calculate the priority score of each agent type; form a vector of the priority scores of all agent types, apply the Softmax function to normalize them, and use the normalized result as the priority score matrix; multiply the original probability by the score of the corresponding type in the priority score matrix element by element to obtain the adjusted probability distribution; normalize the adjusted distribution, generate random numbers and sample specific actions to form an action vector.

2. The intelligent reinforcement learning training method based on self-game according to claim 1 is characterized in that: The acquisition of the original action probability matrix includes the following: Under the guidance of agent type labels, the overall action space is split into multiple independent subspaces; After the split, the actions of each subspace are no longer coupled to each other, and each subspace corresponds to an independent policy network. For the generation of the action probability matrix in each subspace, the global environment state is first input into the shared convolutional layer for feature extraction to obtain the environment feature representation. Then, for each subspace, independent linear mapping weights and bias terms are applied to transform the extracted features, and finally the original action probability matrix of the corresponding subspace is generated.

3. The self-game-based intelligent reinforcement learning training method according to claim 2, characterized in that: The process of obtaining the set of legal action sequences in each subspace is as follows: Generate a dynamic binary mask matrix based on the real-time environment status and the analysis results of the tactical rule library; if an action meets the current environment and tactical requirements, the mask value of the corresponding action is 1; Otherwise, the mask value is 0; the original action probability matrix is ​​multiplied element by element with the dynamic binary mask matrix through the Hadamard product to ensure that only legal actions are retained and illegal actions are set to zero; then, the Softmax function is applied to renormalize the filtered matrix to ensure that the sum of the probabilities of all legal actions is 1; the action probability matrix after the Hadamard product and renormalization is the set of legal action sequences in each subspace; each subspace only retains legal actions under the current environment and tactical rules.

4. The intelligent reinforcement learning training method based on self-game according to claim 1 is characterized in that: The action vectors are fed into the cross-attention network to calculate the tactical relevance between the agents. The specific process is as follows: for the two agents, their action vectors are converted into query vectors, key vectors, and value vectors respectively; the attention weight is obtained by calculating the dot product of the query vector and the key vector; After normalization, the attention weights are weighted and summed with the value vector to generate the tactical relevance vector of a single agent. The tactical relevance vectors of all agents are integrated to form a tactical relevance matrix, in which the elements represent the strength of the tactical relevance between agents.

5. The self-game-based intelligent reinforcement learning training method according to claim 4, characterized in that: Based on the tactical correlation matrix, a synergy gain coefficient is generated through matrix multiplication operation to correct the action value. The specific steps are: multiply the tactical correlation matrix with the action vector matrix to generate a synergy correction vector, in which the elements represent the action adjustment value of the intelligent agent after considering the tactical correlation; then, the synergy correction vector is weightedly fused with the original action value to generate a synergy gain coefficient.

6. The self-game-based intelligent reinforcement learning training method according to claim 5, characterized in that: Input the collaborative gain coefficient into the decision network to generate the final collaborative decision instruction set; The decision network converts the collaborative gain coefficient into an action probability distribution. Its internal parameters include a weight matrix and a bias vector, which are used to adjust the characteristics of the probability distribution. According to the generated action probability distribution, a collaborative decision instruction set is formed by selecting the action with the highest probability.

Citation Information

Patent Citations

  • Multi-agent reinforcement learning transferable method, device and equipment

    CN118627587A

  • Improved DS-PPO reinforcement learning method for intelligent game deduction

    CN119337959A