Dynamic Evolution and Strategy Deduction Optimization Method and System for Game Behaviors in the Space Field
Through deep neural networks and multi-agent reinforcement learning models, analyzing the historical behavior data of participants in the space field, building a dynamic relationship diagram and game behavior model, solving the problem of difficulty in characterizing the dynamic evolution of participants' behavior in the existing technology, and achieving more accurate strategy deduction and optimization.
Patent Information
- Application Number
- CN202510096836.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-01-22
AI Technical Summary
The existing technology is difficult to effectively characterize the dynamic evolutionary characteristics of the behavior of participants in the space field, and the lack of systematic multi-level game scenario modeling, resulting in a large deviation from the actual situation.
By obtaining historical behavioral data of participants in the space field, using deep neural networks to perform timing feature analysis, building multi-dimensional feature vectors and dynamic relationship diagrams, combining multi-agent reinforcement learning models, a game behavior model of adaptive utility functions and dynamic decision rules is generated.
Accurate analysis and strategy optimization of the dynamic evolution characteristics of game behavior in the space field are achieved, and a game behavior model that is closer to reality is generated, which improves the accuracy and efficiency of strategy deduction.
Smart Images

Figure CN119539090B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the optimization technology of game behaviors in the space field, and particularly to a method and system for optimizing the dynamic evolution and strategy deduction of game behaviors in the space field. Background Art
[0002] With the rapid development of space technology and the rise of commercial spaceflight, the number of participants in the space field is increasing day by day, and their strategic game behaviors show high complexity and dynamicity. Traditional game analysis methods are mostly based on static assumptions, which are difficult to effectively characterize the dynamic evolution characteristics of the behaviors of participants. The decision-making behaviors of participants are simplified into fixed strategy selections, ignoring their adaptive adjustments with the changes of time and environment, resulting in a large deviation between the analysis results and the actual situation.
[0003] Existing strategy deduction methods lack systematic modeling of multi-level game scenarios, making it difficult to construct a complete game scenario evolution framework and accurately grasp the overall utility of the decisions of participants. Existing research lacks sufficient depth in mining the historical behavior data of participants and lacks effective data analysis models to extract the deep features and evolution laws of the behaviors of participants. The strategy evaluation system usually evaluates the strategy effects from a single dimension and lacks an evaluation mechanism that comprehensively considers multiple key factors such as technical feasibility, resource constraints, and long-term benefits. In addition, traditional methods mainly rely on expert experience for strategy planning, which is difficult to cope with the rapidly changing space competition environment and cannot achieve dynamic optimization and adjustment during the strategy implementation process.
[0004] Therefore, there is an urgent need to propose a new method that can effectively analyze the dynamic evolution characteristics of game behaviors in the space field and optimize strategies. Summary of the Invention
[0005] Embodiments of the present invention provide a method and system for optimizing the dynamic evolution and strategy deduction of game behaviors in the space field, which can solve the problems in the prior art.
[0006] In the first aspect of the embodiments of the present invention,
[0007] A method for optimizing the dynamic evolution and strategy deduction of game behaviors in the space field is provided, including:
[0008] Obtaining the historical behavior data of the participants in the space field, performing time series feature analysis on the historical behavior data through a deep neural network to obtain the periodic change rules of the behaviors of the participants, constructing a feature extraction model by using a distributed learning framework based on the periodic change rules, generating a multi-dimensional feature vector, constructing a dynamic relationship graph of the participants according to the multi-dimensional feature vector, inputting the dynamic relationship graph into a preset multi-agent reinforcement learning model, and the multi-agent reinforcement learning model adopts an encoder structure and a multi-layer self-attention mechanism to output a game behavior model including an adaptive utility function and a dynamic decision rule;
[0009] Construct a multi - layer game scenario evolution tree according to the game behavior model, where each tree node contains the state information and decision space of the participants. Input the state information and decision space into a dual neural network structure, calculate the immediate payoff value and long - term risk value of the participants in each game scenario respectively. Based on the temporal variation law of the immediate payoff value and long - term risk value, construct an evaluation function, and use a search algorithm with a hierarchical attention mechanism to dynamically update the multi - layer game scenario evolution tree. The hierarchical attention mechanism adaptively adjusts the search strategy according to the evaluation function, determines the sampling weights of the evolution paths, and generates multiple candidate evolution paths with different strategy characteristics;
[0010] Input the candidate evolution paths into a strategy evaluation model constructed by a spatio - temporal graph neural network to generate path evaluation indicators including technology maturity score, resource investment score, and strategy payoff score. Select the optimal strategy path according to the path evaluation indicators, and based on the optimal strategy path, combine with a knowledge reasoning engine to generate a phased implementation strategy deployment plan, and evaluate and optimize the feasibility of the strategy deployment plan through a distributed verification network.
[0011] In an alternative embodiment,
[0012] Construct a dynamic relationship graph of participants according to the multi - dimensional feature vector, and input the dynamic relationship graph into a preset multi - agent reinforcement learning model. The multi - agent reinforcement learning model adopts an encoder structure and a multi - layer self - attention mechanism, and the output game behavior model includes an adaptive utility function and dynamic decision rules:
[0013] Generate an initial dynamic relationship graph according to the multi - dimensional feature vector, where nodes represent participants and graph edges represent the interaction relationships between participants; obtain node representation vectors by feature embedding of nodes through a deep neural network, and extract dynamic pattern features by using a temporal convolutional network for interaction history; combine the node representation vectors and dynamic pattern features to calculate edge weights, and update the topological structure of the dynamic relationship graph according to the edge weights and node representation vectors;
[0014] Input the dynamic relationship graph into a multi - agent reinforcement learning model, and the multi - agent reinforcement learning model includes an encoder network and a policy network; the encoder network encodes the dynamic relationship graph through a node feature layer, a relationship feature layer, and a temporal feature layer to obtain a state representation vector; the policy network generates the action probability distribution of each agent based on the state representation vector;
[0015] Construct a multi-agent value network, and predict the state value and action value of each agent based on the state representation vector and action probability distribution; input the state value and action value into the utility function generation module to construct a utility function including individual utility and group utility;
[0016] Calculate the reward signal of each agent based on the utility function, and input the reward signal into the policy optimization module; the policy optimization module updates the policy network parameters by using the policy gradient method to realize the dynamic optimization of the agent decision rule;
[0017] Adopt an experience replay mechanism to store the interaction data of agents, and synchronously update the value network and utility function parameters based on the interaction data; realize the joint optimization of the policy network, value network and utility function through multi-agent collaborative learning;
[0018] Integrate the optimized policy network, value network and utility function to output a game behavior model including an adaptive utility function and a dynamic decision rule.
[0019] In an alternative embodiment,
[0020] Construct a multi-agent value network, and predict the state value and action value of each agent based on the state representation vector and action probability distribution; input the state value and action value into the utility function generation module to construct a utility function including individual utility and group utility, including:
[0021] Construct a state value network and an action value network. The state value network adopts a three-layer fully connected network structure and inputs the state representation vector to obtain a state value prediction; the action value network adopts a double-branch structure. The first branch processes the state representation vector, and the second branch processes the action probability distribution. After fusing the features of the two branches through an outer product operation, the action value prediction is obtained through a fully connected layer;
[0022] Calculate the temporal difference target value, which is calculated based on the immediate reward value, state value prediction and action value prediction; construct a state value loss function with the difference between the temporal difference target value and the state value prediction, and construct an action value loss function with the difference between the temporal difference target value and the action value prediction; update the value network parameters based on the state value loss function and the action value loss function to obtain the optimized state value and action value;
[0023] Use the optimized action value to represent the individual utility term, and calculate the group utility term through weighted calculation of the optimized state value and the action values of neighboring agents; combine the individual utility term and the group utility term to construct an initial utility function;
[0024] Calculate the gradient value of the initial utility function with respect to the weight coefficients, and update the weight coefficients of the individual utility term and the group utility term by using the gradient descent method based on the gradient value to obtain the final utility function.
[0025] In an alternative embodiment,
[0026] Construct a multi-layer game scenario evolution tree according to the game behavior model, where each tree node contains the state information and decision space of the participating parties, and input the state information and decision space into a dual neural network structure to calculate the immediate payoff value and long-term risk value of the participating parties in each game scenario, including:
[0027] Construct a multi-layer game scenario evolution tree based on the dynamic decision rule in the game behavior model, and store the state vector and decision space information of the participating parties in the tree nodes; calculate the transition probability matrix between adjacent nodes based on the state transition function; perform Monte Carlo sampling according to the transition probability matrix to generate an evolution path from the root node to the leaf node; use the node sequence corresponding to the evolution path and the historical state sequence of the participating parties as features and store them in each tree node;
[0028] Construct a dual neural network structure, including a graph attention network and a hierarchical recurrent neural network. Input the state vector in the tree node into the graph attention network, generate the feature representation matrix of the participating parties through the feature transformation layer, calculate the attention coefficients between the participating parties based on the feature representation matrix, multiply the attention coefficients by the feature representation matrix to obtain the weighted features, perform a residual connection on the weighted features and the feature representation matrix, transform the features after the residual connection through the multi-head attention mechanism to extract the correlation features between the participating parties, and input the correlation features into the adaptive pooling layer of the graph attention network for dimensionality reduction and aggregation to obtain the immediate payoff value reflecting the game state of the participating parties at the current node;
[0029] Group the historical state sequences stored in the tree nodes according to the participating party identifiers, input the grouped sequences into the hierarchical recurrent neural network. The recurrent neural network selectively remembers the grouped sequences through the gating unit, outputs the state features at each time step, calculates the attention weights in the time dimension for the state features to obtain the temporal correlation representation depicting the state evolution law, and at the same time calculates the attention weights in the participating party dimension for the state features to obtain the interactive correlation representation depicting the game relationship; perform a weighted combination of the temporal correlation representation and the interactive correlation representation with the state features, and output the long-term risk value reflecting the cumulative effect of the participating parties on the evolution path through the prediction module of the hierarchical recurrent neural network.
[0030] In an alternative embodiment,
[0031] Based on the temporal variation laws of the immediate return value and the long-term risk value, construct an evaluation function, and use a search algorithm with a hierarchical attention mechanism to dynamically update the multi-layer game scenario evolution tree. The hierarchical attention mechanism adaptively adjusts the search strategy according to the evaluation function, determines the sampling weights of the evolution paths, and generates multiple candidate evolution paths with different strategy characteristics, including:
[0032] Calculate the return change rate according to the immediate return value of the adjacent node game state, and calculate the risk change trend according to the long-term risk value of the historical state sequence; perform weighted combination on the immediate return value, return change rate and risk change trend to construct a temporal evaluation function; calculate the temporal evaluation scores of each tree node in the multi-layer game scenario evolution tree according to the temporal evaluation function;
[0033] Construct a state attention layer based on the temporal evaluation scores of the tree nodes, output the state importance weights of the tree nodes through the state attention layer, and at the same time input the temporal evaluation scores and corresponding state importance weights of adjacent tree nodes into the transfer attention layer to obtain the importance weights of state transitions. Calculate the complete path sampling probability from the root node to the leaf node based on the state importance weights of the tree nodes and the importance weights of state transitions;
[0034] Introduce a temperature parameter to adjust the path sampling probability, detect the gradient change value of the temporal evaluation function, increase the temperature parameter when the gradient change value is less than the convergence threshold, detect the temporal evaluation score of the sampling path, and reduce the temperature parameter when the score exceeds the historical optimum. Use the ratio of the path sampling probability to the temperature parameter as the adjusted sampling probability, and generate candidate evolution paths based on the adjusted sampling probability;
[0035] Maintain a candidate path set with a fixed size, calculate the weighted sum of the temporal evaluation scores of all tree nodes on the newly generated path as the total path evaluation score. When the total evaluation score of the new path is higher than the lowest score in the set, add the new path to the candidate path set and remove the path with the lowest score;
[0036] Calculate the edit distance between the paths in the candidate path set to obtain a path diversity index, dynamically adjust the temperature parameter based on the diversity index, increase the temperature parameter to expand the search range when the diversity index is lower than the target threshold, and reduce the temperature parameter to strengthen local search when the diversity index is higher than the target threshold; calculate the path weights according to the total path evaluation scores and sampling probabilities of each candidate path, and update the multi-layer game scenario evolution tree based on the path weights, and finally obtain multiple candidate evolution paths with different strategy characteristics.
[0037] In an alternative embodiment,
[0038] Input the candidate evolution path into the policy evaluation model constructed by the spatio-temporal graph neural network to generate path evaluation metrics including technology maturity score, resource investment score, and policy benefit score. Selecting the optimal policy path according to the path evaluation metrics includes:
[0039] Construct a spatio-temporal heterogeneous graph structure, use technology nodes and resource nodes as heterogeneous graph nodes, set the technology development sequence relationship as the temporal edge, set the resource dependency relationship as the spatial edge, construct the technology node attribute vector based on technology attribute features and timestamp information, construct the resource node attribute vector based on resource type and capacity information, and combine the technology node attribute vector and the resource node attribute vector to form a node feature matrix;
[0040] Input the node feature matrix and the temporal edge into the causal convolutional network to extract temporal context features, and calculate the feature weights through the temporal attention mechanism to obtain enhanced temporal features; input the node feature matrix and the spatial edge into the graph attention network to extract the interaction features between nodes, and calculate the feature importance through the spatial attention mechanism to obtain enhanced spatial features; input the temporal features and the spatial features into the gated update unit, calculate the gated weights and perform adaptive fusion to obtain a unified feature representation;
[0041] Based on the unified feature representation, use a multi-layer perceptron to match features with an expert knowledge base to obtain a technology feasibility score; calculate the matching degree between the path resource requirements and the template resources through a graph matching network to obtain a resource consumption score; extract policy correlation features through a graph attention network to obtain a policy value score;
[0042] Predict the parameters of the mixture Gaussian distribution based on the unified feature representation, combine the historical evaluation data, and calculate the posterior distribution through variational inference to perform Bayesian calibration on the technology feasibility score, resource consumption score, and policy value score to obtain the calibrated scores;
[0043] Input the calibrated scores into the path optimization objective function, and perform iterative optimization through the beam search algorithm under the constraints of technology feasibility, resource consumption, and policy value, and finally obtain the optimal policy path.
[0044] In an alternative embodiment,
[0045] Based on the optimal policy path, combine a knowledge reasoning engine to generate a phased implementation policy deployment plan, and evaluate and optimize the feasibility of the policy deployment plan through a distributed verification network, including:
[0046] Construct a policy deployment knowledge base, which includes a rule base of technology dependency rules, resource allocation rules, and progress control rules, as well as a fact base of mastered technology lists, current available resource status, and historical deployment case data;
[0047] Construct a knowledge reasoning engine based on the rule base and the fact base. The knowledge reasoning engine extracts the technical dependency relationships of each technical node in the optimal strategy path to obtain a dependency chain, calculates the earliest implementable time of each technical node in combination with the dependency chain and the current system state, matches the earliest implementable time with the progress control rules in the rule base to generate a time execution window, and calculates a resource scheduling plan based on the time execution window and the resource allocation rules to form a phased deployment plan;
[0048] Construct a hierarchical verification architecture. The hierarchical verification architecture includes a first-level verification component for performing global consistency checks, a second-level verification component for performing technical dependency verification, and a third-level verification component for performing resource constraint verification, and issue verification tasks to each level of verification component according to a hierarchical authorization mechanism;
[0049] Input the verification results of each level of verification component into the Byzantine fault-tolerant consensus algorithm, calculate the credibility scores of each verification result based on reputation weighting, sort the verification results according to the credibility scores, and perform structured processing on the verification opinions with high credibility to obtain optimization suggestions;
[0050] Construct a verification index vector according to the optimization suggestions, input the verification data of technical feasibility, resource matching degree, and schedule controllability into the corresponding index calculation models respectively, and generate a quantitative verification evaluation result;
[0051] Input the verification evaluation result into the policy optimization network to generate a targeted optimization adjustment plan, apply the optimization adjustment plan to the execution plan for update, re-enter the updated execution plan into the hierarchical verification architecture for evaluation, and output the final policy deployment plan when the verification evaluation result meets the preset standard, and return the execution plan to the policy optimization network for iterative optimization when the verification evaluation result does not meet the standard.
[0052] In the second aspect of the embodiments of the present invention, a system for dynamically evolving and optimizing strategy deduction of space domain game behaviors is provided, including:
[0053] A first unit for obtaining the historical behavior data of space domain participants, performing temporal feature analysis on the historical behavior data through a deep neural network to obtain the periodic change rules of the participants' behaviors, constructing a feature extraction model based on the periodic change rules by using a distributed learning framework to generate a multi-dimensional feature vector, constructing a dynamic relationship graph of the participants according to the multi-dimensional feature vector, and inputting the dynamic relationship graph into a preset multi-agent reinforcement learning model. The multi-agent reinforcement learning model adopts an encoder structure and a multi-layer self-attention mechanism, and outputs a game behavior model including an adaptive utility function and dynamic decision rules;
[0054] A second unit, configured to construct a multi-layer game scenario evolution tree according to the game behavior model, where each tree node contains the state information and decision space of the participating parties, input the state information and decision space into a dual neural network structure, calculate the immediate revenue value and long-term risk value of the participating parties in each game scenario respectively, construct an evaluation function based on the temporal variation law of the immediate revenue value and long-term risk value, and use a search algorithm with a hierarchical attention mechanism to dynamically update the multi-layer game scenario evolution tree. The hierarchical attention mechanism adaptively adjusts the search strategy according to the evaluation function, determines the sampling weights of the evolution paths, and generates multiple candidate evolution paths with different strategy characteristics;
[0055] A third unit, configured to input the candidate evolution paths into a strategy evaluation model constructed by a spatio-temporal graph neural network, generate path evaluation indicators including technology maturity scores, resource investment scores, and strategy revenue scores, select the optimal strategy path according to the path evaluation indicators, generate a phased implementation strategy deployment plan in combination with a knowledge reasoning engine based on the optimal strategy path, and evaluate and optimize the feasibility of the strategy deployment plan through a distributed verification network.
[0056] The third aspect of the embodiments of the present invention
[0057] Provided is an electronic device, including:
[0058] A processor;
[0059] A memory for storing instructions executable by the processor;
[0060] Wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.
[0061] The fourth aspect of the embodiments of the present invention,
[0062] Provided is a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.
[0063] In this embodiment, by performing temporal feature analysis on historical behavior data through a deep neural network and combining it with a multi-agent reinforcement learning model, the behavior of space domain participants can be predicted more accurately, revealing its periodic change patterns and dynamic relationships, and constructing a game behavior model closer to reality. Based on the game behavior model, a multi-layer game scenario evolution tree is constructed, and combined with a search algorithm of a dual neural network and a hierarchical attention mechanism, the immediate benefits and long-term risks of different strategies can be dynamically evaluated, generating multiple candidate evolution paths with different strategy characteristics, and selecting the optimal strategy path from them. Using a spatio-temporal graph neural network to construct a strategy evaluation model and combining it with a knowledge reasoning engine to generate a phased implementation strategy deployment plan can evaluate and optimize the feasibility of the strategy deployment plan, improving the efficiency and success rate of strategy deployment. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] Figure 1 is a flowchart of the method for optimizing the dynamic evolution and strategy deduction of game behaviors in the space domain according to an embodiment of the present invention;
[0065] Figure 2 is a schematic structural diagram of the system for optimizing the dynamic evolution and strategy deduction of game behaviors in the space domain according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0066] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0067] The technical solutions of the present invention will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.
[0068] Figure 1 is a flowchart of the method for optimizing the dynamic evolution and strategy deduction of game behaviors in the space domain according to an embodiment of the present invention, as Figure 1 shown, the method includes:
[0069] S101. Obtain the historical behavior data of space domain participants, perform time series feature analysis on the historical behavior data through a deep neural network to obtain the periodic change law of the participants' behavior. Based on the periodic change law, use a distributed learning framework to construct a feature extraction model, generate a multi-dimensional feature vector, construct a dynamic relationship graph of participants according to the multi-dimensional feature vector, and input the dynamic relationship graph into a preset multi-agent reinforcement learning model. The multi-agent reinforcement learning model adopts an encoder structure and a multi-layer self-attention mechanism, and outputs a game behavior model including an adaptive utility function and a dynamic decision rule;
[0070] S102. Construct a multi-layer game scenario evolution tree according to the game behavior model, where each tree node contains the state information and decision space of the participant. Input the state information and decision space into a dual neural network structure, calculate the immediate benefit value and long-term risk value of the participant in each game scenario respectively. Based on the time series change law of the immediate benefit value and long-term risk value, construct an evaluation function, and use a search algorithm with a hierarchical attention mechanism to dynamically update the multi-layer game scenario evolution tree. The hierarchical attention mechanism adaptively adjusts the search strategy according to the evaluation function, determines the sampling weight of the evolution path, and generates multiple candidate evolution paths with different strategy characteristics;
[0071] S103. Input the candidate evolution paths into a strategy evaluation model constructed by a spatio-temporal graph neural network, generate path evaluation indicators including technology maturity score, resource investment score, and strategy benefit score. Select the optimal strategy path according to the path evaluation indicators. Based on the optimal strategy path, combine a knowledge reasoning engine to generate a phased implementation strategy deployment plan, and evaluate and optimize the feasibility of the strategy deployment plan through a distributed verification network.
[0072] Among them, the historical behavior data refers to the decision-making records, resource allocation, and behavior patterns of participants in the past period. Through time series feature analysis by a deep neural network, the periodic change law of the participants' behavior can be extracted. Time series feature analysis can identify the periodic fluctuations and trend changes in the behavior pattern. For example, the strategy adjustments or resource allocation changes that may occur to participants within a specific period.
[0073] Based on the extracted periodic change law, use a distributed learning framework to construct a feature extraction model. The distributed learning framework can efficiently process large-scale data sets and map multi-dimensional data to the feature space. The multi-dimensional feature vector generated through this process will be used to describe the behavior characteristics and decision-making environment of participants. Through the multi-dimensional feature vector, a dynamic relationship graph of participants can be constructed, and this relationship graph can reflect the mutual influence and behavior relationship between different participants.
[0074] In an alternative embodiment, a dynamic relationship graph of the participating parties is constructed based on the multi-dimensional feature vector, and the dynamic relationship graph is input into a preset multi-agent reinforcement learning model. The multi-agent reinforcement learning model adopts an encoder structure and a multi-layer self-attention mechanism. The output game behavior model including an adaptive utility function and dynamic decision rules includes:
[0075] Generate an initial dynamic relationship graph based on the multi-dimensional feature vector, where the nodes represent the participating parties and the graph edges represent the interaction relationships between the participating parties; perform feature embedding on the nodes through a deep neural network to obtain node representation vectors, and use a temporal convolutional network to extract interaction history to obtain dynamic pattern features; combine the node representation vectors with the dynamic pattern features to calculate the edge weights, and update the topological structure of the dynamic relationship graph according to the edge weights and node representation vectors;
[0076] Input the dynamic relationship graph into a multi-agent reinforcement learning model, and the multi-agent reinforcement learning model includes an encoder network and a policy network; the encoder network encodes the dynamic relationship graph through a node feature layer, a relationship feature layer, and a temporal feature layer to obtain a state representation vector; the policy network generates the action probability distribution of each agent based on the state representation vector;
[0077] Construct a multi-agent value network, and predict the state value and action value of each agent based on the state representation vector and the action probability distribution; input the state value and the action value into a utility function generation module to construct a utility function including individual utility and group utility;
[0078] Calculate the reward signal of each agent based on the utility function, and input the reward signal into a policy optimization module; the policy optimization module updates the policy network parameters by using a policy gradient method to realize the dynamic optimization of the agent decision rules;
[0079] Adopt an experience replay mechanism to store the interaction data of the agents, and synchronously update the value network and utility function parameters based on the interaction data; realize the joint optimization of the policy network, the value network, and the utility function through multi-agent collaborative learning;
[0080] Integrate the optimized policy network, value network, and utility function to output a game behavior model including an adaptive utility function and dynamic decision rules.
[0081] Exemplarily, first, an initial dynamic relationship graph is generated through multi-dimensional feature vectors. The nodes of this graph represent various parties in the space domain, and the edges in the graph represent the interaction relationships between these parties, such as cooperation, competition, resource sharing, etc. When generating the initial dynamic relationship graph, the weights of the edges can be set according to factors such as the interaction intensity and cooperation frequency between the parties. A deep neural network is used to perform feature embedding on each node, and this process transforms information such as the attributes, behavior patterns, and strategy selections of each party into a vector, called the node representation vector. Then, a temporal convolutional network (TCN) is used to analyze the interaction history and extract dynamic pattern features to help the system identify the behavior change rules of the parties at different time points, such as periodic resource allocation, strategy adjustment, etc. These dynamic pattern features are combined with the node representation vectors, and the topological structure of the dynamic relationship graph is adjusted by calculating the weights of the graph edges, thereby reflecting the changing relationships between the parties.
[0082] After the construction of the dynamic relationship graph is completed, it is input into a multi-agent reinforcement learning model, which includes an encoder network and a policy network. The role of the encoder network is to transform the complex information in the dynamic relationship graph into a state representation vector, and this vector will serve as the basis for the agent's decision-making. The encoder network includes multiple feature layers, and each feature layer processes different types of information respectively. Specifically, the node feature layer is responsible for encoding the attributes and behavior information of the nodes, the relationship feature layer encodes the interaction relationships between the parties, and the temporal feature layer is specifically used to extract the patterns of the parties' behavior changes over time. Through the combination of these layers, the encoder network can generate a comprehensive state representation vector representing the status of each party in the current game scenario.
[0083] Next, the policy network generates the action probability distribution for each agent based on this state representation vector. In the game environment, each agent represents a party, and the action probability distribution represents the probabilities of different strategies that the party may adopt in the current state. The policy network will generate a strategy distribution for each party based on the input state information, determining their behavior in the current game scenario.
[0084] To evaluate the value of each agent adopting different strategies in a specific state, a multi-agent value network is constructed. The role of this network is to predict the state value and action value of each agent based on the state representation vector and the action probability distribution. The state value represents the overall value of the environment where the agent is in a certain state, and the action value is the reward that the agent may obtain by taking a certain action in this state. By calculating the state value and the action value, the value network helps the system evaluate the impact of different strategies on the game process.
[0085] On this basis, a utility function generation module is constructed. The utility function consists of two parts: individual utility and group utility. The individual utility reflects the short-term and long-term interests of each participant in the game, while the group utility evaluates the cooperation and competition effects of all participants. The utility function combines state value, action value, and other environmental factors to provide a reward signal for the agent, which will be used to optimize the decision-making strategy of the agent.
[0086] The reward signal is input into the policy optimization module and optimized using the policy gradient method. The policy gradient method updates the parameters of the policy network, making the action selection of each agent more in line with the optimal policy in different states. As the policy network is continuously optimized, the decision-making rules of the agent will be gradually adjusted dynamically, enabling the game behavior model to more accurately reflect the strategy choices of the participants in complex games.
[0087] To further improve the learning effect, an experience replay mechanism is adopted. The experience replay mechanism stores the interaction data of the agent and allows the model to review and utilize previous experiences in subsequent learning processes. This can prevent the model from falling into local optimal solutions and accelerate the process of multi-agent collaborative learning. Based on this interaction data, the model will synchronously update the parameters of the value network and the utility function, thus achieving parameter optimization in continuous feedback.
[0088] The key to multi-agent collaborative learning lies in jointly optimizing the policy network, value network, and utility function through the interaction and information sharing among multiple agents. This collaborative learning mechanism ensures that the system can consider the needs and interests of different participants from a global perspective and optimize the overall game behavior.
[0089] Finally, through model integration, the optimized policy network, value network, and utility function are integrated to output a game behavior model containing an adaptive utility function and dynamic decision-making rules. This model can not only predict the behavior of the participants in the game process but also adaptively adjust the strategy according to the changes in the game environment, thus providing a scientific basis and decision-making support for multi-party games and strategic decisions in the space field.
[0090] In this embodiment, the dynamic behavior and strategic decision-making of multiple participants in the space field are accurately modeled through deep learning and reinforcement learning technology. Through historical behavior data analysis and multi-dimensional feature vector extraction, the system can capture the complex interactive relationships and behavior patterns between the participants, and reflect their real-time changes through dynamic relationship diagrams. The multi-agent reinforcement learning model effectively predicts the behavior of all parties and optimizes the decision-making process through adaptive decision rules and game behavior modeling, realizing collaborative learning and dynamic optimization in the game process. Through the utility function and strategy optimization module, the system can not only improve the benefits of individual participants, but also optimize the overall benefits of the group. The experience replay mechanism ensures that the model can learn efficiently and avoid local optimality, and promotes the linkage optimization of strategies, value networks and utility functions. This technical solution can provide scientific decision-making basis for complex multi-party games in the space field, support multi-party strategy coordination and long-term planning optimization, improve decision-making efficiency and accuracy, and has broad application prospects.
[0091] In an optional implementation, a multi-agent value network is constructed, and the state value and action value of each agent are predicted based on the state representation vector and the action probability distribution; the state value and action value are input into the utility function generation module, and the utility function including individual utility and group utility is constructed, including:
[0092] Construct a state value network and an action value network. The state value network adopts a three-layer fully connected network structure, and the state representation vector is input to obtain the state value prediction value; the action value network adopts a double-branch structure, the first branch processes the state representation vector, and the second branch processes the action probability distribution. The two branch features are fused through the outer product operation and then the action value prediction value is obtained through the fully connected layer;
[0093] Calculate a time series difference target value, the time series difference target value is calculated based on the immediate reward value, the state value prediction value and the action value prediction value; construct a state value loss function by the difference between the time series difference target value and the state value prediction value, and construct an action value loss function by the difference between the time series difference target value and the action value prediction value; update the value network parameters based on the state value loss function and the action value loss function to obtain optimized state value and action value;
[0094] The optimized action value is used to represent the individual utility term, and the group utility term is obtained by weighted calculation of the optimized state value and the action value of the neighborhood agent; an initial utility function is constructed by combining the individual utility term and the group utility term;
[0095] The gradient value of the initial utility function with respect to the weight coefficient is calculated, and based on the gradient value, the weight coefficients of the individual utility term and the group utility term are updated using the gradient descent method to obtain the final utility function.
[0096] Exemplarily, first, prepare a multi-agent environment and the interaction data between agents. Collect the reward values obtained by each agent after performing various actions in different states, as well as the state transition information of the environment. Next, construct a state value network and an action value network. The state value network is used to predict the long-term value of an agent in a specific state. It adopts a three-layer fully connected network structure. The input is the vector representation of the state, such as the position, speed, surrounding environment information, etc. of the agent, and the output is a numerical value representing the value of this state. The action value network is used to predict the long-term value of an agent performing a specific action in a specific state. It adopts a double-branch structure, one branch processes the state vector, and the other branch processes the action probability distribution. The outputs of the two branches are fused through an outer product operation, and then passed through a fully connected layer. Finally, a numerical value is output, representing the value of performing this action in this state. Assume the state vector dimension is 5 and the action probability distribution dimension is 3, then the feature dimension obtained after the outer product operation is 15, and after passing through a fully connected layer with an output dimension of 1, the action value prediction value is obtained.
[0097] Then, calculate the temporal difference target value. This value is used to evaluate whether the current value prediction is accurate. It combines the current reward, the value of the next state, and the value of performing the action in the current state. For example, if the current reward is 10, the value prediction of the next state is 20, and the value prediction of performing the action in the current state is 15, then the temporal difference target value can be calculated as 10 + 0.9 * 20 - 15 = 13, where 0.9 is the discount factor, which is used to balance the importance of the current reward and future rewards.
[0098] Next, construct and minimize the loss function. Calculate the difference between the temporal difference target value and the state value prediction value, and the difference between the temporal difference target value and the action value prediction value respectively. These two differences constitute the state value loss function and the action value loss function respectively. By minimizing these two loss functions, the parameters of the value network can be updated, thereby improving the accuracy of value prediction.
[0099] After that, construct a utility function. The utility function is used to measure the value obtained by an agent. It includes an individual utility term and a group utility term. The individual utility term is represented by the optimized action value, and the group utility term is calculated by weighting the state value of the current agent and the action values of its neighboring agents. For example, an individual utility value is 15, and the action values of neighboring agents are 10 and 12 respectively, with weights of 0.6 and 0.4 respectively, then the group utility term is 0.6 * 10 + 0.4 * 12 = 10.8. The initial utility function can be defined as a linear combination of the individual utility and the group utility term, such as 0.8 * individual utility + 0.2 * group utility.
[0100] Finally, optimize the utility function. Calculate the gradient value of the initial utility function with respect to the weight coefficients, and then use the gradient descent method to update the weight coefficients of the individual utility term and the group utility term. For example, if the initial value of the weight coefficient of the individual utility term is 0.8, the gradient value is -0.1, and the learning rate is 0.01, then the updated weight coefficient is 0.8 - 0.01 * (-0.1) = 0.801. By continuously iterating and updating the weight coefficients, the final utility function can be obtained.
[0101] In this embodiment, by considering the individual utility and the group utility, the agent is encouraged to promote the achievement of the overall goal while pursuing its own interests, thereby improving the cooperation efficiency of the multi-agent system. By constructing the state value network and the action value network, and combining the temporal difference learning method, the long-term value of the agent performing different actions in different states can be evaluated more accurately. By optimizing the weight coefficients of the utility function using the gradient descent method, this method can adapt to different multi-agent environments and tasks and has stronger generalization ability.
[0102] In an alternative implementation, a multi-layer game scenario evolution tree is constructed according to the game behavior model, where each tree node contains the state information and decision space of the participating parties. Inputting the state information and decision space into a dual neural network structure, calculating the immediate payoff value and long-term risk value of the participating parties in each game scenario respectively includes:
[0103] Construct a multi-layer game scenario evolution tree based on the dynamic decision rules in the game behavior model, and store the state vector and decision space information of the participating parties in the tree nodes; calculate the transition probability matrix between adjacent nodes based on the state transition function; perform Monte Carlo sampling according to the transition probability matrix to generate an evolution path from the root node to the leaf node; use the node sequence corresponding to the evolution path and the historical state sequence of the participating parties as features and store them in each tree node;
[0104] Construct a dual neural network structure, including a graph attention network and a hierarchical recurrent neural network. Input the state vector in the tree node into the graph attention network, generate the feature representation matrix of the participating parties through the feature transformation layer, calculate the attention coefficients between the participating parties based on the feature representation matrix, multiply the attention coefficients by the feature representation matrix to obtain the weighted features, and perform a residual connection on the weighted features and the feature representation matrix. Transform the features after the residual connection through the multi-head attention mechanism to extract the correlation features between the participating parties. Input the correlation features into the adaptive pooling layer of the graph attention network for dimensionality reduction and aggregation to obtain the immediate payoff value reflecting the game state of the participating parties at the current node;
[0105] Group the historical state sequences stored in the tree nodes according to the participant identifiers, and input the grouped sequences into a hierarchical recurrent neural network. The recurrent neural network selectively remembers the grouped sequences through a gating unit, outputs the state features at each time step, calculates the attention weights in the time dimension for the state features to obtain a temporal correlation representation depicting the state evolution law, and simultaneously calculates the attention weights in the participant dimension for the state features to obtain an interaction correlation representation depicting the game relationship; combine the temporal correlation representation and the interaction correlation representation with the state features through weighted combination, and output, through the prediction module of the hierarchical recurrent neural network, a long-term risk value reflecting the cumulative effect of the participants on the evolution path.
[0106] Exemplarily, first, based on a predefined game behavior model, construct a multi-layer game scenario evolution tree. The game behavior model describes the possible decisions of the participants in different states and their probability distributions. Each node of the evolution tree represents a specific game scenario, including the state information of the participants, such as the resource ownership, reputation value, etc., and the decision space in this scenario, such as cooperation, competition, exit, etc. The root node of the tree represents the initial game state, and the child nodes represent the subsequent possible game scenarios. The connection line between the parent node and the child node represents the state transition, and the state transition is determined by the state transition function. The state transition function determines the probability of transitioning from the current state to the next state according to the decisions of the participants and environmental factors. For example, if both parties choose to cooperate, the probability of transitioning to a state where the benefits of both parties increase is relatively high; if one party chooses to compete, the probability of transitioning to a state where the benefit of this party increases but the benefit of the other party decreases is relatively high.
[0107] Next, calculate the transition probability matrix between adjacent nodes. Each element of the transition probability matrix represents the probability of transitioning from one node to another. By analyzing the decision space of the participants and the state transition function, the probability of each state transition can be calculated. For example, assume that at a certain node, participant A has three decision options: cooperate, compete, exit, and participant B also has three decision options: cooperate, compete, exit. Then, the maximum number of child nodes of this node is nine, representing all possible decision combinations. According to the game behavior model, the probability of each decision combination leading to a state transition can be determined, thereby constructing the transition probability matrix.
[0108] Then, the Monte Carlo sampling method is used to generate the evolutionary paths from the root node to the leaf nodes. Monte Carlo sampling is a stochastic simulation method that can simulate the evolutionary process of the game scenario according to the transition probability matrix. Specifically, starting from the root node, a child node is randomly selected as the next state according to the transition probability, and this process is repeated until the leaf node is reached. Repeating the Monte Carlo sampling multiple times can generate multiple evolutionary paths, representing various possibilities of the game process. The node sequences corresponding to these evolutionary paths and the historical state sequences of the participants are stored in the corresponding tree nodes. For example, an evolutionary path may contain three nodes, representing three consecutive game scenarios. Each node stores the state information and decision space of the participants in that scenario, as well as the sequence of state changes of the participants from the root node to that node.
[0109] Subsequently, a dual neural network structure is constructed to evaluate the benefits and risks of the participants. The dual neural network structure includes a graph attention network and a hierarchical recurrent neural network. First, the state vectors in the tree nodes are input into the graph attention network. The graph attention network converts the state vectors into a feature representation matrix through a feature transformation layer. Then, the attention coefficients between the participants are calculated based on the feature representation matrix. The attention coefficients reflect the degree of mutual influence between the participants. The attention coefficients are multiplied by the feature representation matrix to obtain weighted features, and the weighted features are connected with the feature representation matrix through a residual connection, and then transformed through a multi-head attention mechanism to extract the correlation features between the participants. Finally, the correlation features are input into the adaptive pooling layer of the graph attention network for dimensionality reduction and aggregation to obtain the immediate benefit value reflecting the game state of the participants at the current node.
[0110] At the same time, the historical state sequences stored in the tree nodes are grouped according to the participant identifiers, and the grouped sequences are input into the hierarchical recurrent neural network. The recurrent neural network selectively remembers the grouped sequences through a gated unit and outputs the state features at each time step. The attention weights in the time dimension are calculated for the state features to obtain a temporal correlation representation that depicts the state evolution law. At the same time, the attention weights in the participant dimension are calculated for the state features to obtain an interaction correlation representation that depicts the game relationship. The temporal correlation representation and the interaction correlation representation are weighted and combined with the state features, and the long-term risk value reflecting the cumulative effect of the participants on the evolutionary path is output through the prediction module of the hierarchical recurrent neural network.
[0111] For example, assume there are two participating parties, A and B, and the state vector at a certain node is the resource quantities of A and B. The graph attention network can learn the relationship between the resource quantities of A and B, such as a competitive relationship or a cooperative relationship, and output the immediate payoff values of A and B at this node. Meanwhile, the hierarchical recurrent neural network can learn the evolution laws and interaction relationships in the historical state sequences of A and B, such as an increase in the resource quantity of A leading to a decrease in the resource quantity of B, and output the long-term risk values of A and B on this evolution path.
[0112] In this embodiment, by combining the graph attention network and the hierarchical recurrent neural network, it is possible to more comprehensively consider the mutual influences and state evolution laws among the participating parties, thereby more accurately evaluating the payoffs and risks of the participating parties. By generating multiple evolution paths through Monte Carlo sampling, various possibilities in the game process are considered, thus improving the reliability of game prediction. This method can provide more refined game analysis results, helping the participating parties better understand the game situation and formulate more effective game strategies.
[0113] In an alternative implementation manner, based on the temporal variation laws of the immediate payoff values and the long-term risk values, an evaluation function is constructed, and a search algorithm with a hierarchical attention mechanism is used to dynamically update the multi-layer game scenario evolution tree. The hierarchical attention mechanism adaptively adjusts the search strategy according to the evaluation function, determines the sampling weights of the evolution paths, and generates multiple candidate evolution paths with different strategy characteristics, including:
[0114] Calculate the payoff change rate according to the immediate payoff values of adjacent node game states, and calculate the risk change trend according to the long-term risk values of the historical state sequences; perform a weighted combination of the immediate payoff values, the payoff change rate, and the risk change trend to construct a temporal evaluation function; calculate the temporal evaluation scores of each tree node in the multi-layer game scenario evolution tree according to the temporal evaluation function;
[0115] Construct a state attention layer based on the temporal evaluation scores of the tree nodes, output the state importance weights of the tree nodes through the state attention layer, and at the same time input the temporal evaluation scores and the corresponding state importance weights of adjacent tree nodes into the transition attention layer to obtain the importance weights of state transitions. Calculate the complete path sampling probability from the root node to the leaf node based on the state importance weights of the tree nodes and the importance weights of state transitions;
[0116] Introduce a temperature parameter to adjust the path sampling probability, detect the gradient change value of the temporal evaluation function, increase the temperature parameter when the gradient change value is less than the convergence threshold, detect the temporal evaluation scores of the sampled paths, and reduce the temperature parameter when the scores exceed the historical optimum. Take the ratio of the path sampling probability to the temperature parameter as the adjusted sampling probability, and generate candidate evolution paths based on the adjusted sampling probability;
[0117] Maintain a candidate path set of fixed size, calculate the weighted sum of the temporal evaluation scores of all tree nodes on the newly generated path as the total path evaluation score, and when the total evaluation score of the new path is higher than the lowest score in the set, add the new path to the candidate path set and remove the path with the lowest score;
[0118] Calculate the edit distance between paths in the candidate path set to obtain a path diversity index, dynamically adjust the temperature parameter based on the diversity index, increase the temperature parameter to expand the search range when the diversity index is lower than the target threshold, and decrease the temperature parameter to strengthen local search when the diversity index is higher than the target threshold; calculate the path weights based on the total path evaluation scores and sampling probabilities of each candidate path, and update the multi-layer game scenario evolution tree based on the path weights to finally obtain multiple candidate evolution paths with different strategy characteristics.
[0119] Exemplarily, first, the system analyzes adjacent nodes in the multi-layer game scenario evolution tree and calculates the revenue change rate and risk change trend between each pair of nodes. The revenue change rate is reflected by the difference in the immediate revenue values between adjacent nodes, representing the return growth rate of the participating parties at different decision-making stages; the risk change trend is based on the long-term risk values of adjacent nodes to evaluate the long-term uncertainties and risks that the participating parties may face in the game. For example, when the strategy choices of the participating parties change significantly, it may lead to rapid fluctuations in revenue or sharp changes in risk, and the system can identify potential decision-making bottlenecks and turning points through these change patterns.
[0120] Based on the above analysis, the system combines the immediate revenue value, revenue change rate, and risk change trend through weighting to construct a temporal evaluation function. This evaluation function reflects the evolution of the game scenario on the time axis and can comprehensively evaluate the decision-making process of the participating parties and its impact on the overall game. In the multi-layer game scenario evolution tree, each tree node represents a game state, and the temporal evaluation function calculates the temporal evaluation score of each node according to the immediate revenue, revenue change rate, and risk change trend of the node.
[0121] Next, the system constructs a state attention layer and a transition attention layer. The state attention layer is used to calculate the state importance weights of tree nodes, and by analyzing the temporal evaluation scores, it identifies the importance of the current game state to the overall game. For example, at certain decision nodes, the strategy adjustments of some participating parties may have a decisive impact on the entire game process, and these nodes will be assigned higher weights. The transition attention layer further calculates the importance weights of state transitions based on the temporal evaluation scores of adjacent nodes and the corresponding state importance weights. This process helps the system accurately evaluate the transition process from the current node to the next node, thereby providing a more accurate reference for path sampling.
[0122] On this basis, the system adjusts the path sampling probability using temperature parameters to improve the flexibility of the search algorithm. The purpose of adjusting the temperature parameter is to control the exploration scope of the search. When there is no significant change in the temporal evaluation score of path sampling, the temperature parameter will be increased, allowing the system to explore more possible paths and avoid local optima; while when the temporal evaluation score of the sampled path exceeds the historical optimum, the temperature parameter will be decreased to concentrate the search on the local area of the optimal path. Through the dynamic adjustment of the temperature parameter, the system can balance global exploration and local search and optimize the selection of the policy path.
[0123] To maintain path diversity, the system evaluates the path diversity index by calculating the edit distance between each path in the candidate path set. The edit distance measures the degree of difference between two paths. The greater the difference between paths, the more different the policy features are, and the higher the path diversity. The system dynamically adjusts the temperature parameter according to the change of the diversity index. When the diversity index is lower than the target threshold, the temperature parameter increases to expand the search scope; conversely, when the diversity index is higher than the target threshold, the temperature parameter decreases to enhance the accuracy of local search. Through this mechanism, the system can ensure policy diversity while avoiding the situation where path selection becomes overly concentrated on certain similar strategies.
[0124] Each time a new candidate path is generated, the system calculates the weighted sum of the temporal evaluation scores of all tree nodes on the path and uses it as the total evaluation score of the path. If the total evaluation score of the newly generated path exceeds the lowest score in the current candidate path set, the new path is added to the set and the path with the lowest score is removed. This process ensures that the candidate path set always contains the optimal policy path, thus providing high-quality path selection for the game model.
[0125] Finally, the system calculates the path weights based on the total evaluation scores and sampling probabilities of each candidate path and updates the multi-layer game scenario evolution tree according to the path weights. This update process can be continuously carried out to dynamically optimize the game behavior model and generate multiple candidate evolution paths with different policy features, and select the most suitable policy path for the current game environment.
[0126] In this embodiment, by constructing a timing evaluation function and combining a hierarchical attention mechanism, the multi-layer game scenario evolution tree is dynamically updated, effectively optimizing the game behavior model. The system can accurately evaluate the importance of each game node according to the immediate benefits and long-term risk change rules of the participants, dynamically adjust the sampling probability of the strategy path, and avoid premature convergence to the local optimal solution. The adaptive adjustment of the temperature parameter enables the system to balance exploration and exploitation, expands the search scope and improves the path diversity, thereby enhancing the flexibility and accuracy of game decision-making. By maintaining the candidate path set and optimizing it according to the path diversity index, the system ensures that the finally selected strategy path has high diversity and adaptability, and can cope with the complex and changeable game environment. Overall, the system improves the decision-making efficiency and strategy optimization ability of the game model, and can provide effective support and optimization solutions for multi-party games in complex systems.
[0127] In an alternative embodiment, the candidate evolution path is input into a strategy evaluation model constructed by a spatio-temporal graph neural network to generate path evaluation metrics including technology maturity score, resource investment score, and strategy benefit score. Selecting the optimal strategy path according to the path evaluation metrics includes:
[0128] Construct a spatio-temporal heterogeneous graph structure, use technology nodes and resource nodes as heterogeneous graph nodes, set the technology development order relationship as the timing edge, set the resource dependency relationship as the spatial edge, construct a technology node attribute vector based on technology attribute features and timestamp information, construct a resource node attribute vector based on resource type and capacity information, and combine the technology node attribute vector and the resource node attribute vector to form a node feature matrix;
[0129] Input the node feature matrix and the timing edge into a causal convolutional network to extract timing context features, and calculate the feature weights through a timing attention mechanism to obtain enhanced timing features; input the node feature matrix and the spatial edge into a graph attention network to extract the interaction features between nodes, and calculate the feature importance through a spatial attention mechanism to obtain enhanced spatial features; input the timing features and the spatial features into a gated update unit, calculate the gated weights and perform adaptive fusion to obtain a unified feature representation;
[0130] Based on the unified feature representation, use a multi-layer perceptron to perform feature matching with an expert knowledge base to obtain a technology feasibility score; calculate the matching degree between the path resource requirements and the template resources through a graph matching network to obtain a resource consumption score; extract the strategy association features through a graph attention network to obtain a strategy value score;
[0131] Predict the mixture Gaussian distribution parameters based on the unified feature representation, and combine the historical evaluation data. Through variational inference, calculate the posterior distribution to perform Bayesian calibration on the technology feasibility score, resource consumption score, and strategy value score to obtain the calibrated scores;
[0132] Input the optimized objective function of the calibrated scoring input path, and through the beam search algorithm, perform iterative optimization under the constraints of technical feasibility, resource consumption, and policy value, and finally obtain the optimal policy path.
[0133] Exemplarily, first, the system constructs a spatio-temporal heterogeneous graph structure, which includes technology nodes and resource nodes. The technology nodes represent the technology projects or R & D tasks involved in the game scenario, and the resource nodes represent various resources required to implement these technologies, such as funds, equipment, and human resources. The dependency relationships between technology nodes are represented by temporal edges, reflecting the chronological order in the technology development process; while the dependency relationships between resource nodes are represented by spatial edges, indicating the mutual dependencies of different resources. For example, a certain technology node may depend on a specific resource node, and the availability of resources affects the implementation of the technology. Based on these node types, the system constructs a technology node attribute vector for the technology nodes, which includes technical characteristics such as technology maturity, development complexity, and R & D cycle; at the same time, a resource node attribute vector is constructed for the resource nodes, which contains information such as resource type, capacity, and usage period.
[0134] Next, the system combines the attribute vectors of the technology nodes and resource nodes into a node feature matrix, and inputs this matrix together with the temporal edges into a causal convolutional network. The role of the causal convolutional network is to extract temporal context features from time series data, capture the impacts of each technology node and resource node changing over time, and ensure that the system can identify the roles of important nodes at different time points. By introducing a temporal attention mechanism, the system can assign different weights to the features at different time points according to the changes of time nodes, thereby enhancing the feature impacts at important moments and further improving the effectiveness of temporal features.
[0135] In addition, the system inputs the node feature matrix and spatial edges into a graph attention network, which can effectively identify the interactions between technology nodes and resource nodes. Through the spatial attention mechanism, the system assigns different attention weights to them according to the relative importance between nodes, and extracts the most strategically significant node interaction features. This enables the system to focus on the relationships between key resources and technologies in the spatial dimension, thus making accurate predictions about the interactions in the game environment.
[0136] Then, input the obtained temporal features and spatial features into a gated update unit, which adaptively fuses these features by calculating the gating weights, and finally generates a unified feature representation. This unified feature representation is a comprehensive embodiment of temporal features and spatial features, which not only considers the attributes of technology nodes and resource nodes, but also incorporates their dynamic interaction features, and can comprehensively reflect the multi-dimensional information in the game environment.
[0137] Based on this unified feature representation, the system performs feature matching through a multi-layer perceptron and an expert knowledge base to obtain a technical feasibility score. The technical feasibility score measures the feasibility of a certain strategy path in terms of technical implementation, comprehensively considering factors such as the maturity of the technology, development difficulty, and implementation time. At the same time, the system uses a graph matching network to calculate the matching degree between the path resource requirements and the template resources, and obtains a resource consumption score. The resource consumption score reflects the amount of resources required to implement a certain strategy path, helping decision-makers understand the potential resource pressure brought by this path.
[0138] Through the graph attention network, the system further extracts the strategy association features, generates a strategy value score, and evaluates the contribution of the strategy path to achieving the long-term strategy goal. This score takes into account the long-term impact of the strategy, such as factors like market share improvement and technological innovation.
[0139] Subsequently, the system performs Bayesian calibration on the scoring results. First, the system calculates the posterior distribution of the scores through variational inference, and uses historical evaluation data to correct the technical feasibility score, resource consumption score, and strategy value score, making these scores more in line with the actual situation, thereby improving the prediction accuracy. Through Bayesian calibration, the system can eliminate the bias caused by uncertainty and obtain more reliable evaluation results.
[0140] Finally, the system inputs the calibrated scores into the path optimization objective function, and through the beam search algorithm, it performs iterative optimization under the conditions of meeting the constraints of technical feasibility, resource consumption, and strategy value. The beam search algorithm optimizes the strategy selection process through multiple path explorations, and finally obtains the optimal strategy path. This optimal path provides a best implementation plan for decision-makers considering the balance of technical implementation, resource consumption, and strategy value.
[0141] In this embodiment, by combining the construction of a spatio-temporal graph neural network with Bayesian inference, the evaluation and optimization capabilities of strategy paths in multi-party games are effectively improved. Through the fusion of temporal features and spatial features, the system can comprehensively capture the complex relationships among technology, resources, and strategies in the game process, and then optimize the strategy selection. The Bayesian calibration mechanism further improves the accuracy of the scores, making the strategy evaluation more reliable. The beam search algorithm can find the optimal solution among multiple paths, ensuring that the strategy achieves optimal optimization while meeting the balance of technical feasibility, resource consumption, and strategy value. Ultimately, the system can provide more accurate and efficient strategy path selection for decision-makers, enhancing the adaptability and decision-making ability of the system in complex and changing environments, and improving the long-term effect and implementation feasibility of game strategies.
[0142] In an alternative embodiment, based on the optimal policy path, a phased implementation policy deployment plan is generated in combination with a knowledge inference engine, and the feasibility of the policy deployment plan is evaluated and optimized through a distributed verification network, including:
[0143] Construct a policy deployment knowledge base, which includes a rule base of technical dependency rules, resource allocation rules, and schedule control rules, as well as a fact base of mastered technology lists, current available resource status, and historical deployment case data;
[0144] Based on the rule base and the fact base, construct a knowledge inference engine. The knowledge inference engine extracts the technical dependency relationships of each technical node in the optimal policy path to obtain a dependency chain, calculates the earliest implementable time of each technical node in combination with the dependency chain and the current system state, matches the earliest implementable time with the schedule control rules in the rule base to generate a time execution window, and calculates a resource scheduling plan based on the time execution window and the resource allocation rules to form a phased deployment plan;
[0145] Construct a hierarchical verification architecture, which includes a first-level verification component for performing global consistency checks, a second-level verification component for performing technical dependency verification, and a third-level verification component for performing resource constraint verification, and issue verification tasks to each level of verification component according to a hierarchical authorization mechanism;
[0146] Input the verification results of each level of verification component into the Byzantine fault-tolerant consensus algorithm, calculate the credibility scores of each verification result based on reputation weighting, sort the verification results according to the credibility scores, and perform structured processing on the verification opinions with high credibility to obtain optimization suggestions;
[0147] Construct a verification index vector according to the optimization suggestions, input the verification data of technical feasibility, resource matching degree, and schedule controllability into the corresponding index calculation models respectively, and generate a quantitative verification evaluation result;
[0148] Input the verification evaluation result into the policy optimization network to generate a targeted optimization adjustment plan, apply the optimization adjustment plan to the execution plan for update, re-input the updated execution plan into the hierarchical verification architecture for evaluation, and output the final policy deployment plan when the verification evaluation result meets the preset standard, and return the execution plan to the policy optimization network for iterative optimization when the verification evaluation result does not meet the standard.
[0149] Exemplarily, first, a policy deployment knowledge base is constructed. This knowledge base consists of two parts: a rule base and a fact base. The rule base contains technology dependency rules, such as "Technology A must be deployed after Technology B is deployed"; resource allocation rules, such as "Technology C requires Y units of X-type resources"; and progress control rules, such as "Technology D must be deployed before time Z". The fact base contains a list of mastered technologies, such as mastered technologies A, B, and C; the current available resource status, such as having 10 units of X-type resources and 5 units of Y-type resources; and historical deployment case data, such as the average deployment time of Technology A in past cases was 3 days.
[0150] Next, a knowledge inference engine is constructed based on the knowledge base. Taking the deployment of a policy containing technologies A, B, C, and D as an example. First, the inference engine extracts the dependency relationships of each technology node according to the rule base. For example, Technology C depends on Technologies A and B. After obtaining the dependency chain, combined with the current system state (mastered technologies, available resources), calculate the earliest implementable time for each technology node. For example, if the deployment times of Technologies A and B are 1 day and 2 days respectively, then the earliest implementable time for Technology C is the 3rd day. Then, match the earliest implementable time with the progress control rules to generate a time execution window. For example, if Technology C must be deployed before the 5th day, then its time execution window is from the 3rd day to the 5th day. Finally, calculate the resource scheduling plan based on the time execution window and the resource allocation rules. For example, allocate 2 units of X-type resources to Technology C on the 3rd day, and finally form a phased deployment plan. For example, deploy Technologies A and B in the first phase, deploy Technology C in the second phase, and deploy Technology D in the third phase.
[0151] Then, a hierarchical verification architecture is constructed. This architecture contains three levels of verification components: the first level performs global consistency checks, such as checking whether the entire deployment plan meets the overall goal; the second level performs technology dependency verification, such as checking whether the deployment of Technology C satisfies its dependency relationships; the third level performs resource constraint verification, such as checking whether the resource allocation exceeds the limit of available resources. Suppose the first-level verification component discovers a deviation between the deployment plan and the overall goal, the second-level verification component discovers that the dependency relationship of Technology C is not fully satisfied, and the third-level verification component discovers that the resource allocation plan is feasible. These verification tasks will be distributed to the corresponding verification components according to the hierarchical authorization mechanism.
[0152] The verification results of each level of verification components are input into the Byzantine fault-tolerant consensus algorithm. This algorithm calculates the credibility scores of each verification result based on reputation weighting. For example, the reputation of the first-level verification component is 0.8, the second level is 0.9, and the third level is 0.7. Suppose the first-level verification result is "inconsistent", the second level is "partially dependent satisfied", and the third level is "resource satisfied". Then, according to the reputation weighting calculation, the final credibility scores are 0.8*(-1)=-0.8, 0.9*0.5=0.45, and 0.7*1=0.7 respectively. Sort the verification results according to the credibility scores to obtain the sorting of the third level, the second level, and the first level. Structurally process the verification opinions with high credibility to obtain optimization suggestions, such as "prioritize satisfying the dependencies of Technology C".
[0153] Construct a verification index vector according to the optimization suggestions, input the verification data of technical feasibility, resource matching degree, and schedule controllability into the corresponding index calculation models respectively, and generate a quantitative verification evaluation result. For example, the technical feasibility score is 0.8, the resource matching degree score is 0.9, and the schedule controllability score is 0.7.
[0154] Input the verification evaluation result into the policy optimization network to generate a targeted optimization and adjustment plan. For example, advance the deployment time of Technology C to the second day. Apply the optimization and adjustment plan to the execution plan for update, and re-enter the updated execution plan into the hierarchical verification architecture for evaluation. If the verification evaluation result meets the preset criteria (for example, the scores of all indicators are greater than 0.8), then output the final policy deployment plan. If not up to standard, return the execution plan to the policy optimization network for iterative optimization until the preset criteria are met.
[0155] In this embodiment, by automatically generating a deployment plan through the knowledge inference engine and quickly evaluating and optimizing the plan through the distributed verification network, the policy deployment cycle can be significantly shortened. The hierarchical verification architecture can comprehensively verify the deployment plan from multiple dimensions, and the Byzantine fault-tolerant consensus algorithm can effectively identify and filter out incorrect verification results, thereby improving the accuracy and reliability of the deployment plan. The iterative optimization mechanism can dynamically adjust the deployment plan according to the actual situation, so as to better adapt to the complex and changeable deployment environment.
[0156] Figure 2 This is a schematic structural diagram of a dynamic evolution and strategy deduction optimization system for space domain game behaviors according to an embodiment of the present invention, as Figure 2 shown, the system includes:
[0157] The first unit is used to obtain the historical behavior data of space domain participants, perform time series feature analysis on the historical behavior data through a deep neural network to obtain the periodic change law of participants' behaviors, construct a feature extraction model using a distributed learning framework based on the periodic change law, generate a multi-dimensional feature vector, construct a dynamic relationship graph of participants according to the multi-dimensional feature vector, and input the dynamic relationship graph into a preset multi-agent reinforcement learning model. The multi-agent reinforcement learning model adopts an encoder structure and a multi-layer self-attention mechanism, and outputs a game behavior model including an adaptive utility function and dynamic decision rules;
[0158] The second unit is used to construct a multi-layer game scenario evolution tree according to the game behavior model, where each tree node contains the state information and decision space of the participant, input the state information and decision space into a dual neural network structure, calculate the immediate benefit value and long-term risk value of the participant in each game scenario respectively, construct an evaluation function based on the time series change law of the immediate benefit value and long-term risk value, use a search algorithm with a hierarchical attention mechanism to dynamically update the multi-layer game scenario evolution tree, and the hierarchical attention mechanism adaptively adjusts the search strategy according to the evaluation function to determine the sampling weight of the evolution path and generate multiple candidate evolution paths with different strategy characteristics;
[0159] The third unit is used to input the candidate evolution paths into a strategy evaluation model constructed by a spatio-temporal graph neural network, generate path evaluation indicators including technology maturity score, resource investment score and strategy benefit score, select the optimal strategy path according to the path evaluation indicators, generate a phased implementation strategy deployment plan based on the optimal strategy path in combination with a knowledge reasoning engine, and evaluate and optimize the feasibility of the strategy deployment plan through a distributed verification network.
[0160] In the third aspect of the embodiments of the present invention,
[0161] There is provided an electronic device, including:
[0162] A processor;
[0163] A memory for storing instructions executable by the processor;
[0164] Wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.
[0165] In the fourth aspect of the embodiments of the present invention,
[0166] There is provided a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.
[0167] The present invention may be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for performing various aspects of the present invention.
[0168] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A method for dynamic evolution and strategy deduction optimization of game behavior in space, characterized in that: include: Obtain historical behavior data of participants in the space field, the historical behavior data refers to the decision-making records, resource allocation and behavior patterns of the participants in the past period of time, perform time series feature analysis on the historical behavior data through a deep neural network, obtain the periodic change law of the participants' behavior, and build a feature extraction model based on the periodic change law using a distributed learning framework to generate a multi-dimensional feature vector. According to the multi-dimensional feature vector, a dynamic relationship graph of the participants is built, wherein the nodes represent the various participants in the space field, and the edges represent the interaction relationship between the participants, and the interaction relationship includes cooperation, competition and resource sharing. The dynamic relationship graph is input into a preset multi-agent reinforcement learning model, and the multi-agent reinforcement learning model adopts an encoder structure and a multi-layer self-attention mechanism, and outputs a game behavior model including an adaptive utility function and dynamic decision-making rules; A multi-layer game scenario evolution tree is constructed according to the game behavior model, wherein each tree node contains the state information and decision space of the participants, wherein the state information includes the resource ownership and reputation value, and the decision space includes cooperation, competition and exit. The state information and decision space are input into a dual neural network structure, and the immediate benefit value and long-term risk value of the participants in each game scenario are respectively calculated. Based on the temporal variation law of the immediate benefit value and the long-term risk value, an evaluation function is constructed, and a search algorithm with a hierarchical attention mechanism is used to dynamically update the multi-layer game scenario evolution tree. The hierarchical attention mechanism adaptively adjusts the search strategy according to the evaluation function, determines the sampling weight of the evolution path, and generates multiple candidate evolution paths with different strategy characteristics; The candidate evolution path is input into the strategy evaluation model constructed by the spatiotemporal graph neural network to generate path evaluation indicators including technology maturity score, resource input score and strategy benefit score, the optimal strategy path is selected according to the path evaluation indicators, and based on the optimal strategy path, a strategy deployment plan for phased implementation is generated in combination with the knowledge reasoning engine, and the feasibility of the strategy deployment plan is evaluated and optimized through a distributed verification network.
2. The method according to claim 1, characterized in that: A dynamic relationship diagram of the participants is constructed according to the multi-dimensional feature vector, and the dynamic relationship diagram is input into a preset multi-agent reinforcement learning model. The multi-agent reinforcement learning model adopts an encoder structure and a multi-layer self-attention mechanism, and outputs a game behavior model including an adaptive utility function and dynamic decision rules, including: Generate an initial dynamic relationship graph based on the multidimensional feature vector, where nodes represent participants and graph edges represent the interaction relationship between participants; embed nodes with features through a deep neural network to obtain node representation vectors, and use a temporal convolutional network to extract interaction history to obtain dynamic pattern features; combine the node representation vector with the dynamic pattern features to calculate edge weights, and update the topological structure of the dynamic relationship graph based on the edge weights and node representation vectors; The dynamic relationship graph is input into a multi-agent reinforcement learning model, wherein the multi-agent reinforcement learning model comprises an encoder network and a strategy network; the encoder network encodes the dynamic relationship graph through a node feature layer, a relationship feature layer and a time series feature layer to obtain a state representation vector; the strategy network generates an action probability distribution of each agent based on the state representation vector; Constructing a multi-agent value network, predicting the state value and action value of each agent based on the state representation vector and the action probability distribution; inputting the state value and action value into the utility function generation module to construct a utility function including individual utility and group utility; Calculate the reward signal of each agent based on the utility function, and input the reward signal into the policy optimization module; the policy optimization module uses the policy gradient method to update the policy network parameters to achieve dynamic optimization of the agent decision rules; Adopting the experience replay mechanism to store the interaction data of the intelligent agents, and synchronously updating the value network and utility function parameters based on the interaction data; realizing the joint optimization of the strategy network, value network and utility function through multi-agent collaborative learning; The optimized strategy network, value network and utility function are integrated into a model to output a game behavior model that includes an adaptive utility function and dynamic decision rules.
3. The method according to claim 2, characterized in that Constructing a multi-agent value network, predicting the state value and action value of each agent based on the state representation vector and action probability distribution; inputting the state value and action value into the utility function generation module, and constructing a utility function containing individual utility and group utility includes: Construct a state value network and an action value network. The state value network adopts a three-layer fully connected network structure, and the state representation vector is input to obtain the state value prediction value; the action value network adopts a double-branch structure, the first branch processes the state representation vector, and the second branch processes the action probability distribution. The two branch features are fused through the outer product operation and then the action value prediction value is obtained through the fully connected layer; Calculate a time series difference target value, the time series difference target value is calculated based on the immediate reward value, the state value prediction value and the action value prediction value; construct a state value loss function by the difference between the time series difference target value and the state value prediction value, and construct an action value loss function by the difference between the time series difference target value and the action value prediction value; update the value network parameters based on the state value loss function and the action value loss function to obtain optimized state value and action value; The optimized action value is used to represent the individual utility term, and the group utility term is obtained by weighted calculation of the optimized state value and the action value of the neighborhood agent; an initial utility function is constructed by combining the individual utility term and the group utility term; The gradient value of the initial utility function with respect to the weight coefficient is calculated, and based on the gradient value, the weight coefficients of the individual utility term and the group utility term are updated using the gradient descent method to obtain the final utility function.
4. The method according to claim 1, characterized in that: A multi-layer game scenario evolution tree is constructed according to the game behavior model, wherein each tree node contains the state information and decision space of the participants, and the state information and decision space are input into the dual neural network structure to calculate the immediate benefit value and long-term risk value of the participants in each game scenario respectively, including: A multi-layer game scenario evolution tree is constructed based on the dynamic decision rules in the game behavior model, and the state vectors and decision space information of the participants are stored in the tree nodes; the transition probability matrix between adjacent nodes is calculated based on the state transition function; Monte Carlo sampling is performed according to the transition probability matrix to generate an evolution path from the root node to the leaf node; the node sequence corresponding to the evolution path and the historical state sequence of the participants are stored as feature inputs in each tree node; Construct a dual neural network structure, including a graph attention network and a hierarchical recurrent neural network, input the state vector in the tree node into the graph attention network, generate a feature representation matrix of the participants through the feature transformation layer, calculate the attention coefficient between the participants based on the feature representation matrix, multiply the attention coefficient with the feature representation matrix to obtain a weighted feature, and perform a residual connection between the weighted feature and the feature representation matrix, transform the residually connected features through a multi-head attention mechanism, extract the correlation features between the participants, input the correlation features into the adaptive pooling layer of the graph attention network for dimensionality reduction aggregation, and obtain an instant benefit value reflecting the game state of the participants at the current node; The historical state sequences stored in the tree nodes are grouped according to the participant identifiers, and the grouped sequences are input into a hierarchical recurrent neural network. The recurrent neural network selectively memorizes the grouped sequences through a gating unit, outputs the state features of each time step, calculates the attention weight of the time dimension for the state features, and obtains a temporal correlation representation that describes the state evolution law. At the same time, the attention weight of the participant dimension is calculated for the state features to obtain an interactive correlation representation that describes the game relationship. The temporal correlation representation and the interactive correlation representation are weightedly combined with the state features, and the long-term risk value reflecting the cumulative effect of the participants on the evolution path is output through the prediction module of the hierarchical recurrent neural network.
5. The method according to claim 1, characterized in that Based on the temporal variation law of the instant return value and the long-term risk value, an evaluation function is constructed, and a search algorithm with a hierarchical attention mechanism is used to dynamically update the multi-layer game scenario evolution tree. The hierarchical attention mechanism adaptively adjusts the search strategy according to the evaluation function, determines the sampling weight of the evolution path, and generates multiple candidate evolution paths with different strategy characteristics, including: Calculate the rate of change of benefits according to the instant benefit value of the game state of adjacent nodes, and calculate the risk change trend according to the long-term risk value of the historical state sequence; construct a time series evaluation function by weighted combination of the instant benefit value, the rate of change of benefits and the risk change trend; calculate the time series evaluation score of each tree node in the multi-layer game scenario evolution tree according to the time series evaluation function; A state attention layer is constructed based on the temporal evaluation scores of tree nodes. The state importance weights of tree nodes are output through the state attention layer. At the same time, the temporal evaluation scores and corresponding state importance weights of adjacent tree nodes are input into the transfer attention layer to obtain the importance weights of state transfers. The complete path sampling probability from the root node to the leaf node is calculated based on the state importance weights of the tree nodes and the importance weights of state transfers. The temperature parameter is introduced to adjust the path sampling probability, and the gradient change value of the timing evaluation function is detected. When the gradient change value is less than the convergence threshold, the temperature parameter is increased, and the timing evaluation score of the sampling path is detected. When the score exceeds the historical optimal, the temperature parameter is reduced. The ratio of the path sampling probability to the temperature parameter is used as the adjusted sampling probability, and the candidate evolution path is generated based on the adjusted sampling probability. Maintain a fixed-size candidate path set, calculate the weighted sum of the temporal evaluation scores of all tree nodes on the newly generated path as the total evaluation score of the path, and when the total evaluation score of the new path is higher than the lowest score in the set, add the new path to the candidate path set and remove the path with the lowest score; The edit distance between paths in the candidate path set is calculated to obtain the path diversity index, and the temperature parameter is dynamically adjusted based on the diversity index. When the diversity index is lower than the target threshold, the temperature parameter is increased to expand the search range. When the diversity index is higher than the target threshold, the temperature parameter is reduced to strengthen the local search. The path weight is calculated according to the total evaluation score and sampling probability of each candidate path, and the multi-layer game scenario evolution tree is updated based on the path weight, and finally multiple candidate evolution paths with different strategic characteristics are obtained.
6. The method according to claim 1, characterized in that The candidate evolution path is input into the strategy evaluation model constructed by the spatiotemporal graph neural network to generate a path evaluation index including a technology maturity score, a resource input score and a strategy benefit score. The optimal strategy path is selected according to the path evaluation index, including: Constructing a spatiotemporal heterogeneous graph structure, taking technology nodes and resource nodes as heterogeneous graph nodes, setting the technology development sequence relationship as a temporal edge, setting the resource dependency relationship as a spatial edge, constructing a technology node attribute vector based on technology attribute characteristics and timestamp information, constructing a resource node attribute vector based on resource type and capacity information, and combining the technology node attribute vector and resource node attribute vector to form a node feature matrix; The node feature matrix and the temporal edge are input into the causal convolutional network to extract the temporal context features, and the feature weights are calculated through the temporal attention mechanism to obtain enhanced temporal features; the node feature matrix and the spatial edge are input into the graph attention network to extract the node interaction features, and the feature importance is calculated through the spatial attention mechanism to obtain enhanced spatial features; the temporal features and spatial features are input into the gated update unit, the gated weights are calculated, and adaptive fusion is performed to obtain a unified feature representation; Based on the unified feature representation, a multi-layer perceptron and an expert knowledge base are used to perform feature matching to obtain a technical feasibility score; a graph matching network is used to calculate the matching degree between the path resource requirements and the template resources to obtain a resource consumption score; and a graph attention network is used to extract strategy-related features to obtain a strategy value score; Based on the unified feature representation, the mixed Gaussian distribution parameters are predicted, and in combination with the historical evaluation data, the posterior distribution is calculated by variational inference to perform Bayesian calibration on the technical feasibility score, resource consumption score, and strategic value score to obtain a calibrated score; The calibrated score is input into the path optimization objective function, and iterative optimization is performed through a beam search algorithm under the constraints of technical feasibility, resource consumption and strategic value, and finally the optimal strategic path is obtained.
7. The method according to claim 1, characterized in that Based on the optimal strategy path, a strategy deployment plan implemented in stages is generated in combination with a knowledge reasoning engine, and the feasibility of the strategy deployment plan is evaluated and optimized through a distributed verification network, including: Constructing a policy deployment knowledge base, the policy deployment knowledge base including a rule base of technology dependency rules, resource allocation rules and progress control rules, and a fact base of acquired technology lists, current available resource status and historical deployment case data; A knowledge reasoning engine is constructed based on the rule base and the fact base. The knowledge reasoning engine extracts the technical dependency relationship of each technical node in the optimal strategy path to obtain a dependency chain. The earliest implementable time of each technical node is calculated by combining the dependency chain and the current system state. The earliest implementable time is matched with the progress control rule in the rule base to generate a time execution window. A resource scheduling plan is calculated based on the time execution window and the resource allocation rule to form a phased deployment plan. Construct a hierarchical verification architecture, which includes a first-level verification component that performs global consistency checks, a second-level verification component that performs technology dependency verification, and a third-level verification component that performs resource constraint verification, and distributes verification tasks to verification components at each level according to a hierarchical authorization mechanism; The verification results of verification components at all levels are input into the Byzantine Fault Tolerant consensus algorithm, and the credibility score of each verification result is calculated based on the reputation weight. The verification results are sorted according to the credibility score, and the highly credible verification opinions are structured to obtain optimization suggestions. Construct a verification indicator vector based on the optimization suggestions, input the verification data of technical feasibility, resource matching and schedule controllability into the corresponding indicator calculation model, and generate a quantitative verification evaluation result; The verification and evaluation results are input into the policy optimization network to generate targeted optimization and adjustment plans, which are applied to the execution plan for updating. The updated execution plan is re-input into the hierarchical verification architecture for evaluation. When the verification and evaluation results meet the preset standards, the final policy deployment plan is output. When the verification and evaluation results do not meet the standards, the execution plan is returned to the policy optimization network for iterative optimization.
8. A system for dynamic evolution and strategy deduction optimization of game behavior in space, used to implement the method described in any one of claims 1 to 7, characterized in that: include: The first unit is used to obtain historical behavior data of participants in the space field, perform time series feature analysis on the historical behavior data through a deep neural network, obtain the periodic change law of the participants' behavior, and build a feature extraction model based on the periodic change law using a distributed learning framework to generate a multi-dimensional feature vector. A dynamic relationship diagram of the participants is built according to the multi-dimensional feature vector, and the dynamic relationship diagram is input into a preset multi-agent reinforcement learning model. The multi-agent reinforcement learning model adopts an encoder structure and a multi-layer self-attention mechanism to output a game behavior model including an adaptive utility function and dynamic decision-making rules; The second unit is used to construct a multi-layer game scenario evolution tree according to the game behavior model, wherein each tree node contains the state information and decision space of the participants, input the state information and decision space into the dual neural network structure, respectively calculate the immediate benefit value and long-term risk value of the participants in each game scenario, and construct an evaluation function based on the temporal variation law of the immediate benefit value and the long-term risk value. The multi-layer game scenario evolution tree is dynamically updated by using a search algorithm with a hierarchical attention mechanism. The hierarchical attention mechanism adaptively adjusts the search strategy according to the evaluation function, determines the sampling weight of the evolution path, and generates multiple candidate evolution paths with different strategy characteristics; The third unit is used to input the candidate evolution path into the strategy evaluation model constructed by the spatiotemporal graph neural network, generate path evaluation indicators including technology maturity score, resource input score and strategy benefit score, select the optimal strategy path according to the path evaluation indicators, and generate a strategy deployment plan for phased implementation based on the optimal strategy path in combination with the knowledge reasoning engine, and evaluate and optimize the feasibility of the strategy deployment plan through a distributed verification network.
9. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method described in any one of claims 1 to 7.
10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Intelligent decision-making method for military confrontation games under incomplete information conditions
CN112329348A
Network spoofing defense strategy optimization method and system based on intelligent real-time game
CN117220995A