A consensus decision question and answer system based on multi-al intelligent agent game
By using multi-domain information aggregation, interaction effect deduction, policy coherence calibration, and decentralized adaptive mechanisms, policy biases in multi-agent systems are identified and corrected, achieving efficient consensus decision-making in complex environments and improving the system's robustness and adaptability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ZHEJIANG ANYIXIN TECH CO LTD
- Filing Date
- 2025-09-08
- Publication Date
- 2026-04-17
AI Technical Summary
Existing multi-agent intelligent systems are prone to conflict and interference when agents interact in complex environments, leading to the accumulation of policy deviations, inconsistent or invalid system outputs, and a lack of stability and coordination.
Employing a multi-domain information aggregation module, an interaction effect inference module, a strategy coherence calibration unit, a decentralized behavior adaptive mechanism, and an aggregated strategy discrimination module, this system identifies and quantifies non-cooperative strategy biases through a multi-dimensional interaction tensor model and inverse game theory, generates customized correction suggestions, and achieves strategy coordination through gradual adjustments.
It improves the ability to identify and quantify non-cooperative policy biases, enhances the robustness and adaptability of multi-agent systems in dynamic environments, and ensures that the system remains stable and coordinated under external disturbances.
Smart Images

Figure CN121166870B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, specifically to a consensus-based decision-making question-and-answer system based on multi-AI agent game theory. Background Technology
[0002] AI is typically applied to collaborative decision-making and group reasoning in multi-agent intelligent systems. Its fundamental role is to comprehensively utilize the information and decision-making preferences of different individuals through the interaction between multiple intelligent agents. With the development of artificial intelligence, multi-agent collaborative frameworks have been applied in scenarios such as resource scheduling, automatic question answering, and complex task planning, demonstrating their significant value in improving the overall system intelligence level and adapting to dynamic environments.
[0003] However, existing technologies still have limitations in scenarios involving the interaction of multiple agent policies. Especially in complex environments, when potential conflicts or interference arise between agent policies, biases often accumulate, creating vulnerabilities in the overall optimal policy and leading to inconsistent or even invalid system outputs. Traditional methods cannot fully reveal and eliminate the negative impacts caused by agent interaction effects. Therefore, in highly dynamic and uncertain scenarios, existing multi-agent systems struggle to maintain stability and coordination. Summary of the Invention
[0004] To achieve the above objectives, the present invention provides the following technical solution: a consensus-based decision-making question-answering system based on multi-AI agent game theory, comprising:
[0005] A multi-domain information aggregation module is used to receive diverse policy information and external environment situation data from multiple intelligent agents; the policy information and external environment situation data are preprocessed to form an input data stream with a unified format;
[0006] An interaction effect deduction module is used to perform in-depth analysis of the input data stream to identify and quantify the potential mutual influence relationships between the strategies of each agent; by constructing a multidimensional interaction tensor model, it calculates the expected feedback of the combination of agent strategies under different environmental situations, deduces the non-cooperative strategy deviation caused by each agent strategy, and generates an interaction analysis report containing quantitative indicators of the non-cooperative strategy deviation.
[0007] A policy coherence calibration unit receives the interaction analysis report and generates customized cooperative optimization trajectory correction suggestions for each agent based on the non-cooperative policy deviations and their quantitative indicators indicated in the report. This guides the agents to adjust their existing policies to weaken negative interference effects, enhance positive cooperative effects, achieve overall policy coordination, and thus produce the optimal set of correction paths.
[0008] A decentralized behavior adaptation mechanism is logically connected to each agent through network channels, receives collaborative optimization trajectory correction suggestions, and performs incremental policy adjustments locally on each agent, so that the policies of each agent respond to the overall optimization needs of the system and internalize the correction suggestions as part of its own policy generation logic.
[0009] An aggregation strategy discrimination module, logically connected to a distributed behavior adaptation mechanism and a multi-domain information aggregation module, is used to periodically monitor the aggregation status and performance of the adjusted multi-agent strategy and evaluate the current multi-agent aggregation strategy.
[0010] As a further technical solution, the high-order tensor decomposition and causal inference process in the interaction effect inference module, for historical strategy interaction data and dynamic environmental change patterns, firstly reorganizes the multidimensional interaction structure into dynamic subsets according to the similarity of environmental situations, ensuring that environmental constraints remain consistent within each subset; within each subset, constrained tensor decomposition is performed on the interaction tensor containing the strategy combination dimension and the time evolution dimension, so that the strategy combination dimension and the time evolution dimension are represented by mutually independent components, while the environmental situation is retained as an inherent constraint outside the reconstruction error term, to ensure that the strategy action components identified under the same environmental situation are not confused by environmental variables; based on this decomposition, a counterfactual comparison is constructed within the same environmental subset: using the overall system performance measurement of the actual strategy combination as a benchmark, virtual combinations formed by a certain strategy are removed one by one and the overall performance is calculated. The difference is defined as the direct causal contribution of the target policy under a unified scale. The above calculation is repeated in each environmental subset to obtain the contribution trajectory of the same policy under different environmental constraints. Based on the fluctuation amplitude and temporal consistency of the trajectory, the environment-dependent bias and the inherent conflict between policies are separated. Finally, the correlation of the contribution trajectory on the time axis is used to construct a cross-policy transmission chain to quantify the independent proportion of environmental coupling effect and inherent conflict between policies in non-cooperative policy bias. When the transmission chain is monotonically decaying and the residual term is bounded in a given environmental subset, the interaction mode in that environment is considered interpretable and stable for subsequent calibration. This step is limited to multidimensional interaction tensor model, higher-order tensor decomposition, causal inference and environmental subset reorganization, which improves the source of non-cooperative policy bias from statistical correlation to a causally separable structured representation.
[0011] As a further technical solution, the interactive analysis report, while outputting quantitative indicators of non-cooperative strategy deviation, simultaneously provides source tracing entries and risk scenario prediction entries based on counterfactual differences. The source tracing entries are jointly determined by the extreme value positions and extreme value signs of the direct causal contribution of each strategy on different environmental subsets, which are used to define the dominant and passive strategies that cause the deviation. The risk scenario prediction entries are jointly characterized by the maximum propagation amplitude and propagation lag of the cross-strategy transmission chain, which are used to define the secondary deviation amplification path that may occur under given environmental disturbances.
[0012] As a further technical solution, when applying inverse game theory, the policy coherence calibration unit constructs an individual response module for each agent, which includes its current policy state, decision preference weights, and objective function settings. Under the condition of applying preliminary correction suggestions, the optimal response policy of the agent is solved based on the individual response module, and the consistency measure between the solved optimal response policy and the preliminary correction suggestions is used as the acceptance index. The juxtaposition of the optimal response policies of each agent is regarded as a candidate response combination and re-input into the interaction effect inference module to infer new interaction effects and possible secondary non-cooperative policy deviations under the same environmental situation. If the secondary deviation exceeds a preset threshold, the preliminary correction suggestions are updated backtrackingly based on the impact of the deviation on the objective function, and the above solution and inference are repeated until the secondary deviation converges within the threshold, and a stable optimization scheme after predicting the actual response is output.
[0013] As a further technical solution, the individual response module, while keeping the current strategy state, decision preference weights, and objective function settings unchanged, measures the deviation between the optimal response strategy and the initial correction suggestion. The deviation is categorized into three ranges: acceptable, requiring fine-tuning, and unacceptable, as the adjustment range for the collaborative optimization trajectory. When the secondary non-cooperative strategy deviation caused by the candidate response combination in the interaction effect deduction shows a decreasing trend but does not exceed the threshold, the calibration unit only compresses the magnitude of the suggestion and retains the original direction. When the deviation crosses the threshold, the calibration unit, based on the source tracing results in the interaction analysis report, prioritizes reversing the direction or redistributing the intensity of the correction suggestion corresponding to the dominant strategy to accelerate convergence and avoid the continuous amplification of secondary deviations.
[0014] As a further technical solution, the dynamic anti-disturbance evaluation index set introduced by the aggregation strategy discrimination module, under the process of the discrimination module triggering the interaction effect inference module, injecting simulated disturbance signals of preset type and intensity into the current environmental situation data, and constructing a new environmental situation containing disturbance variables, infers and evaluates the strategy combination after adjustment by the distributed behavior adaptive mechanism. The index set is composed of three core dimensions: the maximum offset amplitude of the non-cooperative strategy deviation value during the disturbance injection, the offset stabilization time, and the deviation recovery rate. Specifically, the deviation quantification index before the disturbance injection is used as the benchmark level, and the maximum positive or negative offset of the deviation value relative to the benchmark after the disturbance injection is recorded as the maximum offset amplitude. The time required for the deviation value to enter the stable range from the fluctuation state after first crossing the benchmark is defined as the offset stabilization time. The unit time change of the deviation value from the maximum offset to the benchmark is defined as the deviation recovery rate. The above three dimensions together reflect the maintenance ability and recovery characteristics of the aggregation strategy under the disturbance dynamic situation, and serve as the basis for determining whether the preset aggregation strategy evolution state has been reached.
[0015] As a further technical solution, when the dynamic anti-disturbance evaluation index set is applied, it compares the current three-dimensional index with the preset threshold range or target curve, and makes a comprehensive evaluation in combination with the system's performance in simulated complex situations: when the maximum offset amplitude is within the threshold range and the offset stabilization time does not exceed the upper limit defined by the target curve, and the deviation recovery rate is not lower than the lower limit defined by the target curve, it is determined that the aggregation strategy has reached the preset evolution state under this type of disturbance; when any dimension is not satisfied, the evaluation conclusion is fed back to the strategy coherence calibration unit and the distributed behavior adaptive mechanism, triggering the update of the cooperative optimization trajectory correction suggestion and the continued execution of local progressive adjustment, so that the system gradually reaches the robust range defined by the index set under the same disturbance potential.
[0016] As a further technical solution, the distributed behavior adaptive mechanism adopts a progressive strategy adjustment for the execution of collaborative optimization trajectory correction suggestions: each agent locally follows the trajectory indicated by the correction suggestions, and internalizes the suggestions into part of the strategy generation logic through multiple small-step parameter updates; during the update process, the periodic evaluation results of the aggregation strategy discrimination module are used as the basis for stopping or continuing. When the dynamic anti-disturbance evaluation index set shows that the aggregation strategy has reached the preset evolution state, the local progressive adjustment of the current round is stopped and the monitoring and maintenance phase is entered; when the evaluation fails to meet the standard, the existing adjustment direction is retained and a new round of correction suggestions from the strategy coherence calibration unit is awaited, thereby maintaining a consistent convergence target between the global and local levels.
[0017] This invention provides a consensus-based decision-making question-answering system based on multi-AI agent game theory, which has the following beneficial effects:
[0018] 1. This invention, by setting up a processing mechanism that combines multi-domain information convergence with interaction effect deduction, realizes the characterization of the causal relationship of the interaction of multiple agents in complex environments, improves the ability to identify and quantify potential non-cooperative policy biases, and overcomes the shortcomings of existing technologies that can only perform rough analysis based on statistical correlation and cannot accurately reveal the source of negative interference.
[0019] 2. This invention introduces the joint deduction of inverse game theory and multi-objective optimization into the policy coherence calibration unit, thereby realizing iterative prediction of the agent's behavioral response to the correction proposal and improving the reliability of the consistency between the correction path and the actual behavior.
[0020] 3. By constructing a dynamic anti-disturbance evaluation index set and combining it with an aggregation strategy discrimination module, this invention achieves quantitative monitoring of the stability and recovery capability of multi-agent aggregation strategies under disturbed dynamic potentials. This improves the robustness and adaptability of the system in highly dynamic environments and overcomes the technical problem of existing technologies lacking globally consistent evaluation standards, which leads to the optimal strategy being prone to vulnerabilities under external interference. Attached Figure Description
[0021] Figure 1 This is a schematic diagram of the process of the present invention. Detailed Implementation
[0022] The technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0023] Example 1: Reference Figure 1 The purpose of this invention is to provide a consensus-based decision-making question-and-answer system based on multi-AI agent game theory, which solves the problems of information dispersion, opaque interaction effects, accumulated policy bias, and insufficient stability of aggregated policies in existing multi-agent decision-making processes. This invention designs a multi-domain information aggregation module, an interaction effect inference module, a policy coherence calibration unit, a decentralized behavior adaptive mechanism, and an aggregated policy discrimination module to form a complete system for multi-agent policy generation, adjustment, and evaluation. This enables efficient consensus-based decision-making among different agents in complex environments, thereby significantly improving the overall system's coordination and robustness in the face of dynamic tasks and external disturbances.
[0024] The multi-domain information aggregation module undertakes the core task of unifying the policy information and external environmental situation data from multiple agents. The policy information generated by each agent in different contexts includes behavioral decision sequences, decision preference weights, policy confidence scores, and objective function settings, while the external environmental situation data involves task complexity indicators, available resource constraints, external interference signal strength, and the current perception state of other agents. To avoid the distortion of interactive analysis caused by inconsistencies in the dimension, format, and semantics of data from different sources, this module removes redundant or noisy information through data cleaning, ensures consistent numerical scales through standardization, unifies the data structure through format conversion, and maps different types of data into directly computable high-dimensional vectors through feature vectorization. Finally, it forms an input data stream with unified dimensions and comparable semantics, providing a reliable foundation for subsequent interactive inference. The multi-level preprocessing process ensures the standardization of the input and enhances the overall system's adaptability to complex data environments by reducing information bias.
[0025] In this invention, the interaction effect inference module is used to reveal the potential interaction relationships between agent strategies. Based on the preprocessed input data stream, this module constructs a multidimensional interaction tensor model and captures the coupling effects of different agent strategies in the temporal and spatial dimensions through high-order tensor decomposition and causal inference methods. Specifically, the system first reorganizes historical strategy interaction data and dynamic environmental change patterns into multiple dynamic subsets divided according to the similarity of environmental situations. Within each subset, a constraint tensor decomposition is performed, so that the strategy combination dimension and the time evolution dimension are independently extracted and the constraint features of environmental conditions are preserved, thereby forming an environmentally conditional strategy interaction pattern. On this basis, the module calculates the direct causal contribution of each strategy combination under the same environmental situation through causal inference, further separating environment-dependent biases from inherent strategy conflicts. The inference results not only generate an interaction analysis report containing quantitative indicators of non-cooperative strategy biases, but also locate the source of biases through source tracing analysis and predict potential risk scenarios by combining cross-subset comparative analysis. This inference mechanism significantly improves the interpretability of complex multi-agent interaction relationships, enabling subsequent calibration steps to propose targeted correction paths.
[0026] The strategy coherence calibration unit plays a crucial role in this system by transforming the results of interactive analysis into optimization suggestions. This unit receives interactive analysis reports containing non-cooperative policy biases and their quantitative indicators. It generates collaborative optimization trajectory correction suggestions through a combination of inverse game theory and multi-objective optimization. The inverse game theory component predicts the agent's response behavior after receiving the correction suggestions. By constructing an individual response module that includes the current policy state, decision preference weights, and objective function settings, it simulates the agent's optimal response under correction constraints, thereby predicting the direction of its policy adjustment. The multi-objective optimization component seeks a balance between maximizing overall effectiveness, minimizing non-cooperative biases, and maintaining individual utility, ensuring that the correction path not only resolves conflicts but also maintains the agent's reasonable interests. The final set of correction paths is customized according to the individual characteristics and policy states of different agents, ensuring that each agent receives optimal adjustment suggestions that match its characteristics, avoiding insufficient adaptability caused by a one-size-fits-all correction approach.
[0027] In the specific application of inverse game theory, the strategy coherence calibration unit first uses the deviation quantification data and source tracing information provided in the interaction analysis report to establish a detailed individual response module for each agent. The individual response module not only includes its current policy state but also its decision preferences and objective function. Then, under the condition of applying preliminary correction suggestions, the unit simulates the agent's acceptance of the suggestions and adaptive adjustment behavior, and determines the agent's possible real response by solving for the optimal response strategy. All the optimal response strategies obtained from the simulation are re-inputted into the interaction effect inference module for further inference to detect whether the group of response strategies will produce secondary deviations in the same environment. If the secondary deviation exceeds the threshold, the calibration unit updates the preliminary correction suggestions and continues to iterate until the deviation converges within a preset range. The final output correction suggestion is the stable optimization scheme after predicting the agent's real response. This iterative inverse game simulation ensures the consistency and stability between the correction path and the agent's behavior, thereby significantly improving the reliability of the overall decision.
[0028] The decentralized behavioral adaptation mechanism in this system is responsible for implementing calibration recommendations into the local policy adjustments of the agents. This mechanism connects logically to each agent via a network channel. Upon receiving correction recommendations, it does not immediately force a change in the agent's overall policy, but rather gradually internalizes the recommendations into the agent's own policy generation logic through incremental local adjustments. During this process, the agent adjusts its policy parameters and response rules according to the optimization trajectory in the correction recommendations, enabling the overall policy system to gradually approach its optimum over time. This incremental adjustment mechanism avoids the instability caused by sudden, drastic changes while preserving the agent's flexibility to adapt to environmental changes. Because the adjustments of each agent are performed in parallel and depend on their individual characteristics, the system as a whole exhibits a decentralized collaborative evolutionary process. This mechanism not only enhances the robustness of policy execution but also improves the system's adaptability in dynamic environments.
[0029] The aggregation strategy discrimination module is responsible for periodically evaluating the adaptively adjusted multi-agent strategy in this system to ensure the effectiveness and stability of the overall strategy. This module introduces a dynamic anti-disturbance evaluation index set to quantitatively judge the performance of the strategy under simulated complex situations. Specifically, the discrimination module triggers the interaction effect inference module, injects a preset disturbance signal based on the current environmental situation data to form a new disturbance potential, and inputs the agent strategy combination into the inference process. The inference results pay special attention to the magnitude, rate of change, and recovery time of non-cooperative deviation, thereby generating a dynamic anti-disturbance evaluation index set, including three dimensions: maximum offset magnitude, offset stabilization time, and recovery rate. These indicators can reflect the system's resistance and recovery capability under external disturbances. When the evaluation results show that the aggregation strategy has reached the preset evolution state, the system will recognize the strategy as stable and usable; otherwise, it will prompt the aforementioned calibration and adaptive modules to make further adjustments through a feedback mechanism. This ensures that the entire system has dynamic evaluation and optimization functions, so that the strategy is not only effective under static conditions but can also maintain efficient operation under dynamic conditions.
[0030] Through the close integration of the above modules, this invention achieves consensus-based decision-making and question answering among multiple agents in complex environments. The multi-domain information aggregation module ensures the consistency and integrity of inputs, the interaction effect inference module provides an in-depth understanding of policy coupling relationships, the policy coherence calibration unit implements scientifically reasonable correction suggestions through a combination of theory and optimization, the decentralized behavior adaptive mechanism translates corrections into gradual policy adjustments, and the aggregated policy discrimination module provides continuous dynamic monitoring and evaluation for overall operation. These five components form a complete process of data input, effect inference, path correction, local execution, and global discrimination, enabling the system to continuously evolve and adapt to changes in the external environment. This solves the problems of bias accumulation and instability that are common in existing technologies for multi-agent consensus decision-making.
[0031] Example 2: In a resource scheduling scenario, a consensus-based decision-making question-and-answer system based on multi-AI agent game theory is deployed to achieve efficient allocation of resources from multiple data centers. The system includes a multi-domain information aggregation module, an interaction effect inference module, a strategy fusion calibration unit, a distributed behavior adaptive mechanism, and an aggregation strategy discrimination module.
[0032] The multi-domain information aggregation module collects data center resource usage, task load information, and external environmental factors such as network latency and power costs from various intelligent agents. This data is preprocessed and transformed into a unified format of input data stream. For example, resource utilization is converted from a percentage to a standardized value of 0 to 1, and task load is assigned different weights based on urgency and resource requirements.
[0033] The interaction effect inference module constructs a multidimensional interaction tensor model to analyze the mutual influence between agent policies. Through higher-order tensor decomposition and causal inference, it quantifies non-cooperative policy bias. For example, when an agent tends to over-allocate resources to high-load tasks, it affects the fair acquisition of resources by other agents, leading to a decrease in overall system efficiency. Therefore, an objective function is defined to measure the balance between overall system efficiency and individual utility; the objective function is: Where N is the number of agents, i represents the agent's ID, and a i It is the individual utility weight, U i Let b represent the utility of agent i, b be the weight of noncooperation bias, and s be the weight of agent i. i It is the actual strategy of agent i. It is an optimized equilibrium strategy; in application, a is set. i =0.6 and b=0.4, to balance individual and overall interests.
[0034] Based on the analysis results of the interaction effect inference module, the policy coherence calibration unit generates cooperative optimization trajectory correction suggestions for each agent. For example, when a resource allocation strategy of an agent is detected to cause non-cooperative deviation, the calibration unit will suggest that the agent adjust its strategy to reduce the negative impact on other agents. Therefore, the following thresholds are set to evaluate the effect of policy adjustment: the maximum deviation amplitude does not exceed 20% of the current state, the deviation stabilization time does not exceed 10 seconds, and the recovery rate is not less than 10% of the deviation per second. Through simulation experiments, it was found that when the agents adjust their strategies according to the correction suggestions, the overall system performance is significantly improved, and the non-cooperative deviation is reduced.
[0035] The distributed behavior adaptation mechanism implements calibration recommendations into local policy adjustments for the agent; the agent gradually adjusts its policy parameters based on the optimization trajectory in the correction recommendations; for example, an agent may initially allocate 70% of its resources to high-load tasks, and after adjustment, optimize the resource allocation ratio to 50% to ensure fair use of resources.
[0036] The aggregation strategy discrimination module periodically evaluates the adaptively adjusted multi-agent strategy; it quantifies the strategy's performance in simulated complex scenarios using a dynamic anti-disturbance evaluation index set; for example, after injecting a disturbance signal that increases network latency, it monitors the magnitude, rate of change, and recovery time of the non-cooperative deviation; the results show that the system can quickly recover stability after the disturbance, with a maximum deviation magnitude of 6%, a deviation stabilization time of 8 seconds, and a recovery rate of 0.1 decrease per second.
[0037] Through the implementation examples, the system can identify and correct non-cooperative policy deviations, thereby improving the coordination, stability, and resilience of multi-agent systems in complex environments.
[0038] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A consensus-based decision-making question-answering system based on multi-AI agent game theory, characterized in that, include: A multi-domain information aggregation module is used to receive diverse policy information and external environmental situation data from multiple intelligent agents; The strategy information and external environment situation data are preprocessed to form a unified format input data stream; An interaction effect inference module is used to perform in-depth analysis on the input data stream to identify and quantify the potential mutual influence relationships between the strategies of each agent; By constructing a multidimensional interaction tensor model, the expected feedback of each agent's policy combination under different environmental situations is calculated, the non-cooperative policy deviation caused by each agent's policy is deduced, and an interaction analysis report containing quantitative indicators of the non-cooperative policy deviation is generated. A policy coherence calibration unit receives the interaction analysis report and generates customized cooperative optimization trajectory correction suggestions for each agent based on the non-cooperative policy deviations and their quantitative indicators indicated in the report. This guides the agent to adjust its existing policies to weaken negative interference effects, enhance positive cooperative effects, achieve overall policy coordination, and thus produce the optimal set of correction paths. A decentralized behavior adaptation mechanism is logically connected to each agent through network channels, receives collaborative optimization trajectory correction suggestions, and performs incremental policy adjustments locally on each agent, so that the policies of each agent respond to the overall optimization needs of the system and internalize the correction suggestions as part of its own policy generation logic. An aggregation strategy discrimination module, logically connected to a distributed behavior adaptation mechanism and a multi-domain information aggregation module, is used to periodically monitor the aggregation status and performance of the adjusted multi-agent strategy and evaluate the current multi-agent aggregation strategy.
2. The consensus-based decision-making question-answering system based on multi-AI agent game theory according to claim 1, characterized in that: The multi-domain information aggregation module receives diverse policy information including the behavioral decision sequences, decision preference weights, policy confidence scores, and objective function settings of each agent in different situations; the external environment situation data includes task complexity indicators, available resource constraints, external interference signal strength, and the current perception status of other agents. The preprocessing process includes data cleaning, standardization, format conversion, and feature vectorization to ensure consistency of the input data in terms of dimensionality and semantics.
3. The consensus-based decision-making question-answering system based on multi-AI agent game theory according to claim 1, characterized in that: When constructing a multidimensional interaction tensor model, the interaction effect inference module uses historical policy interaction data and dynamic environmental change patterns to capture the temporal and spatial correlation of each agent's policies and their contribution or weakening effect on the overall system performance through high-order tensor decomposition and causal inference. The interaction analysis report includes not only quantitative indicators of non-cooperative policy deviations, but also source analysis of deviation sources and prediction of potential risk scenarios.
4. The consensus-based decision-making question-answering system based on multi-AI agent game theory according to claim 3, characterized in that: The process of high-order tensor decomposition and causal inference is as follows: In the decomposition stage, based on historical strategy interaction data and dynamic environmental change patterns, the multidimensional interaction structure is reorganized into dynamic subsets according to the similarity of environmental situations; within each subset, a constraint tensor decomposition is performed, forcing the strategy combination dimension and the time evolution dimension to form independent components, while retaining the inherent constraints of the environmental situation, generating an environmentally conditional strategy interaction pattern; in the inference stage, within subsets of the same environmental situation, the direct causal contribution of the strategy is calculated by comparing the difference in overall system performance between the actual strategy combination and the virtual combination after removing a specific strategy; this process is repeated across different environmental subsets to identify the fluctuation of the contribution of the same strategy under changes in environmental constraints, and to separate environmentally dependent biases from inherent policy conflicts. Based on the temporal correlation of contribution fluctuations, a cross-policy transmission chain is constructed to quantify the independent proportion of environmental coupling effects and inherent conflicts between policies in non-cooperative deviations.
5. A consensus-based decision-making question-answering system based on multi-AI agent game theory according to claim 1, characterized in that: The strategy coherence calibration unit generates collaborative optimization trajectory correction suggestions through a combination strategy based on inverse game theory and multi-objective optimization, wherein inverse game theory is used to predict the possible response behavior of each agent after receiving the correction suggestions; Multi-objective optimization seeks the optimal balance between maximizing the overall system performance, minimizing non-cooperative policy bias, and maintaining the individual utility of agents; the set of correction paths includes differentiated correction schemes for different individual agent characteristics and current policy states.
6. A consensus-based decision-making question-answering system based on multi-AI agent game theory as described in claim 5, characterized in that: The process of the inverse game theory is as follows: First, based on the quantified non-cooperative strategy bias and its source tracing results in the interaction analysis report, an individual response module is constructed for each agent, containing its current strategy state, decision preference weights, and objective function settings. Then, given the preliminary correction suggestions generated by the policy coherence calibration unit, the agent's potential acceptance and adaptive adjustment behavior based on its individual response module to the correction suggestions is simulated. This process is completed by solving the agent's optimal response strategy under the constraints of the correction suggestions. The simulated optimal response strategies of each agent are re-input into the interaction effect inference module to infer the new interaction effects of the response strategy combination under the same environmental situation and the possible secondary non-cooperative strategy biases. The preliminary correction suggestions are iteratively updated based on the inference results until the secondary non-cooperative strategy biases caused by the simulated response strategy combination converge within a preset threshold. At this point, the output correction suggestions are the stable optimization schemes after predicting the actual response behavior of each agent.
7. A consensus-based decision-making question-answering system based on multi-AI agent game theory as described in claim 1, characterized in that: When evaluating an aggregation strategy, the aggregation strategy discrimination module incorporates a dynamic anti-disturbance evaluation index set to determine whether the preset aggregation strategy evolution state has been reached. It performs a comprehensive evaluation by comparing the current index set with the preset threshold range or target curve, and by combining the system's performance in simulated complex scenarios.
8. A consensus-based decision-making question-answering system based on multi-AI agent game theory as described in claim 7, characterized in that: The dynamic disturbance rejection assessment index set is used to quantify the maintenance capability and recovery characteristics of a multi-agent aggregation strategy when facing simulated disturbance situations after progressive policy adjustment is performed by a distributed behavioral adaptive mechanism. The dynamic disturbance rejection assessment index set is generated and applied through the following unique process: the aggregation policy discrimination module triggers the interaction effect inference module, injects simulated disturbance signals of preset type and intensity based on the current environmental situation data, and constructs a new environmental situation containing disturbance variables; the adjusted combination of each agent's policy is input into the interaction effect inference module to infer the expected feedback of the policy combination under the disturbance potential, and monitors the change magnitude, change rate, and time required for the non-cooperative policy deviation calculated by the interaction effect inference module to recover to the pre-disturbance level; the dynamic disturbance rejection assessment index set is composed of three core dimensions: the maximum offset magnitude, offset stabilization time, and deviation recovery rate of the non-cooperative policy deviation during the disturbance injection period, which directly reflects the dynamic resistance and adaptive recovery efficiency of the aggregation strategy to external disturbances after policy coherence calibration and adaptive adjustment.
Citation Information
Patent Citations
Hierarchical multi-agent game confrontation and collaborative decision-making algorithm based on federated learning
CN119443312A
Multi-agent causal reasoning dynamic sand table decision-making method, system and equipment and storage medium
CN120450391A