Material research system and method based on multi-agent game
By using a multi-agent game system for materials research and development, the problems of multi-objective optimization and the disconnect between theory and practice have been solved, enabling efficient materials design and experimental verification, and improving the feasibility and success rate of materials research and development.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN INST OF ADVANCED TECH
- Filing Date
- 2026-04-08
- Publication Date
- 2026-05-08
AI Technical Summary
Existing technologies in materials research and development suffer from problems such as difficulty in dynamic coordination of multi-objective optimization, lack of substantial countermeasures in scheme evaluation mechanisms, and disconnect between theoretical prediction and physical experimental verification, resulting in low feasibility and success rate of design schemes.
A materials research system based on multi-agent game theory is adopted, including a human-computer interaction module, a scheme generation module, a review module, an arbitration module, a scheme execution module, and a knowledge base module. Through cross-review and dynamic game among multiple agents, multi-objective collaborative decision-making and virtual-real closed-loop feedback are achieved.
It improved the comprehensiveness of material design schemes and the success rate of experimental transformation, significantly reduced R&D costs and cycles, and enhanced the practical feasibility and R&D efficiency of the schemes.
Smart Images

Figure CN121998102A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer system technology, and in particular to a materials research system and method based on multi-agent game theory. Background Technology
[0002] The research paradigm in materials science is undergoing a transformation from a "trial and error" approach to a "data-driven" and "knowledge-driven" approach. Traditional materials research and development involves tedious steps such as literature review, scheme design, synthesis experiments, and characterization testing. It is not only time-consuming and costly, but also highly dependent on the personal experience and intuition of researchers.
[0003] In recent years, the emergence of Large Language Models (LLMs) has brought new opportunities to materials science. Based on massive amounts of text training, LLMs possess powerful logical reasoning, code generation, and tool invocation capabilities. Combining LLMs with automated laboratories to create "AI scientists" that automate the entire materials research and development process has become a cutting-edge research hotspot. This can not only liberate humans from repetitive labor but also explore the vast chemical space that is difficult for human intuition to access.
[0004] In existing technologies, the application of artificial intelligence to materials research and development generally suffers from the following limitations, resulting in low feasibility, reliability, and engineering implementation success rates of design solutions: (1) The problem of dynamic coordination in multi-objective optimization. Material design schemes need to simultaneously weigh multiple often conflicting objectives such as performance, cost, safety and process feasibility. Existing methods mostly use fixed weights for linear weighting or sequential optimization, which cannot achieve dynamic and adaptive weighing and decision-making based on the specific content of the scheme. The optimization results are often rigid and cannot reflect the flexible compromises required in real engineering scenarios.
[0005] (2) The lack of substantial confrontation in the scheme evaluation mechanism leads to insufficient depth of defect detection. Currently, most multi-agent systems adopt sequential execution or collaborative verification modes, with excessively high consistency between the goals of the evaluation module and the generation module, lacking effective checks and balances and challenge mechanisms. This makes it easy for the evaluation to remain at the level of superficial compliance checks, making it difficult to deeply question and stress test the basic assumptions, internal logic and potential risks of the scheme, and errors and "illusions" in the generated content are not easily detected and corrected in a timely manner.
[0006] (3) The disconnect between theoretical prediction and physical experimental verification, and the weak effectiveness of closed-loop feedback. Existing virtual-real hybrid systems usually use experimental data only for indirect optimization of model parameters, failing to construct it as a strong constraint and decisive verification link for virtual design schemes. This leads the system to tend to prioritize fitting historical or simulation data, while being insensitive to failure results that deviate from expectations in real experiments, and unable to effectively drive the design strategy to be fundamentally adjusted in a direction that is engineering-feasible.
[0007] In summary, this application proposes an innovative adversarial multi-agent architecture based on an arbitration center, aiming to create a next-generation intelligent materials R&D platform with self-correction capabilities, multi-objective collaborative decision-making capabilities, and virtual-real closed-loop evolution capabilities. Summary of the Invention
[0008] This application provides a materials research system and method based on multi-agent game theory. This method transforms the multi-objective conflicts, knowledge uncertainty, and gap between theory and practice in materials research and development into a structured and computable dynamic game process, thereby systematically improving the comprehensiveness of scheme generation, the rigor of review, and the success rate of experimental transformation.
[0009] To address the aforementioned technical problems, in a first aspect, embodiments of this application provide a materials research system based on multi-agent game theory, comprising a human-computer interaction module, a scheme generation module, a review module, an arbitration module, a scheme execution module, and a knowledge base module: the human-computer interaction module generates task instructions based on user needs; the scheme generation module generates multiple candidate schemes corresponding to different technical routes based on the task instructions; the review module performs multi-perspective cross-review and risk interception on the candidate schemes to obtain the review results of the multi-agents; the arbitration module handles the conflict of opinions among the multi-agents based on the review results and outputs a final decision; the scheme execution module executes experiments and collects data based on the final decision; and the knowledge base module provides static knowledge support and a closed-loop feedback of dynamic experimental data for the system.
[0010] In some exemplary embodiments, the human-computer interaction module includes a receiving unit, a conversion unit, a display unit, and a human intervention interface. The receiving unit receives R&D requirements input by the user. The R&D requirements include performance indicators, cost budgets, and process constraints. The conversion unit converts the unstructured natural language and parameters in the R&D requirements into standardized task instructions that can be recognized within the system. The display unit visualizes the system's decision-making process, presenting the interaction of review opinions and the iterative path of solutions among multiple agents in the review module, thus making the R&D process transparent. The human intervention interface allows the user to make a final strategic decision or adjust parameters through human intervention when a consensus cannot be reached within the system.
[0011] In some exemplary embodiments, the scheme generation module includes an initial scheme generation unit and a screening unit; the initial scheme generation unit is used to design candidate material schemes including raw material formulations, reaction paths and process parameters by reasoning and combining based on existing scientific principles and knowledge reserves; the screening unit is used to perform preliminary screening of the rationality of the chemical structure of the candidate material schemes to ensure that the generated candidate schemes have basic validity at the theoretical level. In some exemplary embodiments, the review module includes multiple evaluation units with specific domain perspectives, each responsible for independently verifying candidate solutions from the aspects of scientific principles, engineering feasibility, and safety compliance; each evaluation unit not only outputs a quantitative score, but also proposes modification suggestions or objections for specific defects found.
[0012] In some exemplary embodiments, the arbitration module includes a discrimination unit; when different evaluation units give conflicting evaluation results for the same solution, the discrimination unit automatically discriminates based on preset decision logic or weighting strategy and selects the solution with the best overall benefits; for complex situations involving key risks or disagreements exceeding the system threshold, the arbitration module triggers a suspension mechanism to route the decision-making power to the human-computer interaction module to request manual confirmation.
[0013] In some exemplary embodiments, the scheme execution module can dynamically select an appropriate execution mode based on the complexity of the experimental scheme, the operational precision requirements, and the current laboratory hardware resource configuration.
[0014] In some exemplary embodiments, the knowledge base module is responsible for the storage, management and service of knowledge for the entire system; it integrates static scientific knowledge as well as experimental data dynamically generated during system operation; the knowledge base module provides retrieval services for the generation and review modules through indexing technology, enabling them to refer to relevant historical experience when making decisions; by continuously incorporating new experimental feedback data, the knowledge base module can continuously update its data reserves, thereby supporting the self-correction and capability improvement of the system's decision-making model.
[0015] Secondly, this application also provides a materials research method based on multi-agent game theory, applied to the materials research system based on multi-agent game theory described above, including the following steps: Step S1: The user inputs the target requirements for materials research and development through a human-computer interaction module; the system parses the user's instructions and generates task instructions; Step S2: Based on the task instructions, The solution generation module initiates the design process, generating multiple candidate solutions corresponding to different technical routes. In step S3, the various agents in the review module review the candidate solutions, identifying potential vulnerabilities and risks from different dimensions, thus initiating a high-intensity adversarial verification. During the review process, the arbitration module acts as a "game manager," responsible for collecting scattered adversarial opinions and aggregating and deduplicating them. In step S4, when the solution generation module submits the iterated solution, the arbitration module first makes a preliminary judgment based on the scoring model. If the main indicators of the solution have not yet converged to the preset threshold, it indicates that the game is not yet balanced, and the system will continue iterating in step S3. When the arbitration module determines that the solution is nearing perfection and the game process shows a convergence trend, it initiates the "final review confirmation" mechanism to verify whether the system has truly reached a Nash equilibrium state. The arbitration module resubmits the solution to all review agents for review. Only when all adversaries no longer raise key objections and unanimously give a "feasible" conclusion is the system considered the game complete. Once a consensus is reached, the arbitration module locks in the solution and moves it to the next stage. In step S5, the solution execution module receives the final decision locked through the game and selects automated execution or generates standard operating procedures to guide manual execution based on the experimental conditions. Data and final results during the experiment are collected, formatted in a standard way, and then sent back to the knowledge base module. Real physical data during the experiment will be used as new prior knowledge to correct the value networks of each agent and improve the system's reasoning accuracy in future games. In step S6, the final experimental results are fed back to the user through the human-computer interaction interface, and the user makes the final value judgment. If the user confirms that the research and development goal has been achieved based on the measured data, the task officially ends. If the user believes that the result is not as expected, they can input new feedback. The system injects these new feedbacks as new constraints into the game environment, reactivates the solution generation module, jumps back to step S2, thereby breaking the current equilibrium state and starting a new round of adversarial optimization loop until the optimal solution that meets the user's needs is obtained.
[0016] In some exemplary embodiments, in step S2, the solution generation module, as the "generator" in the game, initiates the design process according to the requirements of the task instructions. The solution generation module first performs a wide-area search of the knowledge base, reviewing existing historical experimental cases, as well as unstructured knowledge content such as academic literature, professional books, and patent materials. Subsequently, the generation module uses a large model to perform comprehensive reasoning on the above information and constructs one or more initial candidate solutions within the limited search space. The initial candidate solutions, as the "objects" of the first round of the game, are submitted to the arbitration module.
[0017] In some exemplary embodiments, in step S3, the arbitration module introduces candidate solutions into the review environment. Each agent in the review module acts as an "adversary," attempting to identify potential vulnerabilities and risks in the solutions from different dimensions, thereby initiating a high-intensity adversarial verification of the solutions. The arbitration module plays the role of a "game manager" in this process, responsible for collecting scattered adversarial opinions and aggregating and deduplicating them. When there are conflicts in the review results of different dimensions, the arbitration module balances the interests of all parties according to the strategy weights, forming a unified correction feedback. The solution generation module then supplements and optimizes the solutions based on the feedback and puts the new solutions back into the review environment, thus forming a dynamic game cycle in which "the generator continuously improves its defense and the adversary continuously searches for vulnerabilities."
[0018] The technical solutions provided in this application have at least the following advantages.
[0019] This application provides a materials research system and method based on multi-agent game theory. The system includes a human-computer interaction module, a scheme generation module, a review module, an arbitration module, a scheme execution module, and a knowledge base module. The human-computer interaction module generates task instructions based on user needs. The scheme generation module generates multiple candidate schemes corresponding to different technical routes based on the task instructions. The review module performs multi-perspective cross-review and risk interception on the candidate schemes to obtain the review results of the multi-agents. The arbitration module handles the conflict of opinions among the multi-agents based on the review results and outputs the final decision. The scheme execution module executes experiments and collects data based on the final decision. The knowledge base module provides static knowledge support and a closed-loop feedback of dynamic experimental data for the system.
[0020] On the one hand, this application, through a dynamic weighting strategy and a convergence determination mechanism based on the rate of change in the arbitration module, enables the system to adjust the focus of optimization in real time based on the specific feedback from each professional agent in each round of review. For example, when the cost agent consistently gives low scores and provides market data evidence, the system will automatically emphasize cost optimization in subsequent iterations. This dynamic adjustment capability based on real-time interaction ensures that the final output solution is a balanced solution reached after sufficient game theory under multiple constraints, avoiding the rigidity of static methods and significantly improving the practical feasibility of the solution.
[0021] On the other hand, each agent in the review module of this application has the ability to directly call external professional computing software, real-time databases, and rule bases for verification. For example, the theoretical agent can quantitatively assess material stability by calling first-principles calculations, and the safety agent can automatically check whether the scheme complies with the latest safety specifications. This in-depth review mechanism based on objective tools and data can discover deep-seated, interdisciplinary risks (such as thermodynamic instability, supply chain bottlenecks, and compliance loopholes) that are difficult to detect by traditional methods, thereby effectively filtering high-risk schemes before investing in expensive experiments, significantly reducing R&D costs and time.
[0022] Finally, this application constructs a standardized and automated end-to-end digital R&D pipeline, improving R&D efficiency and process quality. Through a microservice architecture and a unified data interface, this application seamlessly integrates task parsing, intelligent design, multi-dimensional review, arbitration decision-making, experiment execution, and data feedback, forming a highly automated digital workflow. This significantly reduces delays and errors caused by manual intervention, ensures a high degree of consistency and repeatability in the R&D process, and makes all data, decisions, and status changes traceable throughout the entire R&D process, laying a solid foundation for refined and continuous improvement of R&D management. Attached Figure Description
[0023] One or more embodiments are illustrated by way of example with reference to the accompanying drawings. These illustrations do not constitute a limitation on the embodiments, and unless otherwise stated, the figures in the drawings are not to be limited by scale.
[0024] Figure 1 This is a schematic diagram of the structure of a materials research system based on multi-agent game theory, provided as an embodiment of this application.
[0025] Figure 2 This is a flowchart illustrating a materials research method based on multi-agent game theory, provided as an embodiment of this application. Detailed Implementation
[0026] As can be seen from the background technology, when applying artificial intelligence to materials research and development, there are generally problems such as difficulty in dynamic coordination of multi-objective optimization, insufficient depth of defect detection, disconnect between theoretical prediction and physical experimental verification, and weak closed-loop feedback effectiveness.
[0027] To address the aforementioned technical problems, this application provides a materials research system and method based on multi-agent game theory. The system includes a human-computer interaction module, a scheme generation module, a review module, an arbitration module, a scheme execution module, and a knowledge base module. The human-computer interaction module generates task instructions based on user needs. The scheme generation module generates multiple candidate schemes corresponding to different technical routes based on the task instructions. The review module performs multi-perspective cross-review and risk interception on the candidate schemes to obtain the review results from multiple agents. The arbitration module handles conflicts of opinion among the agents based on the review results and outputs a final decision. The scheme execution module executes experiments and collects data based on the final decision. The knowledge base module provides static knowledge support and a closed-loop feedback of dynamic experimental data for the system. The specific objectives of this application's materials research system and method based on multi-agent game theory are as follows: (1) Construct a game framework in which a cluster of professional domain agents work collaboratively to overcome the knowledge limitations of a single model. By deploying specialized agents that are proficient in material design, cost analysis, process feasibility and safety assessment, a deep integration of interdisciplinary knowledge is achieved. Game rules are designed to drive each agent to actively contribute knowledge and discover potential problems from its own professional perspective, thereby improving the overall reliability of the solution through collaboration and checks and balances.
[0028] (2) Establish a parallel review mechanism based on adversarial debate and arbitration to break the error amplification effect of linear processes. By introducing a cluster of intelligent agents with opposing roles, namely "proposers" and "reviewers," and forcing multiple rounds of adversarial debate and cross-review from multiple perspectives before the implementation of the solution, this mechanism allows serious defects in any dimension to be proactively identified and challenged by reviewers with corresponding professional knowledge, and then adjudicated and fed back by the arbitration module. This effectively cuts off the error chain in advance during the virtual stage, achieving proactive reinforcement of the solution and minimization of risks.
[0029] (3) To achieve an evolutionary closed loop with physical experiments as strong constraints and final arbiters, bridging the gap between theory and practice. The automated experimental platform is positioned as the "fact arbitrator" in the game environment, and its experimental results (especially failure data) will directly and significantly influence the strategies and payoffs of relevant agents in the game as strong constraint signals. This forces virtual design agents to prioritize and satisfy the hard constraints of engineering feasibility in the game, thereby driving the optimization direction of the entire system from "theoretically optimal" to "engineerably feasible," improving the actual success rate of research and development.
[0030] The embodiments of this application will now be described in detail with reference to the accompanying drawings. However, those skilled in the art will understand that many technical details have been provided in the embodiments of this application to facilitate a better understanding of the application. However, the technical solutions claimed in this application can be implemented even without these technical details and various variations and modifications based on the following embodiments.
[0031] refer to Figure 1 This application provides a materials research system based on multi-agent game theory, comprising: a human-computer interaction module, a scheme generation module, a review module, an arbitration module, a scheme execution module, and a knowledge base module. The human-computer interaction module generates task instructions based on user needs. The scheme generation module generates multiple candidate schemes corresponding to different technical routes based on the task instructions. The review module performs multi-perspective cross-review and risk interception on the candidate schemes to obtain the review results of the multi-agents. The arbitration module handles the conflict of opinions among the multi-agents based on the review results and outputs a final decision. The scheme execution module executes experiments and collects data based on the final decision. The knowledge base module provides static knowledge support and a closed-loop feedback of dynamic experimental data for the system.
[0032] This application proposes a materials research system based on multi-agent game theory. This technical solution constructs a cluster of expert agents with different knowledge backgrounds, introduces adversarial review and human-machine collaborative arbitration mechanisms, and combines feedback from automated physical experiments to achieve closed-loop autonomous evolution from materials design to verification. The system disclosed in this application is a modular and scalable materials research system based on multi-agent game theory. Logically, the system is decoupled into six core modules, each with independent interfaces and functions, supporting flexible configuration and expansion for different materials fields (such as metals, catalysts, and polymers).
[0033] See Figure 1 After the user assigns a task to the human-computer interaction module, the module decomposes the task and transmits the generated task instructions to the solution generation module. The solution generation module generates candidate solutions and submits them to the arbitration module, which then transmits the submitted solutions to the review module for evaluation. The review module conducts multi-perspective cross-reviews and risk interception on the submitted solutions, feeding back the multi-agent review results to the arbitration module. Based on the multi-agent review results, the arbitration module handles conflicts of opinion among the multi-agents and transmits the final decision to the solution execution module, while simultaneously feeding the final decision back to the solution generation module. Based on the final decision, the solution execution module executes the experiment and collects data, then feeds the data back to the knowledge base module. The knowledge base module provides static knowledge support and a closed-loop feedback of dynamic experimental data for the entire system.
[0034] In some embodiments, the human-computer interaction module includes a receiving unit, a conversion unit, a display unit, and a human intervention interface. The receiving unit receives R&D requirements input by the user. The R&D requirements include performance indicators, cost budgets, and process constraints. The conversion unit converts the unstructured natural language and parameters in the R&D requirements into standardized task instructions that can be recognized within the system. The display unit visualizes the system's decision-making process, presenting the interaction of review opinions and the iterative path of solutions among multiple agents in the review module, thus making the R&D process transparent. The human intervention interface allows the user to make the final strategic decision or adjust parameters through human intervention when a consensus cannot be reached within the system. Specifically, the human-computer interaction module acts as a communication bridge between the user and the intelligent agent cluster. This module is primarily responsible for facilitating bidirectional communication between user intent and system execution. On one hand, it receives user-inputted R&D requirements, such as specific performance indicators, cost budgets, or process constraints, and is responsible for converting these unstructured natural language or parameters into standardized task instructions recognizable within the system. On the other hand, this module is responsible for visually representing the system's decision-making process. Furthermore, when the system cannot reach a consensus, this module provides a human intervention interface, allowing users to make adjustments.
[0035] In some embodiments, the scheme generation module includes an initial scheme generation unit and a screening unit; the initial scheme generation unit is used to design candidate material schemes including raw material formulations, reaction paths and process parameters through reasoning and combination based on existing scientific principles and knowledge reserves; the screening unit is used to perform preliminary screening of the rationality of the chemical structure of the candidate material schemes to ensure that the generated candidate schemes have basic theoretical validity.
[0036] Specifically, the scheme generation module is primarily responsible for exploring the chemical space and constructing initial schemes. Under given task constraints, this module explores the possibilities of material design and builds initial schemes. Its core function is to design candidate material schemes, including raw material formulations, reaction pathways, and process parameters, based on existing scientific principles and knowledge reserves through reasoning and combination. This module possesses a certain degree of self-verification capability, enabling preliminary screening of the chemical structural rationality of the schemes before output, ensuring that the generated candidate schemes have basic theoretical validity. Simultaneously, this module supports multi-path generation strategies, providing multiple alternative schemes with different technical routes, offering ample sample space for subsequent screening and optimization.
[0037] In some embodiments, the review module includes multiple evaluation units with specific domain perspectives, each responsible for independently verifying candidate solutions from the aspects of scientific principles, engineering feasibility, and safety compliance. Each evaluation unit not only outputs a quantitative score but also proposes modification suggestions or objections for specific defects found. Specifically, the review module is mainly responsible for conducting multi-dimensional adversarial reviews and risk interception of candidate solutions. This module is responsible for multi-dimensional review and evaluation of the generated candidate solutions. In this module, multiple evaluation units with specific domain perspectives are pre-defined, each responsible for independently verifying the solution from the aspects of scientific principles (such as thermodynamic stability), engineering feasibility (such as raw material costs and supply chain), and safety compliance (such as reaction risk). Each evaluation unit not only outputs a quantitative score but also proposes modification suggestions or objections for specific defects found. Through this multi-perspective cross-review mechanism, the system can identify and correct potential logical flaws or engineering risks before the solution enters the experimental stage.
[0038] In some embodiments, the arbitration module includes a discrimination unit. When different evaluation units give conflicting evaluation results for the same solution, the discrimination unit automatically discriminates based on preset decision logic or weighting strategies, selecting the solution with the best overall benefit. For complex situations involving key risks or disagreements exceeding system thresholds, the arbitration module triggers a suspension mechanism, routing the decision-making power to the human-computer interaction module for manual confirmation. Specifically, the arbitration module is mainly responsible for handling conflicts of opinion among multiple agents and outputting the final decision. This module is responsible for handling disagreements that arise during the review process, ensuring the smooth progress of the decision-making process. When different evaluation units give conflicting evaluation results for the same solution (e.g., performance evaluation is excellent but cost evaluation is poor), this module automatically discriminates based on preset decision logic or weighting strategies, selecting the solution with the best overall benefit. For complex situations involving key risks or disagreements exceeding system thresholds, this module triggers a suspension mechanism, routing the decision-making power to the human-computer interaction module for manual confirmation. This module also records all decision-making basis for subsequent system optimization and auditing.
[0039] In some embodiments, the scheme execution module can dynamically select an appropriate execution mode based on the complexity of the experimental scheme, the required operational precision, and the current laboratory hardware resource configuration. Specifically, the scheme execution module is mainly used to bridge the gap between virtual design and physical reality, execute experiments, and collect data. This module is responsible for converting the digital design scheme into physical verification operations and for collecting and integrating experimental data. This module has flexible task scheduling capabilities and can dynamically select an appropriate execution mode based on the complexity of the experimental scheme, the required operational precision, and the current laboratory hardware resource configuration, as shown below: a. Automated execution mode: For standardized processes (such as pipetting and weighing), the plan is converted into equipment control code to drive automated hardware to complete the experiment.
[0040] b. Manual or human-machine collaboration mode: For highly complex or non-standard experimental procedures, generate human-readable standard operating procedures (SOPs) or guides to guide experimenters to complete the operation manually and receive manually entered experimental data through the corresponding interactive interface. It should be noted that, regardless of the mode adopted, the solution execution module is responsible for uniformly cleaning, aligning and formatting the generated experimental data (including automatically collected sensor data and manually uploaded records) to ensure that the data flowing back to the knowledge base is consistent.
[0041] In some embodiments, the knowledge base module primarily provides static knowledge support and a closed-loop feedback mechanism for dynamic experimental data. This module is responsible for the storage, management, and service of knowledge across the entire system. It integrates static scientific knowledge (such as physicochemical property data and literature) and dynamically generated experimental data (including success stories and failure records) during system operation. This module provides retrieval services to the generation and review modules through indexing technology, enabling them to reference relevant historical experience when making decisions. By continuously incorporating new experimental feedback data, the knowledge base module can continuously update its data reserves, thereby supporting the self-correction and capability improvement of the system's decision-making model.
[0042] The following detailed description of the material research system and method based on multi-agent game theory provided in this application is based on specific embodiments.
[0043] Please continue reading. Figure 1 This application proposes a modular, loosely coupled materials research system based on multi-agent game theory. From a technical implementation perspective, the system is built using a distributed microservice architecture, with each core module operating as an independent service unit. Data interaction and collaborative operation are achieved through a standardized communication protocol. The arbitration module, acting as the system's central scheduler and message bus, is responsible for coordinating the entire game process and maintaining system state consistency. Detailed module connections and data interaction mechanisms are as follows: (1) External and internal interfaces of the human-computer interaction module.
[0044] The human-computer interaction module, acting as the system's front-end portal, provides a visual interface for web or desktop use, receiving unstructured user input (such as R&D requirements described in natural language). Internally, this module parses the user's intent through the API gateway and routes the processed structured task instruction data package (containing standard fields such as target parameters and constraints, typically in JSON or Protocol Buffers format) to the solution generation module, thereby triggering the lifecycle. Simultaneously, this module subscribes to the arbitration module's status update channel, obtaining real-time game progress and review data, and visualizing this information on the interface.
[0045] (2) The flow of core game data around the arbitration module.
[0046] The arbitration module is the core of the entire architecture, which realizes high-frequency game interaction between generation and review through a high-throughput message queue mechanism (such as Kafka or RabbitMQ).
[0047] Solution submission path (solid line): After the solution generation module completes the initial solution construction, it serializes the candidate solution data object containing the complete formula, process path and metadata, and submits it asynchronously to the arbitration module's pending queue through the internal service interface.
[0048] Submission and Feedback Path (solid line): The arbitration module, acting as the scheduler, distributes the proposed solutions in parallel to each independent evaluation unit (agent services such as principle, cost, and safety) within the review module. After completing their reasoning, each evaluation unit returns a structured review opinion data package containing quantitative scores, risk labels, and specific modification suggestions to the arbitration module. After aggregating the opinions, the arbitration module, based on its judgment, either feeds back a unified correction instruction data package to the solution generation module to trigger iteration, or locks the final solution for transfer to the next level.
[0049] (3) Data loop between scheme execution and the physical world.
[0050] Once the arbitration module finalizes the solution, it sends the executable solution instruction set to the solution execution module via a dedicated device control interface or SOP generation service. The solution execution module, while driving the physical experiment, integrates a multimodal sensor data acquisition system. It standardizes, cleans, and aligns the real-time time-series data (such as temperature curves) and the final experimental result data (such as product performance indicators) generated during the experiment, forming standardized experimental data records. These records are then fed back to the knowledge base module for persistent storage via a dedicated data channel.
[0051] (4) The underlying support of the knowledge base module (dashed line).
[0052] like Figure 1As shown by the dashed line, the knowledge base module, as the underlying data infrastructure, does not dominate the sequence of business processes, but provides real-time knowledge support for each decision point in the process through high-concurrency, low-latency data services. This module constructs a hybrid index based on vector databases and graph databases. When performing reasoning calculations, the solution generation module and the review module obtain relevant scientific principles, historical cases, or literature knowledge data in real time as context input by calling the semantic retrieval API or graph query interface provided by the knowledge base module, thereby ensuring the scientific nature and accuracy of the agent's decisions.
[0053] Specifically, the detailed technical implementation of each module is as follows: (1) Human-computer interaction module: intent understanding and visual interaction front end.
[0054] As a bridge for communication between users and complex intelligent systems, the key technical implementation of this module lies in the accurate understanding of natural language intent and the intuitive presentation of the system's internal state.
[0055] The LLM-based Natural Language Understanding (NLU) engine technically integrates a large language model, finely tuned by instructions, as its backend. When users input unstructured R&D requirements via text or voice (e.g., "Design a new type of refractory material that can withstand temperatures exceeding 800 degrees Celsius and costs 20% less than existing solutions"), the NLU engine performs in-depth semantic analysis. It uses intent recognition technology to determine the user's core objective and extracts key parameter entities from the text using slot-filling technology (e.g., key indicators "temperature resistance > 800℃" and constraints "cost better than the benchmark"). Finally, this scattered information is automatically assembled into a well-structured JSON-formatted task book object that conforms to the system's internal definition and is then passed to downstream modules.
[0056] A real-time state stream and dynamic visualization front-end based on WebSocket: To allow users to transparently control the complex game process, this module's front-end uses a modern reactive framework (such as React or Vue.js) and maintains real-time communication with the back-end arbitration module via the WebSocket long-connection protocol. Every submission of review comments, score updates, and iteration of the solution within the system is instantly pushed to the front-end in the form of an event stream, triggering dynamic updates to the interface. Technically, this enables the intuitive presentation of the collaboration and "confrontation" process between multiple agents in the form of dynamic dialogue streams, interactive dashboards, or version evolution tree diagrams, and provides a context-rich operation interface for human intervention in arbitration.
[0057] (2) Solution generation module: Search-enhanced and constrained generative design engine.
[0058] This module is the "creative brain" of the system. Its core technology is to build a composite generation process of "retrieval-enhanced generation + rule-constrained filtering", which aims to efficiently find initial solutions that meet the boundary conditions in a huge chemical space.
[0059] The retrieval-enhanced generative framework first runs a high-efficiency retrieval engine that retrieves highly relevant historical cases and synthetic literature fragments from the knowledge base's vector index based on keywords in the task statement. Subsequently, this retrieved knowledge is formatted as contextual hints and input into the core generative large-scale model along with task constraints. This large-scale model is typically a professional model pre-trained on massive amounts of materials science text and structural data. It possesses powerful cross-modal reasoning capabilities and can generate preliminary draft schemes containing raw material proportions (text / tabular data), molecular structures (SMILES / SELFIES encoding), and synthesis steps.
[0060] Built-in rule engine and self-verification mechanism: To reduce the computational burden caused by invalid solutions flowing into subsequent stages, this module incorporates a lightweight chemical rule engine as the first line of defense. Before outputting a solution, the engine performs basic validity checks on the generated molecular structure (such as valence balance checks and ring structure rationality analysis) and filters out obvious errors based on basic chemical common sense (such as incompatible reactant identification). Only solutions that pass this self-verification are packaged into formal candidate solution objects and submitted to the arbitration module.
[0061] (3) Review module: Multidimensional independent intelligent agent cluster and professional toolchain.
[0062] In terms of technical implementation, this module is not a single-node evaluator, but a cluster of intelligent agents with specialized domain knowledge capable of calling external tools for quantitative verification. Each agent is fine-tuned based on a knowledge base specific to its domain, and its core feature is that it is endowed with powerful "tool calling capabilities"—that is, during the inference process, it can autonomously call external professional computing software, database APIs, or internal rule engines for quantitative deep verification, rather than simply relying on the general knowledge of a large model for qualitative judgment.
[0063] Specifically, the review module includes a theoretical feasibility agent, an experimental feasibility agent, a cost feasibility agent, and a safety feasibility agent. Among them, the theoretical feasibility agent is mainly used to examine whether the scheme has fundamental defects in terms of thermodynamic stability, kinetic energy barrier, electronic structure, etc., based on basic scientific principles.
[0064] The technical implementation of the theoretically feasible intelligent agent is as follows: In addition to retrieving phase diagrams and theoretical literature from the knowledge base, the agent integrates interfaces with first-principles calculation software or molecular dynamics simulation software. When encountering an unknown new structure, it automatically constructs an input file and submits it to a high-performance computing cluster to verify its binding energy or band gap using the DFT method, providing quantitative calculation results to support its review opinions.
[0065] The experimental feasibility agent focuses on the operability and reproducibility of the synthesis route in a real-world laboratory environment. It examines reaction conditions (temperature and pressure), post-processing complexity, and the availability of required equipment. This agent interfaces with the current laboratory's hardware resource database and historical process parameter library. It analyzes whether the synthesis steps in the plan depend on extreme or unavailable experimental conditions, assesses whether the stability of intermediates allows for further operations, and identifies potential obstacles at the experimental operational level by comparing the success rates of similar historical synthesis cases.
[0066] The cost feasibility analysis agent assesses the raw material costs, process energy consumption, and potential supply chain risks of a proposed solution from an economic perspective. This agent interfaces with external API (Indexed Price Index) data for chemical raw materials and industrial-grade process cost estimation models. It can capture real-time guidance prices and supply status of key precursors, estimate total costs based on the length of steps in the synthesis route, and output a detailed economic analysis report.
[0067] The safety feasibility agent is responsible for veto-level assessments of environmental, health, and safety risks. This agent incorporates an authoritative chemical safety database and a reaction matrix compatibility rule base. It systematically scans the toxicity levels and flammability / explosive properties of raw materials, solvents, and expected byproducts used, analyzing potential runaway risks or environmental compliance loopholes in the reaction pathway.
[0068] (4) Arbitration module: a strategy-based conflict decision-making and game management engine.
[0069] Unlike the review module, which focuses on a broad, professional perspective, the arbitration module's technical focus is on conflict convergence, strategy scheduling, and decision governance. It incorporates a well-defined quantitative calculation process for handling conflicts and determining the game's state. Its specific technical implementation includes the following three core steps: 1. Logic for calculating comprehensive scores based on dynamic weights.
[0070] The arbitration module maintains a strategy configuration table that defines the weight vector for the current task. These correspond to four dimensions: performance, cost, safety, and experimental feasibility. .
[0071] In each round of the game, each reviewing agent... In addition to outputting textual opinions, it is also necessary to output a normalized quantitative score. And a Boolean "veto label" (1 indicates the existence of an unacceptable hard defect, such as an explosion risk).
[0072] The arbitration module calculates the overall score of the proposed solutions in this round according to the following logic. : Step A (Safety Circuit Breaker): First, check the veto tags of all agents. If any... (Especially for security intelligent agents), then directly determine The plan was marked as "high risk" and forced to be returned for modification.
[0073] Step B (Weighted Calculation): If there is no veto label, then execute the weighted summation formula:
[0074] (Example: Set the performance weight to 0.5 and the cost weight to 0.2. If a solution scores 90 in performance but 40 in cost and has no safety risks, then the overall score contribution is 0.5*90 + 0.2*40 = 53 points, and so on to calculate the total score.)
[0075] 2. Game convergence determination logic based on difference threshold.
[0076] The system does not pursue infinite iterations, but rather determines convergence based on the principle of diminishing returns. The arbitration module records the sequence of comprehensive scores from each iteration. Convergence is determined to be achieved if the following two conditions are met simultaneously: Condition A (Qualification Threshold): Current Overall Score It must be greater than the preset threshold. (For example, 85 points).
[0077] Condition B (Game Equilibrium / Payback Convergence): The rate of change of the score in two consecutive iterations is less than the convergence threshold. (For example, 5%), which satisfies the formula:
[0078] This formula shows that even with further modifications, the improvement in the quality of the solution is very small, indicating that the game has reached a stable state near the Nash equilibrium point, and the system should stop internal friction and lock in the solution.
[0079] 3. "Infinite loop" circuit breaker and human-machine takeover mechanism.
[0080] To prevent the system from falling into an infinite loop of "modification-rejection-modification" due to insufficient capabilities of the generation module or overly stringent constraints, the arbitration module is equipped with a maximum iteration counter. (For example, set to 10 rounds).
[0081] The technical logic is as follows: Each iteration, .
[0082] when If the above convergence conditions are still not met, the arbitration module triggers a "SuspendInterrupt" signal.
[0083] This signal will activate a pop-up window in the human-computer interaction module, displaying the current impasse to the user (e.g., "Performance and cost cannot be reconciled; the highest historical score is 78 points"), requesting the user to make a strategic decision. The user can choose to "lower the acceptance criteria," "modify the constraints," or "force the current highest-scoring solution to pass," and the system will break the deadlock accordingly.
[0084] (5) Solution execution module: virtual-real mapping and dual-mode execution engine.
[0085] This module serves as the mapping interface between the logical world and the physical world. Technically, it comprises a core task scheduling middleware and two physical execution subsystems, designed to adapt to experimental scenarios of varying complexity.
[0086] Automated Execution Subsystem (for Standard Processes): For standardized operations such as liquid handling, weighing, and high-throughput screening, this subsystem interfaces with laboratory automation hardware (such as pipetting workstations and robotic arms) via standard device interface protocols (e.g., SiLA, OPC UA). Technically, it includes an instruction compiler that automatically translates structured experimental protocols issued from the upper layer into low-level control code (e.g., G-code or specific scripts) recognizable by specific hardware devices, achieving unattended automated experimental closed-loop operation.
[0087] Human-Machine Collaborative Execution Subsystem (for non-standard / complex processes): For precision operations involving complex solid-phase processing that are difficult to automate, this subsystem focuses on enhanced human guidance. It utilizes natural language generation technology to transform abstract experimental plans into clearly structured, detailed, human-readable standard operating procedure documents. Furthermore, combined with augmented reality interfaces or tablet PC interfaces, the system can guide operators step-by-step and automatically record execution parameters for key steps through voice input or image recognition, ensuring the standardized collection of data from manual operations.
[0088] (6) Knowledge base module: Multi-source heterogeneous data fusion and knowledge graph construction.
[0089] The knowledge base module serves as the intelligent foundation of the entire system. Its key technological implementation lies in establishing a complete data processing pipeline, from the acquisition of multi-source heterogeneous data to structured knowledge services.
[0090] The module draws on a wide range of data sources, including but not limited to: publicly available academic literature databases (PDF / XML format), various materials science databases, patent documents, industry standard documents, and historical experimental records generated by the system itself. To effectively manage this data, the module constructs a vast domain knowledge graph. Technically, it uses Natural Language Processing (NLP) and information extraction techniques to identify key entities (such as material names and properties) and their relationships (such as "synthesized in," "possesses... properties") from unstructured text, building an entity relationship network. Simultaneously, it utilizes an embedding model to transform text and graph nodes into high-dimensional vectors stored in a vector database, thereby supporting efficient semantic-based retrieval and complex graph reasoning queries, providing precise knowledge enhancement services for upper-layer intelligent agents.
[0091] Based on the aforementioned modular system architecture, such as Figure 2 As shown in the embodiments, this application also provides a materials research method based on multi-agent game theory, applied to the materials research system based on multi-agent game theory described above. This method integrates traditional discrete experimental steps into a self-evolving closed loop through relay and game among modules. This method is a fully closed-loop materials research and development method. Through collaborative game among multiple agents and continuous feedback from human-computer interaction, it achieves self-evolution from requirements to empirical evidence. The specific operation steps are as follows: Step S1: Task initialization and intent parsing.
[0092] Users input their material research and development goals through the human-computer interaction module. The system parses the user's instructions and, in conjunction with industry standards in the knowledge base, transforms the requirements into a structured task specification. This task specification clearly defines the target parameters of the experiment (such as temperature resistance and strength), raw material limitations, and process constraints, and is then sent to the solution generation module.
[0093] Specifically, the process begins with the user inputting unstructured natural language requirements through the human-computer interaction module. The NLU engine in the system's backend first performs semantic analysis on the input text, extracting key entities (such as "high-entropy alloy" and "corrosion resistance > 100h"). Subsequently, the system combines pre-built industry standard templates in the knowledge base to parameterize these discrete requirements, transforming them into a computer-readable JSON format standard task specification. This task specification not only defines the target performance indicators but also clarifies hard constraints such as cost budget, time nodes, and a blacklist of prohibited raw materials, serving as the fundamental guideline for all subsequent actions of the intelligent agent.
[0094] Step S2: Generating a solution based on knowledge retrieval.
[0095] The solution generation module, acting as the "generator" in the game, initiates the design process according to the requirements of the task specification. This module first performs a broad search of the knowledge base, reviewing existing historical experimental cases, as well as unstructured knowledge content such as academic literature, professional books, and patent materials. Subsequently, the generation module uses a large model to comprehensively reason about the above information, constructing one or more initial candidate solutions within the limited search space. These solutions, as the "objects" in the first round of the game, are submitted to the arbitration module.
[0096] Specifically, after receiving the task assignment, the solution generation module initiates a design process based on search-enhanced generation. First, the search engine searches in parallel within the vector space of the knowledge base for relevant synthetic literature, patents, and historical experimental data to construct a high-dimensional context. Next, the generative large model, combining the context and task constraints, performs reasoning within a defined chemical space to generate an initial set of candidate solutions containing raw material formulations, reaction pathways, and process parameters. Before output, the built-in rule engine performs preliminary screening of these solutions based on chemical validity (such as valence equilibrium) to ensure that the solutions submitted to the game theory stage are theoretically sound.
[0097] Step S3: Multi-agent adversarial game and iteration.
[0098] This step constitutes the core interactive link in the multi-agent adversarial game. The arbitration module introduces candidate solutions into the review environment. The various agents in the review module (principles, costs, security, etc.) act as "adversaries," attempting to identify potential vulnerabilities and risks in the solutions from different dimensions, thus initiating a high-intensity adversarial verification of the solutions. The arbitration module plays the role of "game manager" in this process, responsible for collecting scattered adversarial opinions and aggregating and deduplicating them. When review results from different dimensions conflict (e.g., performance meets standards but security is compromised), the arbitration module balances the interests of all parties according to strategy weights, forming a unified corrective feedback. The solution generation module then supplements and optimizes the solutions based on the feedback and puts the new solutions back into the review environment, thus forming a dynamic game cycle of "the generator continuously improving its defenses, and the adversaries continuously searching for vulnerabilities."
[0099] Specifically, step S3 achieves the core closed loop of multi-agent adversarial game and iteration. Specifically, the arbitration module distributes the initial solution to four independent agents (theoretical, experimental, cost-related, and safety-related) in the review module. Each agent uses its own toolchain (such as DFT calculation, price API, and MSDS library) to conduct a comprehensive and in-depth "fault-finding" review of the solution, outputting review opinions including quantitative scores and specific defect descriptions. After collecting these opinions, the arbitration module generates correction instructions based on a weighted strategy and feeds them back to the solution generation module. The solution generation module then makes targeted improvements or modifications to the solution (e.g., replacing expensive raw materials, adjusting reaction temperature) and resubmits the modified solution for review. This process constitutes the internal small closed loop in the diagram, namely the adversarial iteration of "generation-review-modification".
[0100] Step S4: Game convergence determination and final confirmation.
[0101] During the aforementioned adversarial process, the arbitration module continuously monitors the convergence status of the entire game system. Whenever the generation module submits an iterated solution, the arbitration module first makes an initial judgment based on the scoring model: if the main indicators of the solution have not yet converged to the preset threshold, it indicates that the game is not yet balanced, and the system will continue iterating in step S3. When the arbitration module determines that the solution is nearing perfection and the game process shows a convergence trend, it will initiate a "final confirmation" mechanism to verify whether the system has truly reached a Nash equilibrium. The arbitration module resubmits the solution to all review agents for review. Only when all adversaries no longer raise key objections and unanimously give a "feasible" conclusion, does the system determine that the game has reached a consensus, and the arbitration module then locks the solution and moves it to the next stage.
[0102] Specifically, during the adversarial iteration process, the arbitration module continuously monitors the state of the game. It determines whether the game has reached a convergence state based on a preset mathematical model (such as the variance of scores being less than a threshold for three consecutive rounds).
[0103] Branch decision (as shown by the diamond in the figure): If the system fails to converge or there is a critical expert objection (such as a veto by the security agent), the process will follow the "No" branch path and force the system to return to step S3 to continue the adversarial iteration.
[0104] If the system is determined to have converged and meets the final review requirements (i.e., all key indicators meet the standards and the expert group reaches a consensus), the process will follow the "yes" branch path, with the arbitration module digitally signing and locking the final solution, and then moving to the next stage.
[0105] Step S5: Implementation and Data Feedback.
[0106] The solution execution module receives the final solution determined through game theory and, based on the experimental conditions, chooses to execute it automatically or generate a standard operating procedure (SOP) to guide manual execution. Data and final results during the experiment are collected, formatted in a standardized manner, and then transmitted back to the knowledge base module. This real-world physical data will serve as new prior knowledge, used to refine the value networks of each agent and improve the system's reasoning accuracy in future games.
[0107] Specifically, the scheme execution module receives the locked digital scheme and selects the execution mode based on the task attributes: for standardized steps, it compiles them into equipment control code to drive the automated workstation to execute; for non-standard steps, it generates SOPs to guide the experimenters. Regardless of the mode, all sensor data generated during the experiment and the final physicochemical test results are automatically collected and cleaned by the system into standardized experimental data records (EDRs). This real-world physical data then flows back to the knowledge base module through the data channel to update the system's prior knowledge, thus completing the closed loop of "virtual design - physical verification - knowledge update" at the micro level.
[0108] Step S6: Manual acceptance and task closure.
[0109] The final experimental results are fed back to the user through a human-computer interaction interface, allowing the user to make the final value judgment. If the user confirms that the research and development goals have been achieved based on the measured data, the task officially ends. If the user believes the results are not as expected, they can input new feedback. The system injects this new feedback as new constraints into the game environment, reactivating the solution generation module (jumping back to step S2), thereby breaking the current equilibrium state and starting a new round of adversarial optimization loop until the optimal solution that meets the user's needs is obtained.
[0110] The final experimental results are presented to users through a visual interface.
[0111] Branch decision (e.g.) Figure 2 (As shown in the middle diamond shape): Users make value judgments on the achievement of R&D goals based on actual test data.
[0112] If the user confirms that the goal has been achieved (path "Yes"), then this research and development task is officially completed.
[0113] If the user believes the expected results are not met (path "No"), the user can input specific feedback (e.g., "Intensity meets the standard, but color is incorrect"). This feedback will be injected into the system as new constraints, and the process will jump back to step S2 according to the large loop path. This is not just a simple retry, but rather uses the failed data as negative samples to reactivate the solution generation module to find a better solution in the new constraint space, thereby breaking the current local equilibrium and driving the system to evolve towards the global optimum.
[0114] This application provides a materials research system and method based on multi-agent game theory, the main core of which is: On the one hand, this application proposes a method for classifying agent roles and constructing a game framework based on conflicting interests. This application redefines function-oriented agents in traditional collaborative systems as game participants with differentiated and partially conflicting objective functions (payoff functions). Specifically, it divides them into "proposer agents" (whose core payoff lies in the theoretical performance of the solution) and multiple "adversary agents" (whose core payoffs lie in controlling costs, ensuring security, and meeting process requirements, respectively). This application defines at least one proposer agent and at least two adversary agents with different objective functions; wherein the objective functions of the adversary agents are configured to be independent of each other and have a predetermined conflicting relationship with the objective function of the proposer agent.
[0115] On the other hand, this application proposes an arbitration rule engine that drives multi-round adversarial debate and dynamic compromise. The arbitration module of this application is not a simple summary scorer, but rather a game theory rule executor containing a set of algorithmic rules driving game convergence. Its core features are: a) Triggering multiple rounds of iteration: When any opposing party's objection exceeds a threshold, a forced modification and re-examination of the proposed solution is triggered; b) Dynamic trade-offs: Based on the specific defects of the proposed solution in this round, the weighting factors of different objectives (performance, cost, security) in the arbitration are dynamically adjusted; c) Convergence determination: A convergence condition based on game theory (such as the intensity of all opposing parties' objections being below a threshold, or the rate of change of the proposed solution approaching zero) is used to determine whether a Nash equilibrium has been reached, rather than a fixed number of rounds. The arbitration module of this application can collect the proposer's solution and the review opinions of all parties; dynamically calculate and apply a trade-off matrix to generate modification instructions based on the strength and type of objections; determine whether to end the iteration based on a game theory convergence model; and output the final solution when convergence is determined.
[0116] Furthermore, this application proposes a policy update mechanism that uses physical experiment results as "strong constraint signals." This application treats the failure results of physical experiments as high-weight negative feedback (penalty) to relevant agents (especially proposers), directly and significantly reducing the payoff estimate of their policy models. This forces agents to prioritize avoiding features of solutions that previously led to experimental failures (such as a precursor or a process window) in subsequent games, thereby achieving "adversarial evolution" towards engineering feasibility. Specifically, experimental failure records are associated with the features of the solutions that led to the failure; in subsequent games, when a candidate solution generated by the agent contains these features, a penalty term is directly applied to its objective function, or the valuation of the corresponding action in its policy model is lowered.
[0117] Furthermore, this application proposes a closed-loop game theory workflow of "proposal-confrontation-arbitration." This integrates all the aforementioned points into a complete, automatically running workflow. The process begins with task analysis, proceeds through multiple cycles of "generation → adversarial review → arbitration and feedback → iterative modification" until the game converges, and then enters the experimental verification and strategy update phase. This workflow itself is a creative methodological invention.
[0118] Finally, this application also proposes a modular system architecture to support game-theoretic interaction. The core of the specific system architecture designed to implement the aforementioned game-theoretic process lies in the fact that the functions and interaction relationships of each module are specifically designed to support high-frequency, asynchronous adversarial interactions. In particular, the arbitration module acts as a central process controller, the message bus serves as the interaction backbone, and the knowledge base provides support as a real-time data service layer.
[0119] Compared with existing technologies, the material research system and method based on multi-agent game theory provided in this application have the following advantages: 1. It achieves dynamic intelligent trade-offs in multi-objective optimization, improving the overall balance and engineering practicality of the solution.
[0120] Existing technologies use fixed weights for optimization, which cannot flexibly handle conflicts between different objectives in specific solutions. This application, through a dynamic weight strategy in the arbitration module and a convergence determination mechanism based on the rate of change, enables the system to adjust the optimization focus in real time based on the specific feedback from each professional agent in each round of review. For example, when the cost agent consistently gives low scores and provides market data evidence, the system will automatically emphasize cost optimization in subsequent iterations. This dynamic adjustment capability based on real-time interaction ensures that the final output solution is a balanced solution reached after sufficient game theory under multiple constraints, avoiding the rigidity of static methods and significantly improving the practical feasibility of the solution.
[0121] 2. Through in-depth verification using integrated professional tools, early and accurate identification and interception of solution defects were achieved.
[0122] Existing methods often lack sufficient depth in their scheme review. In this application's review module, each agent possesses the ability to directly invoke external specialized computing software, real-time databases, and rule bases for verification. For example, the theoretical agent can quantitatively assess material stability by invoking first-principles calculations, while the safety agent can automatically verify whether a scheme complies with the latest safety standards. This in-depth review mechanism based on objective tools and data can uncover deep-seated, interdisciplinary risks (such as thermodynamic instability, supply chain bottlenecks, and compliance loopholes) that are difficult to detect using traditional methods. This effectively filters out high-risk schemes before investing in expensive experiments, significantly reducing R&D costs and timelines.
[0123] 3. A closed-loop learning mechanism centered on real experimental feedback has been established to drive R&D directions to continuously align with engineering practice.
[0124] Existing closed-loop systems combining virtual and real-world approaches offer relatively weak feedback for correcting optimization directions. This application standardizes the results of physical experiments, particularly failure data, and directly links them to the knowledge base and the decision-making logic of the agent. When a certain feature of a solution (such as using a specific precursor) is experimentally proven infeasible, the relevant agent will impose a significant negative evaluation or rejection on solutions containing that feature in subsequent games. This allows the system's overall optimization strategy to gradually be "shaped" by real-world physical and engineering constraints, moving beyond purely "theoretical performance optimization," ensuring that R&D activities always proceed along an engineeringally feasible path and fundamentally improving the success rate of R&D.
[0125] 4. It provides a transparent and controllable interaction process, realizing efficient and reliable human-machine collaboration.
[0126] Existing automated systems often compromise user trust due to a lack of transparency in their processes. This application addresses this by using a human-computer interaction module to display the complete game process and decision-making basis in real time. When the system reaches an iterative deadlock, it proactively requests human intervention for strategic decision-making. This allows R&D experts to fully utilize the efficiency of automated iteration while clearly understanding the underlying logic of system decisions, and to inject human experience and macro-level judgment at critical junctures. This collaborative model, where humans and machines each perform their respective roles and complement each other's strengths, enhances the reliability and usability of the entire system and makes it easier to integrate into actual R&D workflows.
[0127] 5. A standardized and automated end-to-end digital R&D pipeline has been built, improving R&D efficiency and process quality.
[0128] Traditional R&D processes are fragmented and heavily reliant on manual operation and data transfer. This application seamlessly integrates task parsing, intelligent design, multi-dimensional review, arbitration decision-making, experiment execution, and data feedback through a microservice architecture and a unified data interface, forming a highly automated digital workflow. This significantly reduces delays and errors caused by manual intervention, ensures a high degree of consistency and repeatability in the R&D process, and makes all data, decisions, and status changes traceable throughout the entire R&D process, laying a solid foundation for refined and continuous improvement of R&D management.
[0129] Based on the above technical solutions, this application provides a materials research system and method based on multi-agent game theory. The system includes a human-computer interaction module, a solution generation module, a review module, an arbitration module, a solution execution module, and a knowledge base module. The human-computer interaction module generates task instructions based on user needs; the solution generation module generates multiple candidate solutions; the review module conducts multi-perspective cross-reviews and risk interception of candidate solutions; the arbitration module handles conflicts of opinion among multiple agents; the solution execution module executes experiments and collects data; and the knowledge base module provides static knowledge support and a closed-loop feedback of dynamic experimental data for the system. This application transforms the multi-objective conflicts, knowledge uncertainty, and gap between theory and practice in materials research and development into a structured, computable dynamic game process, thereby systematically improving the comprehensiveness of solution generation, the rigor of review, and the success rate of experimental transformation.
[0130] Those skilled in the art will understand that the above-described embodiments are specific examples of implementing this application, and in practical applications, various changes in form and detail may be made without departing from the spirit and scope of this application. Any person skilled in the art can make their own modifications and alterations without departing from the spirit and scope of this application; therefore, the scope of protection of this application should be determined by the scope defined in the claims.
Claims
1. A materials research system based on multi-agent game theory, characterized in that, It includes a human-computer interaction module, a solution generation module, a review module, an arbitration module, a solution execution module, and a knowledge base module. The human-computer interaction module is used to generate task instructions according to user needs; The scheme generation module is used to generate multiple candidate schemes corresponding to different technical routes according to the task instructions; The review module is used to perform multi-perspective cross-review and risk interception of the candidate solutions to obtain the review results of the multi-agent system. The arbitration module is used to process the conflicting opinions among multiple agents based on the review results of the multiple agents and output the final decision. The scheme execution module is used to execute experiments and collect data based on the final decision; The knowledge base module is used to provide the system with static knowledge support and a closed-loop feedback of dynamic experimental data.
2. The materials research system based on multi-agent game theory according to claim 1, characterized in that, The human-computer interaction module includes a receiving unit, a conversion unit, a display unit, and a human intervention interface; The receiving unit is used to receive R&D requirements input by the user; R&D requirements include performance indicators, cost budgets, and process constraints; The conversion unit is used to convert unstructured natural language and parameters in the R&D requirements into standardized task instructions that can be recognized by the system. The display unit is used to visualize the system's decision-making process, presenting to users the interaction of review opinions and the solution iteration path among multiple agents in the review module, thus making the R&D process transparent. The manual intervention interface is used to allow users to make final strategic decisions or adjust parameters when a consensus cannot be reached within the system.
3. The materials research system based on multi-agent game theory according to claim 1, characterized in that, The scheme generation module includes an initial scheme generation unit and a filtering unit; The initial scheme generation unit is used to design candidate material schemes, including raw material formulations, reaction paths, and process parameters, based on existing scientific principles and knowledge reserves, through reasoning and combination. The screening unit is used to perform preliminary screening of the chemical structure rationality of candidate material schemes to ensure that the generated candidate schemes have basic theoretical validity.
4. The materials research system based on multi-agent game theory according to claim 1, characterized in that, The review module includes multiple evaluation units with specific domain perspectives, each responsible for independently verifying candidate solutions from the aspects of scientific principles, engineering feasibility, and safety compliance. Each evaluation unit not only outputs a quantitative score, but also provides suggestions for improvement or objections regarding the identified defects.
5. The materials research system based on multi-agent game theory according to claim 1, characterized in that, The arbitration module includes a discrimination unit; When different evaluation units give conflicting evaluation results for the same solution, the discrimination unit automatically makes a judgment based on the preset decision logic or weighting strategy and selects the solution with the best overall benefits. In complex situations involving critical risks or disagreements exceeding system thresholds, the arbitration module triggers a suspension mechanism, routing decision-making power to the human-computer interaction module for manual confirmation.
6. The materials research system based on multi-agent game theory according to claim 1, characterized in that, The execution module can dynamically select the appropriate execution mode based on the complexity of the experimental scheme, the required precision of operation, and the current configuration of laboratory hardware resources.
7. The materials research system based on multi-agent game theory according to claim 1, characterized in that, The knowledge base module is responsible for the storage, management and service of knowledge for the entire system; it integrates static scientific knowledge as well as experimental data dynamically generated during system operation. The knowledge base module provides retrieval services to the generation and review modules through indexing technology, enabling them to refer to relevant historical experience when making decisions. By continuously incorporating new experimental feedback data, the knowledge base module can continuously update its data reserves, thereby supporting the self-correction and capability improvement of the system's decision-making model.
8. A materials research method based on multi-agent game theory, applied to the materials research system based on multi-agent game theory as described in any one of claims 1 to 7, characterized in that, Includes the following steps: Step S1: The user inputs the target requirements for material research and development through the human-computer interaction module. The system parses the user's instructions and generates task instructions. Step S2: Based on the task instructions, the solution generation module starts the design process and generates multiple candidate solutions corresponding to different technical routes. Step S3: Each agent in the review module reviews the candidate solutions, identifies potential vulnerabilities and risks from different dimensions, and initiates high-intensity adversarial verification of the solutions. During the review process, the arbitration module acts as a game manager, responsible for collecting scattered opposing opinions and aggregating and deduplicating them; Step S4: When the solution generation module submits the iterative solution, the arbitration module first makes a preliminary judgment based on the scoring model. If the main indicators of the solution have not converged to the preset threshold, it indicates that the game has not yet reached equilibrium, and the system will continue to iterate in step S3. When the arbitration module determines that the solution has become perfect and the game process shows a convergence trend, the final review confirmation mechanism is initiated to verify whether the system has truly reached the Nash equilibrium state. The arbitration module resubmits the solution to all review agents for review. Only when all opposing parties no longer raise key objections and unanimously give feasible conclusions, the system determines that the game has reached a consensus, and the arbitration module then locks the solution and moves it to the next stage. Step S5: The scheme execution module receives the final decision locked by the game and selects automatic execution or generates standard operating procedures to guide manual execution based on the experimental conditions. The data and final results during the experiment will be collected and formatted in a standard way, and then sent back to the knowledge base module. The real physical data during the experiment will be used as new prior knowledge to correct the value network of each agent and improve the reasoning accuracy of the system in future games. Step S6: The final experimental results are fed back to the user through the human-computer interaction interface, and the user makes the final value judgment; if the user confirms that the research and development goal has been achieved based on the actual test data, the task is officially over. If the user feels that the result is not as expected, they can input new feedback. The system will inject these new feedbacks as new constraints into the game environment, reactivate the solution generation module, jump back to step S2, thereby breaking the current equilibrium state and starting a new round of adversarial optimization loop until the optimal solution that meets the user's needs is obtained.
9. The material research method based on multi-agent game theory according to claim 8, characterized in that, In step S2, the solution generation module, acting as the generator in the game, initiates the design process according to the requirements of the task instructions. The solution generation module first performs a wide-area search of the knowledge base, reviewing existing historical experimental cases, as well as unstructured knowledge content such as academic literature, professional books and patent materials; Subsequently, the generation module uses the large model to perform comprehensive reasoning on the above information and constructs one or more initial candidate solutions within the limited search space; the initial candidate solutions are submitted to the arbitration module as the objects of the first round of the game.
10. The material research method based on multi-agent game theory according to claim 8, characterized in that, In step S3, the arbitration module introduces the candidate solutions into the review environment. Each agent in the review module acts as an adversary, attempting to identify potential vulnerabilities and risks in the solutions from different dimensions, thereby initiating a high-intensity adversarial verification of the solutions. In this process, the arbitration module acts as a game manager, responsible for collecting scattered opposing opinions and aggregating and deduplicating them. When there are conflicts in the review results of different dimensions, the arbitration module balances the interests of all parties according to the strategy weights and forms a unified correction feedback. The solution generation module then supplements and optimizes the solution based on the feedback and puts the new solution back into the review environment, thus forming a dynamic game cycle in which the generating party continuously improves its defense and the opposing party continuously looks for loopholes.
Citation Information
Patent Citations
Multi-agent cooperation system and method for material science
CN118737346A
Large model material formula design and evaluation method and system based on multi-agent collaboration and storage medium
CN120878009A
Formula process optimization method, device and equipment based on multiple agents
CN120911121A
Multi-system collaborative data cross validation management system
CN121389162A
Multi-agent-based material performance prediction and synthesis method and system
CN121506290A