Large model material formula design and evaluation method and system based on multi-agent collaboration and storage medium
By employing a large-scale model approach based on multi-agent collaboration, the problem of lack of data governance and mechanism constraints in material formulation design is solved, enabling efficient and accurate formulation generation and evaluation, and filling the application gap of multi-agent systems in material formulation design scenarios.
Patent Information
- Application Number
- CN202511408661.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-29
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2045-09-29
AI Technical Summary
Existing material formulation development relies on experimental trial and error and experience, which is time-consuming and costly. A single large model cannot simultaneously complete multi-source data cleaning and quality control. The lack of mechanistic constraints leads to unreliable generated formulations. Multi-agent systems lack full-process integration in material formulation design.
A large-scale model approach with multi-agent collaboration is adopted. Data agents clean multi-source data, generating agents generate initial formulas, mechanism agents screen compliant formulas, evaluation agents perform performance estimation and risk assessment, and decision agents output the optimal formula. By combining an explicit mechanism knowledge base and a lightweight prediction model, a closed-loop collaboration between data governance and model generation is achieved.
It improves the efficiency and accuracy of material formulation design, reduces invalid experimental verification, ensures that the formulation meets environmental protection standards and performance requirements, supports pluggable large model adaptation, and improves R&D efficiency.
Smart Images

Figure CN120878009A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of artificial intelligence-assisted material development and multi-agent technology, specifically to a method, system, and storage medium for large-scale model material formulation design and evaluation based on multi-agent collaboration. Background Technology
[0002] In the current field of materials formulation research and development, the traditional model relies heavily on experimental trial and error and experience accumulation. Not only does the research and development cycle take several months to several years, but it is also accompanied by high time and economic costs. With the penetration of artificial intelligence technology, single large models have begun to be used for materials formulation generation, becoming a new direction to improve research and development efficiency. However, the existing technology system has not yet broken through the core limitations of the traditional model.
[0003] Existing technologies suffer from three core problems: First, data governance and model generation are disconnected. A single large model cannot simultaneously clean and control the quality of multi-source material data, and the high noise in the input data leads to unreliable basic data for generating formulations. Second, the lack of mechanistic constraints results in illusory formulations. A single large model generates formulations based solely on statistical data patterns without incorporating the physicochemical mechanisms of materials science, easily producing formulations that defy common sense and require extensive manual post-processing, leading to low efficiency. Third, there is a gap in multi-agent applications. Although multi-agent systems have been successfully applied in software development, games, and other fields, in material formulation design scenarios, there is a lack of a computer architecture that organically integrates the entire process of data generation, screening, evaluation, and decision-making. Existing solutions mostly involve single models and manual post-processing, failing to form a closed-loop intelligent collaborative mechanism. Therefore, this paper proposes a method, system, and storage medium for large-scale material formulation design and evaluation based on multi-agent collaboration to address these problems. Summary of the Invention
[0004] Technical problems to be solved To address the shortcomings of existing technologies, this invention provides a method, system, and storage medium for designing and evaluating large-scale material formulations based on multi-agent collaboration, thus solving the problems mentioned in the background section.
[0005] Technical solution To achieve the above objectives, this invention provides the following technical solution: a large-scale model material formulation design and evaluation method based on multi-agent collaboration, such as... Figure 1 As shown, it includes the following steps: S1. Receive the target material performance parameters and constraints input by the user through a computer interactive interface; Specifically, the computer interface is a visual operation interface deployed on a server or in the cloud. It supports users to submit requirements through form input, drop-down selection, etc. The interface has a built-in parameter validation function to verify the numerical range and format of the input content to avoid invalid input. The target material performance parameters are the core performance indicators that the user expects the material to achieve, and the constraints are the limiting rules that must be followed during the formulation design process. After both are input, they will be automatically converted into standardized JSON format and pushed to the message bus of the intelligent agent running framework via API, providing a clear basis for subsequent formulation design.
[0006] S2. The data intelligence agent accesses multi-source material data through a preset data interface, cleans and structures the multi-source material data, generates vector data, and stores it in shared memory; Specifically, the data agent is deployed on designated computer hardware nodes, with preset data interfaces covering database, file, and spectral interfaces, capable of accessing different types of multi-source material data. The data cleaning process follows a workflow of "outlier detection—missing value completion—format unification." Outlier detection uses the IQR method, calculating the interquartile range (ICR) and removing data outside the range. Missing value completion uses the mean of similar formulations. Format unification converts parameters from different units to standard units, and spectral data is converted into numerical matrices. Structured processing transforms the cleaned data into fixed-dimensional vectors, with dimensions including component content, process parameters, spectral features, and measured performance. The generated vector data is ultimately stored in shared memory, providing data support for the generation of the agent.
[0007] S3. The generated intelligent agent calls the multimodal large model or language large model through the message bus, and generates several initial material formulas based on the vector data, target material performance parameters and constraints. The message bus employs a distributed message queue. Generating agents subscribe to specific topics to obtain vector data, target material performance parameters, and constraints from shared memory. The multimodal or language-based large-scale models invoked by the generating agents can be accessed via local interfaces or cloud APIs. During the invocation process, the historical sample information contained in the vector data, the specific requirements of the target material performance parameters, and the constraint rules are combined to generate several initial material formulations according to preset logic. The generated initial material formulations are stored in shared memory in JSON format, and relevant notifications are pushed through the message bus to trigger subsequent processing by the mechanistic agent.
[0008] S4. The mechanism intelligent agent calls domain mechanism rules from the explicit mechanism knowledge base to perform compliance screening on the several initial material formulations and delete illegal formulations that do not conform to common sense of physicochemicals; The explicit mechanism knowledge base stores physicochemical mechanism rules for the materials domain in a specific format. The knowledge base supports both manual and automatic updates to ensure the timeliness and accuracy of the rules. After obtaining initial material formulations from the message bus, the mechanism agent sequentially calls the domain mechanism rules corresponding to the material type in the explicit mechanism knowledge base, performing a line-by-line matching and verification for each initial material formulation. If a formulation violates any domain mechanism rule, it is determined to be an illegal formulation and deleted; if a formulation meets all domain mechanism rules, it is determined to be a compliant formulation and retained in shared memory. After the screening is completed, the mechanism agent pushes a notification of the screening results through the message bus, providing a basis for evaluating the agent's work. S5. The evaluation agent performs performance estimation and risk assessment on the remaining formulations after screening, and outputs the performance prediction value, risk level and confidence score of each formulation. The evaluation agent can invoke the corresponding performance estimation logic based on the material type, achieving cross-material adaptation: For coating materials: Simultaneously estimate decorative and functional parameters such as gloss, leveling, and adhesion; For rubber-based materials: Simultaneously estimate mechanical parameters such as abrasion resistance index, elastic modulus, and compression set; For plastic materials: Simultaneously estimate physical parameters such as heat distortion temperature, flexural modulus, and impact strength; For ceramic materials: Simultaneously estimate performance parameters such as compressive strength, acid and alkali resistance, and sintering shrinkage rate; Performance prediction values must be clearly labeled with units, such as rubber abrasion resistance index 168cm. 3 / 1.61km, plastic heat distortion temperature 125℃, ceramic compressive strength 320MPa, stored in shared memory for later retrieval; VOC content is a specific performance parameter for volatile materials such as coatings, adhesives, and inks, namely the content of volatile organic compounds, expressed in g / L. The lower the value, the more environmentally friendly the material. For non-volatile materials such as rubber, plastics, and ceramics, this parameter can be replaced with specific environmental protection parameters, such as the harmful heavy metal content of rubber materials (mg / kg), the phthalate content of plastic materials (%), and the lead and cadmium leaching of ceramic materials (mg / L). All environmental protection parameters comply with the latest national and industry standards.
[0009] The data agent accesses multi-source material data through a pre-defined data interface. The multi-source material data includes historical formula data, process parameter data, experimental record data, etc. The multi-source material data is cleaned (including outlier detection, missing value completion, and format unification) and structured, and finally structured vector data containing component content, process parameters, spectral features, and measured performance is generated. This data is stored in shared memory as the basis for evaluating the predicted value of the agent's computing performance. The evaluation agent retrieves the structured vector data from shared memory and divides it into training and validation sets in a 7:3 ratio. Both datasets serve as the training foundation for the lightweight prediction model. The evaluation agent uses lightweight models such as gradient boosting trees or linear regression, with "recipe vector + process parameters" as input features and "measured performance from experimental records" as output labels. The model is trained using the training set. During training, the model parameters are continuously adjusted using the validation set to ensure goodness of fit (R²). 2 The accuracy of the model prediction is ensured by setting the value ≥0.85. After training, the model is saved in a lightweight format (such as ONNX format) and deployed on the local node of the evaluation agent for easy and quick access later. After the mechanistic agent completes the screening of illegal formulations, the evaluation agent reads the structured vector data (including key information such as component content and process parameters) of the remaining compliant formulations from shared memory and inputs this data into a pre-trained lightweight prediction model. Based on the input formulation vector features, the model automatically outputs the performance prediction value of the corresponding formulation (e.g., "anti-wear index prediction value 168cm"). 3 / 1.61km" and "VOC content prediction value 42g / L"); the performance prediction values must be clearly labeled with units and then stored in shared memory to provide data support for subsequent risk assessment and credibility calculation. If the historical data sample size for a certain type of formulation is too small (e.g., sample size < 50 records, which cannot support the training of a lightweight prediction model), then an empirical formula is used to calculate the performance prediction value. The empirical formula is derived from fitting a large amount of experimental data in the field of materials (e.g., the empirical formula for rubber abrasion resistance index: abrasion resistance index = 120 + 2 × carbon black content (phr) - 5 × sulfur content (phr). The parameters (coefficients, constants, etc.) in the formula are stored in the "empirical formula parameter table" in shared memory and will be updated regularly with the experimental records added by the data agent to ensure the accuracy of the empirical formula calculation.
[0010] Risk assessment classifies risk levels based on the deviation rate between predicted performance values and target parameters. Based on the magnitude of the deviation rate, risk levels are categorized into three types: low risk, medium risk, and high risk. The specific steps are as follows: Based on the deviation rate, the risk level is divided into three categories: when the deviation rate is ≤10%, it is judged as low risk (the performance prediction value is highly consistent with the target requirements); when the deviation rate is 10% < ≤20%, it is judged as medium risk (the performance prediction value is close to the target requirements, but potential fluctuations need to be monitored); when the deviation rate is >20%, it is judged as high risk. The risk level of each formula is associated with the corresponding deviation rate calculation results and stored in shared memory for subsequent decision-making agents to call.
[0011] The confidence score is calculated using an uncertainty quantification algorithm. This algorithm is based on the historical error rate and data sample size of the prediction model. First, the historical error rate and data sample size are standardized. Then, the uncertainty is calculated by weighted summation. Finally, the uncertainty is converted into a confidence score to achieve accurate quantification of the confidence of the formula.
[0012] The credibility score is obtained as follows: In the formula, This indicates the credibility score. The standardized historical error rate is calculated based on the historical error rate of the prediction model and the model's historical maximum error rate. The data comes from historical experimental records after cleaning by the data agent. The standardized sample size is calculated based on the current task-matched sample size and adjustment coefficients. The data comes from structured vector data filtered by the data agent. and These are the historical weighting factor and the sample weighting factor, respectively. and The sum is 1, for example Indicates 0.6 and 0.4 represents a preset fixed parameter, reflecting the influence of historical error rate and sample size on reliability.
[0013] S6. The decision-making agent reads the performance prediction value, risk level and credibility score from the shared memory and performs a comprehensive calculation to output the top-ranked candidate recipes; if the recipe with the highest comprehensive score does not reach the preset threshold, the generating agent is triggered to re-execute step S3 to generate a new recipe. After extracting performance predictions, risk levels, and credibility scores from shared memory, the decision-making agent uses a combination of rule trees and simple reinforcement learning to calculate a comprehensive score. The specific process is as follows: The rule tree stage prioritizes basic filtering of recipes. Recipes are verified layer by layer in the order of risk level, credibility, and core performance: recipes with risk levels exceeding preset standards are directly eliminated; for example, recipes marked as high-risk are filtered out due to potential excessive experimental bias. Recipes with credibility scores below a set threshold are also excluded because their prediction results have insufficient reference value. Recipes whose core performance does not meet the target baseline are also eliminated, ensuring that the remaining recipes meet the basic needs of R&D. Each round of filtering generates a clear elimination list, recording the IDs and specific reasons for elimination, making the filtering process traceable. Recipes filtered through the rule tree enter the reinforcement learning scoring stage. A scoring model is built using performance achievement rate, cost control rate, and credibility score as core indicators to calculate the comprehensive score. The specific method for obtaining the overall score is as follows: In the formula, This represents the overall score, used for recipe ranking. It outputs the top-ranked candidate recipes. If the highest score does not reach a preset threshold, the generating agent is triggered to regenerate. This represents the performance compliance rate, calculated based on the degree of match between the predicted performance value and the target requirement, and retrieved from shared memory. This represents the cost control rate, calculated based on the degree of matching between the formula cost and the cost threshold, and retrieved from shared memory. The confidence score represents the result of evaluating the agent's output, which is directly retrieved from shared memory. , as well as These are performance weighting factors, cost weighting factors, and credibility weighting factors, respectively, and their sum is 1, for example... , as well as The values are 0.4, 0.3, and 0.3 respectively.
[0014] Table 1. Overall Score Evaluation Table Formula ID Performance compliance rate Cost control rate Credibility score Overall score 01 98 95 85 93.2 021 98 80 80 87.2 025 85 98 78 86.8 Based on the table above, select 3 different formulas for comprehensive score evaluation, rank the formulas according to the comprehensive score evaluation results, and output the top-ranked candidate formulas.
[0015] Reasons for selection: 01. In terms of performance, the performance compliance rate reaches 98%, exceeding the target value by 5 percentage points. Key indicators such as the wear resistance index are better than the industry average. In terms of cost, the cost control rate is 95%, which is 3% lower than the preset threshold, and production costs can be reduced during mass production. In terms of reliability, the score reaches 8.5 points, and the success rate of historical experiments on similar formulas reaches 90%, indicating high reliability of the results. Regarding 025, its performance compliance rate is only 85%, failing to meet the core performance targets set in the R&D. If it's a rubber-based anti-wear formula, its predicted anti-wear index is only 143 cubic centimeters per 1.61 kilometers, lower than the R&D target of no less than 150 cubic centimeters per 1.61 kilometers. If it's a plastic heat-resistant formula, its predicted heat resistance temperature is only 115 degrees Celsius, failing to meet the basic standard of no less than 120 degrees Celsius set in the R&D. The inability to meet core functional requirements and substandard performance directly renders subsequent experimental verification meaningless. Even if the experiments are successful, the finished product will not be suitable for end-user scenarios. In terms of cost, although the cost control rate reaches 98%, the actual estimated cost is 25.5 yuan per kilogram, which is lower than the preset threshold of 26 yuan per kilogram, giving it a certain cost advantage. However, this advantage is based on substandard performance. If this formula is forcibly adopted, additional resources need to be invested to adjust the composition to improve performance, such as increasing the carbon black content to improve abrasion resistance and adding heat-resistant additives to increase heat resistance temperature. After the adjustment, the cost will exceed the threshold, with the estimated cost per kilogram rising to 28 yuan, ultimately causing the cost advantage to disappear. Therefore, this cost advantage has no practical application value.
[0016] In terms of reliability, the reliability score of 78 is the lowest among the three formulation groups. This score is derived from the uncertainty quantification algorithm used to evaluate the agent in the document. On the one hand, the prediction model has a high historical error rate for formulations with performance close to the benchmark, with an average absolute error of 5.6, higher than 4.2 for formulation 01 and 4.8 for formulation 021. On the other hand, the number of historical samples matching this formulation for the current task is only 90, fewer than 180 for formulation 01 and 120 for formulation 021, indicating insufficient data support. Both factors contribute to the low reliability of the prediction results for this formulation. The success rate of historical experiments with similar formulations is only 75%, lower than 90% for formulation 01 and 85% for formulation 021. If experimental verification is conducted, the trial-and-error risk is significantly higher than that of the first two formulation groups.
[0017] Reasons for selection: For example, the performance compliance rate of formula 021 is the same as that of the preferred formula, but the cost control rate is only 80%, which exceeds the threshold by 5%, resulting in a lower overall score of 0.8 points; Formula 025 has a clear cost advantage, but the performance compliance rate is only 85%, which does not meet the core performance requirements, so it is listed as the second choice. At the same time, a preset threshold for the overall score is set. If the overall score of the top-ranked recipe does not reach the threshold, a regeneration instruction is sent to the generating agent through the message bus to trigger the generating agent to re-execute step S3 to generate a new recipe, ensuring that the output recipe meets the R&D standards.
[0018] S7. The top-ranked candidate formulas are returned to the user through a computer interface. Optionally, the user's manual confirmation instruction can be received, and the confirmed formulas can be pushed to the experimental docking interface. The top-ranked candidate formulations output by the decision-making agent will be displayed to the user in tabular form through a computer interface. The table includes information such as the formulation's ingredients, predicted performance values, risk level, confidence score, and overall score. Users can view the formulation details. If they choose manual confirmation, the interface will record the user's confirmation information and generate an experimental task sheet. The experimental task sheet includes formulation details, experimental equipment requirements, and testing standards, and is pushed to the laboratory management system through the experimental interface. If the user does not perform manual confirmation, the interface will automatically save the candidate formulations to the history module for later retrieval and access.
[0019] Specifically, the target material performance parameters are adapted according to the material research and development needs. Examples include at least one of the following for rubber materials: abrasion resistance index, hardness, tensile strength, impact strength, elongation at break, and anti-aging properties; VOC content and gloss for coating materials; heat resistance temperature and flexural modulus for plastic materials; and compressive strength and corrosion resistance for ceramic materials. Constraints include at least one of the following: ingredient prohibition rules are dynamically updated based on industry regulations; cost thresholds support user-defined units and ranges; and process compatibility requirements are adapted to existing injection molding processes.
[0020] Specifically, the multi-source material data includes historical formula data, process parameter data, spectral data, and experimental record data. The data agent employs a traceable data version control mechanism, adding a source identifier, processing timestamp, and version number to each processed data entry to ensure training data quality. Historical formula data represents the material composition ratios used in past R&D or production; process parameter data represents production process conditions matching the formula; spectral data represents the material's structural analysis data; and experimental record data represents the measured performance results corresponding to the formula. The data agent uses a traceable data version control mechanism, adding a source identifier, processing timestamp, and version number to each processed data entry. The source identifier clearly identifies the data acquisition device or database table name; the processing timestamp is accurate to milliseconds, recording the specific time of data processing; the version number uses the "major version.minor version.revision number" rule to distinguish different versions of the data. This mechanism ensures training data quality and facilitates subsequent data traceability.
[0021] Specifically, the generated agent supports a pluggable large model adaptation mechanism, allowing it to call private large models via a local interface or public large models via a cloud API. It also features an automatic prompt word optimization module that iteratively adjusts the prompt word logic based on historical generation results. This mechanism includes a built-in model adaptation interface abstract class, defining three abstract methods: loading the model, generating the recipe, and releasing resources. Different types of large model calling classes inherit from this abstract class and implement the specific methods. When calling a private large model via the local interface, the model file deployed on the local GPU server is loaded; when calling a public large model via the cloud API, the API key is initialized and a request is sent. The generated agent has an automatic prompt word optimization module that maintains a "prompt word template - compliance rate" mapping table. After generating a certain number of recipes, the compliance rate of each template is calculated, templates with higher compliance rates are retained, and the template content is adjusted based on the defects of new recipes. Through iterative optimization, the quality of prompt words is improved, thereby enhancing the effectiveness of recipe generation.
[0022] Specifically, the domain mechanism rules include basic general rules and material-specific rules, and the explicit mechanism knowledge base supports rule expansion; The basic general rules include solubility parameter matching rules based on Hansen theory, component compatibility rules to avoid chemical reaction conflicts, and rules prohibiting harmful components that comply with at least one of the national or industry standards. Examples of material-specific rules include at least one of the following: vulcanization reaction ratio rules for rubber materials, solvent volatility rules for coating materials, polymerization reaction temperature rules for plastic materials, and sintering aid ratio rules for ceramic materials; The explicit mechanism knowledge base supports importing new rules through rule templates, including rule logic, parameter thresholds, and applicable material types. The solubility parameter matching rule is based on Hansen's solubility parameter theory, which determines compatibility by calculating the difference in solubility parameters between the polymer and the solvent. The vulcanization reaction ratio rule is based on the chemical equation of the vulcanization reaction, which determines the reasonable ratio range of vulcanizing agent and accelerator to ensure that the crosslinking density of materials such as rubber meets the requirements. The solvent prohibition rule is based on national environmental protection and safety standards, which clearly defines the list of prohibited solvents to avoid the presence of harmful solvents in the formulation.
[0023] Specifically, the evaluation agent calculates a confidence score using an uncertainty quantification algorithm, which is determined based on the historical error rate of the prediction model and the data sample size. The evaluation agent calculates a confidence score using an uncertainty quantification algorithm. The core parameters of this algorithm are the historical error rate of the prediction model and the data sample size. The historical error rate is calculated using the mean absolute error, and the data sample size is the total number of historical samples matching the current formulation design task. The calculation process first standardizes the historical error rate and data sample size, then calculates the uncertainty through weighted summation, and finally converts the uncertainty back into a confidence score, achieving a quantitative representation of confidence.
[0024] Large-scale model material formulation design and evaluation system based on multi-agent collaboration, such as Figure 2 As shown, it includes: The intelligent agent operation framework includes a message bus and shared memory. The message bus is used to realize data interaction between intelligent agents and supports the addition of new material type topics such as ceramic formulas and metal alloy formulas. Each intelligent agent realizes data interaction by subscribing to the corresponding topic. The shared memory is used to store intermediate data, and is divided into a general data area to store rule tree thresholds and dynamic weight models, and a material-specific data area to store the mechanism rules and empirical formula parameters of each material. When adding a new material, only the dedicated data area needs to be expanded. The framework also reserves an interface for adding intelligent agents, and life cycle assessment intelligent agents and process optimization intelligent agents can be added as needed. They can be integrated into the existing process by registering a message bus topic. The data intelligence agent is deployed on computer hardware nodes and is used to uniformly access multi-source material data through a preset data interface, clean and structure the multi-source material data, and generate vector data. A smart agent is generated and communicates with the data smart agent through a message bus. It is used to call a multimodal large model or a language large model to generate several initial material formulations based on vector data, target performance parameters and constraints. Mechanistic intelligent agents, with a built-in explicit mechanism knowledge base, are used to call domain mechanism rules to screen initial material formulations for compliance and delete illegal formulations; The evaluation agent is used to perform performance estimation and risk assessment on the screened formulations using lightweight prediction models or empirical formulas, and outputs performance prediction values, risk levels and confidence scores. The decision-making agent reads the evaluation results from shared memory, uses a combination of rule trees and simple reinforcement learning to perform comprehensive calculations, and outputs the top-ranked candidate recipes or triggers regeneration instructions. The user interaction module is used to receive target parameters and constraints input by users, display candidate formulation results, and provide an experimental API interface.
[0025] Specifically, the data agent also employs a traceable data version control mechanism to manage data, and the generated agent supports pluggable large model adaptation and automatic optimization of prompt words.
[0026] A computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of a large-scale model material formulation design and evaluation method based on multi-agent collaboration. The computer-readable storage medium includes various types such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, and optical discs. USB flash drives have moderate capacity, suitable for local small-batch deployment, and support high-speed data transfer; portable hard drives have larger capacity, suitable for storing large amounts of historical data and model files; ROM is used to store the core program during system startup, the data is immutable, and it has high stability; RAM is used to temporarily store intermediate data while running the computer program, with fast read and write speeds; magnetic disks are low-cost and suitable for long-term storage of historical data and experimental records; optical discs are easy to archive and distribute, and suitable for storing immutable program installation packages and standard knowledge bases. The computer program stored on the computer-readable storage medium is developed using a specific version of the Python language. Its core dependencies include libraries related to data processing, deep learning, message bus communication, and shared memory interaction. The program structure is divided into a "core agent module," a "runtime framework module," and an "interaction interface module," each corresponding to an independent code package containing code files implementing specific functions. When the processor executes the computer program, it first initializes the agent runtime framework, starts the message bus client, and connects to shared memory. Then, it receives the target parameters input by the user through the API module. Each agent is then triggered sequentially to execute tasks according to the process flow, completing operations such as data processing, recipe generation, screening, evaluation, and decision-making. Finally, the candidate recipe is returned to the user through the interactive interface or pushed to the experimental interface. The entire process is logged and stored in a log file at a specific path for easy troubleshooting. The processor must support a specific architecture; the minimum configuration must meet the program's requirements. If running a large local model, a specific GPU model is required to ensure smooth program execution.
[0027] Beneficial effects The present invention has the following beneficial effects: (1) The large model material formulation design and evaluation method, system and storage medium based on multi-agent collaboration effectively solves the problems of frequent violations of rubber vulcanization ratio, low efficiency of manual screening and omission of prohibited solvents in coatings in existing monomer models and manual post-processing schemes by setting up a generating agent to output multiple initial formulations and combining a mechanism agent to eliminate formulations that do not meet environmental protection standards. Furthermore, the evaluation agent can simultaneously estimate the rationality of the formulation, and the decision agent outputs the optimal formulation based on the comprehensive score. This not only reduces the experimental verification of invalid formulations, but also ensures that the final formulation meets the requirements for use and production. At the same time, it solves the core problems of data governance and model generation disconnect and lack of mechanism constraints in the existing technology.
[0028] (2) The method, system and storage medium for material formulation design and evaluation based on multi-agent collaboration of large models can accurately ensure the reasonable ratio of vulcanizing agent and crosslinking density in rubber anti-wear formulation by setting an explicit mechanism knowledge base built into the mechanism agent. This effectively solves the problem of improper vulcanizing agent dosage when generating formulations using traditional single large models. In addition, the uncertainty quantification method of the evaluation agent outputs a confidence score, allowing engineers to intuitively judge the reliability of the formulation and reduce the cost of manual verification. At the same time, the large model adaptation mechanism supported by the generated agent can be plugged in flexibly and a better model can be switched according to the rubber formulation development needs, further improving the efficiency of formulation design and filling the application gap of multi-agent system in material formulation design scenarios.
[0029] Of course, any product implementing this invention does not necessarily need to achieve all of the advantages described above at the same time. Attached Figure Description
[0030] Figure 1 This is a flowchart of the large-scale material formulation design and evaluation method based on multi-agent collaboration of the present invention; Figure 2 This is a structural diagram of the large-scale model material formulation design and evaluation system based on multi-agent collaboration of the present invention. Detailed Implementation
[0031] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0032] Example 1 Rubber anti-wear formulation design examples This embodiment focuses on the design of rubber anti-wear formulations to verify the effectiveness of a large-scale model-based material formulation design and evaluation method based on multi-agent collaboration. The specific process is as follows: Target parameter input: The user inputs the target material performance parameters and constraints through the computer interactive interface. The target performance parameters are "abrasion resistance index ≥150, hardness 60±5 Shore A" and the constraints are "cost ≤25 yuan / kg, toxic vulcanization accelerators are prohibited". The input information is pushed to the message bus of the intelligent agent operation framework after being standardized.
[0033] Data Preparation and Processing: The data agent accesses multi-source material data in the rubber field (including 1000+ historical abrasion-resistant formulation data, corresponding vulcanization process parameter data, rubber crosslinking density spectrum data, and experimental record data) through preset database and file interfaces; the data is cleaned (abnormal experimental data with abrasion resistance index > 300 are removed using the IQR method, missing sulfur content parameters are supplemented with the average value of similar formulations, and the hardness unit is uniformly converted to ShoreA), and then structured into 32-dimensional vector data containing "carbon black content, sulfur content, vulcanization temperature, vulcanization time, crosslinking density spectrum characteristics, measured abrasion resistance index, and measured hardness", and stored in shared memory.
[0034] Initial formulation generation: The generating agent obtains the above vector data and user target parameters through the message bus, calls the locally deployed rubber formulation-specific large model (based on LLaMA-27B fine-tuning), and combines the template generated by the prompt word automatic optimization module ("Generate rubber formulations that meet the requirements of abrasion resistance index ≥150, hardness 60±5ShoreA, cost ≤25 yuan / kg, and prohibition of toxic vulcanization accelerators, including component names and contents (phr), generate 8 initial formulations"), outputs 8 initial rubber abrasion resistance formulations, and stores them in shared memory in JSON format.
[0035] Compliant Formulation Screening: The mechanism agent retrieves rubber-related mechanism rules from the explicit mechanism knowledge base (including vulcanization reaction ratio rules such as "the mass ratio of sulfur to non-toxic accelerator (e.g., CZ) is 1:0.3-0.5" and solubility parameter matching rules such as "the difference in solubility parameters between the rubber matrix and carbon black is ≤1.2 (cal / cm³)"). 3 The ratio of sulfur to accelerator was checked against 0.5”, and each of the eight initial formulas was verified. Two formulas with a sulfur-accelerator ratio that did not meet the requirements were eliminated (one with a ratio of 1:0.2 and the other with a ratio of 1:0.6), and six compliant formulas were retained.
[0036] Performance Estimation and Reliability Assessment: The assessment evaluates the agent's ability to read vector data of six compliant recipes from shared memory and input them into a pre-trained lightweight XGBoost prediction model (model fit R²). 2 =0.88), output the predicted abrasion resistance index values for each formulation (152-168cm). 3 The predicted values for hardness (58-63 Shore A) and hardness (1.61 km) are calculated. The risk level (all formula deviation rates ≤8%, judged as low risk) and confidence score (based on the model's historical error rate MAE=4.2 and the current task's matching sample size N=180, the confidence score is calculated to be 8.2-8.9 points) are stored in shared memory.
[0037] Optimal formulation decision: The decision-making agent reads the performance prediction value, risk level, and credibility score, and calculates the optimal formulation using the comprehensive score formula (comprehensive score = 0.4 × performance compliance rate + 0.3 × cost control rate + 0.3 × credibility score), where the performance compliance rate and cost control rate are both 100 points (meeting the target requirements) and 100 points (cost ≤ 25 yuan / kg). The top-3 formulations are sorted by comprehensive score, and their corresponding credibility scores (8.9, 8.6, and 8.2 respectively) are output. The selection criteria of the mechanism agent are also noted (e.g., "sulfur 2.4 phr, accelerator 0.8 phr, ratio 1:0.33"). If there are only 40 (<50) historical data entries for rubber abrasion-resistant formulations, the empirical formula 'Abrasion Resistance Index = 120 + 2 × Carbon Black Content (phr) - 5 × Sulfur Content (phr)' will be used to calculate the predicted performance value. The formula parameters will be retrieved from the shared memory 'Empirical Formula Parameter Table'.
[0038] Results Return and Experiment Integration: The system displays detailed components, predicted performance values, risk levels, and credibility of the Top-3 formulations to the user through a computer interface. After user confirmation, the interface generates an experimental task sheet (including formulation details, recommended vulcanizing equipment model XK-160, and abrasion resistance testing standard GB / T1689-2014), which is pushed to the laboratory management system via the experimental integration API to complete the rubber abrasion resistance formulation design process.
[0039] Example 2 This embodiment focuses on the design of low-VOC formulations for water-based coatings to further verify the applicability of the method. The specific process is as follows: Target parameter input: Users input target performance parameters "VOC content <50g / L, gloss ≥80" through the interactive interface, and constraints "drying time ≤2h, benzene solvents are prohibited". The input information is standardized and then pushed to the message bus.
[0040] Data preparation and processing: The data intelligence agent accesses multi-source data in the field of waterborne coatings (including 800+ historical formulas, coating process parameters, gloss spectra of paint films, and VOC detection experimental records of low-VOC coatings). After cleaning, it is structured into 28-dimensional vector data containing "waterborne resin content, environmentally friendly solvent content, additive content, drying temperature, drying time, gloss spectrum characteristics, measured VOC, and measured gloss", and stored in shared memory.
[0041] Initial formula generation: The generating agent calls the cloud-based Tongyi Qianwen big model (accessed through a pluggable adaptation mechanism) to generate 10 initial water-based coating formulas based on vector data and target parameters, and stores them in shared memory.
[0042] Compliant formulation screening: The mechanism agent calls the solvent prohibition rules ("prohibit benzene, toluene, xylene") and drying mechanism rules ("drying time ≤ 2h when water-based resin content ≥ 40%) in the explicit mechanism knowledge base, and removes 3 formulations containing toluene, retaining 7 compliant formulations.
[0043] Performance estimation and credibility assessment: The evaluation agent used a lightweight prediction model to estimate the VOC content (42-48 g / L), gloss (81-86%), and drying time (1.5-1.9 h) of seven formulations and determined the risk level (low risk); based on the model's historical error rate MAE=2.1 and the current task's matched sample size N=150, the credibility score was calculated to be 7.8-8.5.
[0044] Optimal Formulation Decision: The decision-making agent is ranked according to its comprehensive score, taking into account gloss (weight 0.4), VOC content (weight 0.3), and confidence (weight 0.3), and outputs two optimal formulations (one with VOC 42g / L, gloss 86%, and confidence 8.5; the other with VOC 45g / L, gloss 84%, and confidence 8.3). The gloss estimation basis of the evaluation agent is marked as "based on resin content and solvent type, in line with the logic of gloss prediction model".
[0045] If the Top-1 recipe output by the decision agent has a comprehensive score of 62 (preset threshold of 65), then a regeneration instruction is sent to the generating agent via the message bus to regenerate 8 initial recipes. Results return and experiment integration: After the user views the results and confirms the experiment, the system pushes the experiment task sheet (including the formula, recommended testing equipment (TVOC-3000 VOC detector, HG60 gloss meter), and testing standard GB18582-2020) to the laboratory management system, completing the low VOC formula design process for water-based coatings.
[0046] Example 3 The application scenario is the research and development of plastics for electronic device housings. It must meet the following requirements: heat resistance temperature not lower than 120 degrees Celsius, flexural modulus not lower than 2000 MPa, cost per kilogram not exceeding 30 yuan, and fluorine-containing additives are prohibited.
[0047] Researchers input the target parameters: heat resistance temperature not lower than 120 degrees Celsius, flexural modulus not lower than 2000 MPa, cost per kilogram not exceeding 30 yuan, and fluorine-containing additives are prohibited.
[0048] The multi-agent collaborative process is as follows: The data intelligence agent retrieves multi-source data from the plastics field, covering over 600 historical heat-resistant plastic formulations, injection molding process parameters, thermogravimetric analysis spectra, and performance verification records. After cleaning, it generates structured data containing 28 categories of key parameters, including resin type, glass fiber ratio, heat-resistant additive content, and processing temperature.
[0049] The generated agent invokes a large plastics-specific model to generate 12 initial formulations based on the aforementioned data. For example, one formulation contains 70 parts polypropylene resin, 30 parts glass fiber, 4 parts heat-resistant additives, and 1 part antioxidant.
[0050] The mechanism-based intelligent agent screened formulations according to specific rules in the plastics industry: First, it verified the glass fiber to resin mass ratio rule, ensuring that the ratio was within the range of 1:3 to 1:5. This ratio ensures that the plastic's flexural modulus meets the standard and does not become brittle, complying with industry design specifications for plastic reinforcement materials. In this formulation, the ratio was 1:2.33, which did not meet the rule requirements and was therefore rejected. Next, it verified the rules for the addition of heat-resistant additives, ensuring that the addition amount did not exceed 5 parts. Finally, it checked the rules for the prohibition of fluorinated additives, confirming that the formulation contained no fluorinated additives. After screening, 3 formulations that did not meet the rules were rejected, and the remaining 9 entered the evaluation stage.
[0051] The evaluation agent used a random forest model to predict plastic properties, achieving an accuracy of 85% in predicting the heat resistance temperature and flexural modulus of the plastics. The prediction results showed that the remaining formulations had heat resistance temperatures between 122 and 135 degrees Celsius and flexural moduli between 2050 and 2200 MPa, and were classified as low-risk, with deviation rates all below 12%. Furthermore, the confidence scores for each formulation were calculated to be between 8.0 and 8.6.
[0052] The decision-making agent ranks the evaluated formulations based on a weighted calculation of performance compliance rate, cost control rate, and reliability score (performance compliance rate accounts for 40%, cost control rate for 30%, and reliability score for 30%). For example, a formulation with a heat resistance temperature of 135 degrees Celsius, a flexural modulus of 2150 MPa, and a performance compliance rate of 100%; an actual cost of 28 yuan per kilogram, a cost control rate of 100%; and a reliability score of 8.5, resulting in a comprehensive score of 93.5, is the optimal formulation. The decision-making agent outputs the formulation's advantages, explaining that both heat resistance and mechanical properties exceed the targets, costs are controllable, and predictions are reliable. It also explains the options for secondary formulations, such as a formulation with a heat resistance temperature of 128 degrees Celsius and a flexural modulus of 2080 MPa, which has a slightly lower performance compliance rate and a slightly lower comprehensive score, thus ranking it as a secondary option.
[0053] Laboratory injection molding experiments verified that the optimal formula has an actual heat resistance temperature of 132 degrees Celsius and a flexural modulus of 2120 MPa, with a deviation rate of approximately 2.2% from the predicted results, which meets the design requirements.
[0054] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.
[0055] The preferred embodiments of the present invention disclosed above are merely illustrative of the invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the content of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the invention, thereby enabling those skilled in the art to better understand and utilize the invention. The invention is limited only by the claims and their full scope and equivalents.
Claims
1. A method for designing and evaluating material formulations in large-scale models based on multi-agent collaboration, characterized in that, Includes the following steps: S1. Receive the target material performance parameters and constraints input by the user through a computer interactive interface; S2. The data intelligence agent accesses multi-source material data through a preset data interface, cleans and structures the multi-source material data, generates vector data, and stores it in shared memory; S3. The generated intelligent agent calls the multimodal large model or language large model through the message bus, and generates several initial material formulas based on the vector data, target material performance parameters and constraints. S4. The mechanism intelligent agent calls domain mechanism rules from the explicit mechanism knowledge base to perform compliance screening on the several initial material formulations and delete illegal formulations that do not conform to common sense of physicochemicals; S5. The evaluation agent performs performance estimation and risk assessment on the remaining formulations after screening, and outputs the performance prediction value, risk level and confidence score of each formulation. S6. The decision-making agent reads the performance prediction value, risk level and confidence score from the shared memory and performs comprehensive calculations to output the top-ranked candidate formulas; S7. Return the top-ranked candidate formulas to the user through the computer interface, select to receive the user's manual confirmation instruction, and push the confirmed formula to the experimental docking interface.
2. The method for designing and evaluating large-scale material formulations based on multi-agent collaboration as described in claim 1, characterized in that: The target material performance parameters include: hardness, tensile strength, elongation at break, and anti-aging properties in rubber materials; at least one of abrasion resistance index, VOC content, and gloss parameters in coatings; and constraints include at least one of ingredient prohibition rules and cost thresholds.
3. The method for designing and evaluating large-scale material formulations based on multi-agent collaboration as described in claim 1, characterized in that: The multi-source material data includes historical formula data, process parameter data, spectral data, and experimental record data. The data agent adopts a traceable data version control mechanism, adding a source identifier, processing timestamp, and version number to each processed data to ensure the quality of training data.
4. The method for designing and evaluating large-scale material formulations based on multi-agent collaboration as described in claim 1, characterized in that: The generated intelligent agent supports a pluggable large model adaptation mechanism, which can call a private large model through a local interface or a public large model through a cloud API. It also has an automatic prompt word optimization module that iteratively adjusts the prompt word logic based on historical generation results.
5. The method for designing and evaluating large-scale material formulations based on multi-agent collaboration as described in claim 1, characterized in that: The aforementioned domain mechanism rules include at least one of the following: solubility parameter matching rules, sulfidation reaction ratio rules, and solvent prohibition rules.
6. The method for designing and evaluating large-scale material formulations based on multi-agent collaboration as described in claim 1, characterized in that: The evaluation agent calculates a confidence score using an uncertainty quantification algorithm, which is determined based on the historical error rate of the prediction model and the data sample size.
7. A large-scale model material formulation design and evaluation system based on multi-agent collaboration, used to implement the large-scale model material formulation design and evaluation method based on multi-agent collaboration as described in any one of claims 1-6, characterized in that, include: An intelligent agent operation framework, comprising a message bus and shared memory, wherein the message bus is used to realize data interaction between intelligent agents and the shared memory is used to store intermediate data; The data intelligence agent is deployed on computer hardware nodes and is used to uniformly access multi-source material data through a preset data interface, clean and structure the multi-source material data, and generate vector data. A smart agent is generated and communicates with the data smart agent through a message bus. It is used to call a multimodal large model or a language large model to generate several initial material formulations based on vector data, target performance parameters and constraints. Mechanistic intelligent agents, with a built-in explicit mechanism knowledge base, are used to call domain mechanism rules to screen initial material formulations for compliance and delete illegal formulations; The evaluation agent is used to perform performance estimation and risk assessment on the screened formulations using lightweight prediction models or empirical formulas, and outputs performance prediction values, risk levels and confidence scores. The decision-making agent reads the evaluation results from shared memory, uses a combination of rule trees and simple reinforcement learning to perform comprehensive calculations, and outputs the top-ranked candidate recipes or triggers regeneration instructions. The user interaction module is used to receive target parameters and constraints input by users, display candidate formulation results, and provide an experimental API interface.
8. The large-scale model material formulation design and evaluation system based on multi-agent collaboration as described in claim 7, characterized in that: The data agent also employs a traceable data version control mechanism to manage data, and the generated agent supports pluggable large model adaptation and automatic optimization of prompt words.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the large-scale material formulation design and evaluation method based on multi-agent collaboration as described in claim 1.
Citation Information
Patent Citations
Rubber formula optimization method and device and storage medium
CN115862781A
Enterprise collaborative office system and method thereof
CN117787892A
Task integrity judgment method for multi-device cooperative work under complex constraint conditions
CN119578817A
Multi-agent collaboration method based on large language model
CN120338035A
Construction material assessment method and systems
US20210063336A1
Cited By
Evaluation method and system for Chinese herbal medicine-feed compatibility interaction effect
CN121329249A
Material research system and method based on multi-agent game
CN121998102A