Reasoning model training method, recommendation method, device, medium and product
Patent Information
- Application Number
- CN202611146867.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-30
- Publication Date
- 2026-09-18
AI Technical Summary
[0004]本发明提供一种推理模型训练方法、推荐方法、设备、介质及产品,用以解决现有技术中推理过程难以追溯且预测的性能预测值可信度低,导致材料设计方案推荐结果的可靠性较差的缺陷,实现有效提升材料设计方案推荐结果的可靠性
[0020]The reasoning model training method, recommendation method, device, medium, and product provided by this invention obtain a target literature set based on the first target material type and the first target performance parameters, and identify evidence data and mechanistic knowledge data from it. Then, the causal relationship of the mechanistic knowledge data is extracted and fused with the evidence data to obtain the actual reasoning path. The pre-trained language model is then trained based on the sample reasoning problem, the actual reasoning path, and the reasoning judgment label obtained from the reasoning conclusion. This enables the trained material reasoning model to transform the scattered scientific judgment process in the literature into reusable reasoning ability. While outputting the judgment result, it provides a reasoning path that is supported by evidence and is logically traceable. This effectively overcomes the problems of difficult-to-trace reasoning process and low credibility of output performance prediction values in related technologies, and improves the reliability of material design scheme recommendations based on this.
Smart Images

Figure CN122779296A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of materials informatics technology, and in particular to a reasoning model training method, recommendation method, device, medium, and product. Background Technology
[0002] In the development of functional materials, researchers need to comprehensively assess the combination of building blocks, the feasibility of structural formation, the physicochemical properties, and the potential of target performance of candidate materials in order to identify material design schemes worthy of synthesis and testing from a vast candidate pool. As the pool of candidate material combinations continues to expand, how to efficiently and accurately make scientifically based screening judgments on candidate material design schemes has become a crucial issue that the industry urgently needs to address.
[0003] Existing technologies typically train inference models using design parameters and performance labels corresponding to existing material design schemes. The trained inference model is then used to map the performance predictions of candidate design schemes based on their design parameters. Based on these performance predictions, candidate design schemes are ranked to achieve material design scheme recommendations. However, the inference model trained in this way can only learn the mapping relationship between the design parameters and performance labels of material design schemes. The inference process is difficult to trace, and the output performance predictions have low reliability, resulting in poor reliability of the material design scheme recommendations obtained accordingly. Summary of the Invention
[0004] This invention provides a reasoning model training method, recommendation method, device, medium, and product to address the shortcomings of existing technologies, such as the difficulty in tracing the reasoning process and the low reliability of predicted performance values, which leads to poor reliability of material design scheme recommendation results. This invention effectively improves the reliability of material design scheme recommendation results.
[0005] This invention provides a method for training a reasoning model, comprising: Based on the first target material type and the first target performance parameters, a target literature set is obtained, and from the target literature set, evidence data and mechanism knowledge data corresponding to each sample material design scheme under the first target material type are identified; Based on the first target performance parameters and the design information of each of the sample material design schemes, generate sample reasoning questions corresponding to each of the sample material design schemes; The causal relationship is extracted from the mechanistic knowledge data to obtain at least one causal unit. The evidence data and the at least one causal unit are fused to obtain the actual reasoning path corresponding to each of the sample material design schemes. Based on the reasoning conclusions in the actual reasoning path, obtain the reasoning judgment label of the sample reasoning problem; Based on the sample reasoning problem, the actual reasoning path, and the reasoning judgment label, the pre-trained language model is trained to obtain the material reasoning model.
[0006] According to a reasoning model training method provided by the present invention, the step of fusing the evidence data and the at least one causal unit to obtain the actual reasoning path corresponding to each of the sample material design schemes includes: For each sample material design scheme, in the evidence data and at least one causal unit corresponding to the sample material design scheme, the evidence context information associated with the sample reasoning problem corresponding to the sample material design scheme is retrieved to obtain the evidence set of the sample material design scheme. From the evidence set, identify the reasoning information of each reasoning node in the actual reasoning path corresponding to the sample material design scheme; the reasoning information includes the reasoning conclusion, evidence identifier, direction of influence, and confidence level; According to the reasoning order corresponding to each reasoning node, the reasoning information of each reasoning node is combined to obtain the actual reasoning path corresponding to the sample material design scheme; By traversing each of the aforementioned sample material design schemes, the actual reasoning path corresponding to each of the aforementioned sample material design schemes is obtained.
[0007] According to a reasoning model training method provided by the present invention, the step of extracting causal relationships from the mechanistic knowledge data to obtain at least one causal unit includes: For each sample material design scheme, the mechanism knowledge data corresponding to the sample material design scheme is divided into at least one knowledge segment according to the preset relation words; Causal relationships are extracted from each of the knowledge fragments to obtain at least one causal unit corresponding to the sample material design scheme; the causal unit includes an antecedent entity, an action relationship, a consequence entity, applicable conditions, and evidence identifiers; the action relationship is used to characterize the direction of influence of the antecedent entity on the consequence entity; By iterating through each of the sample material design schemes, at least one causal unit corresponding to each sample material design scheme is obtained.
[0008] According to a reasoning model training method provided by the present invention, the step of obtaining the reasoning judgment label of the sample reasoning problem based on the reasoning conclusion in the actual reasoning path includes: For each sample material design scheme, in the reasoning conclusions of the actual reasoning path corresponding to the sample material design scheme, identify the target reasoning conclusions associated with each judgment task in the sample reasoning problem corresponding to the sample material design scheme. Based on the target reasoning conclusion, obtain the reasoning judgment sub-labels corresponding to each judgment task; The reasoning judgment sub-labels corresponding to each of the judgment tasks are combined to obtain the reasoning judgment labels of the sample reasoning problem corresponding to the sample material design scheme. By iterating through each of the aforementioned sample material design schemes, the reasoning judgment labels for the sample reasoning problems corresponding to each of the aforementioned sample material design schemes are obtained; Among them, the judgment tasks configured in the sample reasoning problem corresponding to each of the sample material design schemes include multiple tasks such as judging the feasibility of structure formation, judging the physicochemical properties, judging the performance level, judging the recommendation status, judging the recommendation priority level, judging the reasoning confidence, and judging the risk of chemical synthesis.
[0009] According to a reasoning model training method provided by the present invention, the step of training a pre-trained language model based on the sample reasoning question, the actual reasoning path, and the reasoning judgment label to obtain a material reasoning model includes: Based on the sample reasoning problem, actual reasoning path and reasoning judgment label corresponding to each of the sample material design schemes, construct training samples corresponding to each of the sample material design schemes; Perform consistency verification on each of the training samples and obtain the target training sample that has passed the verification; The pre-trained language model is trained based on the target training samples to obtain the material reasoning model; The consistency verification includes verifying whether the sample reasoning problem in each training sample is related to the reasoning conclusion in the actual reasoning path, whether each reasoning node in the actual reasoning path is associated with a valid evidence identifier, whether there is a logical jump between the reasoning conclusions of adjacent reasoning nodes in the actual reasoning path, whether each reasoning conclusion in the actual reasoning path is consistent with the reasoning judgment label, and whether the performance value in the actual reasoning path is associated with a valid evidence identifier.
[0010] According to a reasoning model training method provided by the present invention, the step of training the pre-trained language model based on the target training samples to obtain the material reasoning model includes: The sample reasoning problem in the target training sample is input into the pre-trained language model to obtain the predicted reasoning path and reasoning judgment result corresponding to the target training sample output by the pre-trained language model; The target loss value is obtained based on the inference loss value between the predicted inference path and the actual inference path in the target training sample, and the inference loss value between the inference judgment result and the inference judgment label in the target training sample. Based on the target loss value, the pre-trained language model is iteratively trained until the performance of the pre-trained language model converges and / or the maximum number of iterations is reached. The material reasoning model is obtained based on the pre-trained language model obtained from each iteration of training.
[0011] The present invention also provides a recommended method, characterized in that it includes: Based on the second target material type and the second target performance parameters, obtain a set of candidate material design schemes; Based on the second target performance parameters and the design information of each candidate material design scheme in the candidate material design scheme set, generate the target reasoning problem corresponding to each candidate material design scheme; The target reasoning problem is input into the material reasoning model to obtain the predicted reasoning path and reasoning judgment result corresponding to each candidate material design scheme output by the material reasoning model. Based on the prediction reasoning path and reasoning judgment result corresponding to each candidate material design scheme, at least one material design scheme to be recommended is selected from the set of candidate material design schemes. The material reasoning model is trained based on the reasoning model training method described in any of the above-mentioned methods.
[0012] According to a recommendation method provided by the present invention, the step of selecting at least one material design scheme to be recommended from the set of candidate material design schemes based on the prediction reasoning path and reasoning judgment result corresponding to each candidate material design scheme includes: Based on the reasoning and judgment results corresponding to each candidate material design scheme, candidate material design schemes that are in the recommended state and whose reasoning confidence is higher than the confidence threshold are selected from the candidate material design scheme set to obtain the first material design scheme subset; The constraint tool component is invoked to perform constraint verification on each candidate material design scheme in the first material design scheme subset, and the constraint verification results corresponding to each candidate material design scheme in the first material design scheme subset are obtained; the verification results include the verification status identifier and verification score output by each constraint tool in the constraint tool component; Based on the verification status identifier, candidate material design schemes that pass the constraint verification are selected from the first material design scheme subset to obtain the second material design scheme subset; Based on the verification score and the reasoning result, at least one material design scheme to be recommended is selected from the second material design scheme subset.
[0013] According to a recommendation method provided by the present invention, the step of selecting at least one material design scheme to be recommended from the second material design scheme subset based on the verification score and the reasoning judgment result includes: Based on the reasoning and judgment results, obtain the reasoning score corresponding to each candidate material design scheme in the second material design scheme subset; Based on the structural description information in the design information of each candidate material design scheme in the second material design scheme subset, the second material design scheme subset is divided into multiple material design scheme groups, and the structural similarity between each candidate material design scheme in each material design scheme group is higher than the similarity threshold. Based on the verification score and the reasoning score, obtain the comprehensive score of each candidate material design scheme in each material design scheme group; Based on the comprehensive score, candidate material design schemes in each of the material design scheme groups are screened to obtain at least one material design scheme to be recommended.
[0014] According to a recommendation method provided by the present invention, after selecting at least one material design scheme to be recommended, the method further includes: Obtain the experimental results corresponding to each of the proposed material design schemes; Based on the experimental results, material design schemes that failed in the experiment were selected as negative examples from the at least one material design scheme to be recommended. The prediction inference path, inference judgment result, and constraint verification result corresponding to the negative example sample are compared with the experimental results corresponding to the negative example sample, and the negative example type label corresponding to the negative example sample is determined based on the comparison result. Based on the negative example type label, determine the target update operation among multiple update operations, and execute the target update operation; The multiple update operations include fine-tuning the training of the material reasoning model, updating the constraint rule information in the constraint tool component, and updating the output weights of each constraint tool in the constraint tool component.
[0015] According to a recommended method provided by the present invention, determining a target update operation among multiple update operations based on the negative example type label, and executing the target update operation, includes: When the negative example type label belongs to the label of abnormal inference information, the experimental results corresponding to the negative example sample are compared one by one with the inference information of each inference node in the prediction inference path corresponding to the negative example sample. Based on the comparison results, the root cause node is determined in the prediction inference path corresponding to the negative example sample. Based on the experimental results corresponding to the negative example sample, the inference information corresponding to the root cause node, the inference information corresponding to the inference node after the root cause node, and the inference judgment result corresponding to the negative example sample in the prediction inference path corresponding to the negative example sample are updated to obtain the corrected inference path and the corrected inference judgment result corresponding to the negative example sample. Based on the target reasoning problem, corrected reasoning path and corrected reasoning judgment result corresponding to the negative example sample, construct the negative example training sample corresponding to the negative example sample; The material reasoning model is fine-tuned based on the negative training samples and the negative type labels.
[0016] According to a recommendation method provided by the present invention, obtaining a set of candidate material design schemes based on a second target material type and a second target performance parameter includes: Based on the second target material type and the second target performance parameters, obtain the generation constraint information; Based on the generated constraint information, the set of candidate material design schemes is obtained by filtering from the candidate material design scheme library.
[0017] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the inference model training method or recommendation method as described above.
[0018] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the inference model training method or recommendation method as described above.
[0019] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements any of the inference model training methods or recommendation methods described above.
[0020] The reasoning model training method, recommendation method, device, medium, and product provided by this invention obtain a target literature set based on the first target material type and the first target performance parameters, and identify evidence data and mechanistic knowledge data from it. Then, the causal relationship of the mechanistic knowledge data is extracted and fused with the evidence data to obtain the actual reasoning path. The pre-trained language model is then trained based on the sample reasoning problem, the actual reasoning path, and the reasoning judgment label obtained from the reasoning conclusion. This enables the trained material reasoning model to transform the scattered scientific judgment process in the literature into reusable reasoning ability. While outputting the judgment result, it provides a reasoning path that is supported by evidence and is logically traceable. This effectively overcomes the problems of difficult-to-trace reasoning process and low credibility of output performance prediction values in related technologies, and improves the reliability of material design scheme recommendations based on this. Attached Figure Description
[0021] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0022] Figure 1 This is one of the flowcharts illustrating the inference model training method provided by the present invention.
[0023] Figure 2 This is a schematic diagram of the process for constructing training samples provided by the present invention.
[0024] Figure 3 This is the second flowchart of the inference model training method provided by the present invention.
[0025] Figure 4 This is one of the flowcharts of the recommended method provided by the present invention.
[0026] Figure 5 This is the second flowchart of the recommended method provided by the present invention.
[0027] Figure 6 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0028] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0029] It should be noted that, in the data processing stage, the technical solution of this application has strictly limited the scope of data collection to the minimum necessary to achieve the technical objectives, preventing the acquisition of irrelevant information. For any user information to be collected, the data subject will be clearly informed and their consent obtained. Furthermore, technologies such as encrypted storage and access control are employed to strengthen data security and ensure the security and compliance of the entire data processing process. The technical model and decision-making mechanism are based on objective technical parameters and do not introduce unnecessary parameters such as gender or age that may lead to discrimination, resolutely eliminating algorithmic discrimination and upholding public order and good morals. In addition, the specification fully describes the technical implementation methods, application scenarios, and compliance protection details. The claims are consistent with the content of the specification, key compliance designs are clear and verifiable, and the overall technical design is guided by the protection of public interests and adherence to social ethics, without any circumstances that harm public interests or violate public order and good morals.
[0030] The development of functional materials typically involves multiple stages, including candidate structure generation, determination of synthesis or preparation routes, structural characterization, physicochemical property analysis, target performance testing, and experimental feedback optimization. For complex materials, such as porous framework materials, organic functional materials, catalytic materials, adsorption and separation materials, and energy storage materials, their target performance is often not determined by a single structural parameter, but rather by a combination of factors, including building blocks, connection methods, spatial structure, crystallinity, pore structure, electronic properties, interfacial reactions, preparation conditions, and testing systems.
[0031] For example, covalent organic frameworks (COFs) are formed by organic building blocks connected by covalent bonds. They possess characteristics such as tunable building blocks, designable connection methods, tunable pore structure, predictable topology, and modifiable functional groups, making them suitable for applications such as photocatalytic hydrogen evolution, carbon dioxide reduction, gas adsorption and separation, electrochemical energy storage, sensing, and drug delivery. The performance of these materials varies depending on the application. In photocatalytic hydrogen evolution, their performance may be influenced by factors such as the electronic structure of the building blocks, donor-acceptor matching, bond type, topology, crystallinity, pore structure, specific surface area, band gap, band edge position, carrier separation efficiency, cocatalyst loading, sacrificial agent system, and interfacial proton reduction capability. Therefore, even if some building blocks appear advantageous in their local electronic structure, the actual material may still perform poorly due to insufficient framework formation, unreasonable pore structure, severe charge recombination, or insufficient interfacial active sites.
[0032] Therefore, how to efficiently and accurately make scientifically based screening judgments on candidate material design schemes has become an important issue that the industry urgently needs to address.
[0033] In traditional technologies, materials research and development heavily relies on researchers' experience-based judgment after reviewing literature. This means researchers typically need to comprehensively assess whether candidate materials are worth synthesizing or testing, based on existing papers, database records, and experimental experience, according to target performance requirements. This approach is effective for small-scale explorations, but when the candidate space reaches tens of thousands or even larger scales, human experience is insufficient to process each candidate combination individually. Furthermore, it heavily depends on researchers' comprehensive judgment of reaction feasibility, structure formation, physicochemical properties, and application scenarios. Different researchers may not have entirely consistent understandings of the relationship between literature evidence, design information, and performance parameters, making it difficult to provide reliable recommendations for candidate materials.
[0034] In response, some technologies propose using methods such as literature data extraction, materials databases, machine learning property prediction, and language model-assisted question answering to assist materials research and development. These technologies typically utilize natural language processing, rule templates, or language models to extract fields such as material name, composition, preparation conditions, characterization results, and performance values from materials literature, and organize them into tables, databases, or knowledge graphs to facilitate subsequent retrieval of performance prediction values for candidate materials. Alternatively, they may use language models for literature question answering, experimental scheme generation, candidate material recommendation, or material category determination, and can combine search enhancement, prompt templates, or domain corpora to improve the quality of responses. Another approach is to directly train machine learning models based on existing structure-performance data, mapping the performance prediction values corresponding to the design information of candidate materials, and then ranking and recommending candidate materials based on these performance prediction values. These technologies can improve the efficiency of materials data processing and property prediction to a certain extent.
[0035] However, the truly valuable part of the literature reflecting expert judgment lies not in isolated end-point fields such as material names, preparation conditions, and performance values, but rather in the analysis, comparison, and interpretation made by experts regarding the relationships between building block selection, structure formation, characterization evidence, property changes, control strategies, and performance results—that is, the scientific judgment process in materials design. For example, for complex systems such as porous framework materials and organic functional materials, prediction scores alone are insufficient to determine whether a candidate is worth preparing. It is also necessary to judge whether the building blocks are matched, whether the reaction or assembly is feasible, whether the structure can be stably formed, whether the key physicochemical properties are reasonable, and whether the target application is supported by structural and interface factors. Therefore, the above-mentioned scientific judgment process often reflects the laws of materials design more effectively than a single performance value, but it is usually not recognized as an independent data object for processing and modeling in related technologies. This makes it difficult for a large amount of valuable mechanistic knowledge in the literature to be effectively learned and reused by the model.
[0036] Furthermore, when a model only learns the mapping relationship between material design information and final performance labels, its reasoning process is difficult to trace, and it is prone to mistaking certain surface features as sufficient conditions for determining material performance. For example, strong visible light absorption, donor-acceptor structures, or large conjugated frameworks do not necessarily correspond to high catalytic activity; a comprehensive judgment must still be made in conjunction with structural order, charge transport, interfacial reactions, and practical preparation feasibility. Moreover, it may also generate unfounded performance values, material entity confusion, or incorrect chemical associations, resulting in low reliability of the output performance predictions, making them unsuitable as a basis for experimental decisions, and ultimately affecting the reliability of the recommended material design schemes.
[0037] In view of this, this application provides a reasoning model training method. The core idea is to transform the scattered scientific judgment processes in materials literature into computable materials design reasoning data. The actual reasoning paths and reasoning judgment labels in this data are then used as supervisory signals to train a pre-trained language model, thereby obtaining a materials reasoning model capable of reasoning based on the process chain of building blocks – preparation conditions – structure – properties – performance. Thus, the obtained materials reasoning model, while outputting the reasoning judgment results of materials design schemes, can also output predictive reasoning paths supported by literature evidence and logically traceable. Furthermore, based on the reasoning judgment results and predictive reasoning paths, a comprehensive selection is made to output candidate materials recommendation results that are supported by evidence, experimentally feasible, and reliable.
[0038] Figure 1 This is one of the flowcharts illustrating the inference model training method provided by the present invention.
[0039] This method can be deployed in electronic devices with data processing and computing capabilities. For example, the electronic device can be a cloud computing server, edge computing node, laboratory workstation, enterprise R&D platform, or terminal device with corresponding computing power, etc., and this invention does not specifically limit its application. The electronic device can also communicate with material databases, knowledge graph libraries, property prediction tools, and laboratory information management systems. This method can be used in various material development scenarios, including but not limited to photocatalytic hydrogen evolution material development scenarios, as well as other material development scenarios, such as porous framework materials, organic functional materials, catalytic materials, adsorption and separation materials, energy storage materials, or materials with other target properties. This embodiment does not specifically limit its application in these scenarios.
[0040] like Figure 1 As shown, the method includes steps 110, 120, 130, 140 and 150.
[0041] Step 110: Based on the first target material type and the first target performance parameters, obtain the target literature set, and identify the evidence data and mechanism knowledge data corresponding to each sample material design scheme under the first target material type from the target literature set.
[0042] Optionally, before training the model, it is necessary to first determine the material scope and performance target to be trained in this inference model, that is, the first target material type and the first target performance parameters, and then obtain the corresponding target literature set based on this, and identify the evidence data and mechanism knowledge data used to construct the training data.
[0043] The first target material type refers to the category of material system targeted in this training. It can be one or more complex material systems such as porous framework materials, organic functional materials, catalytic materials, adsorption and separation materials, or energy storage materials. For example, it can be a covalent organic framework material or a metal organic framework (MOF) material. This embodiment does not make specific limitations on this.
[0044] The primary target performance parameter refers to the performance indicators that are of interest to the type of material in the target application scenario. For example, in the photocatalytic hydrogen evolution scenario, the primary target performance parameter may be the hydrogen evolution rate; in the gas adsorption and separation scenario, the primary target performance parameter may be the adsorption capacity or separation selectivity; and in the electrochemical energy storage scenario, the primary target performance parameter may be the specific capacity, etc.
[0045] The target literature set refers to a collection of literature that provides evidence for material design and is associated with the primary target material type and primary target performance parameters. This literature set can include at least one of the following: journal articles, review articles, literature discussion information, supplementary literature information, figure and table descriptions, experimental data tables, and database records. The literature discussion information and supplementary literature information may include the authors' explanations of the relationship between structure, properties, and performance; analyses of the causes of low performance; descriptions of reaction conditions and testing systems; and empirical judgments regarding factors such as pH, sacrificial agents, co-catalysts, light source conditions, sample activation, and stability.
[0046] Figure 2 This is a schematic diagram of the training sample construction process provided by the present invention; as shown below. Figure 2As shown, when obtaining the target literature set, the literature in the literature library can be initially screened according to the data quality requirements configured for this inference model training to obtain an initial literature set. For example, if the configured data quality requirements require obtaining literature with original experimental data, complete preparation information, effective performance test data, and effective mechanism knowledge data, then literature containing original experimental data, complete preparation information, effective performance test data, and effective mechanism knowledge data can be screened from the literature library to form the initial literature set.
[0047] Furthermore, literature matching the first target material type and the first target performance parameters is selected from the initial literature set to form a target literature set.
[0048] Furthermore, during the acquisition of the target literature set, the literature in the target literature set can be stored locally, and the literature in the target literature set can be identified by source and located by evidence. This allows for tracing back to the corresponding text paragraphs, tables, figure captions, supplementary information, or database entries when extracting evidence data and mechanistic knowledge data. It should be noted that for copyrighted literature, only the factual fields, reasoning summaries, evidence locations, and literature index information extracted from the literature can be stored locally, without storing the full text of the literature. This allows for the formation of a traceable target material literature library while meeting copyright requirements, providing stable and traceable source support for the identification of evidence data and mechanistic knowledge data.
[0049] For example, taking covalent organic framework materials as the first target material type and photocatalytic hydrogen evolution performance as the first target performance parameter, we can first screen the literature library to obtain an initial set of literature related to covalent organic framework materials according to the data quality requirements, and then screen the literature related to photocatalytic hydrogen evolution of covalent organic framework materials from the initial set of literatures to construct a literature library of photocatalytic hydrogen evolution of covalent organic framework materials as the target literature set.
[0050] After obtaining the target literature set, the evidence data and mechanism knowledge data corresponding to the material design schemes of each sample under the first target material type can be identified.
[0051] The sample material design scheme refers to the design scheme of a specific material belonging to the first target material type recorded in the target literature collection. Each sample material design scheme corresponds to a material that has been studied, characterized and tested in the literature.
[0052] Evidence data refers to structured field data that objectively reflects the composition, preparation, structure, physicochemical properties, and performance of the sample material design. For example, evidence data may include material name, abbreviation, literature source, material category, target reaction, or target application; it may also include building block name, functional groups, number of connection sites, donor / acceptor properties, rigidity or flexibility, catalyst, solvent, temperature, time, atmosphere, and post-treatment conditions; furthermore, it may include bond type, topology or connection mode, crystallinity, specific surface area, pore structure, morphology, color, band gap, band edge position, co-catalyst, sacrificial agent, light source, target performance values, and stability. Through the identification of this evidence data, inconsistent statements from different literature can be converted into a unified material-level evidence record.
[0053] Mechanistic knowledge data refers to knowledge-based data in literature that reflects the scientific judgment process of experts and is used to explain the structure-property-performance relationship. This data can originate from the authors' explanations of building block selection, framework formation, electronic structure changes, structural control strategies, and differences in target performance. It can also come from general knowledge in the materials field, materials databases, knowledge graphs, or supporting information from property prediction tools. For example, explanations of the relationship between structure, properties, and performance, analyses of the causes of low performance, and empirical judgments regarding factors such as pH, sacrificial agents, co-catalysts, light source conditions, sample activation, and stability, all found in literature discussion and supplementary information, can all be considered mechanistic knowledge data.
[0054] It should be noted that when identifying evidence data and mechanistic knowledge data from the target literature collection, at least one of the following methods can be used: natural language processing technology, rule templates, or language models. This method can be used to extract and parse the main text, tables, figure captions, and supplementary information of the literature, so as to identify the content describing the composition, preparation, structure, physicochemical properties, and performance of materials in different literature as evidence data, and to identify the mechanistic content characterizing relationships such as promotion, inhibition, causation, dependence, constraint, or conditional establishment as mechanistic knowledge data.
[0055] It should be noted that, in order to control data noise, after identifying and obtaining evidence data and mechanistic knowledge data, if a set of data has abnormalities such as unfounded performance values, confusion of material entities, misunderstanding of abbreviations, missing evidence sources, conflict between reasoning conclusions and evidence, or incorrect labeling, it can be deleted or corrected, thereby providing an accurate and reliable data source for the construction of reasoning data.
[0056] Furthermore, after obtaining the evidence data and mechanism knowledge data corresponding to each sample material design scheme, the following steps 120, 130 and 140 can be executed to convert the evidence data and mechanism knowledge data corresponding to each sample material design scheme into a data form that can be learned by the pre-trained language model, that is, material design reasoning data containing triple data of sample reasoning questions, actual reasoning paths and reasoning judgment labels.
[0057] Step 120: Based on the first target performance parameters and the design information of each of the sample material design schemes, generate sample reasoning problems corresponding to each of the sample material design schemes.
[0058] Optionally, after identifying the evidence data and mechanism knowledge data corresponding to each sample material design scheme, a sample reasoning problem corresponding to each sample material design scheme can be generated by combining the first target performance parameters and the design information of the sample material design scheme, so as to serve as the input problem during model training.
[0059] Among them, the design information of the sample material design scheme refers to the descriptive information of the composition and structure of the sample material design scheme, which can be obtained from the evidence data corresponding to the sample material design scheme. Specifically, it may include at least one of the following: building unit identifier, building unit combination method, connection bond type and metric relationship, topology and topology type, and recommended preparation condition range.
[0060] Sample reasoning problems refer to problems used to describe the material design judgment task to be analyzed. They describe whether the combination of candidate building blocks is suitable for preparing the target material, or what level the target performance might be. For example, the generated sample reasoning problem can at least include information such as the candidate building block identifiers, target material type, target performance parameters, and judgment task type for each sample material design scheme. The judgment task types include, but are not limited to, tasks for judging the feasibility of structure formation, tasks for judging physicochemical properties, tasks for judging performance levels, tasks for judging recommendation status, tasks for judging recommendation priority, tasks for judging reasoning confidence, and tasks for judging the risks of chemical synthesis, among others.
[0061] It should be noted that when generating sample reasoning problems, the first target performance parameters and design information can be input into the existing language model, and the sample reasoning problems can be dynamically generated by the existing language model. This embodiment does not make specific limitations on this.
[0062] Step 130: Extract causal relationships from the mechanistic knowledge data to obtain at least one causal unit; fuse the evidence data and the at least one causal unit to obtain the actual reasoning path corresponding to each of the sample material design schemes.
[0063] Optionally, after obtaining the sample reasoning problem, in order to transform the scattered mechanistic knowledge in the literature into computable and reusable structured reasoning data, it is necessary to extract causal relationships from the mechanistic knowledge data to obtain at least one causal unit, and then integrate the extracted causal unit with the evidence data to obtain the actual reasoning path that can reflect the material design judgment process.
[0064] Causal relationship extraction refers to the process of identifying content representing relationships such as promotion, inhibition, causation, dependence, constraint, or conditionality from the text corresponding to mechanistic knowledge data and representing it in a structured form. A causal unit is a data unit obtained after causal relationship extraction, used to structurally represent a set of causal relationships. For example, the mechanistic knowledge that "increased rigidity of building units promotes improved framework orderliness" can be represented as a causal unit with corresponding antecedent entities, action relationships, consequence entities, applicable conditions, and evidence identifiers, rather than simply storing natural language fragments.
[0065] The actual reasoning path refers to the path that organizes the reasoning information of several reasoning nodes according to a preset judgment order for a sample material design scheme, and is used to reflect the judgment process from the construction unit to the target performance. The actual reasoning path may include multiple reasoning nodes arranged in sequence, and each reasoning node corresponds to a judgment step in the material design judgment process.
[0066] In one possible implementation, the order of the inference nodes may include, in turn, inference nodes for judging building units and preparation conditions, inference nodes for judging frame or topology formation, inference nodes for judging structure and physical properties, inference nodes for judging control strategies, and inference nodes for judging target performance, thereby preserving the cross-scale intermediate judgment chain from building units, preparation conditions, structural features, physical properties to target performance.
[0067] It should be noted that when fusing evidence data and at least one causal unit to obtain the actual reasoning path, content related to the sample reasoning problem can be retrieved from the evidence data and causal units, and the retrieved content can be organized into sequentially arranged reasoning nodes according to the above-mentioned preset judgment order, thereby obtaining the actual reasoning path corresponding to the sample material design scheme.
[0068] Step 140: Obtain the reasoning judgment label of the sample reasoning problem based on the reasoning conclusion in the actual reasoning path.
[0069] Optionally, after obtaining the actual reasoning path, the judgment result corresponding to the sample reasoning problem can be determined based on the reasoning conclusion in the actual reasoning path, and the judgment result can be used as the reasoning judgment label.
[0070] In this context, the reasoning conclusion refers to the local judgment result corresponding to each reasoning node in the actual reasoning path. For example, the conclusion of a reasoning node regarding "the combination of building units has a reasonable possibility of frame formation," and the conclusion at the end of the reasoning path regarding the target performance. The reasoning judgment label refers to the label corresponding to the sample reasoning problem, used to characterize the final judgment result of the sample material design scheme. The reasoning judgment label can be jointly determined by the conclusion at the end of the reasoning path, the target performance threshold, the completeness of evidence, and risk nodes, and output in a structured form.
[0071] Step 150: Train the pre-trained language model based on the sample reasoning problem, the actual reasoning path, and the reasoning judgment label to obtain the material reasoning model.
[0072] Optionally, after obtaining the sample reasoning problem, the actual reasoning path, and the reasoning judgment label, the sample reasoning problem can be used as input, and the actual reasoning path and the reasoning judgment label can be used as supervision targets to train the pre-trained language model. This allows the pre-trained language model to learn the reasoning link information between building units, preparation conditions, structural information, physical and chemical properties, control strategies, and target performance from the literature, thereby obtaining a material reasoning model.
[0073] The pre-trained language model refers to an existing large language model (LLM), which is a natural language processing (NLP) model with a large number of parameters, such as the Bidirectional Encoder Representations from Transformers (BERT) model, generative pre-trained Transformers, text-to-text Transformers, and human-computer interaction models deployed in various interactive systems for human-computer interaction. This embodiment does not specifically limit this type of pre-trained language model. Furthermore, the number of model parameters and / or the complexity of the model structure exceed a preset threshold. This model is pre-trained on a large amount of text data and possesses high semantic understanding and natural language generation capabilities.
[0074] It should be noted that when training the pre-trained language model, the sample reasoning questions corresponding to each sample material design scheme can be used as training input, and the corresponding actual reasoning paths and reasoning judgment labels can be used as training objectives. The loss value between the output of the pre-trained language model and the training objective is calculated, and the parameters of the pre-trained language model are iteratively updated based on this loss value until a preset training stopping condition is met, thus obtaining the material reasoning model. Therefore, the obtained material reasoning model can not only output the reasoning judgment results of the material design scheme, but also output a logically traceable predictive reasoning path supported by literature evidence, structural judgment, physicochemical property judgment, and target performance judgment.
[0075] The method provided in this embodiment obtains a set of target literature based on the first target material type and the first target performance parameters, identifies evidence data and mechanistic knowledge data from it, extracts causal relationships from the mechanistic knowledge data and merges it with the evidence data to obtain the actual reasoning path, and then trains a pre-trained language model based on the sample reasoning problem, the actual reasoning path and the reasoning judgment label obtained from the reasoning conclusion. This enables the trained material reasoning model to transform the scattered scientific judgment process in the literature into reusable reasoning ability. While outputting the judgment result, it provides a reasoning path that is supported by evidence and is logically traceable. This effectively overcomes the problems of difficult-to-trace reasoning process and low credibility of output performance prediction values in related technologies, and improves the reliability of material design scheme recommendations based on this.
[0076] Based on the above embodiments, as an optional embodiment, the step of fusing the evidence data and the at least one causal unit to obtain the actual reasoning path corresponding to each of the sample material design schemes includes the following steps: First, for each sample material design scheme, in the evidence data and at least one causal unit corresponding to the sample material design scheme, the evidence context information associated with the sample reasoning problem corresponding to the sample material design scheme is retrieved to obtain the evidence set of the sample material design scheme.
[0077] Optionally, for each sample material design scheme, before extracting the evidence context information, the evidence data corresponding to the sample material design scheme can be standardized to convert the field information contained in different evidence data, such as material name, building unit, functional group, preparation conditions, structural characterization, physicochemical properties, target performance and test conditions, into a unified field mapping space, and record the source literature and extraction method for each field information.
[0078] In addition, entity alignment and condition alignment can be performed on each causal unit corresponding to the sample material design scheme. That is, the entities in each causal unit, such as material name, abbreviation, construction unit and performance index, are mapped to the normative entities in the standardized evidence data. At the same time, the applicable conditions are unified, such as temperature, time, light source, sacrificial agent, co-catalyst, pH value and performance unit.
[0079] Then, focusing on the sample reasoning problem corresponding to the sample material design scheme, evidence context information is retrieved from the standardized evidence data and at least one aligned causal unit corresponding to the sample material design scheme, and the retrieved evidence context information is assembled into an evidence set for the sample material design scheme.
[0080] Among them, evidence context information refers to the evidence data and causal units associated with the sample reasoning problem and used to support the judgment of each reasoning node; the evidence set refers to the set of all evidence context information retrieved for a sample material design scheme.
[0081] For example, for each sample material design scheme, when retrieving evidence context information, at least one of the following can be considered as a retrieval condition: entity relevance to the sample reasoning problem, target performance relevance, experimental condition similarity, and evidence source reliability. This allows the standardized evidence data corresponding to the sample material design scheme, along with the structured fields and causal units that meet the retrieval conditions from at least one aligned causal unit, to form the evidence set for that sample material design scheme. Each piece of evidence context information in the evidence set retains a unique evidence identifier for subsequent inference nodes to reference that evidence, thus ensuring the traceability of the reasoning conclusion.
[0082] Then, from the evidence set, the reasoning information of each reasoning node in the actual reasoning path corresponding to the sample material design scheme is identified; the reasoning information includes the reasoning conclusion, evidence identification, direction of influence, and confidence level.
[0083] Optionally, after obtaining the evidence set, the reasoning information corresponding to each reasoning node constituting the actual reasoning path is identified from the evidence set. Here, reasoning information refers to information used to describe the judgment content and basis of a reasoning node, specifically including the reasoning conclusion, evidence identifier, direction of influence, and confidence level.
[0084] The reasoning conclusion here is the local judgment result of the reasoning node; the evidence identifier is used to uniquely identify the evidence cited by the reasoning conclusion so that the reasoning conclusion can be traced back to the corresponding evidence data or causal unit in the evidence set; the direction of influence is used to characterize the direction of the effect of the factors involved in the reasoning node on subsequent judgments or target performance, such as promoting or inhibiting; the confidence level is used to characterize the credibility of the reasoning conclusion, and its value can be determined jointly based on the reliability of the source of the evidence cited by the reasoning conclusion, the completeness of entity alignment and condition alignment. By introducing confidence levels into the reasoning information of each reasoning node, when constructing the actual reasoning path, reasoning information with confidence levels below a pre-set confidence threshold that cannot determine the entity's identity or experimental conditions can be avoided from being directly used to generate affirmative reasoning conclusions. This prevents low-confidence evidence from being incorporated into the actual reasoning path and introducing noise. Furthermore, since confidence levels are written into the actual reasoning path as a component of the reasoning information and participate in training as a supervised target, the pre-trained language model can learn the correlation between reasoning conclusions and their confidence levels while learning to generate reasoning paths. Consequently, it can output correspondingly lower confidence levels for reasoning conclusions with insufficient evidence support during the reasoning stage, avoiding the model from giving overly affirmative outputs for judgments lacking sufficient evidence support, thus improving the reliability of the output results of the obtained material reasoning model.
[0085] Then, according to the reasoning order corresponding to each reasoning node, the reasoning information of each reasoning node is combined to obtain the actual reasoning path corresponding to the sample material design scheme.
[0086] Optionally, after identifying the reasoning information of each reasoning node, the reasoning information of each reasoning node is combined according to the reasoning order corresponding to each reasoning node, so as to obtain the actual reasoning path corresponding to the sample material design scheme.
[0087] The reasoning order refers to the arrangement of each reasoning node in the actual reasoning path. In one possible implementation, the reasoning sequence can be as follows: reasoning nodes for judging building units and preparation conditions, reasoning nodes for judging framework or topology formation, reasoning nodes for judging structure and physicochemical properties, reasoning nodes for judging control strategies, and reasoning nodes for judging target performance. This allows the combined actual reasoning path to retain the cross-scale intermediate judgment chain from building units, preparation conditions, structural features, physicochemical properties to target performance. That is, the actual reasoning path can include the cross-scale intermediate judgment chain of analyzing building units, structure formation, structural characterization, physicochemical properties, control strategies, and performance differences. For example, it can consider whether the building unit has suitable reaction sites, whether it is conducive to forming a rigid framework, and whether it has donor or acceptor characteristics; analyze whether the bonding supports dynamic error correction, whether the topology is reasonable, and whether the interlayer stacking is likely to be ordered; analyze the structure and physicochemical properties such as specific surface area, pore structure, morphology, band gap, band position, and charge separation; analyze control strategies such as introducing electron acceptors, extending the conjugated framework, adjusting the pore structure, changing the bonding, or optimizing preparation conditions; and analyze the reasons for performance differences or low performance of materials in the same series.
[0088] By traversing each of the aforementioned sample material design schemes, the actual reasoning path corresponding to each of the aforementioned sample material design schemes is obtained.
[0089] The method provided in this embodiment constructs an evidence set by retrieving evidence context information around the sample reasoning problem, and then combines the reasoning information containing reasoning conclusions, evidence identifiers, influence directions and confidence levels from the evidence set according to the reasoning order. This makes each reasoning node in the obtained actual reasoning path associated with a valid source of evidence and a clear direction of influence, further enhancing the traceability and logical consistency of the actual reasoning path.
[0090] Based on the above embodiments, as an optional embodiment, the step of extracting causal relationships from the mechanistic knowledge data to obtain at least one causal unit includes the following steps: First, for each sample material design scheme, the mechanism knowledge data corresponding to the sample material design scheme is divided into at least one knowledge segment according to preset relation words.
[0091] Optionally, for each sample material design scheme, the mechanism knowledge data corresponding to the sample material design scheme is divided according to preset relational terms to obtain at least one knowledge fragment corresponding to the sample material design scheme.
[0092] Among them, predefined relational terms refer to keywords used to indicate causal relationships, such as words that characterize relationships like promotion, inhibition, cause, dependence, constraint, or conditionality; knowledge fragments refer to text segments in mechanistic knowledge data that contain expressions of causal relationships. By dividing mechanistic knowledge data using predefined relational terms, the mechanistic knowledge data can be segmented into several knowledge fragments, each containing expressions of causal relationships, so that causal relationships can be extracted from each knowledge fragment separately.
[0093] Then, causal relationships are extracted from each of the knowledge fragments to obtain at least one causal unit corresponding to the sample material design scheme; the causal unit includes antecedent entity, action relationship, consequence entity, applicable conditions, and evidence identifier; the action relationship is used to characterize the direction of influence of the antecedent entity on the consequence entity.
[0094] Optionally, causal relationships are extracted from each of the knowledge segments obtained from the division, and the causal relationship represented by each knowledge segment is represented as a structured causal unit, thereby obtaining at least one causal unit corresponding to the design scheme of the sample material.
[0095] The causal unit comprises antecedent entity, action relationship, consequential entity, applicable conditions, and evidence identifier. In other words, each causal unit can be represented as a structured form: antecedent entity - action relationship - consequential entity - applicable conditions - evidence identifier. The antecedent entity is the entity that produces the effect, the consequential entity is the entity that is affected, the action relationship characterizes the direction of the antecedent entity's influence on the consequential entity (e.g., promoting or inhibiting), the applicable conditions define the conditions required for the causal relationship to be established (e.g., specific temperature, time, light source, sacrificial agent, co-catalyst, or pH value), and the evidence identifier uniquely identifies the source of evidence for the causal unit.
[0096] By iterating through each of the sample material design schemes, at least one causal unit corresponding to each sample material design scheme is obtained.
[0097] The method provided in this embodiment divides mechanistic knowledge data into knowledge fragments according to preset relation words, and extracts causal relationships from each knowledge fragment to obtain structured causal units containing antecedent entities, action relationships, consequence entities, applicable conditions, and evidence identifiers. This transforms mechanistic knowledge expressed in natural language in the literature into computable, reusable structured data with applicable conditions and evidence sources.
[0098] Based on the above embodiments, as an optional embodiment, obtaining the reasoning judgment label of the sample reasoning problem according to the reasoning conclusion in the actual reasoning path includes the following steps: First, for each sample material design scheme, in the reasoning conclusions of the actual reasoning path corresponding to the sample material design scheme, identify the target reasoning conclusions associated with each judgment task in the sample reasoning problem corresponding to the sample material design scheme.
[0099] Optionally, for each sample material design scheme, the corresponding sample reasoning problem may be configured with one or more judgment tasks, each judgment task corresponding to a type of content that needs to be judged. When obtaining the reasoning judgment label, firstly, among the reasoning conclusions of the actual reasoning path corresponding to the sample material design scheme, the target reasoning conclusion associated with each judgment task in the sample reasoning problem is identified. Here, the target reasoning conclusion refers to the reasoning conclusion in the actual reasoning path that can respond to a certain judgment task.
[0100] The judgment tasks configured in the sample reasoning problems corresponding to each of the aforementioned sample material design schemes include multiple tasks such as judging the feasibility of structure formation, judging physicochemical properties, judging performance level, judging recommendation status, judging recommendation priority, judging reasoning confidence, and judging chemical synthesis risks. For example, for the task of judging the feasibility of structure formation, the target reasoning conclusion associated with this task can be identified from the reasoning conclusions corresponding to the reasoning nodes for judging frame or topology formation in the actual reasoning path; for the task of judging performance level, the target reasoning conclusion associated with this task can be identified from the reasoning conclusions corresponding to the reasoning nodes for judging target performance in the actual reasoning path.
[0101] Then, based on the target reasoning conclusion, the reasoning judgment sub-labels corresponding to each judgment task are obtained.
[0102] Optionally, after identifying the target inference conclusions associated with each judgment task, a reasoning judgment sub-label corresponding to each judgment task is determined based on the target inference conclusions corresponding to each judgment task. Here, the reasoning judgment sub-label refers to the judgment result label determined for a single judgment task. For example, for a task used to judge the feasibility of structural formation, a sub-label indicating whether the sample material design scheme has structural formation feasibility can be determined based on the corresponding target inference conclusion; for a task used to judge the performance level, a performance level sub-label can be determined based on the corresponding target inference conclusion; for a task used to judge the risk of chemical synthesis, a sub-label indicating whether the sample material design scheme has synthesis risk can be determined based on the corresponding target inference conclusion.
[0103] Then, the reasoning judgment sub-labels corresponding to each judgment task are combined to obtain the reasoning judgment labels for the sample reasoning problem corresponding to the sample material design scheme.
[0104] Optionally, after obtaining the reasoning judgment sub-labels corresponding to each judgment task, the reasoning judgment sub-labels corresponding to each judgment task are combined to obtain the reasoning judgment label for the sample reasoning problem corresponding to the sample material design scheme. Thus, the reasoning judgment label can comprehensively reflect the judgment results of the sample material design scheme in multiple aspects, such as structural feasibility, physicochemical properties, performance level, recommendation status, recommendation priority level, and chemical synthesis risk.
[0105] By iterating through each of the sample material design schemes, the reasoning judgment labels for the sample reasoning problems corresponding to each sample material design scheme are obtained.
[0106] The method provided in this embodiment identifies the target reasoning conclusions associated with each judgment task in the reasoning conclusions of the actual reasoning path, and obtains the reasoning judgment sub-labels corresponding to each judgment task accordingly, and then combines them to obtain reasoning judgment labels. This allows the reasoning judgment labels to cover multiple dimensions of judgment requirements such as structural formation feasibility, physical and chemical properties, performance level, recommendation status, recommendation priority level, and chemical synthesis risk. Moreover, each reasoning judgment sub-label corresponds to the corresponding reasoning conclusion in the actual reasoning path, thereby ensuring the consistency between the reasoning judgment labels and the reasoning process.
[0107] Based on the above embodiments, as an optional embodiment, training the pre-trained language model to obtain the material reasoning model according to the sample reasoning question, the actual reasoning path, and the reasoning judgment label includes the following steps: First, training samples corresponding to each of the sample material design schemes are constructed based on the sample reasoning problem, actual reasoning path, and reasoning judgment label.
[0108] Optionally, for each sample material design scheme, the sample reasoning problem corresponding to that sample material design scheme is used as the input part of the training sample, and the actual reasoning path and reasoning judgment label corresponding to that sample material design scheme are used as the supervision target part of the training sample, thus constructing the training sample corresponding to that sample material design scheme. Here, the training sample refers to the data unit used to train the pre-trained language model.
[0109] Then, consistency verification is performed on each of the training samples, and the target training sample that passes the verification is obtained.
[0110] Optionally, to control the noise of the training data, after constructing each training sample, the consistency of each training sample can be verified first.
[0111] Consistency verification refers to the process of verifying the logical consistency among the sample reasoning questions, actual reasoning paths, and reasoning judgment labels in the training samples. This consistency verification includes verifying several aspects, such as whether the sample reasoning questions in each training sample are related to the reasoning conclusions in the actual reasoning path; whether each reasoning node in the actual reasoning path is associated with valid evidence identifiers; whether there are logical jumps between the reasoning conclusions of adjacent reasoning nodes in the actual reasoning path; whether each reasoning conclusion in the actual reasoning path is consistent with the reasoning judgment labels; and whether the performance values in the actual reasoning path are associated with valid evidence identifiers.
[0112] Furthermore, training samples that pass the consistency verification are used as target training samples in the training set; training samples that fail the consistency verification, i.e., those with unfounded performance values, confusing material entities, misinterpreted abbreviations, missing sources of evidence, conflicting inferences with evidence, or incorrect labels, can be corrected, downweighted, or converted into negative samples.
[0113] It should be noted that there are multiple target training samples here.
[0114] Then, based on the target training samples, the pre-trained language model is trained to obtain the material reasoning model.
[0115] Optionally, after obtaining the verified target training samples, the pre-trained language model is trained based on the target training samples to obtain the material reasoning model.
[0116] The method provided in this embodiment constructs training samples based on sample reasoning questions, actual reasoning paths, and reasoning judgment labels. After verifying the consistency of the training samples, the model is trained using the verified target training samples. This effectively filters out low-quality samples with missing evidence, entity confusion, reversed conclusions, or unfounded numerical values, thereby reducing the interference of training data noise on the model and improving the reliability of the output results and the consistency of reasoning paths of the obtained material reasoning model.
[0117] Figure 3 This is the second flowchart of the inference model training method provided by the present invention.
[0118] like Figure 3 As shown, based on the above embodiments, as an optional embodiment, training the pre-trained language model according to the target training samples to obtain the material reasoning model includes the following steps: First, the sample reasoning problem in the target training sample is input into the pre-trained language model to obtain the predicted reasoning path and reasoning judgment result corresponding to the target training sample output by the pre-trained language model.
[0119] Optionally, during training, the reasoning questions from the target training samples can be input into a pre-trained language model. The pre-trained language model then performs reasoning on the sample reasoning questions and outputs the predicted reasoning path and reasoning judgment result corresponding to the target training sample. Here, the predicted reasoning path refers to the reasoning path predicted by the pre-trained language model for the sample reasoning question; the reasoning judgment result refers to the judgment result predicted by the pre-trained language model for the sample reasoning question.
[0120] Then, the target loss value is obtained based on the inference loss value between the predicted inference path and the actual inference path in the target training sample, and the inference loss value between the inference judgment result and the inference judgment label in the target training sample.
[0121] Optionally, after obtaining the predicted inference path and the inference judgment result, the inference loss value between the predicted inference path and the actual inference path in the target training sample, and the inference loss value between the inference judgment result and the inference judgment label in the target training sample are determined respectively, and the target loss value is obtained accordingly.
[0122] Specifically, the standard Supervised Fine-Tuning (SFT) paradigm can be used for model training, with the sample reasoning problems in the target training samples as training input, which can be denoted as... The complete target sequence, composed of the actual inference path and the inference judgment label in the target training sample, is taken as the supervised output and can be denoted as... ,in For the first in the target sequence The target word element at each position, that is, the component of the reasoning information node or reasoning judgment label in the actual reasoning path; Let be the length of the target sequence. At this point, the inference loss between the predicted inference path and the actual inference path, as well as the inference loss between the inference judgment result and the inference judgment label, are all treated as structured text signals within the same target sequence. These are learned jointly by a unified word-level language modeling objective, thereby optimizing the model through standard autoregressive cross-entropy loss to obtain the target loss value. The formula for calculating this target loss value is: ; in, The target loss value; For the first in the target sequence The sequence of target words preceding each position serves as the context for autoregressive generation. These are the parameters of the pre-trained language model; In order to train input and already generated Under these conditions, the pre-trained language model is Time output The conditional probability value.
[0123] Then, based on the target loss value, the pre-trained language model is iteratively trained until the performance of the pre-trained language model converges and / or the maximum number of iterations is reached.
[0124] Optionally, after obtaining the target loss value, the parameters of the pre-trained language model are iteratively updated based on the target loss value to minimize it, until the performance of the pre-trained language model converges and / or the maximum number of iterations is reached, at which point iterative training stops.
[0125] Among them, model performance convergence means that the validation metrics of the pre-trained language model on the validation set no longer improve after training.
[0126] In actual training, training, validation and test sets can be divided according to material system, performance level and negative example type, and the same material or highly similar building unit combination should be avoided to ensure the objectivity of model evaluation.
[0127] The validation metrics may include at least one of the following: judgment accuracy, macro average F1 score (i.e., macro average F1 score), precision and recall of the target high-performance category, evidence citation accuracy, ranking correlation coefficient, and inference path consistency ratio.
[0128] Then, based on the pre-trained language model obtained from each iteration of training, the material reasoning model is obtained.
[0129] Optionally, after the iterative training stops, a material reasoning model is obtained based on the pre-trained language models obtained from each iteration. For example, the pre-trained language model obtained from the last iteration can be used as the material reasoning model, or the pre-trained language model that has the best validation metrics on the validation set among the pre-trained language models obtained from each iteration can be used as the material reasoning model. This embodiment does not specifically limit this.
[0130] The method provided in this embodiment obtains predicted reasoning paths and reasoning judgment results by inputting sample reasoning problems into a pre-trained language model. Based on the reasoning loss value between the predicted reasoning path and the actual reasoning path, as well as the reasoning loss value between the reasoning judgment result and the reasoning judgment label, a target loss value is obtained. Then, iterative training is performed with the goal of minimizing the target loss value. This allows the pre-trained language model to simultaneously learn the generation of reasoning paths and the prediction of final judgments under a unified lexical-level language modeling objective. As a result, a material reasoning model that can output logically traceable predicted reasoning paths and accurate reasoning judgment results is obtained, effectively improving the interpretability and judgment accuracy of the model output.
[0131] Figure 4 This is one of the flowcharts of the recommended method provided by the present invention; such as Figure 4 As shown, based on the material reasoning model trained in the above embodiments, the present invention also provides a recommendation method, which may include steps 410 to 440.
[0132] Step 410: Obtain a set of candidate material design schemes based on the second target material type and the second target performance parameters.
[0133] Optionally, before recommending a material design scheme, it is necessary to first determine the scope of materials and performance objectives to be recommended, namely the second target material type and the second target performance parameters, and then obtain a set of candidate material design schemes based on this.
[0134] Here, the second target material type refers to the material system category targeted in this recommendation, and the second target performance parameter refers to the performance indicators of the targeted material type in the target application scenario. It should be noted that, to ensure the reasoning capability of the material reasoning model can be effectively transferred to the recommendation scenario, the second target material type can be the same as the first target material type, and the second target performance parameter can be the same as the first target performance parameter; furthermore, the second target material type can also be a subtype of the first target material type, and the second target performance parameter can also be a sub-performance parameter of the first target performance parameter, etc., and this embodiment does not impose specific limitations in this regard.
[0135] The candidate material design scheme set refers to a collection of multiple candidate material design schemes that require performance evaluation and recommendation, where each candidate material design scheme corresponds to a design scheme for a material to be evaluated. When obtaining the candidate material design scheme set, it can be obtained by filtering from a pre-built candidate material design scheme library based on the second target material type and the second target performance parameters; alternatively, it can be dynamically generated based on the building unit library corresponding to the second target material type, reaction or assembly rules, and the second target performance parameters, and each candidate material combination in the candidate material combination space can be used as a candidate material design scheme to constitute the candidate material design scheme set. This embodiment does not specifically limit this approach.
[0136] Step 420: Based on the second target performance parameters and the design information of each candidate material design scheme in the candidate material design scheme set, generate the target reasoning problem corresponding to each candidate material design scheme.
[0137] Optionally, after obtaining the set of candidate material design schemes, for each candidate material design scheme, a target reasoning problem corresponding to the candidate material design scheme is generated by combining the second target performance parameters and the design information of the candidate material design scheme.
[0138] The design information of the candidate material design scheme refers to the descriptive information on the composition and structural structure of the candidate material design scheme. This information may include at least one of the following: candidate building unit identifier, building unit combination method, bonding type and metric relationship, topology and topology type, and suggested preparation condition range. It should be noted that for candidate material design schemes where some preparation conditions are not yet determined, suggested preparation condition ranges can be provided as variables to be optimized, along with the allowable value ranges.
[0139] The target reasoning problem refers to a problem used to describe the material design judgment task of the candidate material design scheme to be evaluated. Its generation method can refer to the generation method of the sample reasoning problem. For example, the second target performance parameter and the design information of the candidate material design scheme can be filled into a preset problem template, or input into a language model for dynamic generation. This will not be elaborated further here. The target reasoning problem can at least include information such as the candidate building unit identifier, target material type, target performance parameter, and judgment task type of the candidate material design scheme.
[0140] Step 430: Input the target reasoning problem into the material reasoning model to obtain the predicted reasoning path and reasoning judgment result corresponding to each candidate material design scheme output by the material reasoning model.
[0141] Optionally, after generating the target reasoning questions corresponding to each candidate material design scheme, each target reasoning question is input into the material reasoning model, and the material reasoning model performs reasoning on each target reasoning question to output the prediction reasoning path and reasoning judgment result corresponding to each candidate material design scheme.
[0142] The material reasoning model is trained based on the reasoning model training method described in steps 110-150, which will not be repeated here.
[0143] The predictive reasoning path refers to the reasoning path predicted by the material reasoning model for the target reasoning problem, which is used to reflect the judgment process from the building unit to the target performance. It can also include multiple reasoning nodes arranged in sequence. Each reasoning node contains reasoning information such as reasoning conclusion, evidence identification, influence direction and confidence level. The reasoning judgment result refers to the judgment result predicted by the material reasoning model for the target reasoning problem. It can include at least one of the following: structural feasibility of the candidate material design scheme, target performance level or numerical range, recommendation status, recommendation priority level and chemical synthesis risk.
[0144] Therefore, while outputting the reasoning and judgment results of candidate material design schemes, the material reasoning model can also output a logically traceable predictive reasoning path supported by literature evidence, structural judgment, physicochemical property judgment and target performance judgment.
[0145] Step 440: Based on the prediction reasoning path and reasoning judgment result corresponding to each candidate material design scheme, at least one material design scheme to be recommended is selected from the set of candidate material design schemes.
[0146] Optionally, after obtaining the predictive inference path and inference judgment result corresponding to each candidate material design scheme, a selection process is performed from the candidate material design scheme set based on the predictive inference path and inference judgment result corresponding to each candidate material design scheme to obtain at least one material design scheme to be recommended. The material design scheme to be recommended refers to the candidate material design scheme determined after selection that is worthy of subsequent synthesis or experimental testing.
[0147] The method provided in this embodiment obtains a set of candidate material design schemes based on the second target material type and the second target performance parameters. After generating a target reasoning question based on the design information of each candidate material design scheme, it inputs the question into the material reasoning model to obtain a predictive reasoning path and reasoning judgment result that is supported by evidence and is logically traceable. Based on this, a material design scheme to be recommended is selected. This makes the obtained recommendation result no longer dependent solely on the terminal performance prediction score, but based on the process chain from the building unit to the target performance, thereby improving the traceability and reliability of the material design scheme recommendation result.
[0148] Figure 5 This is a second flowchart illustrating the recommended method provided by the present invention; as shown below. Figure 5 As shown, based on the above embodiments, as an optional embodiment, the step of selecting at least one material design scheme to be recommended from the set of candidate material design schemes according to the prediction reasoning path and reasoning judgment result corresponding to each candidate material design scheme includes the following steps: First, based on the reasoning and judgment results corresponding to each candidate material design scheme, candidate material design schemes that are in the recommended state and whose reasoning confidence is higher than the confidence threshold are selected from the set of candidate material design schemes to obtain the first material design scheme subset.
[0149] Optionally, since the inference judgment results corresponding to each candidate material design scheme may include a recommended state to characterize whether the candidate material design scheme is worth entering the subsequent screening, candidate material design schemes in the recommended state and with an inference confidence level higher than the confidence threshold can be selected from the candidate material design scheme set based on the inference judgment results corresponding to each candidate material design scheme, and these can be formed into a first material design scheme subset. The first material design scheme subset refers to the subset composed of candidate material design schemes in the candidate material design scheme set that are in the recommended state and have an inference confidence level higher than the confidence threshold. Thus, preliminary screening can be completed first through the inference judgment results of the material inference model, retaining candidate material design schemes with potential target performance.
[0150] Then, the constraint tool component is invoked to perform constraint verification on each candidate material design scheme in the first material design scheme subset, and the constraint verification results corresponding to each candidate material design scheme in the first material design scheme subset are obtained; the verification results include the verification status identifier and verification score output by each constraint tool in the constraint tool component.
[0151] Optionally, after obtaining the first material design scheme subset, the constraint tool component is invoked to perform constraint verification on each candidate material design scheme in the first material design scheme subset, thereby obtaining the constraint verification result corresponding to each candidate material design scheme.
[0152] Among them, the constraint tool component refers to the tool component used to verify and supplement the output of the material reasoning model based on explicit chemical constraints. It can include multiple tools such as building block analysis tools, reaction or assembly feasibility tools, physicochemical property tools, and rule screening tools. Accordingly, constraint verification includes building block verification based on building block analysis tools, reaction feasibility or assembly feasibility verification based on reaction or assembly feasibility tools, physicochemical property verification based on physicochemical property tools, and performance constraint verification based on rule screening tools. Building block verification can be used to identify building block names, functional groups, reaction sites, donor and acceptor properties, and steric hindrance risks. Reaction feasibility or assembly feasibility verification can be used to determine functional group matching, the probability of bond formation, the feasibility of dynamic covalent reactions, or the feasibility of structure formation. Physicochemical property verification can be used to analyze or predict color, specific surface area, crystallinity, band gap, band position, and light response. Performance constraint verification can be used to determine the structural rules, property range, reaction conditions, and risk constraints corresponding to the target performance.
[0153] The constraint verification results include the verification status identifier and verification score output by each constraint tool in the constraint tool component. The verification status identifier indicates whether the candidate material design scheme has passed the verification of the corresponding constraint tool, such as passing or failing the verification; the verification score quantifies the performance score of the candidate material design scheme under the corresponding constraint tool.
[0154] Then, based on the verification status identifier, candidate material design schemes that pass the constraint verification are selected from the first material design scheme subset to obtain the second material design scheme subset.
[0155] Optionally, after obtaining the constraint verification results corresponding to each candidate material design scheme, the candidate material design schemes that pass the constraint verification are selected from the first material design scheme subset according to the verification status identifier in the constraint verification results, and these are combined into the second material design scheme subset. In this way, the model output can be transformed into candidate screening results that are reviewable, traceable, and usable for experimental decision-making.
[0156] The second subset of material design schemes refers to the subset of candidate material design schemes that have passed constraint verification from the first subset of material design schemes. For example, for candidate material design schemes that have functional group incompatibility, insufficient connectivity, or inability to close stoichiometric relationships, or whose band positions still cannot cover the target redox potential after considering prediction errors, their verification status indicates that they have failed verification, and they can be eliminated from the first subset of material design schemes, thus obtaining the second subset of material design schemes.
[0157] Then, based on the verification score and the reasoning judgment result, at least one material design scheme to be recommended is selected from the second material design scheme subset.
[0158] Optionally, after obtaining the second subset of material design schemes, further screening is performed on the second subset of material design schemes based on the verification score in the constraint verification results and the reasoning judgment results output by the material reasoning model, so as to obtain at least one material design scheme to be recommended.
[0159] The method provided in this embodiment first selects candidate material design schemes that are in a recommended state and whose inference confidence is higher than a set confidence threshold based on the inference judgment results to obtain a first subset of material design schemes. Then, it calls the constraint tool component to perform constraint verification and selects a second subset of material design schemes based on the verification status identifier. Finally, it combines the verification score and the inference judgment results to select the material design schemes to be recommended. This multi-stage screening mechanism, which first retains potential high-performance candidates and then gradually compresses them through explicit chemical constraint tools, transforms the output of the open material inference model into candidate screening results that are reviewable, traceable, and conform to explicit chemical constraints, thereby improving the verifiability and reliability of the recommendation results.
[0160] Based on the above embodiments, as an optional embodiment, the step of selecting at least one material design scheme to be recommended from the second material design scheme subset according to the verification score and the reasoning judgment result includes the following steps: First, based on the reasoning and judgment results, obtain the reasoning score corresponding to each candidate material design scheme in the second material design scheme subset.
[0161] Optionally, the reasoning score corresponding to each candidate material design scheme in the second material design scheme subset can be determined based on the reasoning judgment result output by the material reasoning model. The reasoning score refers to a score used to quantify the merits of the candidate material design schemes under the reasoning judgment result of the material reasoning model. It can be quantified based on at least one of the following indicators in the reasoning judgment result: structural feasibility label, physicochemical property label, performance level, recommendation priority level, reasoning confidence level, and chemical synthesis risk label. For example, the reasoning score can be obtained by weighted summing the quantified values corresponding to at least one of the following indicators in the reasoning judgment result: structural feasibility label, physicochemical property label, performance level, recommendation priority level, reasoning confidence level, and chemical synthesis risk label.
[0162] Then, based on the structural description information in the design information of each candidate material design scheme in the second material design scheme subset, the second material design scheme subset is divided into multiple material design scheme groups, and the structural similarity between each candidate material design scheme in each material design scheme group is higher than the similarity threshold.
[0163] Optionally, based on the structural description information in the design information of each candidate material design scheme within the second material design scheme subset, the second material design scheme subset is divided into multiple material design scheme groups. Here, structural description information refers to information used to describe the structural characteristics of the candidate material design schemes; a material design scheme group refers to a group composed of structurally similar candidate material design schemes, where the structural similarity between candidate material design schemes in each material design scheme group is higher than a preset similarity threshold. Therefore, grouping the second material design scheme subset by structural similarity allows candidate material design schemes with highly similar structures to be grouped into the same material design scheme group.
[0164] Then, based on the verification score and the reasoning score, the comprehensive score of each candidate material design scheme in each material design scheme group is obtained.
[0165] Optionally, for each candidate material design scheme in each material design scheme group, a comprehensive score is obtained based on the verification score and inference score corresponding to that candidate material design scheme. The comprehensive score refers to the score obtained by comprehensively considering the verification score and inference score, used to characterize the overall quality of the candidate material design scheme. For example, the comprehensive score can be obtained by weighted summing of the verification scores and inference scores output by each constraint tool.
[0166] Then, based on the comprehensive score, the candidate material design schemes in each of the material design scheme groups are screened to obtain at least one material design scheme to be recommended.
[0167] Optionally, after obtaining the comprehensive score of each candidate material design scheme in each material design scheme group, the candidate material design schemes in each material design scheme group are screened according to the comprehensive score to obtain at least one material design scheme to be recommended. For example, 10 material design schemes with high comprehensive scores are output as recommended schemes, or the candidate material design scheme with the highest comprehensive score is selected in each material design scheme group as the recommended material design scheme, so as to cover more diverse high-performance materials while ensuring the performance of the candidate material design schemes.
[0168] The method provided in this embodiment obtains a reasoning score based on the reasoning judgment result, divides the second material design scheme subset into multiple material design scheme groups according to structural similarity based on structural description information, and obtains a comprehensive score by combining the verification score and the reasoning score. Then, it filters among the material design scheme groups. This ensures that the final recommended material design scheme takes into account both the degree of constraint verification and the quality of reasoning judgment, while avoiding the concentration of the screening results on candidate material design schemes with highly similar structures, thereby improving the diversity of the recommendation results and the experimental value.
[0169] Based on the above embodiments, as an optional embodiment, after selecting at least one material design scheme to be recommended, the method further includes the following steps: First, obtain the experimental results corresponding to each of the proposed material design schemes.
[0170] Optionally, after selecting at least one material design scheme to be recommended, wet experiments can be conducted to verify each scheme and obtain the corresponding experimental results. The experimental results refer to the actual feedback data obtained after experimental verification of the recommended material design scheme. These results may include at least one of the following: whether the synthesis was successful, whether the structure meets expectations, crystallinity, specific surface area, band gap, target performance values, stability, and repeatability. Compared to literature containing sampling errors or experimental noise, these results provide authentic, standardized, and highly reliable data. Therefore, they can be used to verify the recommended material design scheme for reverse correction, thereby further dynamically refining the cognitive boundaries and improving the reliability of the recommendations.
[0171] Then, based on the experimental results, material design schemes that failed in the experiment are selected as negative examples from the at least one material design scheme to be recommended.
[0172] Optionally, after obtaining the experimental results corresponding to each material design scheme to be recommended, the material design schemes that failed in the experiment can be selected from at least one of the material design schemes to be recommended as negative examples. Here, a material design scheme that failed in the experiment refers to a material design scheme whose experimental results indicate that it did not achieve the expected results, such as a material design scheme that predicted high performance but performed poorly in the experiment; a negative example sample refers to a sample composed of material design schemes that failed in the experiment.
[0173] In addition, the reasons for experimental failure of negative samples can be obtained, which may include insufficient structure formation, poor crystallinity, unfavorable pore structure, low charge separation efficiency, insufficient active sites, limited interfacial reaction, incorrect entity identification, or mismatch of experimental conditions.
[0174] Then, the prediction inference path, inference judgment result, and constraint verification result corresponding to the negative example sample are compared with the experimental results corresponding to the negative example sample, and the negative example type label corresponding to the negative example sample is determined based on the comparison result.
[0175] Optionally, the prediction results corresponding to the negative example sample, i.e., the prediction inference path, inference judgment result, and constraint verification result, are aligned and mapped with the experimental results corresponding to the negative example sample. The differences between the aligned prediction results and the experimental results are then compared to determine the negative example type label corresponding to the negative example sample. Here, the negative example type label refers to the label used to characterize the failure reason type of the negative example sample.
[0176] In one possible implementation, negative example type labels can be obtained by comparing the results: if the target material or target bonding is not obtained, it can be marked as a reaction or synthesis failure negative example; if the material has been formed but the crystallinity, pore structure or topology does not meet expectations, it can be marked as a structure formation negative example; if the structure meets expectations but intermediate properties such as band gap, band position, hydrophilicity or charge separation do not meet predictions, it can be marked as a physicochemical property negative example; if the intermediate properties basically meet predictions but the target performance is still below the threshold, it can be marked as an interface reaction or performance conversion negative example; if the error originates from abbreviations, entities with the same name, evidence mismatch or numerical data without source, it can be marked as an entity or evidence negative example; if the experimental conditions exceed the scope of the literature, it can be marked as a condition migration negative example.
[0177] Then, based on the negative example type label, a target update operation is determined from multiple update operations and executed; wherein, the multiple update operations include the operation of fine-tuning the training of the material reasoning model, the operation of updating the constraint rule information in the constraint tool component, and the operation of updating the output weights of each constraint tool in the constraint tool component.
[0178] Optionally, after determining the negative example type label corresponding to the negative example sample, at least one of the following update operations is performed based on the negative example type label: fine-tuning the training of the material reasoning model, updating the constraint rule information in the constraint tool component, and updating the output weights of each constraint tool in the constraint tool component. This ensures that different failure reasons are mapped to different update operations, rather than simply retraining with a uniform low-performance label. This constructs a closed loop of literature and experiment collaboration, which can correct the material reasoning model and constraint tool component based on real experimental feedback. This alleviates the idealized judgment caused by literature success bias and text extraction noise, reduces illusions, entity confusion, and unfounded performance predictions, and further improves the adaptability and reliability of recommendation results in real experimental scenarios.
[0179] Based on the above embodiments, as an optional embodiment, determining the target update operation among multiple update operations according to the negative example type label, and executing the target update operation, includes the following steps: First, when the negative example type label belongs to the label of abnormal inference information, the experimental results corresponding to the negative example sample are compared one by one with the inference information of each inference node in the prediction inference path corresponding to the negative example sample. Based on the comparison results, the root cause node is determined in the prediction inference path corresponding to the negative example sample.
[0180] Optionally, if the negative example type label corresponding to the negative example sample belongs to the label of abnormal inference information, the experimental result corresponding to the negative example sample is compared with the inference information of each inference node in the prediction inference path corresponding to the negative example sample one by one, and the root cause node is determined in the prediction inference path corresponding to the negative example sample based on the comparison result.
[0181] Among them, the label of "abnormal reasoning information" refers to the label of negative example type caused by the reasoning information error of the material reasoning model, such as the aforementioned negative examples of structure formation, physicochemical properties, interface reactions, or performance transformation. The root cause node is the first reasoning node in the predictive reasoning path that is inconsistent with experimental facts. When determining the root cause node, the reasoning conclusions of each reasoning node can be compared with experimental characterization item by item in the order of building blocks and preparation conditions, framework formation, physicochemical properties, control strategies, and target performance. The reasoning node with the earliest conflict is determined as the root cause node, while the reasoning conclusions following the root cause node can be marked as affected conclusions and are not directly counted as independent error repetitions.
[0182] Then, based on the experimental results corresponding to the negative example samples, the reasoning information corresponding to the root cause node, the reasoning information corresponding to the reasoning node after the root cause node, and the reasoning judgment result corresponding to the negative example samples in the prediction reasoning path corresponding to the negative example samples are updated to obtain the corrected reasoning path and corrected reasoning judgment result corresponding to the negative example samples.
[0183] Optionally, after determining the root cause node, based on the experimental results corresponding to the negative example sample, the inference information corresponding to the root cause node, the inference information corresponding to the inference nodes following the root cause node, and the inference judgment result corresponding to the negative example sample in the predicted inference path are updated to obtain the corrected inference path and the corrected inference judgment result corresponding to the negative example sample. The corrected inference path refers to the inference path obtained by replacing the inference information of the root cause node and its subsequent inference nodes in the predicted inference path with inference information consistent with the experimental results; the corrected inference judgment result refers to the inference judgment result modified according to the experimental results, such as being modified to an actual performance level, a non-recommended state, or a conditionally recommended state. During the update process, the differences between the original predicted inference path and the corrected inference path can also be recorded.
[0184] Then, based on the target reasoning problem, corrected reasoning path, and corrected reasoning judgment result corresponding to the negative example sample, a negative example training sample corresponding to the negative example sample is constructed.
[0185] Optionally, after obtaining the corrected reasoning path and the corrected reasoning judgment result, the target reasoning problem corresponding to the negative example sample is retained. Using the target reasoning problem as the input and the corrected reasoning path and the corrected reasoning judgment result corresponding to the negative example sample as the supervision target, a negative example training sample corresponding to the negative example sample is constructed. Here, the negative example training sample refers to the training sample constructed based on the corrected result of the negative example sample, used for fine-tuning the material reasoning model.
[0186] Then, the material reasoning model is fine-tuned based on the negative example training samples and the negative example type labels.
[0187] Optionally, the training weight of negative example training samples can be increased according to the negative example type label, and the material reasoning model can be fine-tuned according to the training weight and negative example training samples, so that the material reasoning model focuses on learning the correction of the corresponding failure reasons.
[0188] The method provided in this embodiment compares the experimental results with the reasoning information of each inference node in the predicted inference path when the negative example type label belongs to the label of abnormal reasoning information to determine the root cause node. Based on this, the reasoning information and reasoning judgment results of the root cause node and its subsequent inference nodes are updated to obtain the corrected inference path and corrected inference judgment results. Then, negative example training samples are constructed to fine-tune the material reasoning model, so that the experimental feedback can be accurately located to the root cause node that caused the error and corrected, rather than retraining the entire inference path indiscriminately. This effectively alleviates model illusion, entity confusion and unfounded performance prediction, and further improves the judgment accuracy of the material reasoning model under real experimental conditions.
[0189] Furthermore, the negative example correction process may also include the following operations: updating the inference data construction logic, updating the explicit knowledge filtering rules, updating the training sample weights and active learning queue, and verifying the update results.
[0190] The operation of updating the inference data construction logic can be performed to update the inference data construction logic so that subsequently constructed inference data can avoid corresponding types of failure reasons. Here, the inference data construction logic refers to the logical rules followed when constructing the actual inference path using evidence data and causal units during training sample construction.
[0191] Specifically, for negative examples of entities or evidence, the priority of entity uniqueness verification and evidence source verification can be increased to avoid inference information anomalies caused by abbreviations, entities with the same name, evidence mismatch, or lack of source for numerical values. For negative examples of structure formation, the inference path can be required to have framework formation evidence before entering the inference node for electronic property judgment, thereby avoiding direct inference of electronic properties when the structure has not been effectively formed. For negative examples of physicochemical properties, the evidence weight of intermediate properties can be increased, such as increasing the evidence weight of measured color values or independent prediction results. For negative examples of interface reaction, inference nodes such as active sites, interface kinetics, or reaction conditions can be added before the inference node for target performance judgment to complete the intermediate judgments related to interface reaction. For negative examples of conditional transfer, the applicable conditions of the literature causal rules can be written into the corresponding inference nodes, and direct extrapolation across conditions can be prohibited, thereby avoiding the incorrect application of literature causal rules when they exceed their applicable conditions.
[0192] The explicit knowledge filtering rule update operation can be performed on the explicit knowledge filtering rules in the constraint tool component. Each explicit knowledge filtering rule can be represented by a "suitable condition - judgment condition - weight update" structure, meaning each rule includes its suitable condition, judgment condition, and corresponding weight. During the update, when experimental results support a rule, the absolute value of the rule's positive or negative weight is increased; when experimental results conflict with a rule, the absolute value of the rule's weight is decreased. This allows the explicit knowledge filtering rules to dynamically adjust their strength in constraint verification based on real experimental feedback.
[0193] The training sample weights and active learning queue update operations can be performed on the model training sample weights and the active learning queue. Specifically, the training weights of root cause negative examples that lead to false positives are increased, so that the materials reasoning model focuses on learning to correct the corresponding failure causes in subsequent fine-tuning training. Boundary samples that are close to the performance grading threshold and samples with low confidence in the materials reasoning model are added to the active learning queue, and in the next round of experiments, candidate material designs that can distinguish adjacent decision regions are selected for experimentation, thereby obtaining the most discriminative feedback data on the model's decision boundary with less experimental cost.
[0194] After performing the above update operations, the performance metrics before and after the update, such as high-performance category precision, false positive rate, evidence citation accuracy, and rule false negative rate, can be compared on an independent validation set. The update result is only written to the current version if the updated material reasoning model or explicit knowledge screening rule improves at least one target metric and other key metrics do not exceed the allowable degradation range; otherwise, the original version is retained, and the negative sample is still marked as a sample to be analyzed for further processing. This avoids overall performance degradation caused by introducing new biases in a single update.
[0195] The method provided in this embodiment maps different negative example types to different update nodes, such as the inference data construction logic, explicit knowledge screening rules, and training sample weights, instead of simply retraining the model with a uniform low-performance label. This allows experimental feedback to specifically correct different aspects, such as entity recognition, structure formation, physical property inference, interface performance judgment, and the scope of condition application. At the same time, by verifying the update results on an independent validation set before deciding whether to write them into the current version, it ensures that each update is based on the premise that key indicators do not degrade. This forms a closed loop from literature data preparation, model inference, candidate screening to wet experiment calibration, continuously improving the reliability and accuracy of the material inference model and constraint tool components under real experimental conditions.
[0196] Based on the above embodiments, as an optional embodiment, obtaining a set of candidate material design schemes according to the second target material type and the second target performance parameters includes the following steps: First, based on the second target material type and the second target performance parameters, the generation constraint information is obtained.
[0197] Optionally, generation constraint information is obtained based on the second target material type and the second target performance parameters to constrain the generation of candidate material design schemes. The generation constraint information refers to constraint information used to limit the range of candidate material design scheme generation, which may include target material constraints converted from the second target material type and target performance constraints converted from the second target performance parameters.
[0198] For example, target material constraints may include at least one of the following: allowed bond types, number of building blocks, node connectivity, target dimension, and optional topology; target performance constraints may include at least one of the following: bandgap range, band position, pore structure, hydrophilicity / hydrophobicity, stability, or other properties related to the target response. Furthermore, the constraints included in the generated constraint information can be divided into hard constraints that must be satisfied and soft constraints used for subsequent ranking. Hard constraints refer to those that candidate material design schemes must satisfy, while soft constraints refer to those used for subsequent ranking of candidate material design schemes.
[0199] Then, based on the generated constraint information, the candidate material design scheme set is obtained by filtering from the candidate material design scheme library.
[0200] Optionally, a library of candidate material design schemes can be obtained first.
[0201] Specifically, this can be achieved by establishing a specification record for each building unit in the building unit library, and establishing a functional group compatibility matrix A for each building unit based on reaction or assembly rules. For any two functional group types... and When the two can form a target linkage bond within the preset reaction conditions, A( , A( ) = 1; when the two do not react, competing side reactions occur, or the linkage product does not belong to the target material type, A( , =0. The compatibility matrix simultaneously records catalyst, solvent, temperature, and dynamic reversibility requirements, and for repeated structures, it removes duplicates through standardized structure identifiers to form a candidate material design scheme library.
[0202] The building block specification record includes at least one of the following for each building block in the building block library: record specification identifier, molecular structure representation, functional group type, number of reaction sites, spatial orientation of reaction sites, molecular geometry, rigidity, conjugation characteristics, donor and acceptor properties, molecular weight, steric hindrance description, and optional procurement or synthesis information.
[0203] Each candidate material design scheme in the candidate material design scheme library can be represented as follows: ,in, Let k be the set of building units for the k-th candidate material design scheme. For the k-th candidate material design scheme, the expected bonding and metric relationships are given. For the k-th candidate material design scheme, consider the candidate topology or structure type. The recommended preparation conditions range for the k-th candidate material design scheme.
[0204] Then, based on the generated constraint information, hard constraint verification is performed on each candidate material design scheme, and the hard constraints are verified by marking. Candidate material design schemes are retained to form a set of candidate material design schemes for subsequent material reasoning model processing; while for those that do not meet the hard constraints, i.e. Candidate material design schemes can retain their reasons for elimination and not include them in the candidate material design scheme set, thus preventing them from entering the subsequent model recall.
[0205] It should be noted that in the aforementioned candidate space generation stage, the optimization variables can be decoupled, locking the boundary of the candidate material design scheme at the dimension of building block combination, while removing synthesis condition parameters such as reaction temperature, solvent, and catalyst. This is because the combination of building blocks directly determines the intrinsic topological network and photoelectric properties of the material, and its design space is extremely large; for example, the cross-combination of 100 amino building blocks and more than 100 aldehyde building blocks can reach a scale of tens of thousands. In contrast, the synthesis condition parameters belong to a finite set of variables and can be empirically matched by domain experts as post-parameters in subsequent experimental verification. Therefore, by reducing the dimensionality of the candidate space to the dimension of building block combination, the size of the candidate space can be effectively controlled while ensuring that the intrinsic properties of the candidate material design scheme are adjustable.
[0206] The method provided in this embodiment obtains generation constraint information containing hard and soft constraints based on the second target material type and the second target performance parameters. Based on this, candidate material design schemes that meet the hard constraints are selected from the candidate material design scheme library to form a candidate material design scheme set. This ensures that all candidate material design schemes entering the subsequent material inference model processing meet the target material constraints and target performance constraints. At the same time, by decoupling the optimization variables of the candidate space and locking them to the dimension of building unit combination, the pre-filtering of candidate material design schemes that do not meet the hard constraints and the effective dimensionality reduction of the candidate space are completed in the candidate space generation stage, thereby improving the processing efficiency of the recommendation process.
[0207] In summary, addressing the problems of existing technologies such as over-reliance on human experience, lack of logical reasoning in intermediate processes, poor interpretability of prediction results, and fragmented multi-stage screening, the method provided in this application significantly improves the scientific credibility and interpretability of predictions by extracting the multi-dimensional mapping relationship of the entire process from building blocks to preparation conditions, structural features, physicochemical properties, control strategies, and target performance, and by introducing explicit chemical constraint tools and evidence tracking mechanisms. Furthermore, the multi-stage screening process achieves precise compression from tens of thousands of candidate material combinations to a small number of experimentally prioritized candidates. Moreover, it breaks the traditional isolation between extraction and prediction, forming a closed-loop R&D system encompassing literature evidence construction, model reasoning, tool verification, candidate ranking, and experimental feedback. This system can efficiently output high-performance porous frameworks and other functional materials with practical experimental value from a large candidate space. Simultaneously, the established digital and intelligent rational design paradigm has strong versatility and can directly serve the development of MOFs and other functional materials.
[0208] Figure 6 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 6As shown, the electronic device may include: a processor 610, a communication interface 620, a memory 630, and a communication bus 640, wherein the processor 610, the communication interface 620, and the memory 630 communicate with each other through the communication bus 640. The processor 610 can call logical instructions in the memory 630 to execute an inference model training method, which includes: obtaining a target literature set based on a first target material type and a first target performance parameter, and identifying evidence data and mechanism knowledge data corresponding to each sample material design scheme under the first target material type from the target literature set; generating sample inference questions corresponding to each sample material design scheme based on the first target performance parameter and the design information of each sample material design scheme; extracting causal relationships from the mechanism knowledge data to obtain at least one causal unit, fusing the evidence data and the at least one causal unit to obtain the actual inference path corresponding to each sample material design scheme; obtaining the inference judgment label of the sample inference question based on the inference conclusion in the actual inference path; and obtaining the inference judgment label of the sample inference question based on the sample inference question and the actual inference judgment label of the sample inference question. The material reasoning model is trained using the reasoning path and the reasoning judgment label to obtain a material reasoning model; or a recommendation method is executed, which includes: obtaining a set of candidate material design schemes based on a second target material type and a second target performance parameter; generating a target reasoning question corresponding to each candidate material design scheme based on the design information of each candidate material design scheme in the set of candidate material design schemes, based on the second target performance parameter and the design information of each candidate material design scheme in the set of candidate material design schemes; inputting the target reasoning question into the material reasoning model to obtain the predicted reasoning path and reasoning judgment result corresponding to each candidate material design scheme output by the material reasoning model; and selecting at least one material design scheme to be recommended from the set of candidate material design schemes based on the predicted reasoning path and reasoning judgment result corresponding to each candidate material design scheme; wherein the material reasoning model is trained based on a reasoning model training method.
[0209] Furthermore, the logical instructions in the aforementioned memory 630 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0210] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the inference model training method provided by the above methods. The method includes: obtaining a target literature set based on a first target material type and a first target performance parameter, and identifying evidence data and mechanism knowledge data corresponding to each sample material design scheme under the first target material type from the target literature set; generating sample inference problems corresponding to each sample material design scheme based on the first target performance parameter and the design information of each sample material design scheme; extracting causal relationships from the mechanism knowledge data to obtain at least one causal unit; fusing the evidence data and the at least one causal unit to obtain the actual inference path corresponding to each sample material design scheme; and obtaining the inference conclusion in the actual inference path. The process involves: taking the reasoning judgment label of the sample reasoning problem; training a pre-trained language model based on the sample reasoning problem, the actual reasoning path, and the reasoning judgment label to obtain a material reasoning model; or executing a recommendation method, which includes: obtaining a set of candidate material design schemes based on a second target material type and a second target performance parameter; generating a target reasoning problem corresponding to each candidate material design scheme based on the second target performance parameter and the design information of each candidate material design scheme in the set of candidate material design schemes; inputting the target reasoning problem into the material reasoning model to obtain the predicted reasoning path and reasoning judgment result corresponding to each candidate material design scheme output by the material reasoning model; and selecting at least one material design scheme to be recommended from the set of candidate material design schemes based on the predicted reasoning path and reasoning judgment result corresponding to each candidate material design scheme; wherein the material reasoning model is trained based on a reasoning model training method.
[0211] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements a reasoning model training method provided by the above methods. This method includes: acquiring a target literature set based on a first target material type and a first target performance parameter; identifying evidence data and mechanistic knowledge data corresponding to each sample material design scheme under the first target material type from the target literature set; generating sample reasoning questions corresponding to each sample material design scheme based on the first target performance parameter and the design information of each sample material design scheme; extracting causal relationships from the mechanistic knowledge data to obtain at least one causal unit; fusing the evidence data and the at least one causal unit to obtain an actual reasoning path corresponding to each sample material design scheme; and obtaining a reasoning judgment for the sample reasoning question based on the reasoning conclusion in the actual reasoning path. The method involves: labeling; training a pre-trained language model based on the sample reasoning question, the actual reasoning path, and the reasoning judgment label to obtain a material reasoning model; or executing a recommendation method, which includes: obtaining a set of candidate material design schemes based on a second target material type and a second target performance parameter; generating a target reasoning question corresponding to each candidate material design scheme based on the second target performance parameter and the design information of each candidate material design scheme in the set of candidate material design schemes; inputting the target reasoning question into the material reasoning model to obtain the predicted reasoning path and reasoning judgment result corresponding to each candidate material design scheme output by the material reasoning model; and selecting at least one material design scheme to be recommended from the set of candidate material design schemes based on the predicted reasoning path and reasoning judgment result corresponding to each candidate material design scheme; wherein the material reasoning model is trained based on a reasoning model training method.
[0212] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0213] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0214] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for training a reasoning model, characterized in that, include: Based on the first target material type and the first target performance parameters, a target literature set is obtained, and from the target literature set, evidence data and mechanism knowledge data corresponding to each sample material design scheme under the first target material type are identified; Based on the first target performance parameters and the design information of each of the sample material design schemes, generate sample reasoning questions corresponding to each of the sample material design schemes; The causal relationship is extracted from the mechanistic knowledge data to obtain at least one causal unit. The evidence data and the at least one causal unit are fused to obtain the actual reasoning path corresponding to each of the sample material design schemes. Based on the reasoning conclusions in the actual reasoning path, obtain the reasoning judgment label of the sample reasoning problem; Based on the sample reasoning problem, the actual reasoning path, and the reasoning judgment label, the pre-trained language model is trained to obtain the material reasoning model.
2. The inference model training method according to claim 1, characterized in that, The process of fusing the evidence data and the at least one causal unit to obtain the actual reasoning path corresponding to each of the sample material design schemes includes: For each sample material design scheme, in the evidence data and at least one causal unit corresponding to the sample material design scheme, the evidence context information associated with the sample reasoning problem corresponding to the sample material design scheme is retrieved to obtain the evidence set of the sample material design scheme. From the evidence set, identify the reasoning information of each reasoning node in the actual reasoning path corresponding to the sample material design scheme; the reasoning information includes the reasoning conclusion, evidence identifier, direction of influence, and confidence level; According to the reasoning order corresponding to each reasoning node, the reasoning information of each reasoning node is combined to obtain the actual reasoning path corresponding to the sample material design scheme; By traversing each of the aforementioned sample material design schemes, the actual reasoning path corresponding to each of the aforementioned sample material design schemes is obtained.
3. The inference model training method according to claim 1, characterized in that, The step of extracting causal relationships from the mechanistic knowledge data to obtain at least one causal unit includes: For each sample material design scheme, the mechanism knowledge data corresponding to the sample material design scheme is divided into at least one knowledge segment according to the preset relation words; Causal relationships are extracted from each of the knowledge fragments to obtain at least one causal unit corresponding to the sample material design scheme; the causal unit includes an antecedent entity, an action relationship, a consequence entity, applicable conditions, and evidence identifiers; the action relationship is used to characterize the direction of influence of the antecedent entity on the consequence entity; By iterating through each of the sample material design schemes, at least one causal unit corresponding to each sample material design scheme is obtained.
4. The inference model training method according to any one of claims 1-3, characterized in that, The step of obtaining the reasoning judgment label of the sample reasoning problem based on the reasoning conclusion in the actual reasoning path includes: For each sample material design scheme, in the reasoning conclusions of the actual reasoning path corresponding to the sample material design scheme, identify the target reasoning conclusions associated with each judgment task in the sample reasoning problem corresponding to the sample material design scheme. Based on the target reasoning conclusion, obtain the reasoning judgment sub-labels corresponding to each judgment task; The reasoning judgment sub-labels corresponding to each of the judgment tasks are combined to obtain the reasoning judgment labels of the sample reasoning problem corresponding to the sample material design scheme. By iterating through each of the aforementioned sample material design schemes, the reasoning judgment labels for the sample reasoning problems corresponding to each of the aforementioned sample material design schemes are obtained; Among them, the judgment tasks configured in the sample reasoning problem corresponding to each of the sample material design schemes include multiple tasks such as judging the feasibility of structure formation, judging the physicochemical properties, judging the performance level, judging the recommendation status, judging the recommendation priority level, judging the reasoning confidence, and judging the risk of chemical synthesis.
5. The inference model training method according to any one of claims 1-3, characterized in that, The step of training a pre-trained language model based on the sample reasoning question, the actual reasoning path, and the reasoning judgment label to obtain a material reasoning model includes: Based on the sample reasoning problem, actual reasoning path and reasoning judgment label corresponding to each of the sample material design schemes, construct training samples corresponding to each of the sample material design schemes; Perform consistency verification on each of the training samples and obtain the target training sample that has passed the verification; The pre-trained language model is trained based on the target training samples to obtain the material reasoning model; The consistency verification includes verifying whether the sample reasoning problem in each training sample is related to the reasoning conclusion in the actual reasoning path, whether each reasoning node in the actual reasoning path is associated with a valid evidence identifier, whether there is a logical jump between the reasoning conclusions of adjacent reasoning nodes in the actual reasoning path, whether each reasoning conclusion in the actual reasoning path is consistent with the reasoning judgment label, and whether the performance value in the actual reasoning path is associated with a valid evidence identifier.
6. The inference model training method according to claim 5, characterized in that, The step of training the pre-trained language model based on the target training samples to obtain the material reasoning model includes: The sample reasoning problem in the target training sample is input into the pre-trained language model to obtain the predicted reasoning path and reasoning judgment result corresponding to the target training sample output by the pre-trained language model; The target loss value is obtained based on the inference loss value between the predicted inference path and the actual inference path in the target training sample, and the inference loss value between the inference judgment result and the inference judgment label in the target training sample. Based on the target loss value, the pre-trained language model is iteratively trained until the performance of the pre-trained language model converges and / or the maximum number of iterations is reached. The material reasoning model is obtained based on the pre-trained language model obtained from each iteration of training.
7. A recommendation method, characterized in that, include: Based on the second target material type and the second target performance parameters, obtain a set of candidate material design schemes; Based on the second target performance parameters and the design information of each candidate material design scheme in the candidate material design scheme set, generate the target reasoning problem corresponding to each candidate material design scheme; The target reasoning problem is input into the material reasoning model to obtain the predicted reasoning path and reasoning judgment result corresponding to each candidate material design scheme output by the material reasoning model. Based on the prediction reasoning path and reasoning judgment result corresponding to each candidate material design scheme, at least one material design scheme to be recommended is selected from the set of candidate material design schemes. The material reasoning model is obtained by training based on the reasoning model training method as described in any one of claims 1-6.
8. The recommended method according to claim 7, characterized in that, The step of selecting at least one recommended material design scheme from the set of candidate material design schemes based on the prediction reasoning path and reasoning judgment result corresponding to each candidate material design scheme includes: Based on the reasoning and judgment results corresponding to each candidate material design scheme, candidate material design schemes that are in the recommended state and whose reasoning confidence is higher than the confidence threshold are selected from the candidate material design scheme set to obtain the first material design scheme subset; The constraint tool component is invoked to perform constraint verification on each candidate material design scheme in the first material design scheme subset, and the constraint verification results corresponding to each candidate material design scheme in the first material design scheme subset are obtained; the verification results include the verification status identifier and verification score output by each constraint tool in the constraint tool component; Based on the verification status identifier, candidate material design schemes that pass the constraint verification are selected from the first material design scheme subset to obtain the second material design scheme subset; Based on the verification score and the reasoning result, at least one material design scheme to be recommended is selected from the second material design scheme subset.
9. The recommended method according to claim 8, characterized in that, The step of selecting at least one material design scheme to be recommended from the second material design scheme subset based on the verification score and the reasoning judgment result includes: Based on the reasoning and judgment results, obtain the reasoning score corresponding to each candidate material design scheme in the second material design scheme subset; Based on the structural description information in the design information of each candidate material design scheme in the second material design scheme subset, the second material design scheme subset is divided into multiple material design scheme groups, and the structural similarity between each candidate material design scheme in each material design scheme group is higher than the similarity threshold. Based on the verification score and the reasoning score, obtain the comprehensive score of each candidate material design scheme in each material design scheme group; Based on the comprehensive score, candidate material design schemes in each of the material design scheme groups are screened to obtain at least one material design scheme to be recommended.
10. The recommended method according to any one of claims 7-9, characterized in that, After selecting at least one material design scheme to be recommended, the method further includes: Obtain the experimental results corresponding to each of the proposed material design schemes; Based on the experimental results, material design schemes that failed in the experiment were selected as negative examples from the at least one material design scheme to be recommended. The prediction inference path, inference judgment result, and constraint verification result corresponding to the negative example sample are compared with the experimental results corresponding to the negative example sample, and the negative example type label corresponding to the negative example sample is determined based on the comparison result. Based on the negative example type label, determine the target update operation among multiple update operations, and execute the target update operation; The multiple update operations include fine-tuning the training of the material reasoning model, updating the constraint rule information in the constraint tool component, and updating the output weights of each constraint tool in the constraint tool component.
11. The recommendation method according to claim 10, characterized in that, The step of determining the target update operation from multiple update operations based on the negative example type label, and then executing the target update operation, includes: When the negative example type label belongs to the label of abnormal inference information, the experimental results corresponding to the negative example sample are compared one by one with the inference information of each inference node in the prediction inference path corresponding to the negative example sample. Based on the comparison results, the root cause node is determined in the prediction inference path corresponding to the negative example sample. Based on the experimental results corresponding to the negative example sample, the inference information corresponding to the root cause node, the inference information corresponding to the inference node after the root cause node, and the inference judgment result corresponding to the negative example sample in the prediction inference path corresponding to the negative example sample are updated to obtain the corrected inference path and the corrected inference judgment result corresponding to the negative example sample. Based on the target reasoning problem, corrected reasoning path and corrected reasoning judgment result corresponding to the negative example sample, construct the negative example training sample corresponding to the negative example sample; The material reasoning model is fine-tuned based on the negative training samples and the negative type labels.
12. The recommendation method according to any one of claims 7-9, characterized in that, The step of obtaining a set of candidate material design schemes based on the second target material type and the second target performance parameters includes: Based on the second target material type and the second target performance parameters, obtain the generation constraint information; Based on the generated constraint information, the set of candidate material design schemes is obtained by filtering from the candidate material design scheme library.
13. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the inference model training method as described in any one of claims 1 to 6; or, it implements the recommendation method as described in any one of claims 7 to 12.
14. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the inference model training method as described in any one of claims 1 to 6; or, it implements the recommendation method as described in any one of claims 7 to 12.
15. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the inference model training method as described in any one of claims 1 to 6; or, it implements the recommendation method as described in any one of claims 7 to 12.