An archer hierarchical inorganic material synthesis route prediction method
Patent Information
- Application Number
- CN202611002710.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-07
- Publication Date
- 2026-09-25
AI Technical Summary
[0010]本发明的目的是基于现有技术的现状,提供一种层次化无机材料合成路线预测方法,尤其是一种ARCHER层次化无机材料合成路线预测方法,本方法将克服现有方法将合成预测视为单一问题、无法兼顾不同决策层级信息需求的不足,同时在无需目标材料历史合成记录的条件下发现可行前驱体路线
[0029]与现有技术相比,相比于传统的数据驱动方法,本发明无需目标材料的历史合成记录即可发现可行前驱体路线,对未曾被合成报道的新材料具有泛化能力。相比于热力学稳定性计算方法,本发明不仅评估合成可行性,还给出完整的前驱体路线、工艺序列和条件参数。相比于大型语言模型直接生成方法,本发明将合成预测分解为五个决策层级,每个层级使用最适合的模型和特征,且所有条件预测满足物理合理性约束,最大程度消除了语言模型的幻觉缺点。相比于现有文本挖掘系统,本发明通过双模型裁定协议消除了单一模型的标注偏差,并提供了可量化的数据质量评估。此外,本发明通过基于裁定一致性的置信度评估机制,使预测结果附带可靠性指标,用户可据此判断预测结果的可信程度。
Smart Images

Figure CN122822167A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computational materials science and artificial intelligence, and in particular to an ARCHER hierarchical method for predicting synthetic routes of inorganic materials, including precursor prediction, operation sequence prediction, and reaction condition filling, which can be used to accelerate the experimental development of inorganic functional materials. Background Technology
[0002] Predicting how to synthesize an inorganic material is one of the most fundamental and challenging open problems in materials science. Given a target compound (target material), the experimenter needs to decide which starting reagents to use, what sequence of process steps to take, and under what conditions to synthesize it. Despite decades of empirical knowledge accumulated in the synthesis literature, this problem remains unsolved for the vast majority of known inorganic compounds. The Materials Project database contains computational properties for over 150,000 materials (APL Materials, 2013, 1(1): 011002), but only a small fraction of them have documented experimental synthesis protocols.
[0003] Existing calculation methods can be mainly classified into the following categories.
[0004] Data-driven approaches utilize large databases of known synthetic reactions to identify statistical patterns in precursor selection and processing conditions. For example, Kim et al. constructed a large-scale inorganic synthesis dataset based on text mining and analyzed precursor selection trends (Chemistry of Materials, 2017, 29(21): 9436-9444). However, such methods cannot be extrapolated to compounds lacking historical precedents, which are precisely the cases most in need of guidance.
[0005] Thermodynamic stability calculations, including phase diagram analysis and energy distance measurements (Science Advances, 2016, 2(11): e1600225), provide constraints on which phases are available under given conditions, but offer limited insight into the practical choices of precursors and processing conditions.
[0006] Large-scale language models generate synthetic descriptions from scientific literature training data (Nature Communications, 2025, 16: 6530), demonstrating the ability to generate seemingly reasonable synthetic schemes. However, this method leaves syntheticity, methods, and precursors to be predicted by independent models, lacking hierarchical cascade information transmission. Furthermore, it does not predict synthetic methods and conditions, resulting in a lack of chemical reliability in the generated suggestions. It is difficult to identify where the predictions succeed or fail, and it is also difficult to provide specific suggestions for conducting effective experiments.
[0007] Text mining systems extract structured synthesis information from papers (Scientific Data, 2019, 6:203), making it possible to analyze synthesis trends across material families. However, the extraction process faces the problem of labeling ambiguity, where different extractors may assign different method labels to the same synthesis description.
[0008] Furthermore, for non-mainstream synthesis methods (i.e. long-tail systems), existing methods have shortcomings such as complete neglect or inability to provide reliable predictions due to insufficient training data, making it impossible for users to determine whether the prediction results are applicable to their target materials. These are precisely the non-mainstream synthesis methods that need the most guidance.
[0009] Based on the current state of the technology, the inventors of this application intend to provide an ARCHER hierarchical inorganic material synthesis route prediction method, which is applicable to accelerating the experimental development of inorganic functional materials. Summary of the Invention
[0010] The purpose of this invention is to provide a hierarchical inorganic material synthesis route prediction method based on the current state of the technology, and in particular an ARCHER hierarchical inorganic material synthesis route prediction method. This method will overcome the shortcomings of existing methods that treat synthesis prediction as a single problem and cannot take into account the information needs of different decision-making levels. At the same time, it can discover feasible precursor routes without the need for historical synthesis records of the target material.
[0011] To achieve the above objectives, the technical solution of the present invention includes: An ARCHER hierarchical inorganic material synthesis route prediction method, comprising: S1: Obtain the composition or crystal structure information of the target material, generate a set of candidate precursor combinations that meet the element coverage constraints from the pre-constructed reagent library; score each candidate precursor combination in the attribute prediction embedding space, and sort the candidate routes by fusing several (at least two) complementary scoring signals, and output at least one precursor route (top-k) determined by ranking. S2: For the precursor route, a feature representation is constructed based on the target material embedding vector, the precursor embedding vector and its residual vector. The precursor route is classified as dry synthesis or wet synthesis, and the obtained synthesis method classification result is used as the scheduling signal for the generation of the process sequence. S3: Using the target material, the precursor combination in the precursor route, and historical synthesis cases retrieved from the synthesis knowledge base as input, an ordered process operation sequence is generated through a retrieval-enhanced language model. S4: For each heat treatment stage in the process sequence generated in step three, the target material, precursor combination, predicted sequence and retrieved historical cases with conditional information are used as inputs to fill in the temperature, time and atmosphere condition parameters through the language model enhanced by retrieval. S5: For the target material, based on the system category label assigned by the annotation model and the chemical characteristics of the target material, determine whether it belongs to the long-tail synthesis system; if it is determined to be a long-tail system, further output the specific category, and generate a synthesis hypothesis scheme based on the polymerization information recorded in the same system in the synthesis knowledge base; at the same time, use the consistency of the adjudication between annotation models as a confidence index, directly output high confidence predictions, and mark low confidence predictions as uncertain and prompt the user to review them.
[0012] In this invention, ARCHER is a five-module-level joint prediction framework, where ARCHER stands for Arithmetic Retrosynthesis and Cascaded Hierarchy for Experimental Routes. The core idea is to discover feasible routes in the attribute prediction embedding space using embedding arithmetic residuals or similarities with byproduct corrections, decomposing the synthetic prediction into five decision levels. Each level uses the model and features best suited to its information needs, and a reliable synthetic knowledge base is constructed through a dual-model decision protocol.
[0013] A further improvement of the present invention is that, in step S1, the process of generating a set of candidate precursor combinations that satisfy the element coverage constraint from the pre-constructed reagent library specifically includes: Compounds are selected from a pre-constructed reagent library, which is commercially available and has an energy distance from the convex hull below a preset threshold. The composition of the compounds, The value ranges from 0 to 100 millielectron volts per atom, with an optimal value of 50 millielectron volts per atom. For each required element in the target material, defined as all elements except hydrogen, carbon, and oxygen, the candidate set must contain at least one precursor covering that element; The generated candidate set is chemically rationalized to exclude candidates that violate stoichiometric constraints or contain incompatible reagent combinations. 50 to 200 candidate combinations (candidate precursor combinations) are generated for each target material.
[0014] The method for scoring each candidate precursor combination in the attribute prediction embedding space is the embedding arithmetic residual method with byproduct correction, which includes: for candidate routes It contains Precursor The target material is The set of candidate byproducts generated by the reaction is The byproducts are automatically identified from the products on the right side of the reaction equation, excluding the target material, and the embedded arithmetic residual score is defined as follows: in For the first Embedding vectors of precursors For the first Embedding vectors of each byproduct Target material The embedding vector, where the material Embedded vector This corresponds to the embedded representation of the property prediction output of the embedded model for this material. For the first The non-negative fitting coefficients corresponding to each precursor embedding vector. For the first The non-negative fitting coefficients corresponding to the embedding vectors of each by-product; the smaller the embedding arithmetic residual, the stronger the embedding compatibility between the candidate precursor route and the target material; Alternatively, the method of scoring each candidate precursor combination in the attribute prediction embedding space is a similarity scoring method, which includes: using the similarity between the route embedding obtained by the precursor embedding and the by-product embedding and the target material embedding as a score; the embedding arithmetic residual, the embedding similarity, or a combination of the two is used for candidate route ranking. The embedding vector is generated by a multi-task graphical neural network trained on at least two material property prediction tasks, such as formation energy and band gap, with an output dimension of [missing information]. The dimensions range from 64 to 256, with an optimal value of 128 dimensions. The training process for this model does not use any synthetic data.
[0015] A further improvement of the present invention is that the embedding vector is generated by a multi-task graph neural network, the input of which is a representation of the crystal structure of the material. ,in For a set of atomic nodes, The set consists of chemical bond edges, and the training tasks include at least two material property prediction tasks such as formation energy and band gap. The output dimension is... The value ranges from 64 to 256 dimensions; for compounds Its embedding vector The output of the message passing layer of this multi-task graph neural network is obtained through global pooling.
[0016] A further improvement of the present invention is that, in step S1, the selection range of the complementary scoring signal includes: Embedded arithmetic residual scoring Based on the embedding arithmetic residual with by-product correction, embedding similarity, or a combination of both, the compatibility between candidate routes and target materials is calculated in the attribute prediction embedding space. Thermodynamic stability measurement signal Its expression is: in This represents the normalized co-occurrence frequency of precursor assemblies in the reaction database. This is the negative value of the normalized energy distance of the precursor compound relative to the convex hull. This is a weighting parameter, with a value ranging from 0 to 1; Chemical family consistency signal Its expression is: in The number of precursors. Precursor Chemical family categories, 1 For indicator functions; based on the mathematical form of the formula; The chemical family categories include oxides, carbonates, nitrates, hydroxides, chlorides, sulfides, organic salts, and others; each scoring signal generates an independent ranking, which is then merged into a final ranking using the inverse ranking. in For the route In the Ranking under each signal It is a smoothing constant, with a value ranging from 0.1 to 10.
[0017] A further improvement of the present invention is that, in step S2, the method for constructing the feature representation includes: Feature vector Defined as: in Embed a vector for the target material. For the precursor to be embedded in the center of mass, For the residual vector, The number of precursors. For the number of elements, For the embedding dimension; when the residual vector The closer the precursor assembly is to the zero vector, the smaller the compositional difference between the precursor assembly and the target material. Perform binary classification using at least one classifier, such as gradient boosting decision tree ensemble, support vector machine, or neural network, and output the objective probability. and wet method probability Based on the mathematical form of the formula, it is clear that the closer the residual vector is to zero, the smaller the compositional difference between the precursor assembly and the target material.
[0018] A further improvement of the present invention is that, in step S3, the method for generating the enhanced process sequence is as follows: Independent parameter-efficient adaptation modules are maintained for dry synthesis and wet synthesis respectively; each parameter-efficient adaptation module includes at least one of low-rank adaptation, prefix fine-tuning, or cue fine-tuning. When low-rank adaptation is used, the adaptation matrix... Decomposed into: in The rank is 4, and its value ranges from 4 to 64. The hidden layer dimension of the model; the generated input consists of the target material's chemical formula, precursor combination, and... Composed of retrieved historical synthetic cases, The value ranges from 1 to 10; the historical cases are selected from the synthetic knowledge base based on semantic similarity, and the retrieval pool is deduplicated and quality filtered to exclude degenerate cases containing only a single operation.
[0019] A further improvement of the present invention is that, in step S4, the condition parameter filling method includes: using an independent parameter efficient adaptation module for condition prediction; employing a sequence-aware reordering strategy during retrieval, and candidate historical cases. Reordering weights Defined as: in For the set of operations to predict the process sequence, Candidate cases The set of operations, This is an indicator function; the output is the temperature corresponding to each heat treatment stage. ,time and atmosphere Structured mapping; physical validation of the prediction results: , The value ranges from 0 to 50 degrees Celsius. The value ranges from 1500 to 3000 degrees Celsius. The optimal temperature is 20 degrees Celsius. The optimal temperature is 2500 degrees Celsius. , The value range is from 100 to 2000 hours. The optimal time is 1000 hours.
[0020] A further improvement of the present invention is that, in step S5, the synthesized knowledge base is constructed through a bimodel adjudication protocol, the bimodel adjudication protocol comprising: Synthesis records are extracted from the literature. Each record includes the target material's chemical formula, precursor information, reaction equation, operating conditions, and a synthesis description. A multi-class classification method for synthesis systems is defined, including labels for mainstream synthesis systems and labels for long-tail synthesis systems. Each record is independently labeled with the main system label by two annotation models with different architectures or different training data. Records with consistent labels are directly accepted, while inconsistent records are resolved through a structured adjudication process, where the judgment of the second model serves as the final label and is accompanied by an approval mark. Each record is further... Each dimension of evidence quality is marked. The value ranges from 3 to 10, and the overall quality score is defined as follows: in For the first Assigning values to each dimension, The value range is from 0 to 6.
[0021] A further improvement of the present invention is that, in step S5, the classification method for synthetic systems includes mainstream synthetic system labels and long-tail synthetic system labels: The mainstream synthesis system labels include mainstream solid-phase synthesis and mainstream sol-gel synthesis; The long-tailed synthesis system labels include: flux growth, crystal growth, floating zone method, Czochralski method, melt quenching method, molten salt-assisted synthesis, molten hydroxide-assisted synthesis, high pressure or pressure-assisted synthesis, hydrothermal method, solvothermal method, coprecipitation-dominant synthesis, precipitation-dominant synthesis, Pechini method, polymer complexation route, combustion route, auto-combustion route, mechanochemical driven route, ball mill driven route, electrochemical driven synthesis, potential driven synthesis, ammonium hydrolysis, nitriding method, hydrogen-containing unconventional system, or hydride-based unconventional system, nitrogen-rich unconventional system, polynitride unconventional system, chalcogenide growth system, non-oxide growth system, and unlisted long-tailed system categories.
[0022] A further improvement of the present invention is that, in step S5, the method for generating the synthetic hypothesis scheme includes: Aggregate and determine records with the same system category from the synthetic knowledge base, and extract at least two of the following information: similar material system reference, common precursor family, common process template, typical condition range, and reported risks; The extracted information is integrated into a structured synthetic hypothesis scheme, which includes... Candidate synthetic pathways and their supporting evidence, The values range from 1 to 10; each hypothesis is labeled with its strength of evidence. and directness level .
[0023] A further improvement of the present invention is that, in step S5, the confidence index... Defined as: in and These are two annotation models for the target material. Assigned system category labels; when When the result is high confidence, it is marked as such and output directly; when When the result is determined, it is marked as low confidence, an uncertainty label is added, and the user is prompted to review it.
[0024] In this invention, the training process of the embedding model (multi-task graph neural network) does not use any synthetic data, but only uses the material property prediction task for training, thereby ensuring that the embedding space captures the intrinsic physical properties of the material rather than the statistical bias in the synthetic literature.
[0025] The route sequencing module in step S1 does not require synthesis-specific training for the target material and operates without accessing the synthesis history of the target material, thus making it applicable to new materials that have not been previously reported in synthesis.
[0026] The outputs of steps S1 to S5 form a hierarchical synthesis scheme. The route sorting result of step S1 serves as the input for the method classification in step S2. The method classification result of step S2 serves as the scheduling signal for sequence generation in step S3. The sequence generation result of step S3 serves as the structural framework for condition filling in step S4. The long-tail detection result of step S5 is presented in parallel with the outputs of the previous four steps. When step S5 determines that the target material belongs to a long-tail system, the synthesis hypothesis scheme generated in step S5 is output simultaneously with the mainstream prediction results of the previous four steps for the user to choose from. Unless otherwise specified, the long-tail detection result of step S5 is suggestive and does not replace the mainstream prediction results of steps S1 to S4.
[0027] In this invention, the research system can be either an oxide material or a non-oxide material such as a chalcogenide, nitride, or hydride. For non-oxide materials, the long-tail detection module in step five automatically identifies and labels their unconventional anionic chemical characteristics.
[0028] In this invention, all predicted conditions in step four meet the physical rationality constraints, with the predicted temperature value not lower than 20 degrees Celsius and not exceeding 2500 degrees Celsius, and the predicted time value being positive and not exceeding 1000 hours.
[0029] Compared to existing technologies and traditional data-driven methods, this invention can discover feasible precursor routes without requiring historical synthesis records of the target material, and it has generalization capabilities for novel materials that have not been previously reported in synthesis. Compared to thermodynamic stability calculation methods, this invention not only assesses synthesis feasibility but also provides complete precursor routes, process sequences, and conditional parameters. Compared to methods that directly generate large language models, this invention decomposes synthesis prediction into five decision levels, each using the most suitable model and features, and all conditional predictions satisfy physical rationality constraints, minimizing the illusionary drawbacks of language models. Compared to existing text mining systems, this invention eliminates the annotation bias of a single model through a dual-model adjudication protocol and provides quantifiable data quality assessment. Furthermore, this invention uses a confidence assessment mechanism based on adjudication consistency to attach reliability indicators to the prediction results, allowing users to judge the credibility of the prediction results. Attached Figure Description
[0030] Figure 1 : Schematic diagram of the ARCHER five-module cascaded architecture of the present invention.
[0031] Figure 2 : Schematic diagram of multi-signal fusion ablation experimental results. The horizontal axis represents configuration, and the vertical axis represents accuracy percentage.
[0032] Figure 3 : Schematic diagram of the consistency analysis results of the dual-model adjudication. The horizontal axis represents the consistency status of the labeled model, and the vertical axis represents the percentage of consistency with the adjudication label.
[0033] Figure 4 : A schematic diagram of the confusion matrix for the synthetic method. The columns represent the true category of the behavior, and the predicted category; the color intensity indicates the percentage.
[0034] Figure 5 Example 1: A schematic diagram of the precursor route fusion score for Li2AuH6. The horizontal axis represents the fusion ranking score, and the vertical axis represents the precursor combinations, arranged in descending order of score.
[0035] Figure 6 Example 2: Schematic diagram of TiTe precursor route fusion score. The horizontal axis represents the fusion ranking score, and the vertical axis represents the precursor combination, arranged in descending order of score.
[0036] Figure 7 Example 3: A schematic diagram of the precursor route fusion score for TaFeSe2. The horizontal axis represents the fusion ranking score, and the vertical axis represents the precursor combinations, arranged in descending order of score. Detailed Implementation
[0037] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments. These embodiments are based on the technical solution of the present invention and provide detailed implementation methods and specific operating procedures. However, the scope of protection of the present invention is not limited to the following embodiments.
[0038] Example 1: This embodiment performs precursor route sequencing and long-tail detection for Li2AuH6. The crystal structure information of Li2AuH6 is input using the method of this invention. Embedding Dimension Set to 128, smoothing constant. Set to 1, the weight of the thermodynamic stability metric. Set to 0.5, energy threshold Set at 50 millielectron volts per atom. The top-ranked route is LiH + Au, with a fusion ranking score. The probability of a dry process is 0.905, and the probability of a wet process is 0.095, classifying it as a dry process. The process sequence generated by the dry process branch is: mixing → calcination, with a calcination condition of 800 degrees Celsius. Long-tail detection determines that this target belongs to an unconventional hydride system, with a high warning level, and recommends that the system be classified as mechanochemical. The family tags include complex hydride-related, hydride-related, hydrogen-rich-related, and non-oxide targets. Among the global long-tail hypotheses, the top-ranked hypothesis is that the precursor is LiH+Au+H2, with a process sequence of: mixing → grinding → hydrogenation, and a moderate strength of evidence. The results are as follows. Figure 5 As shown in the example, this embodiment demonstrates that the present invention can successfully recommend feasible precursor routes with complete elemental coverage even without historical synthesis records of Li2AuH6. It also correctly identifies the unconventional hydride system properties and provides users with mechanochemical alternative synthesis hypotheses, achieving a dual-track output of mainstream prediction and long-tail supplementation.
[0039] Example 2: This embodiment performs precursor route sorting and long-tail detection for TiTe. The method of this invention is used to input the composition information of TiTe. Embedding dimension. Set to 128, smoothing constant. Set to 1. The top-ranked route is Te+Ti, with a combined ranking score. The probability of a dry process is 0.918, and the probability of a wet process is 0.082, classifying it as a dry process. The process sequence generated by the dry process branch is: mixing → calcination, with a calcination condition of 950 degrees Celsius. Long-tail detection determines that this target belongs to the chalcogenide growth system, with a high warning level, and suggests that the system be quenched. The family tags include chalcogenide-related, layered chalcogenide-related, and non-oxide targets. Among the global long-tail hypotheses, the top-ranked hypothesis is that the precursor is Te+Ti, with a process sequence of: mixing → melting → quenching, and a moderate strength of evidence. The results are as follows. Figure 6As shown in the figure. This embodiment demonstrates that the present invention can correctly identify the long-tail properties of non-oxide chalcogenide TiTe, and provides a melt-quenching alternative hypothesis while offering the mainstream solid-state solution, overcoming the deficiency of existing methods in providing guidance for non-mainstream synthesis systems.
[0040] Example 3: This embodiment performs precursor route sequencing and long-tail detection for TaFeSe2. The crystal structure information of TaFeSe2 is input using the method of this invention. Embedding Dimension Set to 128, smoothing constant. Set to 1. The top-ranked route is Fe+Se+Ta, with the combined ranking score... The probability of a dry process is 0.913, and the probability of a wet process is 0.087, classifying it as a dry process. The process sequence generated by the dry process branch is: mixing → calcination → grinding → tableting, with a calcination condition of 900 degrees Celsius. Long-tail detection determines that this target belongs to the chalcogenide growth system, with a high warning level, and recommends that the system be quenched. The family tags include chalcogenide-related, layered chalcogenide-related, non-oxide targets, and transition metal chalcogenide-related. Among the global long-tail hypotheses, the top-ranked hypothesis is that the precursor is Fe+Se+Ta, with a process sequence of: mixing → melting → quenching, and a moderate strength of evidence. The results are as follows. Figure 7 As shown in the figure. This embodiment demonstrates that the present invention can simultaneously recommend precursor routes with complete elemental coverage for ternary non-oxide TaFeSe2 and generate process sequences containing multiple heat treatment stages. The long-tail detection module accurately labels multiple material family tags, providing experimenters with a wealth of synthetic strategy references.
[0041] Example 4: like Figure 1 As shown, this embodiment performs an overall performance evaluation. Using the method of this invention, 63,507 route matching records were evaluated and divided into training, validation, and test sets in an 80 / 10 / 10 ratio. The results of the multi-signal fusion ablation experiment are as follows: Figure 2 As shown, the baseline first-order accuracy was 19.6%, which improved to 24.3% after adding a thermodynamic stability metric, 37.2% after adding chemical family consistency, and reached 49.0% after adding an embedded arithmetic residual score; the top-ten accuracy improved from 80.0% at baseline to 96.8%. The confusion matrix for the synthesis method classification is shown below. Figure 4 As shown, the accuracy of dry-to-dry method analysis is 96.6%, and the accuracy of wet-to-wet method analysis is 96.6%. The consistency analysis of the two models is as follows. Figure 3As shown, when the two annotation models are consistent, the consistency rate between the adjudicated label and the consensus is 98.6%; when the two annotation models are inconsistent, the consistency rate between the adjudicated label and the predicted category is only 4.3%. This embodiment demonstrates that the present invention improves the top ten accuracy to 96.8% through multi-signal fusion ranking, an improvement of 16.8 percentage points compared to the baseline; the dual-model adjudication protocol enables the label credibility of consistent samples to reach 98.6%, providing a high-quality data foundation for the synthetic knowledge base; the classification accuracy of the synthetic method reaches 96.6%, verifying the effectiveness of the hierarchical cascade framework.
[0042] Example 5: This embodiment performs independent verification of the embedding arithmetic residuals. To verify that the embedding arithmetic residuals in step S1 do not solely rely on precursor route ordination data, this invention further employs Materials Project decomposition relationships for matching decoy verification. This verification does not use target material-precursor route labels, but instead calculates the residual between the weighted sum of reactant embeddings and their decomposition product embeddings, and compares it with the matching decoy product set. Under the coverage matching control condition, a total of 100,524 valid decomposition relationships satisfy the sampling conditions. The results show that the embedding arithmetic residuals of the real decomposition products are systematically lower than those covering the matching decoy product set, with an AUC of 0.7239 distinguishing between real products and decoy products. The median percentile of the real residuals in the matching decoy distribution is 0.14, and in 76.46% of the reactions, the real residuals are lower than the median decoy residuals. This embodiment demonstrates that the embedding arithmetic residuals can capture non-random chemical relationships and can serve as an independent characterization layer compatibility signal for precursor route ordination.
[0043] Example 6: This embodiment verifies the effectiveness of the retrieval enhancement module. To verify the contribution of retrieval enhancement in steps S3 and S4 to process sequence generation and condition filling, this invention compares settings with and without retrieval. In solid-phase process sequence generation, after removing the retrieval, the exact matching rate decreased from 32.82% to 15.24%, stage F1 decreased from 77.85% to 58.43%, and the normalized edit distance increased from 0.316 to 0.512. In calcination condition prediction, after removing the retrieval, the average absolute temperature error increased from 126.3 degrees Celsius to 174.2 degrees Celsius, the relative time error increased from 0.834 to 1.256, and the atmosphere prediction accuracy decreased from 42.4% to 3.1%.
[0044] Further validation was performed using document time-division: training and retrieval used records from 2014 and earlier, validation used records from 2015 to 2016, and testing used records from 2017 and later. The results showed that under this more stringent setting, the stage F1 score of the solid-phase route time-division retrieval enhancement setting was 67.98%, higher than the 55.27% without the retrieval setting; the stage F1 score of the sol-gel route time-division retrieval enhancement setting was 67.18%, higher than the 54.94% without the retrieval setting. This embodiment demonstrates that the retrieval enhancement module can provide transferable synthetic precedent information, rather than relying solely on nearest-neighbor overlap in random partitioning.
[0045] The preferred embodiments of the present invention have been described in detail above. It should be understood that those skilled in the art can make numerous modifications and variations based on the concept of the present invention without creative effort. Therefore, all technical solutions that can be obtained by those skilled in the art based on the concept of the present invention through logical analysis, reasoning, or limited experimentation on the basis of existing technology should be within the scope of protection defined by the claims.
Claims
1. A method for predicting the synthesis route of ARCHER hierarchical inorganic materials, characterized in that... include: S1: Obtain the composition or crystal structure information of the target material, generate a set of candidate precursor combinations that meet the element coverage constraints from the pre-constructed reagent library; score each candidate precursor combination in the attribute prediction embedding space, and sort the candidate routes by fusing several complementary scoring signals, and output at least one precursor route determined by the ranking. S2: For the precursor route, a feature representation is constructed based on the target material embedding vector, the precursor embedding vector and its residual vector. The precursor route is classified as dry synthesis or wet synthesis, and the obtained synthesis method classification result is used as the scheduling signal for the generation of the process sequence. S3: Using the target material, the precursor combination in the precursor route, and historical synthesis cases retrieved from the synthesis knowledge base as input, an ordered process operation sequence is generated through a retrieval-enhanced language model. S4: For each heat treatment stage in the process sequence generated in step three, the target material, precursor combination, predicted sequence and retrieved historical cases with conditional information are used as inputs to fill in the temperature, time and atmosphere condition parameters through the language model enhanced by retrieval. S5: For the target material, based on the system category label assigned by the annotation model and the chemical characteristics of the target material, determine whether it belongs to the long-tail synthesis system; if it is determined to be a long-tail system, further output the specific category, and generate a synthesis hypothesis scheme based on the polymerization information recorded in the same system in the synthesis knowledge base; at the same time, use the consistency of the adjudication between annotation models as a confidence index, directly output high confidence predictions, and mark low confidence predictions as uncertain and prompt the user to review them.
2. The ARCHER hierarchical inorganic material synthesis route prediction method according to claim 1, characterized in that, Step S1, the process of generating a set of candidate precursor combinations that satisfy the element coverage constraint from the pre-constructed reagent library, specifically includes: Compounds are selected from a pre-constructed reagent library, which is commercially available and has an energy distance from the convex hull below a preset threshold. The composition of the compounds, The value ranges from 0 to 100 millielectronvolts per atom; For each required element in the target material, defined as all elements except hydrogen, carbon, and oxygen, the candidate set must contain at least one precursor covering that element; The generated candidate set is chemically rationalized to exclude candidates that violate stoichiometric constraints or contain incompatible reagent combinations, generating 50 to 200 candidate combinations for each target material.
3. The ARCHER hierarchical inorganic material synthesis route prediction method according to claim 1, characterized in that, In step S1: The method for scoring each candidate precursor combination in the attribute prediction embedding space is the embedding arithmetic residual method with byproduct correction, which includes: for candidate routes It contains Precursor The target material is The set of candidate byproducts generated by the reaction is The byproducts are automatically identified from the products on the right side of the reaction equation, excluding the target material; (Note: The original text contains some inconsistencies and unclear grammatical structure. A more accurate translation would require the full context.) Target material Embedded vector, Let j be the embedding vector of the j-th precursor. For the first Embedding vectors of several byproducts, where the material Embedded vector The embedded representation of the property prediction output of the embedded model for this material is defined as follows: The embedded arithmetic residual score is defined as: ; in, For the first The non-negative fitting coefficients corresponding to each precursor embedding vector. For the first The non-negative fitting coefficients corresponding to the embedding vectors of each by-product; Alternatively, the method of scoring each candidate precursor combination in the attribute prediction embedding space is a similarity scoring method, which includes: using the similarity between the route embedding obtained by the precursor embedding and the by-product embedding and the target material embedding as a score; the embedding arithmetic residual, the embedding similarity, or a combination of the two is used for candidate route ranking.
4. The ARCHER hierarchical inorganic material synthesis route prediction method according to claim 3, characterized in that, The embedding vector is generated by a multi-task graph neural network, the input of which is a representation of the crystal structure of the material. ,in For a set of atomic nodes, The set consists of chemical bond edges, and the training task includes at least two material property prediction tasks such as formation energy and band gap. The output dimension is... The value ranges from 64 to 256 dimensions; for compounds Its embedding vector The output of the message passing layer of this multi-task graph neural network is obtained through global pooling.
5. The ARCHER hierarchical inorganic material synthesis route prediction method according to claim 3, characterized in that, In step S1, the selection range of complementary scoring signals includes: Embedded arithmetic residual scoring Based on the embedding arithmetic residual with by-product correction, embedding similarity, or a combination of both, the compatibility between candidate routes and target materials is calculated in the attribute prediction embedding space. Thermodynamic stability measurement signal Its expression is: ; in This represents the normalized co-occurrence frequency of precursor assemblies in the reaction database. This is the negative value of the normalized energy distance of the precursor compound relative to the convex hull. This is a weighting parameter, with a value ranging from 0 to 1; Chemical family consistency signal Its expression is: ; in The number of precursors. Precursor Chemical family categories, 1 For indicator functions; based on the mathematical form of the formula; The chemical family categories include oxides, carbonates, nitrates, hydroxides, chlorides, sulfides, organic salts, and others; each scoring signal generates an independent ranking, which is then merged into a final ranking using the inverse ranking. ; in For the route In the Ranking under each signal It is a smoothing constant, with a value ranging from 0.1 to 10.
6. The ARCHER hierarchical inorganic material synthesis route prediction method according to claim 1, characterized in that, In step S2, the method for constructing the feature representation includes: Feature vector Defined as: ; in Embed a vector for the target material. For the precursor to be embedded in the center of mass, For the residual vector, The number of precursors. For the number of elements, For the embedding dimension; when the residual vector The closer the precursor assembly is to the zero vector, the smaller the compositional difference between the precursor assembly and the target material. Perform binary classification using at least one classifier, such as gradient boosting decision tree ensemble, support vector machine, or neural network, and output the objective probability. and wet method probability .
7. The method for predicting the synthesis route of ARCHER hierarchical inorganic materials according to claim 1, characterized in that, In step S3, the method for generating the enhanced process sequence is as follows: Independent parameter-efficient adaptation modules are maintained for dry synthesis and wet synthesis respectively; each parameter-efficient adaptation module includes at least one of low-rank adaptation, prefix fine-tuning, or cue fine-tuning. When low-rank adaptation is used, the adaptation matrix... Decomposed into: ; in The rank is 4, and its value ranges from 4 to 64. The hidden layer dimension of the model; The input for generation consists of the target material's chemical formula, precursor combination, and Composed of retrieved historical synthetic cases, The value ranges from 1 to 10; the historical cases are selected from the synthetic knowledge base based on semantic similarity, and the retrieval pool is deduplicated and quality filtered to exclude degenerate cases containing only a single operation.
8. The method for predicting the synthesis route of ARCHER hierarchical inorganic materials according to claim 7, characterized in that, In step S4, the condition parameter filling method includes: using an independent parameter efficient adaptation module for condition prediction; employing a sequence-aware reordering strategy during retrieval, and candidate historical cases. Reordering weights Defined as: ; in For the set of operations to predict the process sequence, Candidate cases The set of operations, This is an indicator function; the output is the temperature corresponding to each heat treatment stage. ,time and atmosphere Structured mapping; physical validation of the prediction results: , The value ranges from 0 to 50 degrees Celsius. The value ranges from 1500 to 3000 degrees Celsius; , The value ranges from 100 to 2000 hours.
9. The ARCHER hierarchical inorganic material synthesis route prediction method according to claim 1, characterized in that, In step S5, the synthetic knowledge base is constructed using a bi-model adjudication protocol, which includes: Synthesis records are extracted from the literature. Each record includes the target material's chemical formula, precursor information, reaction equation, operating conditions, and a synthesis description. A multi-class classification method for synthesis systems is defined, including labels for mainstream synthesis systems and labels for long-tail synthesis systems. Each record is independently labeled with the main system label by two annotation models with different architectures or different training data. Records with consistent labels are directly accepted, while inconsistent records are resolved through a structured adjudication process, where the judgment of the second model serves as the final label and is accompanied by an approval mark. Each record is further... Each dimension of evidence quality is marked. The value ranges from 3 to 10, and the overall quality score is defined as follows: ; in For the first Assigning values to each dimension, The value range is from 0 to 6.
10. The ARCHER hierarchical inorganic material synthesis route prediction method according to claim 9, characterized in that, In step S5, the classification method for synthetic systems includes mainstream synthetic system labels and long-tail synthetic system labels; The mainstream synthesis system labels include mainstream solid-phase synthesis and mainstream sol-gel synthesis; The long-tailed synthesis system labels include: flux growth, crystal growth, floating zone method, Czochralski method, melt quenching method, molten salt-assisted synthesis, molten hydroxide-assisted synthesis, high pressure or pressure-assisted synthesis, hydrothermal method, solvothermal method, coprecipitation-dominant synthesis, precipitation-dominant synthesis, Pechini method, polymer complexation route, combustion route, auto-combustion route, mechanochemical driven route, ball mill driven route, electrochemical driven synthesis, potential driven synthesis, ammonium hydrolysis, nitriding method, hydrogen-containing unconventional system, hydride-based unconventional system, nitrogen-rich unconventional system, polynitride unconventional system, chalcogenide growth system, non-oxide growth system, and unlisted long-tailed system categories.
11. The ARCHER hierarchical inorganic material synthesis route prediction method according to claim 1, characterized in that, In step S5, the method for generating the synthetic hypothesis scheme includes: Aggregate and determine records with the same system category from the synthetic knowledge base, and extract at least two of the following information: similar material system reference, common precursor family, common process template, typical condition range, and reported risks; The extracted information is integrated into a structured synthetic hypothesis scheme, which includes... Candidate synthetic pathways and their supporting evidence, The values range from 1 to 10; each hypothesis is labeled with its strength of evidence. and directness level .
12. The ARCHER hierarchical inorganic material synthesis route prediction method according to claim 1, characterized in that, In step S5, the confidence index Defined as: ; in and These are two annotation models for the target material. Assigned system category labels; when When the result is high confidence, it is marked as such and output directly; when When the result is determined, it is marked as low confidence, an uncertainty label is added, and the user is prompted to review it.