A method for recommending a material synthesis process in the chemical industry
Through multi-step iterative semantic parsing and large-model reasoning, combined with chemical engineering knowledge and database retrieval, a systematic guide for the synthesis of chemical materials is generated, solving the problem of the lack of systematic guidance for chemical material synthesis schemes and realizing efficient and accurate synthesis process recommendations.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-15
- Publication Date
- 2026-03-27
AI Technical Summary
The lack of systematic guidance in the synthesis schemes of chemical materials in existing technologies leads to prolonged research and development cycles and waste of resources, and the reliance on experience and manual operation results in low efficiency.
By acquiring textual information on the synthesis requirements of target chemical materials and performing multi-step iterative semantic parsing, combined with a pre-built professional model for recommending synthesis processes, intelligent reasoning is performed using chemical common sense and reaction laws, database retrieval and reasoning expansion are performed by combining element requirements, material types and synthesis generation constraints, and finally the results are fused to generate accurate and comprehensive synthesis process recommendation results.
It improves the automation level and recommendation efficiency of chemical material synthesis, ensures the accuracy and innovation of recommendation results, optimizes the synthesis process, and reduces costs and time consumption.
Smart Images

Figure CN121054158B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of knowledge graph, in particular to a method for recommending material synthesis process in chemical industry. BACKGROUND
[0002] In the synthesis process of chemical materials, it still mainly depends on the experience accumulation and manual operation of researchers at present, and the formulation of material synthesis scheme often lacks systematic guidance, and a large number of experiments and repeated attempts are needed to gradually explore suitable reaction conditions and process paths, which not only greatly prolongs the research and development cycle, but also increases the experimental cost and human input. The synthesis process of chemical materials is complex, involving various reactants, catalysts and environmental factors, and slight carelessness may lead to failure or failure to obtain ideal performance. Therefore, how to improve the efficiency and accuracy of the synthesis process has become an important direction of research in this field. SUMMARY
[0003] The present application provides a method for recommending material synthesis process in chemical industry, aiming to solve the technical problem that the existing technology relies on experience accumulation and manual operation, resulting in lack of systematic guidance in the formulation of chemical material synthesis scheme, thereby greatly prolonging the research and development cycle and causing resource waste.
[0004] The method for recommending material synthesis process in chemical industry disclosed in the present application comprises: obtaining target chemical material synthesis demand text information for multi-step iterative semantic analysis to determine element demand information, material type demand information and synthesis generation restriction demand information; pre-building a synthesis process recommendation professional large model; combining the element demand information, material type demand information and synthesis generation restriction demand information to retrieve and screen in a chemical synthesis database to obtain a first retrieval and screening result, using the synthesis process recommendation professional large model to deeply screen the first retrieval and screening result to determine a first knowledge recommendation result; using the synthesis process recommendation professional large model to reason and expand chemical materials with similar synthesis processes for the first retrieval and screening result to determine a second knowledge recommendation result; using the synthesis process recommendation professional large model to reason the synthesis process path for the element demand information, material type demand information and synthesis generation restriction demand information to determine a third knowledge recommendation result; and fusing the first knowledge recommendation result, the second knowledge recommendation result and the third knowledge recommendation result to obtain a target knowledge recommendation result.
[0005] One or more technical solutions provided in the present application have at least the following beneficial effects:
[0006] By acquiring target chemical material synthesis demand text information for multi-step iterative semantic analysis, the user's synthesis demand can be automatically extracted and understood, reducing the dependence on manual operation, improving the automation degree, and thus improving the recommendation efficiency; by pre-building a synthesis process recommendation professional model, not only relying on traditional database query, but also based on the chemical knowledge and reaction rules learned in the model training process, intelligent reasoning is carried out to generate innovative synthesis paths, so that the recommendation system can more flexibly and accurately adapt to the diversified needs of users; combined with element demand, material type and synthesis generation restriction information, screening is carried out in the chemical synthesis database, which can effectively narrow the recommendation range and focus only on materials and synthesis paths that meet the user's demand, this process optimizes data retrieval and improves the relevance and pertinence of the results, avoiding irrelevant or inadmissible material and path recommendations; through the reasoning expansion of the first retrieval screening result by the large model, other synthesis paths and chemical materials similar to or potentially similar to the existing synthesis paths can be inferred, which improves the comprehensiveness of the recommendation and identifies excellent candidate materials or synthesis processes that may not be directly retrieved, increasing more selectivity; using the large model to infer the synthesis process path based on element demand information, material type demand information and synthesis generation restriction information can further optimize the synthesis process flow, ensuring that the recommended path not only meets the basic requirements, but also provides more optimal process parameters and conditions in actual operation, which helps users optimize the existing process flow, reduce costs and improve efficiency; by fusing the first knowledge recommendation result, the second knowledge recommendation result and the third knowledge recommendation result, the information from different recommendation sources can be integrated to eliminate the possible bias of single recommendation method, so as to obtain more accurate and comprehensive target knowledge recommendation result, and such fusion ensures that the recommendation system considers multiple dimensions, so that the final result is improved in accuracy, practicality and innovation.
[0007] The above description is only a summary of the technical solutions of the present application. In order to more clearly understand the technical means of the present application, the present application can be implemented according to the content of the specification, and in order to make the above and other purposes, characteristics and advantages of the present application more obvious and easy to understand, the following specific embodiments of the present application are described. BRIEF DESCRIPTION OF DRAWINGS
[0008] Figure 1 A method flowchart for recommending a material synthesis process in the chemical industry is provided for the embodiments of the present application.
[0009] Figure 2 A flowchart for determining the first knowledge recommendation result in the method for recommending a material synthesis process in the chemical industry is provided for the embodiments of the present application.
[0010] Figure 3An exemplary chemical material synthesis mode recommendation flowchart is provided in a method for recommending a material synthesis process in the chemical industry. DETAILED DESCRIPTION
[0011] The embodiments of the present application provide a method for recommending a material synthesis process in the chemical industry, which solves the technical problem that the prior art relies on experience accumulation and manual operation, resulting in a lack of systematic guidance for formulating a chemical material synthesis scheme, which further leads to a significant extension of the research and development period and causes resource waste.
[0012] After introducing the basic principles of the present application, various non-limiting embodiments of the present application will be specifically introduced in combination with the drawings of the specification.
[0013] As shown in the drawings, Figure 1 The embodiments of the present application provide a method for recommending a material synthesis process in the chemical industry, which comprises:
[0014] Obtaining target chemical material synthesis requirement text information for multi-step iterative semantic analysis to determine element requirement information, material type requirement information, and synthesis generation restriction requirement information.
[0015] Further, obtaining target chemical material synthesis requirement text information for multi-step iterative semantic analysis to determine element requirement information, material type requirement information, and synthesis generation restriction requirement information comprises:
[0016] The target chemical material synthesis requirement text information is divided into a semantic sub-information sequence; the semantic sub-information sequence is taken as an explicit information sequence, and a chemical common sense knowledge graph is obtained as implicit information; a semantic analysis unit is initialized, wherein the semantic analysis unit stores a preset synthesis requirement keyword set; based on the semantic analysis unit, the explicit information sequence and the implicit information are subjected to multi-step iterative semantic analysis to determine element requirement information, material type requirement information, and synthesis generation restriction requirement information.
[0017] Receiving target chemical material synthesis requirement text information input by a user, which is usually in natural language form and describes the user's requirements, for example, "I want to synthesize a two-dimensional material containing zinc, and it is best to be carried out under mild conditions using a green solvent". The target chemical material synthesis requirement text information is segmented and analyzed, and is divided into a semantic sub-information sequence, each semantic sub-information representing a specific element or requirement in the text.
[0018] The divided semantic sub-information sequence is taken as an explicit information sequence. The explicit information is the content directly expressed in the user text. They are clear and explicit, and usually contain explicit requirements. For example, the original text is disassembled into sentences or phrases such as “containing zinc”, “two-dimensional material”, “mild condition”, “green solvent” and the like as the explicit information sequence. The chemical industry common sense knowledge graph is obtained as implicit information. The implicit information refers to the information that does not directly appear in the user text in the semantic analysis process, but is deduced through common sense or knowledge base. Here, the implicit information is derived from the chemical industry common sense knowledge graph, that is, a knowledge base containing chemical reactions, material properties, common synthesis methods, chemical rules and the like. By accessing the chemical industry common sense knowledge graph, the information in the user requirement is supplemented and enriched, for example, the green solvent includes water, ethanol and the like.
[0019] The semantic analysis unit is one of the core modules of the entire system, which is responsible for analyzing the requirement text input by the user and identifying the specific synthesis requirement. The semantic analysis unit stores a preset synthesis requirement keyword set, which includes common chemical synthesis related terms, elements, material types, synthesis methods, reaction conditions and the like, which are used to identify the key elements in the text during the analysis process.
[0020] Based on the initialized semantic analysis unit, the combined information of the explicit information sequence (text extracted from the user requirement) and the implicit information (from the chemical industry common sense knowledge graph) is subjected to multi-step iterative semantic analysis. Through repeated interaction process, the corresponding key information is gradually extracted, and each iteration can more carefully identify and separate different requirements in the text. After completing the multi-step iteration, the element requirement information, material type requirement information and synthesis generation restriction requirement information of the target chemical material are clearly extracted, providing comprehensive input for the subsequent synthesis process recommendation process.
[0021] Further, based on the semantic analysis unit, the explicit information sequence and the implicit information are subjected to multi-step iterative semantic analysis to determine the element requirement information, material type requirement information and synthesis generation restriction requirement information, which includes:
[0022] The explicit information sequence is respectively subjected to word vector conversion representation to obtain an explicit information word vector sequence. The first explicit information word vector in the explicit information word vector sequence is extracted, which is sequentially subjected to first-order interaction and second-order interaction with the preset synthesis requirement keyword set and the implicit information to obtain a first iteration semantic analysis unit. According to the order from front to back, the explicit information word vector sequence is subjected to multi-step iterative semantic analysis in combination with the first iteration semantic analysis unit and the implicit information to determine a target iteration semantic analysis unit. The target iteration semantic analysis unit is integrated according to the semantic type to determine the element requirement information, material type requirement information and synthesis generation restriction requirement information.
[0023] To enable the computer to understand the explicit information sequence, each word or phrase of the explicit information sequence is converted into a word vector, which is a representation method that converts words into high-dimensional vectors. Usually, a pre-trained language model such as Word2Vec, GloVe, BERT, etc. is used to obtain the vector representation of each word or phrase. Through word vector conversion, an explicit information word vector sequence is obtained, and each explicit information element is represented as a vector, providing a basis for the subsequent semantic analysis process.
[0024] In the explicit information word vector sequence, the first explicit information word vector is extracted, representing the first key information in the input text. The extracted first explicit information word vector is interacted with a preset synthetic demand keyword set in the first order, which is a pre-stored word set containing all keywords related to chemical material synthesis, such as common elements (hydrogen, oxygen, carbon, etc.), material types (oxides, metals, catalysts, etc.), and synthesis methods. First-order interaction refers to determining the matching degree of the first explicit information word vector and the relevant words in the keyword set by calculating the similarity between the word vectors, for example, if the first explicit information word vector represents "zinc", it is matched with "zinc" in the preset synthetic demand keyword set, and the relevance is determined by similarity measurement. After the first-order interaction, the first explicit information word vector is further interacted with implicit information (from the chemical common sense knowledge graph) in the second order, and the second-order interaction involves more complex semantic reasoning, which can enhance the understanding and reasoning of explicit information by comparing relevant knowledge in implicit information, such as "common conditions of zinc oxidation reaction". According to the results of the first-order and second-order interactions, the first iteration semantic analysis unit is generated, which contains the comprehensive information obtained from the explicit information word vector, the keyword set and the interaction of implicit information.
[0025] According to the order of the input text, the explicit information word vector sequence is parsed step by step, considering the previous parsing results at each step to ensure the coherence and consistency of the information. In the parsing process of each step, the first iteration semantic analysis unit is combined with the subsequent explicit information word vector and implicit information, so that each subsequent explicit information word vector not only depends on the position in the text, but also considers the influence of the previous information and implicit information. For example, if the first word vector represents "zinc" and the next word vector represents "oxidation", combining the relevant knowledge of "zinc" obtained before, it can be inferred that "zinc oxide" may be the target synthesis material of the user. The entire explicit information word vector sequence is subjected to multi-step iteration semantic analysis, i.e. updating the current analysis unit in each iteration step, gradually improving the understanding of element demand, material type demand, synthesis restriction, etc. At the end of the iteration process, the final target iteration semantic analysis unit is generated, which integrates explicit information and implicit information, and accurately reflects the user's demand.
[0026] According to different semantic types, such as element demand, material type demand, synthesis restriction demand, etc., the target iterative semantic parsing unit is integrated, and this integration process will ensure that different types of information are correctly classified and extracted, for example, all information related to elements (such as "zinc") is classified as element demand, and information related to material type (such as "oxide") and reaction conditions (such as "300℃ reaction temperature") is classified separately, and finally element demand information, material type demand information and synthesis generation restriction demand information are determined.
[0027] Further, the first explicit information word vector in the explicit information word vector sequence is extracted, and first-order interaction and second-order interaction are performed with the preset synthesis demand keyword set and the implicit information in turn to obtain a first iterative semantic parsing unit, including:
[0028] The first explicit information word vector and the preset synthesis demand keyword set are interacted in first order by using the cosine similarity formula, and a first explicit information keyword vector is extracted; the implicit information is interacted in second order based on the first explicit information keyword vector, and a first implicit information keyword sentence vector is extracted; the first explicit information keyword vector and the first implicit information keyword sentence vector are added to the semantic parsing unit to obtain the first iterative semantic parsing unit.
[0029] Cosine similarity is a commonly used method to measure the similarity between two vectors. By calculating the cosine similarity between the first explicit information word vector and each word vector in the preset synthesis demand keyword set, the similarity between them is determined. According to the calculated cosine similarity, the most relevant keyword vector is selected, i.e. those with higher similarity. The selected keyword vector is the first explicit information keyword vector, which represents the synthesis keyword that best matches the user's demand.
[0030] The sentence containing the matching keyword in the implicit information is taken as a sentence vector, and the first explicit information keyword vector is matched with the sentence vector. By calculating the similarity between the first explicit information keyword vector and the sentence vector in the implicit information, the sentence most relevant to the keyword is selected to obtain the first implicit information keyword sentence vector. This vector represents the implicit information related to the user's demand.
[0031] The first explicit information keyword vector and the first implicit information keyword sentence vector are integrated into the semantic parsing unit, which is responsible for managing and processing these vectors for subsequent parsing. The obtained first iterative semantic parsing unit provides a preliminary basis for subsequent iterative parsing, which combines key information from user input and implicit knowledge, and provides a core basis for more detailed semantic understanding and recommendation.
[0032] Further, the target chemical material synthesis requirement text information is divided into semantic sub-information sequences, including:
[0033] The target chemical material synthesis requirement text information is subjected to punctuation standardization and stop word removal to obtain preprocessed target chemical material synthesis requirement text information; the preprocessed target chemical material synthesis requirement text information is subjected to step-by-step processing to obtain the semantic sub-information sequences.
[0034] Punctuation standardization refers to unified format processing of punctuation marks in the target chemical material synthesis requirement text information. In different texts, punctuation marks can appear in different forms or be repeatedly used. The punctuation marks are unified into a standard format for subsequent text processing. Stop words refer to words with high frequency of occurrence in the corpus but small contribution to understanding the semantic of sentences. In the stop word removal process, a stop word list is maintained to automatically remove these words in the text, thereby reducing the redundant part of the text and improving the efficiency and accuracy of subsequent processing. After completing punctuation standardization and stop word removal, a cleaner and standardized text version, referred to as preprocessed target chemical material synthesis requirement text information, is obtained.
[0035] Step-by-step processing refers to decomposing the preprocessed target chemical material synthesis requirement text information into multiple small and meaningful parts, each part representing a core information in the text. The semantic sub-information sequences are the results of step-by-step processing of the text, containing different semantic sub-information extracted from the preprocessed target chemical material synthesis requirement text information, and each semantic sub-information represents a specific requirement or description in the text.
[0036] A pre-constructed synthesis process recommendation professional large model is constructed.
[0037] Further, the pre-constructed synthesis process recommendation professional large model includes:
[0038] A data source including at least chemical material name, physicochemical property, and synthesis process text is collected to construct a basic data source set; a text vector of the synthesis process text is generated from the basic data source set using a qwen text-embedding-v3 model, and a material pair set is determined by material pair screening according to similarity Top-k; a teacher model is used to infer the synthesis path similarity of the material pair set according to the physicochemical property in the basic data source set, and the inference text is reviewed to determine a target inference corpus; the target inference corpus is encapsulated, and the basic model is fine-tuned using the encapsulated target inference corpus to construct the synthesis process recommendation professional large model.
[0039] The chemical material name includes the name of chemical elements, compounds and new materials, such as zinc oxide, aluminum sulfate, etc.; the physicochemical property is the physical and chemical characteristics of each chemical material, such as melting point, boiling point, solubility, density, chemical reactivity, etc., which are crucial for selecting the synthesis path; the synthesis process text is the text information describing the synthesis process of the chemical material, including reactants, catalysts, reaction conditions (such as temperature, pressure, etc.) and reaction steps, etc. The collected data is integrated into a basic data source set as the basis for subsequent model training.
[0040] The synthesis process text in the basic data source set is vectorized using the text-embedding-v3 model of qwen, which can convert text information into numerical vectors, enabling computers to understand and process these text data. The vectorized synthesis process text can preserve the semantic information in the text, facilitating similarity calculation and subsequent reasoning.
[0041] Based on vectorization, the similarity between synthesis process text vectors is calculated, and according to the similarity Top-k, those material pairs with similar synthesis paths are selected, where similarity Top-k refers to selecting the top k material pairs with the highest similarity to the target material in the vector space from all materials. For example, if k is set to 5, it means selecting the top 5 material pairs with the highest similarity. These material pairs refer to two chemical materials that have similarities in synthesis path, reaction conditions, etc., and may produce similar synthesis results. The pre-set similarity threshold ensures that only similar material pairs are selected. The selected material pair set contains material pairs with high similarity in synthesis process.
[0042] A teacher model, such as a large pre-trained language model, is used to perform synthesis path similarity reasoning on each pair of materials in the material pair set. The reasoning process includes: based on the physicochemical properties of each chemical material, the model infers the synthesis path they may take. Physicochemical properties provide the behavior of materials under certain reaction conditions, thus helping the model to infer possible synthesis paths. The goal of reasoning is to find out which synthesis paths are similar for these pairs of chemical materials, i.e., they can be synthesized in a similar way.
[0043] During the reasoning process, the reasoning text is reviewed to ensure that the reasoning results conform to chemical common sense and reaction rules. The review includes verifying whether the reactants in the reasoning results are correct, whether the reaction steps are reasonable, and whether the synthesis conditions are consistent with reality, etc. After review, the reasoning text that does not conform to reality is removed to ensure the quality of the target reasoning corpus, which contains the synthesis paths derived based on physicochemical properties and is the core data for subsequent fine-tuning.
[0044] The target reasoning corpus is encapsulated, which means converting the target reasoning corpus into a format that is easy for the model to process, such as a JSON structure, for subsequent use. The base model, which is a large pre-trained language model such as Qwen3-8B, is fine-tuned using the encapsulated target reasoning corpus. The purpose of fine-tuning is to enable the base model to automatically infer the corresponding synthesis process according to the input chemical material name and physical and chemical properties. During the fine-tuning process, the model will learn how to generate accurate synthesis paths based on existing synthesis processes and material characteristics. After fine-tuning, the base model becomes a synthesis process recommendation professional large model, which can intelligently recommend appropriate chemical synthesis paths according to user input.
[0045] Further, the base model is Qwen3-8B, which is fine-tuned by using the llama-factory training framework combined with the encapsulated target reasoning corpus to build the synthesis process recommendation professional large model.
[0046] The base model is Qwen3-8B, which is a powerful pre-trained language model with a large number of parameters, for example, 8B means the model has 80 billion parameters, which can handle complex text generation and reasoning tasks. It is based on the Transformer architecture and is suitable for various natural language processing tasks. The llama-factory training framework is an open-source deep learning training framework specifically designed for efficient distributed training, which can support large-scale model training and fine-tuning. In this step, llama-factory is used to fine-tune the Qwen3-8B model.
[0047] Through the llama-factory training framework, Qwen3-8B can be fine-tuned with the target reasoning corpus, making Qwen3-8B more focused on the task of recommending chemical material synthesis processes. The fine-tuning process includes loading the Qwen3-8B base model, inputting the encapsulated target reasoning corpus into the llama-factory training framework, performing batch processing of training data, and updating the parameters of Qwen3-8B through the backpropagation algorithm, so that Qwen3-8B can perform better in specific chemical synthesis tasks. With the support of a large amount of chemical synthesis path data, Qwen3-8B gradually learns how to generate accurate synthesis process paths based on material names, physical and chemical properties, etc. After fine-tuning is completed, Qwen3-8B is converted into a synthesis process recommendation professional large model, which is specifically used for the task of recommending chemical material synthesis processes.
[0048] Further, the target reasoning corpus is encapsulated into a JSON structure in fine-tuning format.
[0049] In the fine-tuning process, the target reasoning corpus needs to be input into the model in a standardized format. To facilitate model processing, the target reasoning corpus is packaged into a fine-tuning format JSON structure. JSON structure is a commonly used format, which is easy to read and parse, and has flexible structure, suitable for storing complex nested data. This packaging process ensures that the data can be correctly parsed by the Qwen3-8B model, and provides standardized input for model fine-tuning.
[0050] In combination with the element requirement information, material type requirement information and synthesis generation restriction requirement information, the first retrieval and screening result is obtained by searching and screening in the chemical synthesis database. The first knowledge recommendation result is determined by deep screening of the first retrieval and screening result using the synthesis process recommendation professional large model.
[0051] Further, as shown in Figure 2 In combination with the element requirement information, material type requirement information and synthesis generation restriction requirement information, the first retrieval and screening result is obtained by searching and screening in the chemical synthesis database. The first knowledge recommendation result is determined by deep screening of the first retrieval and screening result using the synthesis process recommendation professional large model.
[0052] Based on the element requirement information, the synthesis path and chemical material are retrieved in the chemical synthesis database to determine the matched synthesis path set and matched chemical material set, and the matched synthesis path set and matched chemical material set are used as the first retrieval and screening result. The first knowledge recommendation result is determined by screening the first retrieval and screening result based on the material type requirement information and synthesis generation restriction requirement information using the synthesis process recommendation professional large model.
[0053] In the chemical synthesis database, the retrieval based on the element requirement information is performed to screen out all synthesis paths and materials related to the element requirement information, for example, multiple candidate chemical materials containing nitrogen and oxygen are retrieved in the database and their synthesis paths can also be generated. These paths and materials are the matched synthesis path set and matched chemical material set. These matched synthesis path set and matched chemical material set are used as the first retrieval and screening result for subsequent steps.
[0054] The material type requirement information and the synthesis generation restriction requirement information are input into a synthesis process recommendation professional large model which has been trained to be able to evaluate and screen the synthesis process of the chemical material according to multiple factors. The synthesis process recommendation professional large model further screens the matched synthesis path set and the matched chemical material set of the first search screening result according to the material type requirement information and the synthesis generation restriction requirement information. For example, assuming that the material type requirement information requires the synthesized material to be high conductivity, and the synthesis generation restriction requirement information requires the reaction temperature to be no more than 200 DEG C, if a certain synthesis path involves a material that cannot meet the requirement of high conductivity or cannot complete the reaction below 200 DEG C, the path and the material will be excluded. After screening, the synthesis path and the chemical material that meet all the requirement conditions constitute the first knowledge recommendation result, which is used as the candidate result recommended to the user.
[0055] The synthesis process recommendation professional large model is used to perform inference expansion of the chemical material with a similar synthesis process on the first search screening result to determine the second knowledge recommendation result.
[0056] Further, the synthesis process recommendation professional large model is used to perform inference expansion of the chemical material with a similar synthesis process on the first search screening result to determine the second knowledge recommendation result, including:
[0057] The synthesis process recommendation professional large model is used to perform inference expansion on the first search screening result, and an expanded chemical material set and an expanded synthesis path set are output. The expanded chemical material set and the expanded synthesis path set are input into the synthesis process recommendation professional large model, and screening is performed based on the material type requirement information and the synthesis generation restriction requirement information to determine the second knowledge recommendation result, wherein the second knowledge recommendation result includes a screened expanded chemical material set and a screened expanded synthesis path set.
[0058] The first search screening result, i.e., the matched synthesis path set and the matched chemical material set, is passed as input to the synthesis process recommendation professional large model. Based on the existing synthesis path and material information, the synthesis process recommendation professional large model uses its learned synthesis rules and reasoning ability of chemical reactions to infer possible other chemical materials and synthesis paths similar to the existing path and materials. For example, if a synthesis path contains a specific catalyst or temperature condition, the synthesis process recommendation professional large model can infer other materials or paths that can be synthesized under similar catalysts or temperature conditions. This inference is based on similar chemical reaction mechanisms, similar physical and chemical properties, known reaction reactivity and selectivity of chemical reactions, etc. The materials and synthesis paths obtained by inference expansion form two new components, wherein the expanded chemical material set contains inferred chemical materials similar to existing materials but possibly new; the expanded synthesis path set is the corresponding possible new synthesis path for the expanded chemical material, which has similar synthesis process and conditions to the existing path.
[0059] The expanded chemical material set and the expanded synthesis path set obtained by inference expansion are input into the synthesis process recommendation professional large model. The synthesis process recommendation professional large model further screens the expanded chemical material set and the expanded synthesis path set according to the material type requirement information and the synthesis generation restriction requirement information. The screening process is similar to the first search screening, and for the sake of brevity of the specification, it will not be repeated here. After screening, the expanded chemical materials and the expanded synthesis paths that meet all the requirement conditions form the screened expanded chemical material set and the screened expanded synthesis path set, which constitute the second knowledge recommendation result and are used as candidate results recommended to the user.
[0060] The synthesis process recommendation professional large model is used to infer the synthesis process path based on the element requirement information, the material type requirement information, and the synthesis generation restriction requirement information, to determine the third knowledge recommendation result.
[0061] The element requirement information, the material type requirement information, and the synthesis generation restriction requirement information are input as input data into the synthesis process recommendation professional large model. The synthesis process recommendation professional large model has learned how to generate synthesis process paths through a large amount of training data. Through synthesis process path inference, a new set of synthesis process paths is generated. These synthesis process paths constitute the third knowledge recommendation result, which is obtained through intelligent inference and calculation of the synthesis process recommendation professional large model, and are usually more flexible and innovative than database-based path recommendations.
[0062] The first knowledge recommendation result, the second knowledge recommendation result, and the third knowledge recommendation result are fused to obtain the target knowledge recommendation result.
[0063] The first knowledge recommendation result, the second knowledge recommendation result and the third knowledge recommendation result are fused, and exemplarily, the three recommendation results are weighted and sorted based on the reliability, the innovativeness and the degree of meeting the demand of the synthetic path, for example, for the path in the database, the weighting is performed according to the success rate or the universality in the actual synthesis, and for the path recommended by the model, the weighting is performed according to the innovativeness or the predicted success rate. After the fusion processing, the target knowledge recommendation result is generated, such as Figure 3 As shown in the example chemical material synthesis mode recommendation flowchart, the target knowledge recommendation result obtained by the synthesis path in the existing database and the innovative path inferred by the large model is able to provide more innovative synthesis schemes while ensuring the reliability.
[0064] In summary, the method for recommending a material synthesis process in the chemical industry provided by the embodiments has the following technical effects:
[0065] By obtaining the target chemical material synthesis demand text information for multi-step iterative semantic analysis, the synthesis demand of the user can be automatically extracted and understood, the dependence on manual operation is reduced, the automation degree is improved, and the recommendation efficiency is improved; by pre-building a synthesis process recommendation professional large model, not only the traditional database query is relied on, but also the chemical knowledge and reaction rules learned in the model training process are relied on for intelligent inference to generate innovative synthesis paths, so that the recommendation system can more flexibly and accurately adapt to the diversified needs of the user; in combination with the element demand, the material type and the synthesis generation restriction information, the recommendation range is effectively narrowed down in the chemical synthesis database, and only the materials and synthesis paths meeting the user's demand are focused on, which optimizes the data retrieval and improves the relevance and pertinence of the results, avoiding the recommendation of irrelevant or non-compliant materials and paths; the first retrieval and screening result is inferred and expanded by the large model, which can infer other synthesis paths and chemical materials similar to or potentially similar to the existing synthesis paths, which improves the comprehensiveness of the recommendation and identifies excellent candidate materials or synthesis processes that may not be directly retrieved, increasing the selectivity; the element demand information, the material type demand information and the synthesis generation restriction information are inferred by the large model to further optimize the synthesis process flow, ensuring that the recommended path not only meets the basic demand, but also provides more optimal process parameters and conditions in actual operation, which helps the user to optimize the existing process flow, reduce the cost and improve the efficiency; by fusing the first knowledge recommendation result, the second knowledge recommendation result and the third knowledge recommendation result, the information of different recommendation sources can be integrated, and the possible deviation of a single recommendation method can be eliminated, so as to obtain a more accurate and comprehensive target knowledge recommendation result, and the fusion ensures that the recommendation system is considered from multiple dimensions, so that the accuracy, practicality and innovativeness of the final result are improved.
[0066] The foregoing description of the disclosed embodiments enables a person skilled in the art to make or use the application. Modifications of these embodiments will occur to persons of skill in the art, and, while certain embodiments according to the principles set forth herein are shown and described, it is to be understood that the same are not limiting of the scope of the application as it is set forth in the appended claims, and that various modifications are made within the scope of the appended claims. Therefore, it is contemplated to cover the application in its broadest scope, including all features that can be made or used in the embodiments described herein.
Claims
1. A method for recommending a material synthesis process in the chemical industry, characterized by, The method comprises: obtaining target chemical material synthesis demand text information for multi-step iterative semantic analysis to determine element demand information, material type demand information and synthesis generation restriction demand information; pre-building a synthesis process recommendation professional large model; combining the element demand information, material type demand information and synthesis generation restriction demand information for retrieval and screening in a chemical synthesis database to obtain a first retrieval and screening result, and using the synthesis process recommendation professional large model to perform deep screening on the first retrieval and screening result to determine a first knowledge recommendation result; using the synthesis process recommendation professional large model to perform chemical material reasoning expansion with similar synthesis processes on the first retrieval and screening result to determine a second knowledge recommendation result; using the synthesis process recommendation professional large model to perform synthesis process path reasoning on the element demand information, material type demand information and synthesis generation restriction demand information to determine a third knowledge recommendation result; fusing the first knowledge recommendation result, the second knowledge recommendation result and the third knowledge recommendation result to obtain a target knowledge recommendation result; The pre-built synthesis process recommendation professional large model comprises: collecting data sources including at least chemical material names, physicochemical properties and synthesis process texts to build a basic data source set; using the text-embedding-v3 model of qwen to generate synthesis process text vectors of the basic data source set, and performing material pair screening according to similarity Top-k to determine a material pair set; using a teacher model to perform synthesis path similarity reasoning on the material pair set according to the physicochemical properties in the basic data source set, and performing review on the reasoning text to determine target reasoning corpus; packaging the target reasoning corpus, and fine-tuning a basic model through the packaged target reasoning corpus to build the synthesis process recommendation professional large model; The basic model is Qwen3-8B, and the synthesis process recommendation professional large model is built by fine-tuning the basic model through the llama-factory training framework combined with the packaged target reasoning corpus; The method comprises: dividing the target chemical material synthesis demand text information into semantic sub-information sequences; taking the semantic sub-information sequences as explicit information sequences and obtaining a chemical common sense knowledge graph as implicit information; initializing a semantic analysis unit, wherein the semantic analysis unit stores a preset synthesis demand keyword set; based on the semantic analysis unit, performing multi-step iterative semantic analysis on the explicit information sequences and implicit information to determine element demand information, material type demand information and synthesis generation restriction demand information.
2. A method of recommending a process for the synthesis of a material in the chemical industry as claimed in claim 1, wherein, The element requirement information, the material type requirement information and the synthesis generation restriction requirement information are combined for retrieval screening in the chemical synthesis database to obtain a first retrieval screening result, the first retrieval screening result is screened in depth by using the synthesis process recommendation professional large model to determine a first knowledge recommendation result, including: Based on the element requirement information, a synthesis path and a chemical material are retrieved in the chemical synthesis database to determine a matched synthesis path set and a matched chemical material set, and the matched synthesis path set and the matched chemical material set are taken as the first retrieval screening result; The first retrieval screening result is screened based on the material type requirement information and the synthesis generation restriction requirement information by using the synthesis process recommendation professional large model to determine the first knowledge recommendation result.
3. A method of recommending a process for the synthesis of a material in the chemical industry as claimed in claim 1, wherein, The first retrieval screening result is reasoned and expanded by using the synthesis process recommendation professional large model to obtain a chemical material with a similar synthesis process to determine a second knowledge recommendation result, including: The first retrieval screening result is reasoned and expanded by using the synthesis process recommendation professional large model to output an expanded chemical material set and an expanded synthesis path set; The expanded chemical material set and the expanded synthesis path set are input into the synthesis process recommendation professional large model, and are screened based on the material type requirement information and the synthesis generation restriction requirement information to determine the second knowledge recommendation result, wherein the second knowledge recommendation result includes a screened expanded chemical material set and a screened expanded synthesis path set.
4. The method of recommending a process for the synthesis of a material in the chemical industry of claim 1, wherein, The target reasoning corpus is packaged into a JSON structure in a fine-tuning format.
5. The method of recommending a process for the synthesis of a material in the chemical industry of claim 1, wherein, Based on the semantic analysis unit, the explicit information sequence and the implicit information are iteratively analyzed in multiple steps to determine the element requirement information, the material type requirement information and the synthesis generation restriction requirement information, including: The explicit information sequence is converted into a word vector to represent the explicit information word vector sequence; The first explicit information word vector in the explicit information word vector sequence is extracted, and is sequentially interacted with a preset synthesis requirement keyword set and implicit information in a first-order interaction and a second-order interaction to obtain a first iterative semantic analysis unit; The explicit information word vector sequence is iteratively analyzed in multiple steps in the order from front to back in combination with the first iterative semantic analysis unit and the implicit information to determine a target iterative semantic analysis unit; The target iterative semantic analysis unit is integrated according to the semantic type to determine the element requirement information, the material type requirement information and the synthesis generation restriction requirement information.
6. A method of recommending a process for the synthesis of a material in the chemical industry as claimed in claim 5, c h a r a c t e r i z e d b y The first explicit information word vector in the explicit information word vector sequence is extracted, and is sequentially interacted with a preset synthesis requirement keyword set and implicit information in a first-order interaction and a second-order interaction to obtain a first iterative semantic analysis unit, including: The first explicit information word vector and the preset synthesis requirement keyword set are interacted in a first-order interaction by using a cosine similarity formula to extract a first explicit information keyword vector; The implicit information is interacted in a second-order interaction based on the first explicit information keyword vector to extract a first implicit information keyword vector; The first explicit information keyword vector and the first implicit information keyword vector are added into a semantic parsing unit to obtain a first iteration semantic parsing unit.
7. The method of recommending a process for the synthesis of a material in the chemical industry of claim 1, wherein, The target chemical material synthesis demand text information is divided into semantic sub-information sequences, including: Punctuation standardization and stop word elimination are performed on the target chemical material synthesis demand text information to obtain preprocessed target chemical material synthesis demand text information. The preprocessed target chemical material synthesis demand text information is processed step by step to obtain the semantic sub-information sequences.
Citation Information
Patent Citations
Synthetic route recommendation method and terminal
CN115206450A
Text analysis method and device based on big data
CN115470773A