A processing method and device for generating molecular descriptions based on a chemical large model
By fine-tuning the general large language model with enhanced chemical knowledge, the issues of accuracy and comprehensiveness in generating molecular descriptions were resolved, improving the model's chemical reasoning ability and the chemical accuracy of the generated text.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING DP TECH CO LTD
- Filing Date
- 2026-04-13
- Publication Date
- 2026-07-21
AI Technical Summary
The general-purpose large language model lacks specialized chemical knowledge when generating molecular descriptions, resulting in inaccurate or incomplete descriptions. Furthermore, it lacks chemical reasoning ability, which affects the accuracy of the generated text.
A pre-trained general language model was selected as the chemical model. By configuring chemical question-and-answer instruction templates and molecular description instruction templates, and combining big data collection in the field of chemistry to construct question-and-answer datasets and inference chain datasets, a three-step reinforcement learning fine-tuning was performed to improve the model's chemical knowledge understanding and reasoning ability.
It improves the comprehensiveness and chemical accuracy of the generated molecular descriptions, and enhances the model's ability to understand and analyze specialized chemical knowledge.
Smart Images

Figure CN122436048A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to a processing method and apparatus for generating molecular descriptions based on large chemical models. Background Technology
[0002] Molecular captioning (MCC) generation is the process of generating natural language descriptions from molecular structures (such as the string "SELFIES"). This technology has significant application value in fields such as drug design and materials design. Currently, general-purpose Large Language Models (LLMs), such as Qwen, GPT, and DeepSeek, have made significant progress in a series of general Natural Language Processing (NLP) tasks (such as text generation, translation, question answering, and chain of thought derivation). However, some problems still exist when handling molecular caption generation tasks in specialized chemical fields: 1) The pre-training corpus of general-purpose LLMs and the fine-tuning corpus for NLP tasks do not contain sufficient specialized chemical knowledge, making the generated molecular captions prone to inaccuracies or incompleteness; 2) The chain of thought (CoT) reasoning ability of general-purpose LLMs only possesses general knowledge reasoning ability and lacks specialized chemical knowledge reasoning ability, making the generated molecular captions prone to chemical accuracy deviations. Summary of the Invention
[0003] The purpose of this invention is to address the shortcomings of existing technologies by providing a method, apparatus, electronic device, and computer-readable storage medium for generating molecular descriptions based on a large chemical model. This invention selects a general-purpose large language model that has completed pre-training for both large language models and general NLP tasks as the large chemical model; and configures corresponding chemical question-and-answer instruction templates and molecular description instruction templates for it. A corresponding question-and-answer dataset is constructed through large-scale data collection from publicly available chemical professional knowledge media; and a corresponding inference chain dataset and molecular description dataset are constructed through large-scale data collection from publicly available molecular databases of SELFIES format molecular sequences and their corresponding molecular structures and physicochemical properties. The large chemical model is trained using a three-step reinforcement learning approach: first, a first-stage chemical knowledge reinforcement fine-tuning is performed based on the chemical question-and-answer instruction templates and the question-and-answer dataset; second, a second-stage chemical inference reinforcement fine-tuning is performed based on the molecular description instruction templates and the inference chain dataset; and finally, a third-stage molecular description reinforcement fine-tuning is performed based on the molecular description instruction templates and the molecular description dataset. After model fine-tuning, the user-input SELFIES format molecular sequence is substituted into the molecular description instruction template. The molecular sequence text of the SELFIES segment in the template is configured to obtain the corresponding molecular description instruction. The current molecular description instruction is then input into the large-scale chemical model for molecular description generation. The molecular description text output by the model in this processing is then fed back to the user. This invention, through a three-step progressive reinforcement learning approach, can improve the model's understanding and parsing ability of specialized chemical knowledge, enhance the model's chemical reasoning ability, and improve the comprehensiveness and chemical accuracy of the generated text when processing molecular description generation tasks based on the large-scale chemical model provided by this invention.
[0004] To achieve the above objectives, a first aspect of the present invention provides a processing method for generating molecular descriptions based on a large chemical model, the method comprising: Select a general-purpose language model that has completed pre-training for both large language models and general NLP tasks as the corresponding chemical large model; the general-purpose language model includes at least the Qwen series models, GPT series models, and DeepSeek series models; the general-purpose NLP tasks include at least text generation tasks, translation tasks, question answering tasks, and thought chain deduction tasks; Configure corresponding chemical question-and-answer instruction templates and molecular description instruction templates for the large chemical model; We construct corresponding question-and-answer datasets by collecting big data from publicly available knowledge media in the field of chemistry; and we construct corresponding inference chain datasets and molecular description datasets by collecting big data from publicly available molecular databases of SELFIES format molecular sequences and their corresponding molecular structures and physicochemical properties. First, the chemical big model is fine-tuned in one stage by enhancing chemical knowledge based on the chemical question-and-answer instruction template and the question-and-answer dataset; then, the chemical big model is fine-tuned in two stages by enhancing chemical reasoning based on the molecular description instruction template and the reasoning chain dataset; finally, the chemical big model is fine-tuned in three stages by enhancing molecular description based on the molecular description instruction template and the molecular description dataset. After the model fine-tuning is completed, the user-input SELFIES format molecular sequence is substituted into the molecular description instruction template to configure the molecular sequence text of the SELFIES segment in the template to obtain the corresponding current molecular description instruction; the current molecular description instruction is then input into the large chemical model for molecular description generation task processing; and the molecular description text output by the model in this processing is fed back to the current user.
[0005] Preferably, the chemical question-and-answer instruction template consists of a question-and-answer instruction requirement text and a question text; The question-and-answer instruction requires the text to be a fixed natural language text, which is used to prompt the chemical big model to generate a corresponding answer based on the given question in the question text; The question segment consists of a fixed question segment title and configurable question text; the question segment title defaults to the string "Question:"; the question text is initialized to empty; when the question text is not empty, it is a piece of natural language text used to ask questions about the basic chemical concepts given in the text or a piece of natural language text used to ask questions about one or more types of molecular features given the molecular name or SELFIES molecular sequence in the text. The knowledge scope corresponding to basic chemical concepts includes at least the atomic model and quantum numbers, the principle of electron configuration and structure, the periodic table, the atomic structure and periodic properties of the periodic law, various chemical bond structures and their corresponding force and energy theories, intermolecular forces, chemical thermodynamics theory, chemical reaction theory, acid-base theory, redox theory, functional group theory, polymer theory, organic matter theory, inorganic matter theory, stereochemistry theory, physicochemical property theory, chemical nomenclature rules, quantum chemistry theory, and biological and pharmaceutical chemistry theory. The knowledge scope corresponding to molecular characteristics includes at least basic constituent characteristics, two-dimensional structural characteristics, three-dimensional structural characteristics, physicochemical property characteristics, spectroscopic characteristics, crystallographic characteristics, biological activity characteristics, material performance characteristics, and reaction characteristics; the basic constituent characteristics include at least molecular formula, molecular mass, chemical formula, and isotopic composition; the two-dimensional structural characteristics include at least atomic connection sequence characteristics, chemical bond characteristics, functional group characteristics, and atomic skeleton characteristics; the three-dimensional structural characteristics include at least spatial geometric characteristics, rotational conformation characteristics, chiral characteristics, and geometric isomerism characteristics; the physicochemical property characteristics include at least melting point, boiling point, flash point, solubility, partition coefficient, vapor pressure, polarity, dipole moment, ionization energy, electron affinity, color, optical rotation, chemical stability, and acid / base dissociation constant; the spectroscopic characteristics include at least spectral characteristics and mass spectrometry characteristics. The molecular description instruction template consists of a description instruction requirement section, the SELFIES section, a reasoning step description section, and a molecular description formatting requirement section. The description instruction requires the text segment to be a fixed natural language text, which is used to prompt the chemical big model to perform step-by-step parsing and reasoning according to the SELFIES molecular sequence information given in the SELFIES text segment, the step sequence given in the reasoning step description text segment, and generate a molecular description text that conforms to chemical laws based on the reasoning context. The SELFIES segment consists of a fixed sequence segment title and configurable molecular sequence text; the molecular sequence title defaults to the string "SELFIES:"; the molecular sequence text is initialized to empty; when the molecular sequence text is not empty, its text format is a molecular sequence in SELFIES format; The inference step description is a fixed natural language text consisting of S inference step texts, where S is the preset total number of inference steps. Each inference step text is a step description text for one step of inference, used to prompt the chemical macro-model to perform analysis and inference according to the requirements of this step and the current inference context, and to take the result of this step as a corresponding single-step inference text C. s The output is 1 ≤ index s ≤ S; the current inference context includes the SELFIES molecular sequence information given by the SELFIES text segment, the inference results of all historical steps before the current inference step, and the single-step inference text C corresponding to step S. s=S This is the corresponding molecular description text; The molecular description formatting requirement is a fixed natural language text segment, which is used to prompt the chemical model to encapsulate and output the generated molecular description text according to the preset molecular description text output format; the molecular description text output format is formed by connecting the preset start marker text, the molecular description text generated by the model, and the preset end marker text in sequence by default. The publicly available knowledge media in the field of chemistry include at least publicly available textbooks or theoretical books in the field of chemistry, publicly available journals or papers in the field of chemistry, and publicly available databases or molecular libraries in the field of chemistry. The publicly available molecular databases include at least the PubChem database and the PDB database; The question-and-answer dataset includes a basic chemistry question-and-answer subset and a molecular feature question-and-answer subset; both the basic chemistry question-and-answer subset and the molecular feature question-and-answer subset consist of multiple first data records; each first data record includes a first training question and a first labeled answer; the first training question of the basic chemistry question-and-answer subset is a piece of natural language text used to ask questions about a given basic chemical concept within the text; the first training question of the molecular feature question-and-answer subset is a piece of natural language text used to ask questions about one or more types of molecular features given a molecular name or SELFIES molecular sequence within the text; The inference chain dataset includes multiple second data records; the second data records include a first training sequence and a first label inference chain; the first label inference chain includes S first single-step labels. The first training sequence is a SELFIES molecular sequence. The molecular description dataset includes multiple third data records; each third data record includes a second training sequence and a first label description; the second training sequence is a SELFIES molecular sequence; the first label description is the molecular description text of the corresponding molecule.
[0006] Preferably, the step of constructing a corresponding question-and-answer dataset by collecting big data from publicly available knowledge media in the field of chemistry specifically includes: Multiple knowledge entries are obtained by collecting big data from the publicly available chemical professional knowledge media; and corresponding first training questions and first labeled answers are generated for each knowledge entry through manual or machine questioning and answer labeling to form the corresponding first data record; and all the first data records corresponding to basic chemistry questions and answers are combined into the corresponding basic chemistry question and answer subset through manual or machine classification, and all the first data records corresponding to molecular feature questions and answers are combined into the corresponding molecular feature question and answer subset; and the basic chemistry question and answer subset and the molecular feature question and answer subset are combined to form the corresponding question and answer dataset. Among them, the knowledge domain scope of all the knowledge items obtained through big data collection is greater than or equal to the preset knowledge scope of basic chemical concepts and molecular characteristics; the knowledge scope of basic chemical concepts and molecular characteristics is the combination of the knowledge scope corresponding to the basic chemical concepts and the knowledge scope corresponding to the molecular characteristics.
[0007] Preferably, the step of constructing corresponding inference chain datasets and molecular description datasets through large-scale data collection of SELFIES format molecular sequences and their corresponding molecular structures and physicochemical properties from publicly available molecular databases specifically includes: In the publicly available molecular database, large-scale data collection is performed on molecular sequences in SELFIES format. By querying the molecular database, the basic constituent features, two-dimensional structural features, three-dimensional structural features, physicochemical property features, spectroscopic features, crystallographic features, bioactivity features, material performance features, and reaction features corresponding to each collected molecular sequence are queried. Based on the query results, corresponding molecular feature records are formed. Each collected molecular sequence and its corresponding molecular feature record form a corresponding first collection entry. All the first collection entries form a corresponding first collection library. Each of the first acquisition entries in the first acquisition library is taken as the current entry; and the molecular sequence of the current entry is taken as a set of corresponding first training sequences and second training sequences; and a chemical professional or other specialized chemical large-scale model, based on the molecular feature record of the current entry, generates a corresponding single-step reasoning result as the corresponding first single-step label for each reasoning step text specified in the reasoning step description text of the molecular description instruction template. And the S first single-step labels obtained Form the corresponding first tag inference chain; and set the first single-step tag corresponding to step S. As the current description text; and according to the molecular description formatting requirements specified in the molecular description instruction template, the first label description is formed by sequentially connecting the start marker text, the current description text, and the end marker text; and the second data record is formed by the first training sequence and the first label inference chain corresponding to the current entry, and the third data record is formed by the second training sequence and the first label description corresponding to the current entry; and the inference chain dataset is formed by all the obtained second data records, and the molecular description dataset is formed by all the obtained third data records.
[0008] Preferably, the step of performing a first-stage chemical knowledge enhancement and fine-tuning of the large chemical model based on the chemical question-and-answer instruction template and the question-and-answer dataset specifically includes: Step 51: Use the basic chemical question-and-answer subset of the question-and-answer dataset as the current dataset; Step 52: Divide the current dataset into multiple first data batches based on a preset batch size B1; and take the first first data batch of the current dataset as the current data batch; Each first data batch includes B1 first data records; the first tag answer of each first data record in each first data batch is denoted as the corresponding tag answer. , 1 ≤ index j ≤ B1; each of the stated label answers The total number of word segments is denoted as n. j Each of the aforementioned tagged answers Each word segment is recorded as the corresponding 1 ≤ index k ≤ n j ; Step 53: Substitute the first training question of each first data record in the current data batch into the chemical question-and-answer instruction template, and configure the question text of the question segment in the template to obtain the corresponding current chemical question-and-answer instruction Q. j ; and the current chemical question-and-answer command Q j The chemical model is input for question-answering task processing; the autoregressive text generation process of the chemical model during this processing is recorded; and after completing B1 model processing iterations, the first model loss function L is used as the basis for the calculation. M1 Calculate the corresponding first loss value; Wherein, the first model loss function L M1 It is implemented based on the cross-entropy loss function, specifically as follows: ; θ is used to identify the model parameters of the large chemical model; For the tagged answer The word segmentation sequence preceding the kth word; For the model in the autoregressive text generation process, the current chemical question-answering instruction Q is used. j and word segmentation sequence The k-th word generated for the context is the corresponding word. The probability of; Step 54: Identify whether the first loss value meets the preset first loss value range; if not, then based on the preset first model optimizer, move towards making the first model loss function L... M1 The model parameters of the large chemical model are fine-tuned in the direction that reaches the minimum value, and the process returns to step 53 after this round of fine-tuning is completed; if the condition is met, proceed to step 55. The first model optimizer includes the Adam optimizer, the SGD optimizer, and the AdamW optimizer. Step 55: Identify whether the current data batch is the last first data batch of the current dataset; if not, take the next first data batch of the current dataset as the new current data batch and return to step 53; if yes, identify whether the current dataset is the basic chemical question-and-answer subset; if yes, take the molecular feature question-and-answer subset of the question-and-answer dataset as the new current dataset and return to step 52; otherwise, stop training and confirm that the chemical knowledge reinforcement fine-tuning training of the model has ended.
[0009] Preferably, the step of performing two-stage chemical inference enhancement and fine-tuning on the large chemical model based on the molecular description instruction template and the inference chain dataset specifically includes: Step 61: Divide the inference chain dataset into multiple second data batches based on a preset batch size B2; and use the first second data batch as the current data batch; Each second data batch includes B2 second data records; the first tag inference chain of each second data record in each second data batch is denoted as the corresponding tag chain. , 1 ≤ index q ≤ B2; each of the stated tag chains Each of the first single-step tags Record as the corresponding single-step label Each of the aforementioned single-step labels The total number of word segments is denoted as n. q,s Each of the aforementioned single-step labels Each word segment is recorded as the corresponding 1 ≤ index u ≤ n q,s ; Step 62, set the sequence X of each of the second data records in the current data batch. q Substituting the molecular description instruction template into the template, the molecular sequence text of the SELFIES segment is configured to obtain the corresponding current molecular description instruction X. q ; and the current molecular description instruction X q The chemical large model is input for molecular description generation; the autoregressive text generation process of the chemical large model is recorded during this process; and after completing B2 model processing iterations, a second model loss function L is used as the basis for the process. M2 Calculate the corresponding second loss value; Wherein, the second model loss function L M2 It is implemented based on the cross-entropy loss function, specifically as follows: ; θ is used to identify the model parameters of the large chemical model; It is a sequence of single-step tags formed by concatenating the first s-1 single-step tags in sequence; For the single-step label The word segmentation sequence preceding the u-th word; For the autoregressive generation step of the inference text in step s of the model, the corresponding molecular description instruction X is used. q Single-step label sequence and word segmentation sequence The u-th word generated for the context is the corresponding word. The probability of; Step 63: Identify whether the second loss value meets the preset second loss value range; if not, then based on the preset second model optimizer, move towards making the second model loss function L... M2 The model parameters of the large chemical model are fine-tuned in the direction of reaching the minimum value, and the process returns to step 62 after the fine-tuning is completed. If the condition is met, it is identified whether the current data batch is the last second data batch. If not, the next second data batch is taken as the new current data batch and the process returns to step 62. If the condition is met, training is stopped and the chemical inference enhancement fine-tuning training of the model is confirmed to be completed. The second model optimizer includes the Adam optimizer, the SGD optimizer, and the AdamW optimizer.
[0010] Preferably, the three-stage molecular description enhancement and fine-tuning of the large chemical model based on the molecular description instruction template and the molecular description dataset specifically includes: Step 71: The chemical macro model is replicated to obtain two replicated models, and the current chemical macro model is denoted as the corresponding new strategy model M. new The two replication models are denoted as the corresponding old policy models M. old Reference Model M ref and the reference model M ref The model parameters are solidified; and the new strategy model M is then... new The old strategy model M old The reference model M ref The model parameters are denoted as the corresponding model parameters θ. new Model parameters θ old Model parameters θ ref The molecular description dataset is divided into multiple third data batches based on a preset batch size B3, and the first third data batch is taken as the current data batch. Each of the third data batches includes B3 third data records; the second training sequence of each third data record in each of the third data batches is denoted as the corresponding sequence X. e The first tag description is denoted as the corresponding tag text. , 1 ≤ index e ≤ B3; Step 72, record the sequence X of each of the third data records in the current data batch. e The molecular sequence text of the SELFIES segment in the template is configured by substituting the molecular description instruction template to obtain the corresponding current molecular description instruction, and the current molecular description instruction is input into the old strategy model M consecutively G times. old The molecular description generation task is performed to obtain G formatted molecular description texts, and these G formatted molecular description texts are respectively recorded as the corresponding predicted texts. 1 ≤ index g ≤ G; and each of the predicted texts The total number of word segments is recorded as the corresponding Each of the predicted texts Each word segment is recorded as the corresponding 1≤index t≤ ; and each word segmentation The corresponding predicted probability is denoted as ; Wherein, the number of repetitions G is a preset positive integer; Step 73: Based on the text format check results and the language similarity assessment results between the predicted text and its corresponding tag text, process each of the predicted texts. The corresponding text reward R e,g Perform calculations; Specifically, this involves: processing each of the predicted texts The current predicted text is used as the basis for identification; and the text format of the current predicted text is checked to see if it meets the molecular description formatting requirements specified in the molecular description instruction template. If it does, the corresponding format check result f is set. e,g If the condition is not met, the corresponding format check result f is set to 1. e,g The value is 0; and the format check result f is... e,g If the value is 0, then the corresponding text reward R is set. e,g If the value is 0, then the current predicted text and its corresponding tag text are compared based on a preset six-class text similarity evaluation algorithm. The text similarity was evaluated to obtain six corresponding evaluation values, and the average of the six evaluation values was taken as the corresponding language similarity. And based on the language similarity Calculate the corresponding text reward R e,g ; The six types of text similarity evaluation algorithms include BLEU-2 evaluation algorithm, BLEU-4 evaluation algorithm, METEOR evaluation algorithm, ROUGE-1 evaluation algorithm, ROUGE-2 evaluation algorithm, and ROUGE-L evaluation algorithm. The text reward R e,g The calculation method is as follows: ; Step 74: Generate the G predicted texts corresponding to each of the third data records in the current data batch. Cluster them into groups; and calculate the average reward µ for each group. e and standard deviation σ e Perform calculations; and based on each of the predicted texts in each group. The corresponding text reward R e,g The average value µ e and the standard deviation σ e Calculate the corresponding within-group advantage V e,g ; Wherein, the average value µ e The standard deviation σ e and the aforementioned intra-group advantage V e,g The calculation method is as follows: , ; ; λ1 is a preset small constant used to prevent the denominator from being zero; Step 75, record the sequence X of each of the third data records in the current data batch. e The molecular sequence text of the SELFIES segment in the template is configured by substituting the molecular description instruction template to obtain the corresponding current molecular description instruction, and the current molecular description instruction is input into the reference model M G times consecutively. ref Perform molecular description generation task processing; and during the g-th task processing corresponding to the e-th third data record, match the word probability vector corresponding to the t-th segment generated in this processing with the corresponding predicted text. The t-th segmentation The corresponding word segmentation probabilities are extracted and used as the corresponding prediction probabilities. ; Step 76, record the sequence X of each of the third data records in the current data batch. eThe molecular sequence text of the SELFIES segment in the template is configured by substituting the molecular description instruction template to obtain the corresponding current molecular description instruction, and the current molecular description instruction is input into the new strategy model M G times consecutively. new Perform molecular description generation task processing; and during the g-th task processing corresponding to the e-th third data record, match the word probability vector corresponding to the t-th segment generated in this processing with the corresponding predicted text. The t-th segmentation The corresponding word segmentation probabilities are extracted and used as the corresponding prediction probabilities. ; Step 77, for each of the predicted texts Each of the aforementioned words Importance sampling ratio r e,g,t And truncation sampling ratio Perform calculations; Wherein, the importance sampling ratio r e,g,t and the truncation sampling ratio The calculation method is as follows: , ; λ2 is the preset truncation threshold, and Clip() is the Clip truncation function; Step 78, based on the preset third model loss function L M3 The corresponding third loss value is obtained through calculation; Wherein, the third model loss function L M3 The objective function J based on the group relative strategy optimization algorithm GRPO To achieve this, the objective function J of the group relative strategy optimization algorithm is... GRPO Then, by the policy function term J policy , divergence function term J KL The composition is as follows: , , ; β is the preset divergence coefficient; Step 79: Identify whether the third loss value meets a preset third loss value range; if the third loss value does not meet the third loss value range, then based on the preset third model optimizer, move towards making the third model loss function L... M3 The direction for reaching the minimum value corresponds to the new strategy model M. new The model parameters θ newPerform one round of modulation, and return to step 72 at the end of this round of modulation; if the third loss value meets the range of the third loss value, then based on the model parameter θ new For the old strategy model M old The model parameters θ old Perform a reset and identify whether the current data batch is the last third data batch. If not, use the next third data batch as the new current data batch and return to step 72. If yes, stop training, confirm that the molecular description enhancement fine-tuning training of the model is complete, and set the new strategy model M. new As the latest chemical large model; wherein, the third model optimizer includes the Adam optimizer and the AdamW optimizer.
[0011] A second aspect of the present invention provides an apparatus for implementing the processing method for generating molecular descriptions based on a large chemical model as described in the first aspect above. The apparatus includes: a language model selection module, an instruction template configuration module, a data acquisition module, a three-stage enhancement module, and a model application module. The language model selection module is used to select a general-purpose language model that has completed pre-training for both large language models and general NLP tasks as the corresponding chemical large model; the general-purpose language model includes at least the Qwen series models, GPT series models, and DeepSeek series models; the general-purpose NLP tasks include at least text generation tasks, translation tasks, question answering tasks, and thought chain derivation tasks; The instruction template configuration module is used to configure corresponding chemical question-and-answer instruction templates and molecular description instruction templates for the large chemical model. The data acquisition module is used to construct a corresponding question-and-answer dataset by collecting big data from publicly available knowledge media in the field of chemistry; and to construct a corresponding inference chain dataset and molecular description dataset by collecting big data from publicly available molecular databases of SELFIES format molecular sequences and their corresponding molecular structures and physicochemical properties. The three-stage enhancement module is used to first perform a first-stage chemical knowledge enhancement and fine-tuning of the large chemical model based on the chemical question-and-answer instruction template and the question-and-answer dataset; then perform a second-stage chemical reasoning enhancement and fine-tuning of the large chemical model based on the molecular description instruction template and the reasoning chain dataset; and finally perform a third-stage molecular description enhancement and fine-tuning of the large chemical model based on the molecular description instruction template and the molecular description dataset. The model application module is used to, after the model fine-tuning is completed, substitute the user-input SELFIES format molecular sequence into the molecular description instruction template to configure the molecular sequence text of the SELFIES segment of the template to obtain the corresponding current molecular description instruction; input the current molecular description instruction into the chemical large model for molecular description generation task processing; and feed back the molecular description text output by the model in this processing to the current user.
[0012] A third aspect of the present invention provides an electronic device, including: a memory, a processor, and a transceiver; The processor is used to couple with the memory, read and execute instructions in the memory to implement the steps of the method described in the first aspect above; The transceiver is coupled to the processor, and the processor controls the transceiver to send and receive messages.
[0013] A fourth aspect of the present invention provides a computer-readable storage medium storing computer instructions that, when executed by a computer, cause the computer to perform the instructions described in the first aspect.
[0014] This invention provides a method, apparatus, electronic device, and computer-readable storage medium for generating molecular descriptions based on a large chemical model. As described above, this invention selects a general-purpose large language model that has completed pre-training for both large language models and general NLP tasks as the large chemical model; and configures corresponding chemical question-and-answer instruction templates and molecular description instruction templates for it. A corresponding question-and-answer dataset is constructed by collecting large amounts of publicly available chemical knowledge; and a corresponding inference chain dataset and molecular description dataset are constructed by collecting large amounts of SELFIES format molecular sequences and their corresponding molecular structures and physicochemical properties from publicly available molecular databases. The large chemical model is trained using a three-step reinforcement learning approach: first, a first-stage chemical knowledge reinforcement fine-tuning is performed based on the chemical question-and-answer instruction templates and the question-and-answer dataset; second, a second-stage chemical inference reinforcement fine-tuning is performed based on the molecular description instruction templates and the inference chain dataset; and finally, a third-stage molecular description reinforcement fine-tuning is performed based on the molecular description instruction templates and the molecular description dataset. After model fine-tuning, the user-input SELFIES format molecular sequence is substituted into the molecular description instruction template. The molecular sequence text of the SELFIES segment in the template is configured to obtain the corresponding molecular description instruction. The current molecular description instruction is then input into the large-scale chemical model for molecular description generation. The molecular description text output by the model in this processing is then fed back to the current user. This embodiment of the invention improves the model's understanding and parsing ability of professional chemical knowledge and enhances its chemical reasoning ability through a three-step progressive reinforcement learning approach. Using the large-scale chemical model provided by this embodiment of the invention to process the molecular description generation task can improve the comprehensiveness and chemical accuracy of the generated text. Attached Figure Description
[0015] Figure 1 This is a schematic diagram of a processing method for generating molecular descriptions based on a large chemical model, provided in Embodiment 1 of the present invention. Figure 2 This is a module structure diagram of a processing device for generating molecular descriptions based on a large chemical model, provided in Embodiment 2 of the present invention. Figure 3 This is a schematic diagram of the structure of an electronic device provided in Embodiment 3 of the present invention. Detailed Implementation
[0016] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this invention, and not all embodiments. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.
[0017] Embodiment 1 of this invention provides a processing method for generating molecular descriptions based on large chemical models, such as... Figure 1 The schematic diagram shows a method for generating molecular descriptions based on a large chemical model, as provided in Embodiment 1 of the present invention. This method mainly includes the following steps: Step 1: Select a general-purpose large language model that has completed pre-training for both large language models and general NLP tasks as the corresponding chemical large model.
[0018] Here, the general-purpose large language model in this embodiment of the invention includes at least the Qwen series models, the GPT series models, and the DeepSeek series models. The general-purpose NLP tasks mentioned in this embodiment of the invention include at least text generation tasks, translation tasks, question answering tasks, and thought chain derivation tasks.
[0019] Step 2: Configure the corresponding chemical question-and-answer instruction template and molecular description instruction template for the large chemical model.
[0020] The chemical question-and-answer instruction template of this invention consists of a question-and-answer instruction requirement text and a question text. Wherein: 1) Question and answer instructions require the following text: The question-and-answer instruction in this embodiment of the invention requires the text to be a fixed natural language text, which is used to prompt the chemical big model to generate the corresponding answer based on the given question in the question text.
[0021] It should be noted that the specific text content of the question-and-answer instruction in this embodiment of the invention can be customized based on application requirements or the developer's language habits. For example, "Provide accurate, detailed answers that conform to chemical principles based on the following questions."
[0022] 2) Problematic passage: In this embodiment of the invention, the question segment consists of a fixed question segment title and configurable question text. The question segment title defaults to the string "Question:". The question text is initialized to empty; when the question text is not empty, it is a piece of natural language text used to ask questions about the basic chemical concepts given in the text, or a piece of natural language text used to ask questions about one or more types of molecular features given the molecular name or SELFIES molecular sequence in the text.
[0023] It should be noted that the knowledge scope corresponding to the basic chemical concepts in the embodiments of this invention includes at least: atomic models and quantum numbers, principles of electron configuration and construction, the periodic table, atomic structure and periodic properties of the periodic law, various chemical bond structures and their corresponding force and energy theories, intermolecular forces, chemical thermodynamics theory, chemical reaction theory, acid-base theory, redox theory, functional group theory, polymer theory, organic matter theory, inorganic matter theory, stereochemistry theory, physicochemical property theory, chemical nomenclature rules, quantum chemistry theory, and biological and pharmaceutical chemistry theory.
[0024] It should be noted that the knowledge scope corresponding to the molecular features in the embodiments of the present invention includes at least: basic constituent features, two-dimensional structural features, three-dimensional structural features, physicochemical property features, spectroscopic features, crystallographic features, biological activity features, material performance features, and reaction features. Among these, basic constituent features include at least molecular formula, molecular mass, chemical formula, and isotopic composition. Two-dimensional structural features include at least atomic connection sequence features, chemical bond features, functional group features, and atomic skeleton features. Three-dimensional structural features include at least spatial geometric features, rotational conformation features, chiral features, and geometric isomerism features. Physicochemical property features include at least melting point, boiling point, flash point, solubility, partition coefficient, vapor pressure, polarity, dipole moment, ionization energy, electron affinity, color, optical rotation, chemical stability, and acid / base dissociation constant. Spectroscopic features include at least spectral features and mass spectrometry features.
[0025] The molecular description instruction template of this invention consists of a description instruction requirement section, a SELFIES section, a reasoning step description section, and a molecular description formatting requirement section. Wherein: 1) Description of the instruction requirement text: The description instruction requires the text to be a fixed natural language text, which prompts the chemical large model to perform step-by-step analysis and reasoning based on the SELFIES molecular sequence information given in the SELFIES text, the step-by-step explanation text given in the reasoning steps, and generate a molecular description text that conforms to chemical laws based on the reasoning context.
[0026] It should be noted that the specific text content of the description instruction required in the embodiments of the present invention can be customized based on application requirements or the developer's language habits. For example, "Based on the SELFIES information provided below, perform step-by-step analysis according to the given reasoning steps, and finally generate a natural, concise molecular description text that conforms to chemical laws, and output the obtained molecular description text based on the molecular description formatting requirements."
[0027] 2) Selfies paragraph: The SELFIES passage consists of a fixed sequence passage title and configurable molecular sequence text. The molecular sequence title defaults to the string "SELFIES:". The molecular sequence text is initially empty; when the molecular sequence text is not empty, its text format is a molecular sequence in SELFIES format.
[0028] 3) Passage for explaining reasoning steps: The passage for explaining reasoning steps is a fixed natural language text, consisting of S reasoning step texts, where S is the total number of preset reasoning steps. Each reasoning step text is the step explanation text for one-step reasoning, used to prompt the chemical large model to analyze and reason according to the requirements of this step of reasoning based on the current reasoning context and output the reasoning result of this step as a corresponding single-step reasoning text C s and output, where 1 ≤ index s ≤ S. The current reasoning context includes the SELFIES molecular sequence information given in the SELFIES passage and the reasoning results of all previous historical steps before the current reasoning step. The single-step reasoning text C s=S corresponding to the S-th step is the corresponding molecular description text.
[0029] It should be noted that the specific text content of the passage for explaining the reasoning steps in the embodiments of the present invention can be customized based on application requirements or the developer's language usage habits. For example: the total number of reasoning steps S = 5, and the passage for explaining the reasoning steps is set as "Explanation of reasoning steps: First step, determine the skeleton structure of the molecule; second step, analyze the functional groups of the molecule; third step, analyze the key fragments of the molecule; fourth step, analyze the stereochemical properties of the molecule; fifth step, based on the known derivation results, generate a text describing the molecule, including the molecule name, functional groups, key physical properties, common uses, etc., ensuring that the language is concise and conforms to the common sense of chemical laws".
[0030] 4) Passage for formatting requirements of molecular description: The passage for formatting requirements of molecular description is a fixed natural language text, used to prompt the chemical large model to perform text encapsulation and output on the generated molecular description text according to the preset output format of the molecular description text. The default output format of the molecular description text is sequentially connected by a preset start marker text, the molecular description text generated by the model, and a preset end marker text.
[0031] It should be noted that the specific text content of the passage for formatting requirements of molecular description in the embodiments of the present invention can be customized based on application requirements or the developer's language usage habits. For example, if the start marker text is "<Molecular description>" and the end marker text is "< / Molecular description>", the passage for formatting requirements of molecular description is "Formatting requirements for molecular description: The generated molecular description text must be returned in the following format: <Molecular description>... < / Molecular description>".
[0032] Step 3: Construct a corresponding question-and-answer dataset by collecting big data from publicly available knowledge media in the field of chemistry; and construct a corresponding inference chain dataset and molecular description dataset by collecting big data from publicly available molecular databases of SELFIES format molecular sequences and their corresponding molecular structures and physicochemical properties.
[0033] The current step 3 specifically includes: Step 31: Construct a corresponding question-and-answer dataset by collecting big data from publicly available knowledge media in the field of chemistry.
[0034] Here, the knowledge media in the field of chemistry disclosed in the embodiments of the present invention include at least publicly available textbooks or theoretical books in the field of chemistry, publicly available journals or papers in the field of chemistry, and publicly available databases or molecular libraries in the field of chemistry.
[0035] The question-and-answer dataset of this invention includes a basic chemistry question-and-answer subset and a molecular feature question-and-answer subset; both the basic chemistry question-and-answer subset and the molecular feature question-and-answer subset consist of multiple first data records; the first data record includes a first training question and a first labeled answer; the first training question of the basic chemistry question-and-answer subset is a piece of natural language text used to ask questions about a given basic chemical concept in the text; the first training question of the molecular feature question-and-answer subset is a piece of natural language text used to ask questions about one or more types of molecular features given a molecular name or SELFIES molecular sequence in the text.
[0036] The current step 31 specifically includes: obtaining multiple knowledge entries through big data collection from publicly available chemical professional knowledge media; generating corresponding first training questions and first-label answers to form corresponding first data records based on each knowledge entry through manual or machine questioning and answer labeling; forming a corresponding basic chemistry question and answer subset by manually or machine classifying all the first data records corresponding to basic chemistry questions and answers, forming a corresponding molecular feature question and answer subset by manually or machine classifying all the first data records corresponding to molecular feature questions and answers; and forming a corresponding question and answer dataset by the basic chemistry question and answer subset and the molecular feature question and answer subset.
[0037] It should be noted that the knowledge domain of all knowledge items obtained through big data collection in this embodiment of the invention is greater than or equal to the preset knowledge domain of basic chemical concepts and molecular characteristics; the knowledge domain of basic chemical concepts and molecular characteristics in this embodiment of the invention is the sum of the knowledge domains corresponding to basic chemical concepts and molecular characteristics.
[0038] It should also be noted that when setting the answer for the first tag, you must first ensure that the text content conforms to common sense of chemical laws, and secondly, the natural language expression of the text should be grammatically correct and adopt a concise and refined language style.
[0039] Step 32 involves constructing the corresponding inference chain dataset and molecular description dataset by collecting large amounts of data on SELFIES format molecular sequences and their corresponding molecular structures and physicochemical properties from publicly available molecular databases.
[0040] Here, the publicly disclosed molecular databases in this embodiment of the invention include at least the PubChem database and the PDB database.
[0041] The inference chain dataset of this embodiment includes multiple second data records; each second data record includes a first training sequence and a first label inference chain; the first label inference chain includes S first single-step labels. The first training sequence is a SELFIES molecular sequence.
[0042] The molecular description dataset of this invention includes multiple third data records; each third data record includes a second training sequence and a first label description; the second training sequence is a SELFIES molecular sequence; and the first label description is the molecular description text of the corresponding molecule.
[0043] The current step 32 specifically includes: Step 321: In a public molecular database, large-scale data collection is performed on molecular sequences in SELFIES format. By querying the molecular database, the basic constituent features, two-dimensional structural features, three-dimensional structural features, physicochemical property features, spectroscopic features, crystallographic features, biological activity features, material performance features, and reaction features corresponding to each collected molecular sequence are queried. Based on the query results, corresponding molecular feature records are formed. Each collected molecular sequence and its corresponding molecular feature record form a corresponding first collection entry. All the obtained first collection entries form a corresponding first collection library.
[0044] Step 322: Take each first acquisition entry in the first acquisition library as the current entry; take the molecular sequence of the current entry as a set of corresponding first training sequences and second training sequences; and have professionals in the field of chemistry or other specialized chemical models generate a corresponding single-step reasoning result as the corresponding first single-step label for each reasoning step text specified in the reasoning step description text of the molecular description instruction template, based on the molecular feature record of the current entry. And from the obtained S first single-step labels Form the corresponding first label inference chain; and set the first single-step label corresponding to step S. As the current description text; and in accordance with the molecular description formatting requirements specified in the molecular description instruction template, the molecular description text output format is composed of the start marker text, the current description text, and the end marker text, which are sequentially connected to form the corresponding first label description; and the first training sequence and the first label inference chain corresponding to the current entry form the corresponding second data record, and the second training sequence and the first label description corresponding to the current entry form the corresponding third data record; and all the obtained second data records form the corresponding inference chain dataset, and all the obtained third data records form the corresponding molecular description dataset.
[0045] It should be noted that, in the first single-step label When setting the first tag description, first ensure that the text content conforms to common sense of chemical laws, and secondly, require that the natural language expression of the text be grammatically correct and adopt a concise and refined language style.
[0046] Step 4: First, perform a first-stage fine-tuning of chemical knowledge enhancement on the large chemical model based on the chemical question-and-answer instruction template and question-and-answer dataset; then, perform a second-stage fine-tuning of chemical reasoning enhancement on the large chemical model based on the molecular description instruction template and reasoning chain dataset; finally, perform a third-stage fine-tuning of molecular description enhancement on the large chemical model based on the molecular description instruction template and molecular description dataset.
[0047] The current step 4 specifically includes: Step 41: First, perform a phase of chemical knowledge reinforcement and fine-tuning on the large chemical model based on the chemical question-and-answer instruction template and question-and-answer dataset.
[0048] Specifically, it includes: Step 411: Use the basic chemical question-and-answer subset of the question-and-answer dataset as the current dataset.
[0049] Step 412: Divide the current dataset into multiple first data batches based on the preset batch size B1; and take the first first data batch of the current dataset as the current data batch.
[0050] Here, the batch size B1 in this embodiment of the invention is a pre-set positive integer. Each first data batch includes B1 first data records; the first label answer of each first data record in each first data batch is denoted as the corresponding label answer. 1 ≤ index j ≤ B1; answers for each label The total number of word segments is denoted as n. j Answers under various tags Each word segment is recorded as the corresponding 1 ≤ index k ≤ n j .
[0051] Step 413: Substitute the first training question of each first data record in the current data batch into the chemical question-and-answer instruction template, configure the question text of the template's question segment, and obtain the corresponding current chemical question-and-answer instruction Q. j ; and will the current chemical question-and-answer command Q j The chemical large-scale model is input for question-answering task processing; the autoregressive text generation process of the chemical large-scale model is recorded during this processing; and after completing B1 model processing iterations, the first model loss function L is used as the basis for the calculation. M1 Calculate the corresponding first loss value.
[0052] Here, the first model loss function L in this embodiment of the invention M1 It is implemented based on the cross-entropy loss function, specifically as follows: ; Where θ is used to identify the model parameters of the large chemical model; For tagged answers The word segmentation sequence preceding the kth word; For the model in the autoregressive text generation process, the current chemical question-answering instruction Q is used. j and word segmentation sequence The k-th word generated for the context is the corresponding word. The probability of.
[0053] Step 414: Identify whether the first loss value meets the preset first loss value range; if not, then based on the preset first model optimizer, move towards making the first model loss function L... M1 The direction that reaches the minimum value is used to fine-tune the model parameters of the large chemical model, and after this round of fine-tuning is completed, return to step 413; if satisfied, proceed to step 415.
[0054] Here, the first loss value range in this embodiment of the invention is a pre-set numerical range. The first model optimizer includes the Adam optimizer, the SGD optimizer, and the AdamW optimizer.
[0055] Step 415: Identify whether the current data batch is the last first data batch of the current dataset; if not, take the next first data batch of the current dataset as the new current data batch and return to step 413; if yes, identify whether the current dataset is a subset of basic chemical questions and answers; if yes, take the molecular feature question and answer subset of the question and answer dataset as the new current dataset and return to step 412; otherwise, stop training and confirm that the chemical knowledge enhancement fine-tuning training of the model has ended.
[0056] Step 42: Then, perform a two-stage chemical reasoning enhancement and fine-tuning on the large chemical model based on the molecular description instruction template and the reasoning chain dataset.
[0057] Specifically, it includes: Step 421: Divide the inference chain dataset into multiple second data batches based on the preset batch size B2; and use the first second data batch as the current data batch.
[0058] Here, the batch size B2 in this embodiment of the invention is a pre-set positive integer. Each second data batch includes B2 second data records; the first tag inference chain of each second data record in each second data batch is denoted as the corresponding tag chain. 1 ≤ index q ≤ B2; each tag chain Each first single-step label Record as the corresponding single-step label Each single-step label The total number of word segments is denoted as n. q,s Each single-step label Each word segment is recorded as the corresponding 1 ≤ index u ≤ n q,s .
[0059] Step 422, extract the sequence X of each second data record in the current data batch. q Substituting the molecular description instruction template into the template and configuring the molecular sequence text of the SELFIES segment yields the corresponding current molecular description instruction X. q ; and the current molecular description instruction X q The chemical large model is input for molecular description generation; the autoregressive text generation process of the chemical large model is recorded during this process; and after completing B2 model processing iterations, the pre-set second model loss function L is used. M2 Calculate the corresponding second loss value.
[0060] Here, the second model loss function L in this embodiment of the invention M2 It is implemented based on the cross-entropy loss function, specifically as follows: ; θ is used to identify the model parameters of large chemical models; It is a sequence of single-step tags formed by concatenating the first s-1 single-step tags in sequence; For single-step tags The word segmentation sequence preceding the u-th word; For the autoregressive generation step of the inference text in step s of the model, the corresponding molecular description instruction X is used.q Single-step label sequence and word segmentation sequence The u-th word generated for the context is the corresponding word. The probability of.
[0061] Step 423: Identify whether the second loss value meets the preset range of the second loss value; if not, then based on the preset second model optimizer, move towards making the second model loss function L... M2 The direction that reaches the minimum value is used to fine-tune the model parameters of the large chemical model in one round, and the process returns to step 422 after this round of fine-tuning. If the condition is met, it is determined whether the current data batch is the last second data batch. If not, the next second data batch is used as the new current data batch and the process returns to step 422. If the condition is met, training is stopped and the chemical inference enhancement fine-tuning training of the model is confirmed to be completed.
[0062] Here, the second loss value range in this embodiment of the invention is a pre-set numerical range. The second model optimizer includes the Adam optimizer, the SGD optimizer, and the AdamW optimizer.
[0063] Step 43: Finally, the chemical large model is fine-tuned in three stages to enhance molecular description based on the molecular description instruction template and the molecular description dataset.
[0064] Specifically, this includes: Step 431, replicating the large chemical model to obtain two replicated models, and denoting the current large chemical model as the corresponding new strategy model M. new The two replication models are denoted as the corresponding old policy models M. old Reference Model M ref and the reference model M ref The model parameters are solidified; and the new strategy model M is then implemented. new Old strategy model M old Reference Model M ref The model parameters are denoted as the corresponding model parameters θ. new Model parameters θ old Model parameters θ ref The molecular description dataset is divided into multiple third data batches based on the preset batch size B3, and the first third data batch is used as the current data batch.
[0065] Here, the batch size B3 in this embodiment of the invention is a pre-set positive integer. Each third data batch includes B3 third data records; the second training sequence of each third data record in each third data batch is denoted as the corresponding sequence X. e The first tag description is denoted as the corresponding tag text. , 1≤indexe≤B3.
[0066] Step 432, extract the sequence X of each third data record in the current data batch. e The molecular sequence text of the SELFIES segment in the template is configured by substituting the molecular description instruction template to obtain the corresponding current molecular description instruction. The current molecular description instruction is then input into the old strategy model M G times consecutively. old The molecular description generation task is performed to obtain G formatted molecular description texts, and these G formatted molecular description texts are respectively recorded as the corresponding predicted texts. 1 ≤ index g ≤ G; and each predicted text The total number of word segments is recorded as the corresponding , and the predicted text Each word segment is recorded as the corresponding 1≤index t≤ ; and each word segmentation The corresponding predicted probability is denoted as .
[0067] Wherein, the number of repetitions G is a preset positive integer; Step 433: Based on the text format check results and the language similarity assessment results between the predicted text and its corresponding tag text, perform a process for each predicted text. The corresponding text reward R e,g Perform the calculation.
[0068] Specifically, this involves: processing each predicted text. This serves as the current predicted text; it also identifies whether the text format of the current predicted text meets the molecular description formatting requirements specified in the molecular description instruction template. If it does, it sets the corresponding format check result f. e,g If the condition is not met, the corresponding format check result f is set to 1. e,g The result is 0; and the format check result f is... e,g If the value is 0, then perform recognition; if so, set the corresponding text reward R. e,g If the value is 0, then the current predicted text and its corresponding tag text are compared based on a preset six-class text similarity evaluation algorithm. The text similarity was evaluated to obtain six corresponding evaluation values, and the average of the six evaluation values was taken as the corresponding language similarity. And based on language similarity Calculate the corresponding text reward R e,g .
[0069] Here, the six types of text similarity evaluation algorithms in this embodiment of the invention include BLEU-2 evaluation algorithm, BLEU-4 evaluation algorithm, METEOR evaluation algorithm, ROUGE-1 evaluation algorithm, ROUGE-2 evaluation algorithm, and ROUGE-L evaluation algorithm.
[0070] Text reward R in embodiments of the present invention e,g The calculation method is as follows: .
[0071] Step 434: Generate the G predicted texts corresponding to each third data record in the current data batch. Cluster them into groups; and calculate the average reward µ for each group. e and standard deviation σ e Perform calculations; and based on each predicted text in each group The corresponding text reward R e,g Average value µ e and standard deviation σ e Calculate the corresponding within-group advantage V e,g .
[0072] Here, the average value µ in the embodiments of the present invention e Standard deviation σ e and group advantages V e,g The calculation method is as follows: , ; ; Where λ1 is a preset small constant used to prevent the denominator from being zero.
[0073] It should be noted that the evaluation value output by each type of text similarity evaluation algorithm in this embodiment of the invention is a score between 0 and 1. The average of the six evaluation values is the language similarity. It is also a numerical value between 0 and 1. In the format check result f e,g When = 0, the text reward R e,g The result is 0, in the format check result f e,g When =1, the text reward R is... e,g This is a numerical value between 0.5 and 2. The text reward R given in this embodiment of the invention is based on text format checking and six categories of text similarity evaluation. e,gIt is used to comprehensively reward the text description text for compliance with text format, chemical accuracy of text content, and conciseness of natural language: non-compliant text format will receive 0 points, that is, no reward. If the text format is compliant, but the chemical accuracy or conciseness of the text content is insufficient, the reward score will also be reduced. Only when the text format is compliant and the chemical accuracy and conciseness of the text content are sufficient will a larger reward score be obtained.
[0074] Step 435, extract the sequence X of each third data record in the current data batch. e The molecular sequence text of the SELFIES segment in the template is configured by substituting the molecular description instruction template to obtain the corresponding current molecular description instruction, and the current molecular description instruction is input into the reference model M G times consecutively. ref Perform molecular description generation task processing; and in the g-th task processing corresponding to the e-th third data record, match the word probability vector corresponding to the t-th word segment generated in this step with the corresponding predicted text. The t-th word The corresponding word segmentation probabilities are extracted and used as the corresponding prediction probabilities. .
[0075] Step 436: Sequence X of each third data record in the current data batch. e The molecular sequence text of the SELFIES segment in the template is configured by substituting the molecular description instruction template to obtain the corresponding current molecular description instruction, and the current molecular description instruction is input into the new strategy model M G times consecutively. new Perform molecular description generation task processing; and in the g-th task processing corresponding to the e-th third data record, match the word probability vector corresponding to the t-th word segment generated in this step with the corresponding predicted text. The t-th word The corresponding word segmentation probabilities are extracted and used as the corresponding prediction probabilities. .
[0076] Step 437, for each predicted text Each word Importance sampling ratio r e,g,t And truncation sampling ratio Perform the calculation.
[0077] Here, the importance sampling ratio r of the embodiments of the present invention e,g,t And truncation sampling ratio The calculation method is as follows: , ; Where λ2 is the preset truncation threshold, which takes a value between 0 and 0.5, for example, 0.2. Clip() is the Clip truncation function.
[0078] Step 438, based on the preset third model loss function L M3 The corresponding third loss value is obtained through calculation.
[0079] Here, the third model loss function L in this embodiment of the invention M3 The objective function J based on the group relative strategy optimization algorithm GRPO The objective function J of the group relative strategy optimization algorithm is implemented. GRPO Then, by the policy function term J policy , divergence function term J KL The composition is as follows: , , .
[0080] Where β is the preset divergence coefficient.
[0081] Step 439: Identify whether the third loss value meets the preset range of the third loss value; if the third loss value does not meet the range of the third loss value, then based on the preset third model optimizer, move towards making the third model loss function L... M3 The direction that reaches the minimum value corresponds to the new strategy model M new Model parameters θ new Perform one round of modulation and return to step 432 at the end of this round of modulation; if the third loss value meets the range of the third loss value, then based on the model parameter θ new For the old strategy model M old Model parameters θ old Perform a reset and identify whether the current data batch is the last third data batch. If not, use the next third data batch as the new current data batch and return to step 432. If yes, stop training, confirm that the molecular description enhancement fine-tuning training of the model is complete, and set the new strategy model M. new As the latest large-scale chemical model.
[0082] Here, the third loss value range in this embodiment of the invention is a pre-set numerical range. The third model optimizer includes the Adam optimizer and the AdamW optimizer.
[0083] Step 5: After the model fine-tuning is completed, the user-input SELFIES format molecular sequence is substituted into the molecular description instruction template to configure the molecular sequence text of the SELFIES segment in the template to obtain the corresponding current molecular description instruction; the current molecular description instruction is then input into the large chemical model for molecular description generation task processing; and the molecular description text output by the model in this processing is fed back to the current user.
[0084] Figure 2 This is a module structure diagram of a processing device for generating molecular descriptions based on a large chemical model, provided in Embodiment 2 of the present invention. This device can be a terminal device or server implementing the aforementioned method embodiments, or it can be a device that enables the aforementioned terminal device or server to implement the aforementioned method embodiments. For example, the device can be a device or chip system of the aforementioned terminal device or server. Figure 2 As shown, the device includes: a language model selection module 201, an instruction template configuration module 202, a data acquisition module 203, a three-stage reinforcement module 204, and a model application module 205.
[0085] The language model selection module 201 is used to select a general-purpose large language model that has completed pre-training of a large language model and pre-training of a general NLP task as the corresponding chemical large model; the general-purpose large language model includes at least the Qwen series model, the GPT series model, and the DeepSeek series model; the general-purpose NLP task includes at least the text generation task, the translation task, the question answering task, and the thought chain derivation task.
[0086] The instruction template configuration module 202 is used to configure corresponding chemical question and answer instruction templates and molecular description instruction templates for large chemical models.
[0087] The data acquisition module 203 is used to construct the corresponding question-and-answer dataset by collecting big data from publicly available knowledge media in the field of chemistry; and to construct the corresponding inference chain dataset and molecular description dataset by collecting big data from publicly available molecular databases of SELFIES format molecular sequences and their corresponding molecular structures and physicochemical properties.
[0088] The three-stage enhancement module 204 is used to first perform a first-stage chemical knowledge enhancement and fine-tuning of the large chemical model based on the chemical question-and-answer instruction template and question-and-answer dataset; then perform a second-stage chemical reasoning enhancement and fine-tuning of the large chemical model based on the molecular description instruction template and reasoning chain dataset; and finally perform a third-stage molecular description enhancement and fine-tuning of the large chemical model based on the molecular description instruction template and molecular description dataset.
[0089] The model application module 205 is used to, after the model fine-tuning is completed, substitute the user-input SELFIES format molecular sequence into the molecular description instruction template, configure the molecular sequence text of the SELFIES segment in the template to obtain the corresponding current molecular description instruction; input the current molecular description instruction into the large chemical model for molecular description generation task processing; and feed back the molecular description text output by the model in this processing to the current user.
[0090] The present invention provides a processing device for generating molecular descriptions based on a large chemical model, which can execute the method steps in the above method embodiments. Its implementation principle and technical effect are similar, and will not be repeated here.
[0091] It should be noted that the division of the various modules in the above device is merely a logical functional division. In actual implementation, they can be fully or partially integrated into a single physical entity, or they can be physically separated. Furthermore, these modules can be implemented entirely in software via processing element calls; they can be fully implemented in hardware; or some modules can be implemented by processing element calls to software, while others are implemented in hardware. For example, the language model selection module can be a separate processing element, or it can be integrated into a chip in the above device. Alternatively, it can be stored as program code in the memory of the above device, and called and executed by a processing element of the device. The implementation of other modules is similar. Moreover, these modules can be fully or partially integrated together, or they can be implemented independently. The processing element described here can be an integrated circuit with signal processing capabilities. In the implementation process, each step of the above method or each of the above modules can be completed through integrated logic circuits in the hardware of the processor element or through software instructions.
[0092] For example, these modules can be one or more integrated circuits configured to implement the above methods, such as one or more Application Specific Integrated Circuits (ASICs), one or more Digital Signal Processors (DSPs), or one or more Field Programmable Gate Arrays (FPGAs). As another example, when a module is implemented using processing element scheduler code, the processing element can be a general-purpose processor, such as a Central Processing Unit (CPU) or other processor capable of calling program code. Furthermore, these modules can be integrated together as a System-on-a-Chip (SOC).
[0093] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. This computer program product includes one or more computer instructions. When these computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the foregoing method embodiments are generated. The computer described above can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The aforementioned computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the aforementioned computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, Digital Subscriber Line (DSL)) or wireless (e.g., infrared, wireless, Bluetooth, microwave, etc.) means. The aforementioned computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The aforementioned available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state disks (SSDs)).
[0094] Figure 3 This is a schematic diagram of an electronic device provided in Embodiment 3 of the present invention. This electronic device can be a terminal device or server implementing the methods of the aforementioned embodiments, or it can be a terminal device or server connected to the aforementioned terminal device or server implementing the methods of the aforementioned embodiments. Figure 3 As shown, the electronic device may include: a processor 301 (e.g., CPU), a memory 302, and a transceiver 303; the transceiver 303 is coupled to the processor 301, and the processor 301 controls the transmission and reception operations of the transceiver 303. The memory 302 may store various instructions for performing various processing functions and implementing the processing steps described in the foregoing embodiments. Preferably, the electronic device involved in the embodiments of the present invention further includes: a power supply 304, a system bus 305, and a communication port 306. The system bus 305 is used to realize communication connections between components. The communication port 306 is used for communication between the electronic device and other peripherals.
[0095] exist Figure 3The system bus 305 mentioned can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This system bus can be divided into address bus, data bus, control bus, etc. For ease of representation, it is represented by only one thick line in the figure, but this does not indicate that there is only one bus or one type of bus. The communication interface is used to enable communication between the database access device and other devices (e.g., clients, read-write libraries, and read-only libraries). Memory may include Random Access Memory (RAM) and may also include Non-Volatile Memory, such as at least one disk storage device.
[0096] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), graphics processing units (GPUs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0097] It should be noted that the embodiments of the present invention also provide a computer-readable storage medium storing instructions that, when run on a computer, cause the computer to perform the methods and processes provided in the above embodiments.
[0098] This invention provides a method, apparatus, electronic device, and computer-readable storage medium for generating molecular descriptions based on a large chemical model. As described above, this invention selects a general-purpose large language model that has completed pre-training for both large language models and general NLP tasks as the large chemical model; and configures corresponding chemical question-and-answer instruction templates and molecular description instruction templates for it. A corresponding question-and-answer dataset is constructed by collecting large amounts of publicly available chemical knowledge; and a corresponding inference chain dataset and molecular description dataset are constructed by collecting large amounts of SELFIES format molecular sequences and their corresponding molecular structures and physicochemical properties from publicly available molecular databases. The large chemical model is trained using a three-step reinforcement learning approach: first, a first-stage chemical knowledge reinforcement fine-tuning is performed based on the chemical question-and-answer instruction templates and the question-and-answer dataset; second, a second-stage chemical inference reinforcement fine-tuning is performed based on the molecular description instruction templates and the inference chain dataset; and finally, a third-stage molecular description reinforcement fine-tuning is performed based on the molecular description instruction templates and the molecular description dataset. After model fine-tuning, the user-input SELFIES format molecular sequence is substituted into the molecular description instruction template. The molecular sequence text of the SELFIES segment in the template is configured to obtain the corresponding molecular description instruction. The current molecular description instruction is then input into the large-scale chemical model for molecular description generation. The molecular description text output by the model in this processing is then fed back to the current user. This embodiment of the invention improves the model's understanding and parsing ability of professional chemical knowledge and enhances its chemical reasoning ability through a three-step progressive reinforcement learning approach. Using the large-scale chemical model provided by this embodiment of the invention to process the molecular description generation task can improve the comprehensiveness and chemical accuracy of the generated text.
[0099] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented in hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.
[0100] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A processing method for generating molecular descriptions based on large chemical models, characterized in that, The method includes: Select a general-purpose language model that has completed pre-training for both large language models and general NLP tasks as the corresponding chemical large model; the general-purpose language model includes at least the Qwen series models, GPT series models, and DeepSeek series models; the general-purpose NLP tasks include at least text generation tasks, translation tasks, question answering tasks, and thought chain deduction tasks; Configure corresponding chemical question-and-answer instruction templates and molecular description instruction templates for the large chemical model; We construct corresponding question-and-answer datasets by collecting big data from publicly available knowledge media in the field of chemistry; and we construct corresponding inference chain datasets and molecular description datasets by collecting big data from publicly available molecular databases of SELFIES format molecular sequences and their corresponding molecular structures and physicochemical properties. First, the chemical big model is fine-tuned in one stage by enhancing chemical knowledge based on the chemical question-and-answer instruction template and the question-and-answer dataset; then, the chemical big model is fine-tuned in two stages by enhancing chemical reasoning based on the molecular description instruction template and the reasoning chain dataset; finally, the chemical big model is fine-tuned in three stages by enhancing molecular description based on the molecular description instruction template and the molecular description dataset. After the model fine-tuning is completed, the user-input SELFIES format molecular sequence is substituted into the molecular description instruction template to configure the molecular sequence text of the SELFIES segment in the template to obtain the corresponding current molecular description instruction; the current molecular description instruction is then input into the large chemical model for molecular description generation task processing; and the molecular description text output by the model in this processing is fed back to the current user.
2. The processing method for generating molecular descriptions based on large chemical models according to claim 1, characterized in that, The chemical Q&A instruction template consists of a Q&A instruction requirement text and a question text; The question-and-answer instruction requires the text to be a fixed natural language text, which is used to prompt the chemical big model to generate a corresponding answer based on the given question in the question text; The question segment consists of a fixed question segment title and configurable question text; the question segment title defaults to the string "Question:"; the question text is initialized to empty; when the question text is not empty, it is a piece of natural language text used to ask questions about the basic chemical concepts given in the text or a piece of natural language text used to ask questions about one or more types of molecular features given the molecular name or SELFIES molecular sequence in the text. The knowledge scope corresponding to basic chemical concepts includes at least the atomic model and quantum numbers, the principle of electron configuration and structure, the periodic table, the atomic structure and periodic properties of the periodic law, various chemical bond structures and their corresponding force and energy theories, intermolecular forces, chemical thermodynamics theory, chemical reaction theory, acid-base theory, redox theory, functional group theory, polymer theory, organic matter theory, inorganic matter theory, stereochemistry theory, physicochemical property theory, chemical nomenclature rules, quantum chemistry theory, and biological and pharmaceutical chemistry theory. The knowledge scope corresponding to molecular characteristics includes at least the basic composition characteristics, two-dimensional structural characteristics, three-dimensional structural characteristics, physicochemical property characteristics, spectroscopic characteristics, crystallographic characteristics, biological activity characteristics, material performance characteristics, and reaction characteristics; The basic structural features include at least molecular formula, molecular mass, chemical formula, and isotopic composition; the two-dimensional structural features include at least atomic connection sequence, chemical bond, functional group, and atomic skeleton features; the three-dimensional structural features include at least spatial geometric features, rotational conformation features, chiral features, and geometric isomerism features; the physicochemical properties include at least melting point, boiling point, flash point, solubility, partition coefficient, vapor pressure, polarity, dipole moment, ionization energy, electron affinity, color, optical rotation, chemical stability, and acid / base dissociation constant; the spectroscopic features include at least spectral features and mass spectrometry features. The molecular description instruction template consists of a description instruction requirement section, the SELFIES section, a reasoning step description section, and a molecular description formatting requirement section. The description instruction requires the text segment to be a fixed natural language text, which is used to prompt the chemical big model to perform step-by-step parsing and reasoning according to the SELFIES molecular sequence information given in the SELFIES text segment, the step sequence given in the reasoning step description text segment, and generate a molecular description text that conforms to chemical laws based on the reasoning context. The SELFIES segment consists of a fixed sequence segment title and configurable molecular sequence text; the molecular sequence title defaults to the string "SELFIES:"; the molecular sequence text is initialized to empty; when the molecular sequence text is not empty, its text format is a molecular sequence in SELFIES format; The inference step description is a fixed natural language text consisting of S inference step texts, where S is the preset total number of inference steps. Each inference step text is a step description text for one step of inference, used to prompt the chemical macro-model to perform analysis and inference according to the requirements of this step and the current inference context, and to take the result of this step as a corresponding single-step inference text C. s The output is 1 ≤ index s ≤ S; the current inference context includes the SELFIES molecular sequence information given by the SELFIES text segment, the inference results of all historical steps before the current inference step, and the single-step inference text C corresponding to step S. s=S This is the corresponding molecular description text; The molecular description formatting requirement is a fixed natural language text segment, which is used to prompt the chemical model to encapsulate and output the generated molecular description text according to the preset molecular description text output format; the molecular description text output format is formed by connecting the preset start marker text, the molecular description text generated by the model, and the preset end marker text in sequence by default. The publicly available knowledge media in the field of chemistry include at least publicly available textbooks or theoretical books in the field of chemistry, publicly available journals or papers in the field of chemistry, and publicly available databases or molecular libraries in the field of chemistry. The publicly available molecular databases include at least the PubChem database and the PDB database; The question-and-answer dataset includes a subset of basic chemistry questions and answers and a subset of molecular feature questions and answers; Both the basic chemistry question-and-answer subset and the molecular feature question-and-answer subset consist of multiple first data records; the first data record includes a first training question and a first labeled answer; the first training question of the basic chemistry question-and-answer subset is a piece of natural language text used to ask questions about a given basic chemical concept in the text; the first training question of the molecular feature question-and-answer subset is a piece of natural language text used to ask questions about one or more types of molecular features given a molecular name or SELFIES molecular sequence in the text; The inference chain dataset includes multiple second data records; the second data records include a first training sequence and a first label inference chain; the first label inference chain includes S first single-step labels. The first training sequence is a SELFIES molecular sequence. The molecular description dataset includes multiple third data records; each third data record includes a second training sequence and a first label description; the second training sequence is a SELFIES molecular sequence; the first label description is the molecular description text of the corresponding molecule.
3. The processing method for generating molecular descriptions based on large chemical models according to claim 2, characterized in that, The step of performing a first-stage chemical knowledge enhancement and fine-tuning of the large chemical model based on the chemical question-and-answer instruction template and the question-and-answer dataset specifically includes: Step 31: Use the basic chemical question-and-answer subset of the question-and-answer dataset as the current dataset; Step 32: Divide the current dataset into multiple first data batches based on a preset batch size B1; and take the first first data batch of the current dataset as the current data batch; Each first data batch includes B1 first data records; the first tag answer of each first data record in each first data batch is denoted as the corresponding tag answer. , 1 ≤ index j ≤ B1; each of the stated label answers The total number of word segments is denoted as n. j Each of the aforementioned tagged answers Each word segment is recorded as the corresponding 1 ≤ index k ≤ n j ; Step 33: Substitute the first training question of each first data record in the current data batch into the chemical question-and-answer instruction template, and configure the question text of the question segment in the template to obtain the corresponding current chemical question-and-answer instruction Q. j ; and the current chemical question-and-answer command Q j The chemical model is input for question-answering task processing; the autoregressive text generation process of the chemical model during this processing is recorded; and after completing B1 model processing iterations, the first model loss function L is used as the basis for the calculation. M1 Calculate the corresponding first loss value; Wherein, the first model loss function L M1 It is implemented based on the cross-entropy loss function, specifically as follows: ; θ is used to identify the model parameters of the large chemical model; For the tagged answer The word segmentation sequence preceding the kth word; For the model in the autoregressive text generation process, the current chemical question-answering instruction Q is used. j and word segmentation sequence The k-th word generated for the context is the corresponding word. The probability of; Step 34: Identify whether the first loss value meets the preset first loss value range; if not, then based on the preset first model optimizer, move towards making the first model loss function L... M1 The model parameters of the large chemical model are fine-tuned in the direction that reaches the minimum value, and the process returns to step 33 after this round of fine-tuning is completed; if the condition is met, proceed to step 35. The first model optimizer includes the Adam optimizer, the SGD optimizer, and the AdamW optimizer. Step 35: Identify whether the current data batch is the last first data batch of the current dataset; if not, take the next first data batch of the current dataset as the new current data batch and return to step 33; if yes, identify whether the current dataset is the basic chemical question-and-answer subset; if yes, take the molecular feature question-and-answer subset of the question-and-answer dataset as the new current dataset and return to step 32; otherwise, stop training and confirm that the chemical knowledge enhancement fine-tuning training of the model has ended.
4. The processing method for generating molecular descriptions based on large chemical models according to claim 2, characterized in that, The step of performing a two-stage chemical inference enhancement and fine-tuning of the large chemical model based on the molecular description instruction template and the inference chain dataset specifically includes: Step 41: Divide the inference chain dataset into multiple second data batches based on a preset batch size B2; and use the first second data batch as the current data batch; Each second data batch includes B2 second data records; the first tag inference chain of each second data record in each second data batch is denoted as the corresponding tag chain. , 1 ≤ index q ≤ B2; each of the stated tag chains Each of the first single-step tags Record as the corresponding single-step label Each of the aforementioned single-step labels The total number of word segments is denoted as n. q,s Each of the aforementioned single-step labels Each word segment is recorded as the corresponding 1 ≤ index u ≤ n q,s ; Step 42, set the sequence X of each of the second data records in the current data batch. q Substituting the molecular description instruction template into the template, the molecular sequence text of the SELFIES segment is configured to obtain the corresponding current molecular description instruction X. q ; and the current molecular description instruction X q The chemical large model is input for molecular description generation; the autoregressive text generation process of the chemical large model is recorded during this process; and after completing B2 model processing iterations, a second model loss function L is used as the basis for the process. M2 Calculate the corresponding second loss value; Wherein, the second model loss function L M2 It is implemented based on the cross-entropy loss function, specifically as follows: ; θ is used to identify the model parameters of the large chemical model; It is a sequence of single-step tags formed by concatenating the first s-1 single-step tags in sequence; For the single-step label The word segmentation sequence preceding the u-th word; For the autoregressive generation step of the inference text in step s of the model, the corresponding molecular description instruction X is used. q Single-step label sequence and word segmentation sequence The u-th word generated for the context is the corresponding word. The probability of; Step 43: Identify whether the second loss value meets the preset second loss value range; if not, then based on the preset second model optimizer, move towards making the second model loss function L... M2 The model parameters of the chemical model are fine-tuned in the direction of reaching the minimum value, and the process returns to step 42 after the fine-tuning is completed. If the condition is met, it is identified whether the current data batch is the last second data batch. If not, the next second data batch is taken as the new current data batch and the process returns to step 42. If the condition is met, training is stopped and the chemical inference enhancement fine-tuning training of the model is confirmed to be completed. The second model optimizer includes the Adam optimizer, the SGD optimizer, and the AdamW optimizer.
5. The processing method for generating molecular descriptions based on large chemical models according to claim 2, characterized in that, The three-stage molecular description enhancement and fine-tuning of the large chemical model based on the molecular description instruction template and the molecular description dataset specifically includes: Step 51: The chemical macro model is replicated to obtain two replicated models, and the current chemical macro model is denoted as the corresponding new strategy model M. new The two replication models are denoted as the corresponding old policy models M. old Reference Model M ref and the reference model M ref The model parameters are solidified; and the new strategy model M is then... new The old strategy model M old The reference model M ref The model parameters are denoted as the corresponding model parameters θ. new Model parameters θ old Model parameters θ ref The molecular description dataset is divided into multiple third data batches based on a preset batch size B3, and the first third data batch is taken as the current data batch. Each of the third data batches includes B3 third data records; the second training sequence of each third data record in each of the third data batches is denoted as the corresponding sequence X. e The first tag description is denoted as the corresponding tag text. , 1 ≤ index e ≤ B3; Step 52, record the sequence X of each of the third data records in the current data batch. e The molecular sequence text of the SELFIES segment in the template is configured by substituting the molecular description instruction template to obtain the corresponding current molecular description instruction, and the current molecular description instruction is input into the old strategy model M consecutively G times. old The molecular description generation task is performed to obtain G formatted molecular description texts, and these G formatted molecular description texts are respectively recorded as the corresponding predicted texts. 1 ≤ index g ≤ G; and each of the predicted texts The total number of word segments is recorded as the corresponding Each of the predicted texts Each word segment is recorded as the corresponding 1≤index t≤ ; and each word segmentation The corresponding predicted probability is denoted as ; Wherein, the number of repetitions G is a preset positive integer; Step 53: Based on the text format check results and the language similarity evaluation results between the predicted text and its corresponding tag text, process each of the predicted texts. The corresponding text reward R e,g Perform calculations; Specifically, this involves: processing each of the predicted texts The current predicted text is used as the basis for identification; and the text format of the current predicted text is checked to see if it meets the molecular description formatting requirements specified in the molecular description instruction template. If it does, the corresponding format check result f is set. e,g If the condition is not met, the corresponding format check result f is set to 1. e,g The value is 0; and the format check result f is... e,g If the value is 0, then the corresponding text reward R is set. e,g If the value is 0, then the current predicted text and its corresponding tag text are compared based on a preset six-class text similarity evaluation algorithm. The text similarity was evaluated to obtain six corresponding evaluation values, and the average of the six evaluation values was taken as the corresponding language similarity. And based on the language similarity Calculate the corresponding text reward R e,g ; The six types of text similarity evaluation algorithms include BLEU-2 evaluation algorithm, BLEU-4 evaluation algorithm, METEOR evaluation algorithm, ROUGE-1 evaluation algorithm, ROUGE-2 evaluation algorithm, and ROUGE-L evaluation algorithm. The text reward R e,g The calculation method is as follows: ; Step 54: Generate the G predicted texts corresponding to each of the third data records in the current data batch. Cluster them into groups; and calculate the average reward µ for each group. e and standard deviation σ e Perform calculations; and based on each of the predicted texts in each group. The corresponding text reward R e,g The average value µ e and the standard deviation σ e Calculate the corresponding within-group advantage V e,g ; Wherein, the average value µ e The standard deviation σ e and the aforementioned intra-group advantage V e,g The calculation method is as follows: , ; ; λ1 is a preset small constant used to prevent the denominator from being zero; Step 55, record the sequence X of each of the third data records in the current data batch. e The molecular sequence text of the SELFIES segment in the template is configured by substituting the molecular description instruction template to obtain the corresponding current molecular description instruction, and the current molecular description instruction is input into the reference model M G times consecutively. ref Perform molecular description generation task processing; and during the g-th task processing corresponding to the e-th third data record, match the word probability vector corresponding to the t-th segment generated in this processing with the corresponding predicted text. The t-th segmentation The corresponding word segmentation probabilities are extracted and used as the corresponding prediction probabilities. ; Step 56, record the sequence X of each of the third data records in the current data batch. e The molecular sequence text of the SELFIES segment in the template is configured by substituting the molecular description instruction template to obtain the corresponding current molecular description instruction, and the current molecular description instruction is input into the new strategy model M G times consecutively. new Perform molecular description generation task processing; and during the g-th task processing corresponding to the e-th third data record, match the word probability vector corresponding to the t-th segment generated in this processing with the corresponding predicted text. The t-th segmentation The corresponding word segmentation probabilities are extracted and used as the corresponding prediction probabilities. ; Step 57, for each of the predicted texts Each of the aforementioned words Importance sampling ratio r e,g,t And truncation sampling ratio Perform calculations; Wherein, the importance sampling ratio r e,g,t and the truncation sampling ratio The calculation method is as follows: , ; λ2 is the preset truncation threshold, and Clip() is the Clip truncation function; Step 58, based on the preset third model loss function L M3 The corresponding third loss value is obtained through calculation; Wherein, the third model loss function L M3 The objective function J based on the group relative strategy optimization algorithm GRPO To achieve this, the objective function J of the group relative strategy optimization algorithm is... GRPO Then, by the policy function term J policy , divergence function term J KL The composition is as follows: , , ; β is the preset divergence coefficient; Step 59: Identify whether the third loss value meets a preset third loss value range; if the third loss value does not meet the third loss value range, then based on a preset third model optimizer, move towards making the third model loss function L... M3 The direction for reaching the minimum value corresponds to the new strategy model M. new The model parameters θ new Perform one round of modulation, and return to step 52 at the end of this round of modulation; if the third loss value meets the range of the third loss value, then based on the model parameter θ new For the old strategy model M old The model parameters θ old Perform a reset and identify whether the current data batch is the last third data batch. If not, use the next third data batch as the new current data batch and return to step 52. If yes, stop training, confirm that the molecular description enhancement fine-tuning training of the model is complete, and set the new strategy model M. new As the latest chemical large model; wherein, the third model optimizer includes the Adam optimizer and the AdamW optimizer.
6. An apparatus for performing the processing method based on a large chemical model to generate molecular descriptions as described in any one of claims 1-5, characterized in that, The device includes: a language model selection module, an instruction template configuration module, a data acquisition module, a three-stage reinforcement module, and a model application module; The language model selection module is used to select a general-purpose language model that has completed pre-training for both large language models and general NLP tasks as the corresponding chemical large model; the general-purpose language model includes at least the Qwen series models, GPT series models, and DeepSeek series models; the general-purpose NLP tasks include at least text generation tasks, translation tasks, question answering tasks, and thought chain derivation tasks; The instruction template configuration module is used to configure corresponding chemical question-and-answer instruction templates and molecular description instruction templates for the large chemical model. The data acquisition module is used to construct a corresponding question-and-answer dataset by collecting big data from publicly available knowledge media in the field of chemistry; and to construct a corresponding inference chain dataset and molecular description dataset by collecting big data from publicly available molecular databases of SELFIES format molecular sequences and their corresponding molecular structures and physicochemical properties. The three-stage enhancement module is used to first perform a first-stage chemical knowledge enhancement and fine-tuning of the large chemical model based on the chemical question-and-answer instruction template and the question-and-answer dataset; then perform a second-stage chemical reasoning enhancement and fine-tuning of the large chemical model based on the molecular description instruction template and the reasoning chain dataset; and finally perform a third-stage molecular description enhancement and fine-tuning of the large chemical model based on the molecular description instruction template and the molecular description dataset. The model application module is used to, after the model fine-tuning is completed, substitute the user-input SELFIES format molecular sequence into the molecular description instruction template to configure the molecular sequence text of the SELFIES segment of the template to obtain the corresponding current molecular description instruction; input the current molecular description instruction into the chemical large model for molecular description generation task processing; and feed back the molecular description text output by the model in this processing to the current user.
7. An electronic device, characterized in that, include: Memory, processor, and transceiver; The processor is configured to be coupled to the memory, read and execute instructions in the memory to implement the method according to any one of claims 1-5; The transceiver is coupled to the processor, and the processor controls the transceiver to send and receive messages.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that, when executed by a computer, cause the computer to perform the method described in any one of claims 1-5.