Molecular multi-objective optimization method and device based on language model and evolutionary algorithm
By constructing an initial molecular population and combining iterative optimization with molecular language models and multi-objective optimization algorithms, the problems of insufficient chemical space exploration and poor multi-objective collaborative optimization in existing technologies are solved, and high-quality protein-targeting molecules that meet multiple key performance indicators are generated.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-30
- Publication Date
- 2026-04-03
AI Technical Summary
Existing molecular optimization techniques are insufficient in chemical space exploration and have poor multi-objective synergistic optimization effects, making it difficult to generate high-quality protein-targeting molecules that meet multiple key performance indicators.
By constructing an initial molecular population and conducting multi-objective evaluation, and combining a semantically guided generation mechanism based on a molecular language model and a multi-objective optimization algorithm for evolutionary iteration optimization until the preset termination conditions are met, a molecular population that balances broad-domain exploration of chemical space and goal-oriented convergence is generated.
This study achieved balanced optimization of multiple molecular objectives, generating high-quality protein-targeting molecules that meet several key performance indicators such as binding affinity, drug compatibility, and synthetic feasibility, thus solving the problems of insufficient chemical space exploration and poor multi-objective synergistic optimization.
Smart Images

Figure CN121789762A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of bioinformatics technology, specifically to a molecular multi-objective optimization method and apparatus based on language models and evolutionary algorithms. Background Technology
[0002] Molecular multi-objective optimization is a core technology in fields such as drug development and materials design. Its core requirement is to generate or screen molecular structures that simultaneously meet multiple key indicators such as binding affinity, drug compatibility, and synthetic feasibility in a vast and discrete chemical space, so as to adapt to the requirements of practical applications for the comprehensive performance of molecules.
[0003] In existing molecular optimization techniques, one type of method relies on screening and local modification of existing molecular libraries. Limited by the coverage of the initial library, it struggles to break through known structural boundaries and fully explore potential novel molecular spaces. Another type optimizes through generative models or evolutionary algorithms, but it has shortcomings in achieving multi-objective synergistic balance, often resulting in optimization of a single metric while other metrics become unbalanced, making it difficult to achieve robust and balanced improvements across multiple metrics. Therefore, existing molecular optimization techniques suffer from insufficient chemical space exploration and poor multi-objective synergistic optimization performance. There is an urgent need for a molecular optimization scheme that can balance broad-area exploration with objective convergence to achieve balanced multi-objective optimization.
[0004] The preceding description is intended to provide general background information and does not necessarily constitute prior art. Summary of the Invention
[0005] This application provides a molecular multi-objective optimization method and apparatus based on language models and evolutionary algorithms, which can take into account both the wide-area exploration of chemical space and goal-oriented convergence, and achieve balanced optimization of molecular multi-objectives, thereby generating high-quality protein-targeting molecules that meet multiple key performance indicators.
[0006] In a first aspect, embodiments of this application provide a molecular multi-objective optimization method based on language models and evolutionary algorithms, including: An initial molecular population is constructed, and a multi-objective evaluation is performed on each molecule in the initial molecular population to obtain the multi-objective evaluation vector corresponding to each molecule; Based on the current molecular population and its corresponding multi-objective evaluation vector, an evolutionary iteration optimization is performed by integrating a semantically guided generation mechanism based on a molecular language model and a multi-objective optimization algorithm to obtain a new generation of molecular population; The evolutionary iteration optimization process is repeated until the preset termination condition is met, and the resulting final molecular population is taken as the optimization result.
[0007] Secondly, embodiments of this application provide a molecular multi-objective optimization device based on language models and evolutionary algorithms, comprising: An initial processing module is used to construct an initial molecular population and perform multi-objective evaluation on each molecule in the initial molecular population to obtain a multi-objective evaluation vector corresponding to each molecule. The iterative optimization module is used to perform evolutionary iterative optimization based on the current molecular population and its corresponding multi-objective evaluation vector by integrating a semantic-guided generation mechanism based on molecular language model and a multi-objective optimization algorithm to obtain a new generation of molecular population; The loop control module is used to repeat the evolutionary iterative optimization process until the preset termination condition is met, and the resulting final molecular population is taken as the optimization result.
[0008] This application provides a molecular multi-objective optimization method and apparatus based on language models and evolutionary algorithms. First, an initial molecular population is constructed and multi-objective evaluation is performed to provide starting samples with basic validity and clear performance benchmarks for subsequent optimization. Then, an evolutionary iterative optimization is carried out by integrating a semantic-guided generation mechanism based on molecular language models and a multi-objective optimization algorithm. The semantic-guided generation mechanism can overcome the limitations of traditional reliance on existing molecular libraries or local modifications, expand the exploration boundary of chemical space, and solve the problem of insufficient exploration in existing technologies. The multi-objective optimization algorithm can comprehensively consider multiple key indicators of molecules during the iteration process, avoiding the situation where a single indicator is optimal while other indicators are unbalanced. By repeating the iteration until the termination condition is met, the final generated molecular population can take into account both the wide-area exploration of chemical space and goal-oriented convergence, achieving balanced optimization of molecular multi-objectives. This solves the problems of insufficient chemical space exploration and poor multi-objective synergistic optimization in existing technologies, providing high-quality protein-targeting molecules that meet multiple key performance indicators for practical applications. Attached Figure Description
[0009] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0010] Figure 1 This is an application environment diagram of the molecular multi-objective optimization method based on language models and evolutionary algorithms provided in the embodiments of this application; Figure 2 This is a flowchart illustrating the molecular multi-objective optimization method based on language models and evolutionary algorithms provided in the embodiments of this application; Figure 3 This is a comparative schematic diagram of the molecular representation methods provided in the embodiments of this application; Figure 4This is a schematic diagram of the molecular multi-objective optimization device based on language models and evolutionary algorithms provided in the embodiments of this application; Figure 5 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0011] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of systems and methods consistent with those detailed in the appended claims or with some aspects of this application.
[0012] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover descriptions such as non-exclusive inclusion, so that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, components, features, and elements with the same names in different embodiments of this application may have the same meaning or different meanings, the specific meaning of which must be determined by its interpretation in that specific embodiment or further in conjunction with the context of that specific embodiment.
[0013] It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit this application.
[0014] In the following description, the use of suffixes such as "module," "part," or "unit" to denote elements is solely for the purpose of illustrative purposes and has no specific meaning in itself. Therefore, "module," "part," or "unit" may be used interchangeably.
[0015] To address the aforementioned technical problems and overcome the shortcomings of existing technologies, this application provides a molecular multi-objective optimization method and apparatus based on language models and evolutionary algorithms. This method can balance the broad-area exploration of chemical space with goal-oriented convergence, achieving balanced optimization of molecular multi-objectives, thereby generating high-quality protein-targeting molecules that meet multiple key performance indicators.
[0016] Figure 1 This is an application environment diagram of a molecular multi-objective optimization method based on language models and evolutionary algorithms in one embodiment. (Refer to...) Figure 1This molecular multi-objective optimization method based on language models and evolutionary algorithms is applied to a molecular multi-objective optimization system based on language models and evolutionary algorithms. The system includes a terminal 110 and a server 120. The terminal 110 and server 120 are connected via a network. The terminal 110 can be a desktop terminal or a mobile terminal, specifically a mobile phone, tablet computer, laptop computer, etc. The server 120 can be a standalone server or a server cluster consisting of multiple servers. The server 120 is configured to execute the aforementioned molecular multi-objective optimization method based on language models and evolutionary algorithms, including: constructing an initial molecular population and performing multi-objective evaluation on each molecule in the initial population to obtain a multi-objective evaluation vector for each molecule; based on the current molecular population and its corresponding multi-objective evaluation vector, performing iterative evolutionary optimization by fusing a semantically guided generation mechanism based on molecular language models and a multi-objective optimization algorithm to update and obtain a new generation of molecular population; repeating the iterative evolutionary optimization process until a preset termination condition is met, and using the generated final molecular population as the optimization result.
[0017] Please see Figure 2 , Figure 2 This is a flowchart illustrating a molecular multi-objective optimization method based on language models and evolutionary algorithms according to an embodiment of this application. This embodiment primarily uses the application of this molecular multi-objective optimization method based on language models and evolutionary algorithms to a computer device as an example. Specifically, the molecular multi-objective optimization method based on language models and evolutionary algorithms provided in this embodiment may include the following steps: S1. Construct an initial molecular population and perform multi-objective evaluation on each molecule in the initial molecular population to obtain the multi-objective evaluation vector corresponding to each molecule; Specifically, for step S1, the source of molecular screening is determined by selecting a certain number of molecules from publicly available small molecule databases as candidate molecules. These molecules need to cover different structural types and molecular weight ranges to ensure the diversity of the initial sample. Basic effectiveness screening is performed on the selected candidate molecules, eliminating molecules with obvious chemical structural defects (such as valence bond mismatch, unstable ring structures, etc.) to form an initial molecular population. For example, 100-200 structurally stable molecules without obvious chemical defects are selected to form the initial population. A multi-objective evaluation index system is established, with core indicators including key performance indicators such as molecule-target binding affinity, drug compatibility, and synthetic feasibility. Quantitative evaluation is performed for each indicator. Binding affinity is determined through molecular docking experiments (with docking score as the quantitative result), and quantitative values for drug compatibility (such as QED value) and synthetic feasibility (such as SA value) are calculated using professional algorithms.
[0018] S2. Based on the current molecular population and its corresponding multi-objective evaluation vector, an evolutionary iteration optimization is performed by integrating a semantic-guided generation mechanism based on a molecular language model and a multi-objective optimization algorithm to update and obtain a new generation of molecular population; Specifically, for step S2, a pre-trained molecular language model is invoked. Based on the molecular structural features of the current molecular population, and utilizing the chemical semantic rules learned by the model, a batch of candidate molecules with novel structures and chemical rationality are generated. For example, the model generates new molecules with similar core skeletons but optimized side chain groups, or potentially effective molecules with entirely new skeleton structures, based on the skeletal structures of existing molecules. New candidate molecules generated by the semantically guided generation mechanism are collected and merged with the current molecular population to form a mixed set containing existing high-quality molecules and novel candidate molecules. A multi-objective optimization algorithm is used to comprehensively rank all molecules in the mixed set. This algorithm considers the multi-objective evaluation vectors of each molecule simultaneously, without relying on a single index for ranking. Instead, it is based on the collaborative optimization logic of multiple indices to screen out molecules that perform evenly or have outstanding advantages in multiple dimensions. According to the preset population size, the best few molecules are selected from the ranking results to form a new generation of molecular population, completing one evolutionary iteration. For example, if the population size is set to 150, the top 150 molecules in the ranking are selected as the new generation of population.
[0019] S3. Repeat the evolutionary iteration optimization process until the preset termination condition is met, and take the generated final molecular population as the optimization result; Specifically, for step S3, after completing one evolutionary iteration, it is checked whether the preset termination conditions are met. Common termination conditions include reaching the preset maximum number of iterations (e.g., 20-30 generations) and no significant improvement in the multi-objective evaluation metrics of consecutive generations of populations (i.e., population performance convergence). If the termination conditions are not met, the new generation of molecular population is used as the current molecular population, and the process returns to step 2 to perform semantic-guided generation, multi-objective optimization, and other operations again to start the next round of evolutionary iteration. If the termination conditions are met, the iteration process is stopped, and the molecular population obtained from the last iteration is determined as the final optimization result. This result contains a set of high-quality molecules that have undergone multiple rounds of screening and optimization.
[0020] This embodiment lays the foundation for optimization through initial population construction and multi-objective evaluation, expands the boundaries of chemical space exploration by integrating semantic-guided generation mechanism, and ensures balanced improvement of multiple indicators by using multi-objective optimization algorithm. After multiple iterations, the final generated molecular population takes into account both structural novelty and synergistic optimization of multiple performance indicators, effectively solving the problems of insufficient chemical space exploration and high difficulty in balancing multiple objectives in traditional methods, and providing a set of molecular candidates with better performance for related fields.
[0021] Furthermore, in some embodiments, step S1, "constructing an initial molecular population and performing multi-objective evaluation on each molecule in the initial molecular population to obtain a multi-objective evaluation vector corresponding to each molecule," may specifically include: S11. According to the preset sampling strategy, extract multiple molecules from the molecular database to form an initial molecule set; Specifically, for step S11, the source of the molecular database is determined by selecting a public database containing a large number of small molecule structures. This ensures that the database covers molecules of different chemical types and molecular weight ranges, providing a rich sample base for sampling. The core logic of the pre-set sampling strategy is clearly defined. This strategy must guarantee the diversity and representativeness of the initial molecules, avoiding sample bias towards a single structure or performance characteristic. Taking molecular weight stratified sampling as an example, the molecules in the database are first divided into intervals such as <100Da, 100-150Da, 150-200Da, and 200-250Da according to their molecular weight. Then, molecules are randomly selected from each interval according to a pre-set ratio (e.g., 2:3:4:1). If the plan is to construct an initial set of 200 molecules, then 20, 30, 40, and 10 molecules are selected from each interval respectively, and the results are then combined to form the initial molecule set.
[0022] S12. Perform a chemical validity check on each molecule in the initial molecular set to filter out invalid structures and obtain the initial molecular population; Specifically, for step S12, chemical validity check standards are established, including SMILES format validity verification, atomic valence state rationality check, and ring structure stability check, to ensure that the molecular structure conforms to basic chemical laws. Each molecule in the initial molecular set is checked: for example, molecules with incorrect SMILES string format, mismatched atomic valence bonds (e.g., carbon atoms forming 5 covalent bonds), or ring structures with significant strain that cannot exist stably are removed. All molecules that pass the checks are summarized to form the initial molecular population. Assuming the initial molecular set contains 200 molecules, after checking, 25 invalid molecules with structural defects are filtered out, ultimately resulting in an initial molecular population of 175 molecules.
[0023] S13. Perform multi-objective evaluation on each molecule in the initial molecular population to obtain the corresponding multi-objective evaluation vector; Specifically, for step S13, the core indicators for multi-objective evaluation are determined, including key performance indicators such as the binding affinity between the molecule and the target, drug compatibility, and synthetic feasibility, ensuring that the evaluation dimensions align with practical application needs. Quantitative detection and calculation are performed for each indicator. Binding affinity is measured using molecular docking technology to obtain a docking score (e.g., -6.8 kcal / mol); drug compatibility quantification values (e.g., QED value 0.73) and synthetic feasibility quantification values (e.g., SA value 2.3) are calculated using specialized algorithms. The evaluation results for each molecule are integrated to construct a multi-objective evaluation vector. For example, the evaluation vector for a molecule might be (docking score: -7.1 kcal / mol, QED value: 0.76, SA value: 2.0), with each molecule corresponding to a unique multi-objective evaluation vector that fully reflects its comprehensive performance.
[0024] This embodiment ensures the diversity and representativeness of the initial molecules through targeted sampling, eliminates invalid structures by chemical validity checks, and accurately quantifies the comprehensive performance of molecules through multi-objective evaluation. Finally, it obtains an initial molecular population with valid structures and well-defined performance, providing a high-quality and reliable starting foundation for subsequent evolutionary iteration optimization. This effectively avoids invalid molecules occupying computational resources and improves the efficiency and accuracy of the overall optimization process.
[0025] Furthermore, in some embodiments, the preset sampling strategy is a stratified sampling strategy, and step S11, "according to the preset sampling strategy, extracting multiple molecules from the molecular database to form an initial molecule set," may specifically include: S111. Stratify molecules in a molecular database according to at least one of the following attributes: molecular weight, structural type, or source database; Specifically, for step S111, based on the actual needs of molecular optimization, one or more criteria are selected from molecular weight, structure type, and source database as stratification bases to ensure that molecules in each subset have clear distinguishing characteristics after stratification. First, the molecular weight distribution range of all molecules in the molecular database is statistically analyzed, and stratification is divided according to reasonable intervals. For example, molecular weight is divided into five intervals, each interval being an independent stratum. If structure type is selected as the stratification attribute, molecules are classified according to features such as core functional groups and skeletal structure, for example, into aromatic rings, heterocycles, aliphatic chains, esters, etc., ensuring that molecular structures within the same stratum are similar and that structural differences between different strata are significant. If source database is selected as the stratification attribute, molecules from different public small molecule databases are divided into independent strata to ensure coverage of molecules from different sources. If multiple attributes are combined for stratification, for example, first dividing into three intervals by molecular weight, and then further dividing each molecular weight interval into sub-layers by structure type, a multi-level stratified structure is formed, further improving the representativeness of the samples.
[0026] S112. Extract molecules from each layer according to a preset ratio to form an initial molecule set; Specifically, for step S112, based on the computational resources and iteration efficiency requirements of the subsequent optimization process, the total number of molecules in the initial molecular set is set, for example, 200 molecules are selected to form the initial set. The sampling ratio for each stratum is determined, based on factors such as the number of molecules, structural diversity, and optimization requirements of each stratum. Based on the total sampling size and the preset ratio, the number of molecules to be extracted from each stratum is calculated. Within each stratum, a corresponding number of molecules are selected using random sampling to avoid sample bias caused by human preference and ensure that the molecular characteristics of each stratum are reflected in the initial set. The molecules extracted from each stratum are then integrated to form the initial molecular set.
[0027] This embodiment ensures that the initial molecule set can uniformly cover molecular types with different characteristics by stratifying according to specific attributes and sampling according to a preset ratio, avoiding the concentration of samples in a single attribute range, improving the diversity and representativeness of the initial molecules, providing a comprehensive sample basis for subsequent multi-objective optimization, helping to broaden the scope of chemical space exploration, and improving the reliability and practicality of the final optimization results.
[0028] Furthermore, in some embodiments, step S2, "based on the current molecular population and its corresponding multi-objective evaluation vector, and through evolutionary iterative optimization by fusing a semantically guided generation mechanism based on a molecular language model and a multi-objective optimization algorithm, to update and obtain a new generation of molecular population," may specifically include: S21. Generate a batch of new candidate molecules based on the current molecular population using at least one generation mechanism, including a semantically guided generation mechanism based on a pre-trained molecular language model; Specifically, for step S21, the current molecular population base is defined, using the molecular population obtained through previous optimization as a blueprint. This population contains a certain number of molecules with valid structures and well-defined properties, providing structural references for the generation of new candidate molecules. A pre-trained molecular language model is invoked. This model has learned chemical semantic rules through large-scale molecular data and can identify reasonable connection relationships and structural features between molecular fragments. Semantic-guided generation is then performed. Based on information such as the skeletal structure and functional group combinations of molecules in the current population, the model generates new molecules through semantic reasoning. All generated new molecules are collected to form a preliminary candidate molecule set.
[0029] S22. Perform multi-objective evaluation on new candidate molecules to obtain the corresponding multi-objective evaluation vector, and merge the multi-objective evaluation vector with the current molecule population to form a candidate set; Specifically, for step S22, a multi-objective evaluation standard is set, using core indicators consistent with the initial population evaluation, including binding affinity (docking score), drug compatibility (QED value), and synthetic feasibility (SA value), ensuring consistency and coherence across evaluation dimensions. New candidate molecules are evaluated one by one. The docking score of each new candidate molecule is determined using molecular docking technology, and the QED and SA values are calculated using specialized algorithms, integrated into a multi-objective evaluation vector for that molecule. The evaluation vectors of all new candidate molecules are then integrated with the evaluation vectors of the current molecular population, eliminating duplicate molecules to form a candidate set.
[0030] S23. Based on the multi-objective evaluation vectors of all molecules in the candidate set, a multi-objective optimization algorithm is used to select the next generation of molecular population; Specifically, for step S23, a multi-objective optimization algorithm is initiated. This algorithm uses multi-objective evaluation vectors as its core basis, focusing on the synergistic balance of multiple molecular properties rather than relying on a single index ranking. By analyzing the evaluation vectors of each molecule, molecules with no significant weaknesses in indicators such as binding affinity, drug compatibility, and synthetic feasibility, or those with outstanding performance in some indicators and satisfactory performance in others, are selected. The size of the next-generation population is determined, and individuals with the best overall performance are selected from the selected molecules according to the preset population size to form the next-generation molecular population.
[0031] This embodiment broadens the scope of chemical space exploration through a semantically guided generation mechanism, ensures that the performance of new candidate molecules is quantifiable through multi-objective evaluation, and achieves comprehensive performance optimization through a multi-objective optimization algorithm. The resulting new generation of molecular populations combines structural novelty with balanced performance, effectively improving the targeting and reliability of molecular optimization and laying a high-quality foundation for subsequent iterations.
[0032] Furthermore, in some embodiments, step S21, "generating a batch of new candidate molecules based on the current molecular population using at least one generation mechanism, including a semantically guided generation mechanism based on a pre-trained molecular language model," may specifically include: S211. Segment the molecules in the current molecular population based on chemical reaction rules to obtain a sequence of segments composed of multiple structural segments according to their connection relationships; Specifically, for step S211, the applicable chemical reaction rules are determined, and rules that ensure the feasibility of fragment synthesis are selected. These rules identify stable chemical bond connection sites in the molecule, ensuring that the segmented fragments possess independent structural integrity and subsequent connection potential. Molecular segmentation is performed; for each molecule in the current molecular population, according to the selected chemical reaction rules, it is disassembled into multiple structural fragments, and connection sites are marked for each fragment. The segmented structural fragments are then arranged sequentially according to the original chemical bond connection order in the molecule to form a fragment sequence.
[0033] S212. Perform a masking operation on selected segments in the segment sequence to generate an input sequence containing at least one mask position; Specifically, in step S212, based on the molecular structure optimization requirements, one or more fragments are selected from the fragment sequence as masking objects. These can be side chain fragments, functional group fragments, etc., to avoid masking the core skeleton fragments and causing structural integrity damage. The selected fragments are replaced with preset mask markers, while other fragments and the connection relationship markers between fragments are preserved, generating the input sequence.
[0034] S213. Input the input sequence into the pre-trained molecular language model. The model predicts and fills in the new fragments corresponding to the masked positions based on the chemical context provided by the unmasked fragments, forming a predicted fragment sequence. Specifically, in step S213, the masked input sequence is fed into a pre-trained molecular language model. The model, through learned chemical semantic rules, identifies the structural features and connection site requirements of the unmasked segments, among other contextual information. Based on contextual reasoning, the model generates new segments that are chemically plausible, and these new segments must match the connection site types of the unmasked segments at both ends. The new segments predicted by the model replace the mask markers and are combined with the original unmasked segments according to their connection relationships to form a complete predicted segment sequence.
[0035] S214. According to the predetermined connection rules, chemically bond and connect the fragments in the predicted fragment sequence to reconstruct and generate new candidate molecules; Specifically, for step S214, chemical bonding rules corresponding to the segmentation rules are used to clarify the bonding type of the connection sites between fragments, ensuring that the molecular structure after ligation conforms to the rules of chemical bonding. Following the predicted fragment sequence, the connection sites of adjacent fragments are chemically bonded according to predetermined rules. After ligation, a structurally complete new molecule is formed.
[0036] This embodiment ensures fragment effectiveness through chemical reaction rule segmentation, focuses on optimizing target points through masking operations, and provides novel fragments that match the context through model prediction. Finally, chemically reasonable new molecules are generated through rule-based connection. This not only expands the novelty of molecular structures but also ensures synthetic feasibility, effectively enriching the diversity and quality of candidate molecules.
[0037] Furthermore, in some embodiments, the masking operation employs a dynamic masking strategy. Step S212, "performing a masking operation on selected segments in the segment sequence to generate an input sequence containing at least one mask position," may specifically include: Based on the current evolutionary iteration generation, and according to the preset dynamic scheduling rules, calculate the number or proportion of mask segments required for this iteration. Based on the number or proportion of segments to be masked in this iteration, select the appropriate number of segments from the segment sequence for masking to generate the input sequence.
[0038] Specifically, the key parameters for evolutionary iteration are first determined, including the current iteration generation and the preset maximum iteration generation. A rule is adopted where the mask ratio decreases with increasing iteration generation to ensure that early iterations encourage broad-area exploration while later iterations focus on local refinement. The rule can use a linear decreasing logic, meaning the mask ratio / number decreases linearly with each iteration. The mask number or ratio is calculated by combining the parameters and rules. Based on the calculation results and the actual total number of molecular fragments, the final number or ratio of masked fragments used in this iteration is determined to ensure that the masking operation does not damage the fundamental integrity of the molecular core structure. To avoid disrupting the stability of the molecular core skeleton, non-core fragments further down the fragment sequence are prioritized for masking, while core skeleton fragments are not prioritized. Based on the calculated mask number, a corresponding number of target fragments are selected from the fragment sequence. The selected target fragments are replaced with preset mask markers, while unselected fragments and the connection relationship markers between fragments are retained, generating the input sequence. The input sequence is validated to ensure that the generated input sequence contains at least one masked position and that the unmasked segment provides sufficient chemical context information to provide a valid reference for subsequent model predictions.
[0039] This embodiment adapts to the optimization needs of different iteration stages through dynamic scheduling rules. In the early stage, multiple masked fragments are used to expand the scope of chemical space exploration, while in the later stage, fewer masked fragments are used to focus on local structural refinement. This not only ensures the exploration of novel molecular structures but also guarantees the convergence of the optimization process, effectively balancing exploration and development in molecular optimization and improving the quality and optimization efficiency of candidate molecules.
[0040] Furthermore, in some embodiments, the generation mechanism of step S21 also includes rule mutation based on chemical reaction rules and cross-recombination based on the largest common substructure between molecules. Therefore, step S22, "performing multi-objective evaluation on new candidate molecules to obtain corresponding multi-objective evaluation vectors, and merging the multi-objective evaluation vectors with the current molecular population to form a candidate set," may specifically include: S221. Check the rationality of atomic valence state, bonding mode and ring structure of all generated candidate molecules, and filter out molecules that do not conform to chemical rules; Specifically, for step S221, the chemical rationality check standards are clearly defined. For example, for the atomic valence state check, based on the common bonding rules of each element (e.g., carbon atoms can form a maximum of 4 covalent bonds, oxygen atoms can form a maximum of 2 covalent bonds), the number of bonds formed by all atoms in the molecule is verified to conform to chemical rules. For the bonding mode check, it is confirmed that the connection of chemical bond types (single bonds, double bonds, triple bonds) in the molecule is reasonable, and there are no abnormal bonds (e.g., triple bonds formed between saturated carbon atoms, unstable cumulative double bonds in aromatic rings). For the ring structure rationality check, the ring strain is assessed to determine whether it is within a stable range, avoiding excessive crowding, abnormal bond angles (e.g., large substituents in three-membered rings causing excessive strain), or ring atom bonding conflicts. The check is performed molecule by molecule. Taking the candidate molecule "1-hydroxy-2-methyl-3-pentavalent carbon benzene" as an example, the check found that one carbon atom formed 5 covalent bonds, violating the atomic valence state rules; another candidate molecule "cyclopropanenaphthalene" has excessive ring strain due to the direct fusion of the three-membered ring and the naphthalene ring, which does not meet the requirements for ring structure rationality. All candidate molecules were subjected to the above three checks one by one, and molecules with any violation were recorded. Molecules that were confirmed to be non-compliant with chemical rules were eliminated, and only candidate molecules with normal atomic valence states, reasonable bonding modes, and stable ring structures were retained.
[0041] S222. Compare candidate molecules that have passed chemical filtering based on normalized molecular representation to remove duplicate molecular structures; Specifically, in step S222, a unified molecular structure encoding method (such as normalized SMILES strings) is used to process candidate molecules that have passed the chemical filtering, transforming different representations of the same molecule into a unique standardized code. The normalized representations of all candidate molecules are compared one by one, identifying molecules with completely identical codes as repetitive structures. For each group of repetitive structures, only one is retained (either randomly selected or the one generated earliest), and the remaining repetitive individuals are discarded.
[0042] This embodiment filters out invalid molecules that do not conform to chemical rules by checking atomic valence state, bonding mode and ring structure; by normalizing representation comparison, it removes repetitive structures, effectively improving the chemical validity and uniqueness of candidate molecules, reducing the invalid computational burden in the subsequent multi-objective evaluation and optimization process, and providing high-quality, non-redundant molecular samples for the subsequent selection steps.
[0043] Furthermore, in some embodiments, step S13, "performing a multi-objective evaluation on each molecule in the initial molecular population to obtain its corresponding multi-objective evaluation vector," may specifically include: S131. Based on the multi-objective evaluation vectors of all molecules in the candidate set, perform non-dominated sorting to divide the molecules into several Pareto front levels; Specifically, for step S131, the core of the non-dominant ranking is to identify molecules that are not dominated by other molecules. If all evaluation metrics (such as docking score, QED value, SA value) of molecule X are not inferior to those of molecule Y, and at least one metric is superior to Y, then X dominates Y. Molecules not dominated by any other molecules are classified into the same Pareto front level, and ranked in descending order of dominance (such as F1, F2, F3, etc.), with F1 being the optimal front. The complete evaluation vector of each molecule is extracted from the candidate set to ensure that the metric dimensions are consistent (such as "docking score-QED value-SA value"). Comparing the evaluation vectors of all molecules, molecule A has the best docking score and the best SA value, and no other molecule can dominate it, so it is classified into F1; molecule B has the best QED value, but its docking score is inferior to A, and no other molecule dominates it, so it is classified into F1; all metrics of molecule C are dominated by A or B, so it is classified into F2; and so on, dividing all molecules into several Pareto front levels according to their dominance relationships.
[0044] S132. For molecules within the same Pareto front order, calculate the crowding distance in the multi-object space; Specifically, for step S132, the definition of crowding distance is clarified. Crowding distance is used to measure the sparsity of molecules within the same frontier in the multi-object space. The larger the distance, the fewer molecules are around the molecule. Retaining the molecule can improve the diversity of the solution set.
[0045] Sort by evaluation metric individually: Taking the F1 frontier as an example, first sort by docking score from smallest to largest, then by QED value from largest to smallest, and finally by SA value from smallest to largest. Set the crowding degree of the molecules at the top and bottom of each metric ranking (boundary molecules) to infinity (meaning they must be retained to cover the extreme values of the metric). For non-boundary molecules, calculate the crowding degree contribution of each metric dimension according to the formula, and sum them to obtain the total crowding degree. Summarize the crowding degree results and output the crowding degree distance of all molecules within the same frontier.
[0046] S133. Select a predetermined number of molecules in descending order of Pareto front hierarchy, and in descending order of crowding distance within the same hierarchy, to form a new generation of molecular population. Specifically, for step S133, the size of the next-generation population is determined, and a preset population size is set as the upper limit for the final number of molecules selected. Selection begins with the optimal frontier F1. If the number of molecules in F1 is less than or equal to the population size, all are included; if the number of molecules in F1 is insufficient, selection then begins with F2, and so on. For frontiers requiring partial selection (such as F2), molecules are sorted from largest to smallest by their crowding distance, prioritizing molecules with high crowding to ensure a uniform distribution within the same frontier. The molecules selected from each frontier are then aggregated to ensure the total number equals the preset population size, forming a next-generation molecular population that combines optimal performance with diverse distribution.
[0047] This embodiment selects the optimal molecular set for multiple objectives through non-dominated sorting and ensures the diversity of molecular distribution within the same frontier by using crowding distance. The final selected new generation of molecular population not only takes into account the optimal performance of various indicators, but also avoids the homogenization of structure and performance, providing a high-quality and highly diverse foundation for subsequent iterative optimization and improving the robustness and effectiveness of multi-objective optimization.
[0048] Furthermore, in some embodiments, step S3, "repeating the evolutionary iterative optimization process until a preset termination condition is met, and taking the generated final molecular population as the optimization result," may specifically include: S31. After each iteration, determine whether the preset termination condition is met. The termination condition includes reaching the maximum number of iterations or the population performance converges. Specifically, for step S31, a post-iteration judgment node is defined. After each complete process of "generating candidate molecules—multi-objective evaluation—selecting the next generation population," a termination condition check is immediately initiated to ensure that the optimization process is not prematurely interrupted or over-iterated. Based on optimization requirements, computational resources, and expected improvements in molecule performance, a maximum number of iterations is preset as a hard termination boundary. Convergence judgment rules are set based on the statistical values of the population's multi-objective evaluation vectors. If the average value of key indicators (such as average docking score, average QED value, and average SA value) of the population across multiple consecutive generations is less than a preset threshold, then the population performance is considered to have converged. Each condition, either "reaching the maximum number of iterations" or "population performance convergence," is checked. If either condition is met, the process proceeds to the result output stage; otherwise, the next iteration continues.
[0049] S32. When the preset termination condition is met, stop the iteration and output the molecular population obtained in the last iteration as the optimization result. The optimization result shall include at least the molecules located at the Pareto front and their corresponding multi-objective evaluation vectors. Specifically, for step S32, once any termination condition is met, the subsequent iteration process is immediately stopped, the next-generation molecular population generated in the last iteration is frozen, and no further candidate molecule generation or selection operations are performed. Molecules located at the Pareto front (i.e., the optimal set of molecules with undominated multi-objective performance) are extracted from the last-generation population to ensure that the output focuses on core high-quality molecules. The Pareto front molecules are associated one-to-one with their corresponding multi-objective evaluation vectors to form a structured output. The output is presented in a standardized file format (such as tables or structured text) to ensure that the results can be directly used for subsequent experimental verification, patent layout, or further analysis without additional processing.
[0050] This embodiment avoids insufficient optimization or redundant iterations by using explicit termination conditions, which ensures that molecular performance is fully improved while avoiding waste of computational resources. The output results focus on high-quality molecules and quantitative performance data at the Pareto frontier, accurately meeting the screening needs for cost-effective molecules in practical applications, and improving the practicality and implementation efficiency of the optimization results.
[0051] To facilitate understanding of the molecular multi-objective optimization method based on language models and evolutionary algorithms provided in this embodiment, this embodiment also provides a specific implementation of the molecular multi-objective optimization method based on language models and evolutionary algorithms. Taking the method executed in a molecular multi-objective optimization system based on language models and evolutionary algorithms as an example, the system includes an input and configuration module, a mutation module guided by FragMLM inference execution semantics, a rule base mutation and crossover module, a chemovalidity and basic filtering module, a multi-objective scoring module, a multi-objective optimization and elite retention module, a dynamic mask scheduling module, and a data management module. The specific steps are as follows: Task initialization: Read protein pockets and docking parameters; set the optimization objective vector form (uniformly minimized); determine population size. Maximum Algebra Candidate amplification fold Random seed and parallel thread configuration. Constructing the initial population. Stratified random sampling from the ZINC database Individual molecules; perform chemical and basic filtering on the initial population. Initial scoring and normalization: calculate docking score, QED, and SA for all individuals in generation 0 and convert them into target vectors that are as small as possible; perform initial non-dominated sorting and crowding calculation using NSGA-II.
[0052] Iterate through the main loop ( =1 to ) 1) Dynamic mask update: The scheduler sets the mask ratio for the current generation. 1) Sampling temperature T, top-k. 2) Candidate generation (three-way parallel generation of candidate molecules): a) Semantic variation: for the current population according to... 1) Masking: Call FragMLM for reconstruction and sample to generate candidates; b) Rule mutation: Call reaction / replacement rules; c) Crossover recombination: Pair from current elite or highly diverse individuals and generate new skeletons based on MCS crossover. 3) Filtering and deduplication: Perform validity, substructure, and simple similarity (such as Tanimoto threshold) filtering to maintain candidate diversity. 4) Multi-objective evaluation: Batch docking scoring, calculate QED and SA, and form target vectors. 5) NSGA-II optimization: Perform non-dominated ranking and crowding calculation on "current population ∪ candidate pool" and select the top candidates. 6) Recording and auditing: Store the parent-child relationships, operator labels, score distribution, and convergence curve of the current generation (which can be used to plot the "optimized forest" and Pareto trajectory later).
[0053] Termination and result summary: The termination condition is reaching the maximum algebra. The goal is to stop early or exhaust the budget. The final output includes a Pareto frontier set (including SMILES / structure, docking sub-segments, QED, SA), a Top-k candidate list, and a source path and synthesizability summary for each candidate.
[0054] The input and configuration module receives target protein information (pocket coordinates / grid, docking parameters), an initial molecular library, optimization target settings (minimizing docking score, maximizing QED, minimizing SA, etc., all unified as "the smaller the better" target vector), and a search budget (algebraic). Group size (Including the number of parallel threads, etc.). Outputs a standardized task description and a global hyperparameter table. For the FragMLM inference module, it is used to mask and reconstruct the FragSeq format of a given SIMLES, and outputs the conditional distribution of fragment replacement / sponging; it supports sampling control such as temperature and top-k. It provides a dynamic mask ratio interface (driven by the scheduler) to achieve an adaptive balance between exploration and development.
[0055] To facilitate differentiation from existing technologies in this field, the FragMLM module is implemented as follows: First, the molecule is fragmented using the BRICS (Bond-Retrosynthesis-In-Combination-Strategy) rule, decomposing the original SMILES representation into a one-dimensional fragment sequence FragSeq. Each fragment corresponds to a structural unit that meets synthetic feasibility requirements and is marked with a connection site. After this processing, the original molecular graph structure is transformed into an ordered sequence of fragments from left to right, facilitating the language model's learning of the contextual dependencies between fragments.
[0056] The training data, model structure, and inference process of FragMLM are as follows: 1. Training data construction: First, the original small molecule is fragmented based on the BRICS rule, breaking the molecule down into ordered fragment sequences frag1, frag2, ..., frag n The fragments are concatenated using the segment separator [SEP], and special markers [BOS] and [EOS] are added to the beginning and end of the sequence to form the training sample string: [BOS]frag1[SEP]frag2[SEP]…frag n[EOS]. All samples are stored in HuggingFace datasets format, containing two splits: a training set (train) and a validation set (validation). After loading, they are encoded into discrete token sequences using SMILES / fragment-specific tokenizers for autoregressive modeling.
[0057] 2. Model Structure: FragMLM uses a one-way decoder generated by a Transformer-based generative model. (1) The word embedding layer maps fragment tokens to =768-dimensional continuous vector space; (2) Stack L=12 TransformerBlocks, each Block is internally composed of multi-head causal self-attention ( =12) and feedforward network (4× extension) are connected in series and fused with LayerNorm through residual connection; (3) the output is mapped to vocabulary size |V| through linear layer, and then softmax is connected to obtain the conditional distribution of the next token; (4) during training, standard cross-entropy is used as the loss function to perform autoregressive prediction on the entire segment sequence: given x0… predict .
[0058] 3. Training process: (1) Distributed data parallelism (DDP) is used to train on multiple GPUs. The samples in the batch are first padded to a uniform length and then divided according to "input sequence x" and "label y shifted one bit to the right". (2) The optimizer uses AdamW and adopts mixed precision (AMP) training, gradient clipping and optional warmup + cosine annealing learning rate scheduling in each step, and periodically saves the checkpoint weights.
[0059] 4. Inference and Molecular Generation Logic: During online inference, the current parent molecule is first fragmented and masked using BRICS (see point 2) to obtain the prefix sequence: [BOS]frag1[SEP]…frag_k[SEP]. Only the unmasked prefix fragments are retained. This prefix is encoded as a token sequence and input into FragMLM. Autoregressive sampling is used to generate subsequent fragments: using temperature T, top Parameters such as k-truncation and repetition penalty rp control sampling diversity; generation stops when [SEP] or [EOS] is encountered, producing a complete fragment sequence. The generated fragment sequence is parsed back into a fragment SMILES list, and then reassembled into a complete molecule using a BRICS-based recombination algorithm; only molecules that are valid according to RDKit are retained and written back to the GA population as the third "generation operator" output generated by the Transformer-based generative model.
[0060] Building upon this, FragMLM preferentially employs a generative pre-trained language model based on the Transformer architecture for unsupervised training on a large number of FragSeq sequences of molecules. The training objective can be expressed as: given a prefix fragment sequence... Predict the current location segment By studying the conditional probability distribution of molecules, we can learn the statistical laws governing the overall molecular structure. Let the molecular... If a sequence consists of L segments, its joint probability can be expressed as: in These are the model parameters for FragMLM. Through pre-training on a large-scale molecular corpus, FragMLM learns "which segments reasonably co-occur in what context," forming a chemical semantic prior that can be used for mutation operators. During training, the parameters are learned by minimizing the negative log-likelihood of all molecules in the training set. : in To train the corpus, Indicates molecule In position The FragMLM learns the co-occurrence patterns of fragments in different contexts and the statistical distribution of molecular structures by minimizing the aforementioned loss.
[0061] In a preferred embodiment, the training corpus is derived from publicly available small molecule databases (such as ZINC, UniChem, etc.). First, the original SMILES are pre-screened (e.g., molecular weight limited to 150-750, element type limited to common organic elements, metal complexes and obviously unstable groups removed). Random sampling is then performed on approximately ten million SMILES after filtering, and each fragment is segmented using the BRICS rule to obtain approximately ten million valid FragSeq fragment sequences. The resulting FragSeq sequences are divided into a training set (95%) and a validation set (5%) to monitor model convergence and implement early stopping strategies.
[0062] The FragMLM preferably employs a decoder-only Transformer architecture, including sub-modules such as a fragment embedding layer, a position embedding layer, a multi-head self-attention layer, and a feedforward network layer. Specifically, each fragment is first mapped to... A continuous vector of dimension is added to the position embedding of the same dimension, and then sequentially passed through... A multi-head self-attention block is stacked; each layer consists of a multi-head self-attention sublayer and a feedforward network sublayer connected in series, and combined with residual connections and layer normalization to extract high-order fragment-level semantic information. Preferably, it can be set as =768, Set to 12, the number of attention heads is set to 12, and the corresponding number of parameters is approximately The scale allows for both expressive power and reasoning efficiency.
[0063] During the model training phase, adaptive optimization algorithms such as AdamW can be used, combined with a learning rate scheduling strategy of piecewise linear preheating and cosine annealing. Unsupervised pre-training is performed under large batch conditions (such as 512 to 1024 FragSeqs) until early stopping is triggered when the perplexity of the validation set no longer decreases significantly, thereby obtaining a well-converged fragmented molecular language model.
[0064] In the FragEvo inference (online mutation) stage of this model, the preferred calling process of FragMLM includes the following steps: (1) Fragmentation and mask position determination: Perform BRICS fragmentation on the current individual molecule to obtain the fragment sequence FragSeq; the dynamic mask module calculates the number of mask fragments for the current generation based on the current generation g. Select from the maskable locations excluding the first segment. Each segment is replaced with a special marker [MASK] to form a masked FragSeq. (2) Calculation of conditional probability distribution and temperature recalibration: Input the masked FragSeq into FragMLM and calculate the conditional probability distribution for each [MASK] position. To adjust for diversity, logits are adjusted according to the temperature corresponding to the current generation. Scaling is performed by dividing the log probability of each candidate segment by... The higher the temperature, the flatter the distribution and the more exploratory it is. (3) Top-k truncation and random sampling: Perform top-k truncation on the probability distribution after temperature recalibration. (4) Sequential generation of multiple mask positions: When there are multiple mask positions, it is preferable to fill them in order from left to right; after each new fragment is generated, it is immediately written back to FragSeq and used as the context condition for subsequent mask positions until all mask positions are replaced by new fragments. (5) Stitching and chemical verification: The generated FragSeq is "stitched" into complete SMILES according to the BRICS connection rules, and chemical rationality checks (including valence state checks, ring structure checks, etc.) are performed using RDKit or similar tools. If it fails, the candidate molecule is discarded; otherwise, it enters the subsequent filtering and multi-target scoring module.
[0065] Meanwhile, during execution, to achieve a balance between exploration and optimization in the vast molecular space, this invention employs a masking approach for control. Combining this masking mechanism, the semantic mutation process guided by FragMLM in the FragMLM module preferably includes the following steps: 1. Fragment selection and masking: The current individual molecule is fragmented into FragSeq representation using BRICS. The dynamic masking module automatically adjusts the number of masked fragments according to the current generation. In the early stages, more fragments are masked to encourage backbone transitions, while in the later stages, only a few fragments are masked for local refinement.
[0066] 2. Semantic Condition Generation: Input the masked fragment sequence into FragMLM. Given the unmasked fragment as context, obtain the candidate fragment probability distribution for each masked position. In the implementation, temperature parameters and top-k truncation can be combined to recalibrate and truncate the distribution to balance diversity and reliability.
[0067] 3. Fragment Replacement and Reconstruction: At each mask location, a new target fragment is sampled from the candidate distribution, replacing the original fragment. The fragment is then "stitched" according to the BRICS connection rules to restore a complete FragSeq, which is then converted back to valid SMILES and molecular structures. The generated molecule subsequently enters the chemical validity check and multi-objective scoring module.
[0068] Through the above mechanism, the FragMLM module no longer serves as an "offline candidate molecule generator," but directly acts as a fragment-level semantic mutation operator in the genetic algorithm. In each generation iteration, it continuously injects chemical semantic priors from a large-scale corpus into the population, significantly improving the effectiveness, novelty, and diversity of newly generated molecules.
[0069] For the rule base and cross-operation module, standard chemical variation / reaction rules and cross-operations based on the maximum common substructure (MCS) are provided, ensuring chemical valence and basic feasibility. A set of rule candidates (replacement, growth, truncation, etc.) is returned for the input molecule.
[0070] Specifically, this embodiment introduces 94 organic reaction rules compiled in AutoGrow 4.0 for rule variation, including 36 click chemistry rules from the AutoClickChemRxn rule set and 58 robust organic reaction rules from the RobustRxn rule set. These rules are pre-coded in the form of "reaction template + action site marker," enabling local structural adjustments such as addition, substitution, ring closure, bond breaking, and functional group modification to the parent molecule, ensuring that the resulting offspring conforms to chemical common sense in terms of valence state, bonding mode, and basic synthetic feasibility.
[0071] Crossover operations preferably employ a skeleton recombination strategy based on MCS (Multi-Structure Recombination). Specifically, two parent molecules are selected from the current population. and (Preferably drawn from individuals ranked high in NSGA-II or with high diversity), the MCS between the two is calculated using the RDKit library; if no valid MCS is found, the parent pair is discarded and a new parent pair is selected. If a valid MCS is obtained, the common substructure is used as the anchor point, and the two parent molecules are divided into two parts along the MCS boundary: "common backbone + respective side chains". By exchanging or recombinating the side chains, one or more daughter molecules are generated according to the SMILES merging step in AutoGrow 4.0. The generated daughter molecules undergo the following processes in sequence: 1. Chemical Sanitization: Global checks and repairs of atomic valence states, bonding modes, and aromaticity based on RDKit; 2. Valence and Charge Checks: Filtering out obviously incorrect valence bond combinations or charge distributions; 3. Protonation / Neutrition: Adjusting the protonation state of specific functional groups as needed to obtain structures under more reasonable neutral or physiological pH conditions. If a progeny molecule fails the checks in any of the above steps, it is discarded and does not proceed to the subsequent multi-objective evaluation process. By using MCS as a crossover anchor point and combining it with rigorous chemical validity checks, this invention can efficiently recombine key structural units of different molecules while maintaining the rationality of the parent skeleton, generating progeny molecules with new skeleton combinations, thereby significantly expanding the reachable chemical space. Throughout the candidate generation stage, regular mutation, MCS crossover, and FragMLM semantic mutation work synergistically. Regular mutation is responsible for local, synthetically friendly fine-tuning; MCS crossover is responsible for recombination between known skeletons; and FragMLM introduces entirely new fragments and skeletons from the global semantic space. The combined use of these three methods ensures chemical feasibility and significantly reduces the risk of premature convergence in the search.
[0072] For the semantically guided candidate generation module, three operators are used to fuse at each generation stage: a) FragMLM semantic variation: Predicts and replaces / assembles fragment sites according to the current mask ratio; (detailed above) b) Rule variation: Calls the rule base for local reactive editing; c) Crossover (recombination): Performs MCS crossover on individuals to generate a new backbone. Outputs a candidate molecule pool (after deduplication).
[0073] The specific execution logic of the three-way operator is as follows: First, in FragEvo, there are three types of operators for generating new candidate molecules: (1) FragMLM generation operator, (2) crossover operator, and (3) mutation operator. Their collaborative process and proportion control are as follows: 1. Triggering and output scale of each operator Each generation Execute the following sequence: (1) From the previous generation of the father Pure SMILES were extracted from the docking results and used as the basic population of GA. (2) For the parent molecule Fragmentation and dynamic masking are performed, and FragMLM is called to generate new molecules, which serve as "progeny offspring". By default, each mask prefix is sampled a certain number of times. That is, for a given molecule, the FragMLM model generates three different new molecules by masking it. This value is adjustable; in a preferred embodiment, this parameter is set to 1. Therefore, the number of candidates generated by FragMLM is approximately [number missing]. ≈ c· ; (3) Combine the parent generation + FragMLM products into the GA input pool; (4) Call the crossover operator on the input pool and generate approximately according to the configured crossover_attempts target. Crossover offspring; (5) Use the same input pool to call the mutation operator and generate approximately according to the configured mutation_attempts target. A mutant offspring.
[0074] The `crossover_attempts` and `mutation_attempts` parameters allow for explicit control over the triggering ratio of crossover / mutation relative to the FragMLM generation operator. : : This balances the contributions of "learning-based generation" and "evolutionary recombination." Furthermore, it's worth noting that in this FragEvo model, crossover and mutation are parallel processes; that is, the crossover operator and the mutation operator can be executed in parallel to reduce model evolution time.
[0075] 2. Rules for resolving conflicts between operators (repeated numerators) 1) First-level deduplication: After crossover and mutation, the two child files are merged into a unified child pool and deduplicated according to the SMILES string; if the same molecule is generated multiple times by different operators, only 1 entry is kept.
[0076] 2) Second-level deduplication: Before making a selection, the "parent docking result + child docking result" are merged; if the same SMILES exists in both the parent and child generations, the docking scores are compared, and only the record with the better docking score (smaller value) is retained.
[0077] By employing a strategy of "two-level deduplication based on SMILES + selective retention based on docking score", we ensure that the effective diversity of the population is not weakened by duplicate individuals occupying slots among different operators.
[0078] 3. Fusion and synergistic strategies for candidate molecules 1) At the generation level: FragMLM uses dynamic masking to extract 5 "structural prefixes" from the parent, and the generated molecules are directly incorporated into the GA input pool, so that crossover / mutation can work together in the same structural space (i.e., MLM first generates a new backbone, and GA then uses it as material for recombination and modification).
[0079] 2) At the evaluation level: crossover and mutation offspring are first merged, and the entire offspring pool is scored uniformly. 3) At the selection level: The next generation parent is always selected from the unified docking results of "parent generation + all offspring (including FragMLM, crossover, and mutation paths)" through multi-target NSGA. II or Comprehensive A global selection process is performed to ensure that the products of the three operators compete within the same Pareto framework, thereby naturally forming a collaborative optimization.
[0080] The chemical validity and basic filtering module is used for SMILES validity verification, valence bond and ring structure checking, and initial screening of toxicity / PAINS substructures (optional). The output is a set of candidates that pass the filter.
[0081] The multi-objective scoring module is used to calculate the score for each objective for each candidate molecule: a) Docking Score DS: Docking and scoring based on a given protein pocket (values are negative, the smaller the value, the better the docking effect); b) QED: Drug compatibility index (the higher the value, the better the properties of the molecular drug); c) SA: Synthesis feasibility score (the smaller the value, the easier the molecule is to synthesize).
[0082] The Multi-Objective Optimization and Elite Preservation Module (NSGA-II) is used to perform non-dominated sorting and crowding distance calculation, conduct selection, elite preservation, and next-generation population construction; maintaining the diversity and robustness of the Pareto front. It outputs the next-generation population and the current Pareto solution set.
[0083] The specific rules of the elite retention mechanism are as follows: FragEvo in different selection modes (NSGA-II and comprehensive) The strategy employed a two-tiered elite retention approach: one tier was "intergenerational elite retention" (combining parent and child selection), and the other tier was "docking". "Scores are explicitly reserved for elite players."
[0084] 1. Intergenerational elite retention (applicable to NSGA) II. Multi-target / Single-target selection) The input for each generation's selection phase is not simply the offspring, but a combined set of "parent generation docking results + the entire offspring docking results": First, the docking results of the parent and offspring are merged according to SMILES. If the same molecule exists in both, the record with the better docking score (smaller value) is retained. Then, single-objective selection (Rank / Roulette / Tournament) or NSGA is applied to this merged set. II. Multi-objective selection: The selected N_select molecules are directly used as the next generation's parents. Since the "next generation's parents" are always generated from the entire set of "parents + children", and the higher-scoring individuals are always retained for duplicates, any high-quality molecules that have appeared will not be completely eliminated in the selection process as long as they still maintain a relatively good index in the current generation. This forms an implicit elite retention at the algorithmic level.
[0085] 2. Comprehensive The retention of elites under this model, in comprehensive In the scoring selection strategy, the standardized docking score is first calculated for the merged parent + child samples. Composite score and comprehensive ; = · · Then make a selection. Highest Each molecule is retained.
[0086] 3. Elite Population Assessment and Recording: For each selected parent generation (i.e., the elite population), the system will run a separate statistical evaluation script to calculate and output statistics such as DS, QED, and SA for that generation of elites. The "best overall score" molecule will be recorded in a summary file for subsequent analysis and visualization. Although this step does not participate in the optimization decision-making, it demonstrates the interpretability and traceability of elite group behavior in this invention.
[0087] Termination and Output: Reaching After determining whether to replace or stop early, export the Top candidates, along with the source path and scoring details for each individual.
[0088] For the dynamic mask scheduling module, used by algebraic It dynamically adjusts the FragMLM mask ratio to achieve a balance between wide-area exploration in the early stage and local convergence in the later stage; it can be linked to temperature and top-k.
[0089] The parallelization and resource scheduling module is used for batch parallelization of computation and property evaluation; queued scheduling of GPU / CPU resources; failure retry and timeout recycling.
[0090] The data management and auditing module is used to record each generation of the population, parent-child relationships (source tree), operator type, random seed, scoring details, Pareto curves, and key intermediates; this facilitates reproducibility and compliance review.
[0091] In a specific embodiment, the molecular multi-objective optimization method based on language models and evolutionary algorithms provided in this embodiment is as follows: Task initialization includes reading protein pockets and docking parameters, setting the target vector format and direction (NSGA-II does not require manual weighting, but the target direction must be set), and determining the population size. Maximum Algebra Candidate amplification fold Random seed, parallel threads.
[0092] The initial population was constructed by stratified random sampling from the ZINC molecular library. The initial population consists of 10 individuals; initial scores and normalization are performed, and docking score, QED, and SA are calculated for all individuals in generation 0, and converted into a "smaller is better" target vector; results are cached. The target vector is uniformly converted to a "smaller is better" transformation logic (NSGA-II). The original indicators for optimization in the multi-objective selection stage include: docking score DS, drug-likeness score QED, and synthetic feasibility score SA. Among them, DS and SA are themselves "smaller is better" (DS itself is a negative value), while QED is "larger is better". To uniformly use the "smaller is better" target vector, the following transformation is actually performed on each candidate molecule: 1. Define the original goal: DS: Molecular-receptor docking score (usually negative, the smaller the value, the better the binding); QED∈[0,1]: Drug similarity score, the larger the value, the better; SA: Synthesis feasibility score, the smaller the value, the easier it is to synthesize.
[0093] 2. The unified minimized objective vector Let the transformed unified objective vector be . ,but: (Maintain the original orientation of the docking score and minimize it directly); (Maximize) Convert to "minimize" (”); (Keep Original direction, directly minimize).
[0094] 3. General Explanation For any metric that "needs to be maximized" ,use Transformation; for any metric that "needs to be minimized". ,Keep Unchanged; both multi-objective ranking and congestion calculation are based on the unified version. Completed, thus ensuring NSGA Algorithms such as II can unambiguously compare dominance relationships within the same "smallest possible" target space. Ultimately, multiple targets are unified into a "smallest possible" target vector: .
[0095] Iterate through the main loop ( = 1… ): 1. Dynamic mask update: The scheduler sets the mask ratio for the current generation. Sampling temperature , top-k.
[0096] 2. Candidate generation: (Three parallel paths generate candidate molecules) a) Semantic variation: For the current population, according to... Masking, calling FragMLM for reconstruction, sampling to generate candidates; calling FragMLM for fragment-level reconstruction of the current population according to the current generation masking strategy, sampling to generate candidate molecules, the specific process can be found in the semantic mutation steps above.
[0097] b) Rule mutation: Invoke the reaction / replacement rule; In this embodiment, rule variation is preferably selected from a pre-organized library of 94 organic reaction rules. A one-step reaction transformation is performed on a specific functional group or fragment of the parent molecule, such as site-specific aromatic ring substitution, amide bond formation, or heterocyclic ring closure. For each rule, it is first determined whether the parent molecule contains a matching reaction center and a dissociative group. If a match is successful, one or more candidate product molecules are generated based on the reaction template. The product molecules need to pass SMILES validity, valence state, and ring structure checks; those that fail are directly eliminated.
[0098] c) Crossover recombination: Select individuals from the current population for pairing and generate a new skeleton based on MCS crossover.
[0099] Specifically, based on the NSGA-II ordination results, elite individuals or those with high diversity are paired up as parents. For each parent pair, the RDKit is used to calculate the maximum common substructure (MCS). If a valid MCS is found, the two parents are divided along the MCS into "common backbone + respective side chains." New composite structures are then constructed by exchanging side chains or selectively retaining some substituents, and the SMILES merging process in AutoGrow 4.0 is used to generate daughter molecules. The generated daughter molecules undergo chemical correction and valence state checks; those that fail are discarded, while those that pass proceed to filtering and deduplication steps. This crossover strategy can introduce new substitution patterns while retaining favorable backbones, enhancing cross-backbone jumping and structural recombination capabilities.
[0100] 3. Filtering and Deduplication: Perform validity, substructure and simple similarity filtering to maintain candidate diversity. 4. Multi-objective evaluation: Batch docking scoring, calculation of QED and SA, and formation of target vectors.
[0101] 5. NSGA-II Optimization: Perform non-dominated sorting and crowding calculation on the "current population ∪ candidate pool", select the top N as the next generation population, and output / update the Pareto front.
[0102] After obtaining the unified objective vector of each candidate molecule, this invention adopts non-dominated sorting multi-objective selection based on crowding degree to maintain the diversity of solutions on the Pareto front. The crowding degree calculation process is as follows.
[0103] Non-dominated sorting first involves sorting all candidate molecules non-dominated according to a unified target vector, resulting in several Pareto fronts F1, F2, ..., where individuals within each front have the same dominance level.
[0104] Frontal congestion calculation. For a given frontal F containing K individuals and with a target dimension of M, this invention initializes the congestion degree for each individual. And accumulate the crowding level for each target dimension using the following steps: (1) Sort the individuals in the front edge according to the value of the target dimension from smallest to largest to obtain the sorted index sequence; (2) Set the crowding degree of the minimum and maximum value individuals in the target dimension to infinity as the boundary solution; (3) If the maximum and minimum values of the target dimension are not equal, then for each intermediate individual, normalize and accumulate the difference between its previous and next neighbors on the target axis into the crowding degree, that is, measure the local density in the form of "difference between adjacent target values / target interval width".
[0105] When it is necessary to select only a subset of individuals from a given frontier, this invention prioritizes individuals with higher crowding density, thereby preserving a more evenly distributed and comprehensive solution set in the multi-objective space, effectively avoiding the concentration of solutions in a localized region. The resulting crowding density... The larger the value, the sparser the "neighborhood" of the solution in the multi-objective space. Therefore, when it is necessary to select a small number of individuals from the current frontier to fill the next generation population, individuals with high crowding will be retained first, thereby maintaining the coverage and diversity of the entire Pareto frontier.
[0106] 6. Recording and Auditing: Store the current generation's parent-child relationships, operator labels, score distribution, and convergence curve (which can be used to plot "optimized forest" and Pareto trajectory later).
[0107] Termination and result summary: The termination condition is reaching the maximum algebra. The goal is to stop early or exhaust the budget. Output the final Pareto frontier set (including SMILES / structure, docking score, QED, SA), the Top-k candidate list, and the source path and synthesizability summary for each candidate.
[0108] The hyperparameter settings in this embodiment can use the recommended default values or be adjusted within a reasonable range to suit the needs of different targets, computational budgets, or project stages. Typical settings and optional implementations include: 1. Population size The initial value should be between 100 and 200. When computational resources are plentiful and a richer Pareto solution set is desired, the value can be increased. Take around 200; when resources are limited or there is a greater focus on rapid iteration, you can... Take 80-120.
[0109] 2. Maximum number of iterations Limited to 20-25 generations For early-stage exploration or scenarios requiring rapid results, it is advisable to... =10~15 (used in the main experiment); for fine-tuning of important targets, it can be... Extended to generations 30 and above.
[0110] The dynamic masking module automatically schedules the number of fragment masks based on the algebra to balance early wide-area exploration with later local refinement. For any algebra g∈[1, The number of mask fragments in the current generation. It can be calculated using the following formula: in: The number of masks represents the first generation population. The optimal strategy is to retain the first segment for each molecule and mask all other segments. Set the number of iterations to execute synchronously (maximum number of generations); This indicates the number of segments masked for the last generation of the population. The default value is 1, meaning only the last segment of each molecule is masked. < Schedule factor ; In practice, the same masking strategy described above is applied to all molecules.
[0111] Specifically, the quantization details of the dynamic mask scheduling used in this invention are as follows: 1. Mask ratio and calculation formula For each parent molecule, it is first fragmented to obtain a length of The goal of dynamic masking is to "mask more fragments" in the early stages of evolution (giving FragMLM more generation space) and "mask fewer fragments" in the later stages (fixing more existing structures and making local optimizations).
[0112] Let the current algebra be... The total algebra is The number of mask segments that are linearly decaying for a single molecule is: Initial mask fragment number (Only the first segment is retained as the starting point); Final mask fragment count (Only mask the last segment, retaining most of the structure); The formula for calculating the number of mask segments in each generation can be found in the previous explanation.
[0113] 2. Fragment Sequence Construction and Linkage Parameters In calculation Then, the list of fragments Employ a "tail mask" strategy: The visible fragment is ; The constructed input prefix is .
[0114] Therefore, the number of mask fragments With the total number of segments Algebra The interaction of these three factors determines the length of the visible context, forming a dynamic mask curve that changes linearly with the evolution process: the earlier the stage, the shorter the preserved fragments and the larger the generation space; the later the stage, the more preserved fragments and the more the generation is biased towards local modification.
[0115] 3. Scheduling Triggering Conditions When `dynamic_masking.enable = true` and `max_generations > 1` in the configuration, the system will call the above dynamic mask calculation in each generation's "decomposition and masking" phase. If `dynamic_masking` is disabled, it will fall back to fixed mask mode, where `n_fragments_to_mask` or `(initial_mask_fragments, final_mask_fragments)` in the configuration will use the same linear interpolation formula to provide a constant mask at the generation level. This parameter is set to adjust the masking strategy. The introduction of this control parameter explicitly represents the dynamic masking strategy of the model.
[0116] If N ≤ 1 for a single molecule, it is determined to be an invalid mask, and the shortest prefix [BOS][SEP] is returned directly to avoid anomalies.
[0117] The introduction of dynamic masking mechanism allows the number of mask fragments to approach [a certain value] in the early stages of the search (when g is small]. This involves masking and reconstructing large segments of a molecule to encourage the generation of new skeletons and achieve wide-area exploration; in the later stages (when g approaches G), the number of masked segments gradually approaches... By replacing only a small number of key segments, local refinement and property fine-tuning are achieved, thereby realizing an adaptive balance of "broad first and narrow later" throughout the evolutionary process.
[0118] Regarding FragMLM sampling control, to further control the diversity and quality of new fragments provided by FragMLM, this invention can synchronize the following sampling parameters with the algebra: 1. Temperature In the early stages of the search, a higher temperature (e.g., T=1.0~1.2) is used to enhance randomness and structural diversity; in the later stages, the temperature is gradually reduced to 0.7~0.9, so that the generated results are more concentrated in high-probability fragment combinations with more stable chemical semantics.
[0119] 2. Sampling: In the early stages, a larger top-k (e.g., 50-100) can be used to encourage "jumping out" of the current skeleton; as the iteration progresses, the top-k is gradually narrowed to a smaller one (e.g., 10-20) to improve the determinism of generated fragments and optimization efficiency.
[0120] Typical NSGA-II parameter settings in the multi-objective optimization module are as follows: 1. Definition of target vector Docking Score: Directly used as the target to be minimized; QED: Unified as the "smaller the better" target by taking the negative form; SA: Directly used as the target to be minimized.
[0121] 2. Crossover and mutation probabilities The selection operations within the NSGA-II framework are shared with the three types of operators (semantic mutation, regular mutation, and crossover recombination) of this invention, typically set as follows: crossover probability 0.8–0.9, mutation probability 0.1–0.2.
[0122] 3. Crowded Distance and Deduplication Strategies After performing non-dominated sorting in each generation, individuals with larger crowding distances are retained first to maintain a uniform distribution of the Pareto front in the target space. After merging the parent and offspring generations, SMILES are normalized to remove duplicates. If duplicates are found, individuals with better pairing scores are retained to reduce invalid evaluations.
[0123] Regarding parallelism and caching mechanisms: 1. Parallel evaluation: Dock computation and QED / SA computation can be submitted to the CPU / GPU cluster for parallel execution in batches, with a typical batch size of 32 to 256.
[0124] 2. Result caching: For repeated molecules or substructures whose structures have not changed, historical scores can be reused to avoid repeated docking and property calculations, thereby reducing the total number of evaluations and computational costs.
[0125] 3. Retry on failure and timeout recovery: If a single docking task fails or times out due to an error, it will automatically retry 1 to 3 times. Molecules that fail to dock multiple times can be marked as "non-docking" and will not be retried in subsequent generations to improve overall stability.
[0126] Regarding optional extension targets: In addition to the current three targets of Docking / QED / SA, this invention also reserves an extension interface that can be connected according to the actual project needs: ADMET-related predictive indicators (such as LogP, solubility, and preliminary toxicity prediction); Specific physicochemical properties (such as polarizability, coordination ability, glass transition temperature, etc.); Accessibility score or laboratory feasibility score for a specific reaction pathway.
[0127] In this embodiment, 10 representative protein targets were selected, including various GPCR and kinase targets from the DUD-E benchmark library, as well as the SARS-CoV-2 main protease Mpro, to verify the applicability and robustness of the present invention in a multi-target scenario. In addition, AutoDock vina was used to dock small molecules and target proteins.
[0128] 1. Initial Setup Approximately 250,000 drug pattern molecules were selected from the ZINC database and stratified into molecular weight ranges of <100, 100–150, 150–200, and 200–250 Da. A candidate library was constructed by random sampling in a ratio of 2:3:4:1. N=200 molecules were randomly selected from the candidate library as the generation 0 population, and SMILES validity verification, valence state check, and basic PAINS / toxicity substructure filtering were performed. A docking grid was generated for each target, with a uniform pocket size of (15Å, 15Å, 15Å). The docking engine used was AutoDock Vina. The docking score, QED, and SA were calculated for all molecules in generation 0 and uniformly converted into a target vector that is "smaller is better" as the initial input for NSGA-II.
[0129] 2. Dynamic Masking and Candidate Generation Set the maximum number of iterations to 20, and calculate the number of mask segments in each generation according to the aforementioned formula. ; Semantic variation: For each individual in the current population, with The FragSeq sequence is masked, and the masked sequence is fed into a pre-trained FragMLM at a temperature... Replacement fragments are generated by top-k constraint downsampling; Rule mutation: Apply the 94 reaction rules in AutoGrow 4.0 (including the AutoClickChemRxn and RobustRxn rules) to the same batch of individuals, and perform local growth, truncation or replacement operations; Crossover recombination: Individuals with high diversity or high current NSGA-II ranking are extracted in pairs and crossover recombination is performed based on MCS to generate progeny molecules with new backbones.
[0130] 3. Filtering, evaluation and selection Candidates generated by semantic variation, rule variation, and crossover are merged, and SMILES validity checks, valence checks, loop structure checks, structural deduplication, and similarity deduplication based on a Tanimoto similarity threshold (e.g., 0.9) are performed. Candidate molecules that pass the filtering are batch-docked and scored, and QED and SA are calculated to form the candidate score set for the current generation. The current parent population is merged with the candidate set, and NSGA-II non-dominated sorting and crowding calculation are performed according to the three-objective combination of "minimum Docking Score, maximum QED (negative), and minimum SA". The top N individuals are selected to form the next generation population, and the Pareto front of the current generation is updated.
[0131] like Figure 3 As shown, Figure 3This is a comparative diagram of the molecular representation methods in this embodiment. The left side shows the SMILES representation of the same molecule and the corresponding FragSeq fragment sequence representation. The right side is a schematic diagram of the overall structure of the FragEvo molecule generation and optimization process provided in this embodiment.
[0132] In summary, compared with the prior art, the molecular multi-objective optimization method based on language model and evolutionary algorithm provided in this embodiment, by integrating the semantic-guided generation mechanism based on molecular language model and the multi-objective optimization algorithm for evolutionary iterative optimization, can take into account both the wide-area exploration of chemical space and the goal-oriented convergence, and achieve balanced optimization of molecular multi-objectives, thereby generating high-quality protein-targeting molecules that meet multiple key performance indicators.
[0133] To facilitate better implementation of the molecular multi-objective optimization method based on language models and evolutionary algorithms in the embodiments of this application, this application also provides a molecular multi-objective optimization device based on language models and evolutionary algorithms, which is based on the aforementioned molecular multi-objective optimization method based on language models and evolutionary algorithms. The meanings of the terms used are the same as in the aforementioned molecular multi-objective optimization method based on language models and evolutionary algorithms, and specific implementation details can be found in the descriptions in the method embodiments.
[0134] Please see Figure 4 , Figure 4 This is a schematic diagram of the molecular multi-objective optimization device based on language models and evolutionary algorithms provided in an embodiment of this application. Specifically, this molecular multi-objective optimization device may include an initial processing module 201, an iterative optimization module 202, and a loop control module 203, as follows: The initial processing module 201 is used to construct an initial molecular population and perform multi-objective evaluation on each molecule in the initial molecular population to obtain a multi-objective evaluation vector corresponding to each molecule. The iterative optimization module 202 is used to perform evolutionary iterative optimization based on the current molecular population and its corresponding multi-objective evaluation vector by integrating a semantic-guided generation mechanism based on molecular language model and a multi-objective optimization algorithm to update and obtain a new generation of molecular population. The loop control module 203 is used to repeat the evolutionary iteration optimization process until the preset termination condition is met, and the generated final molecular population is taken as the optimization result.
[0135] For specific limitations regarding the molecular multi-objective optimization device based on language models and evolutionary algorithms, please refer to the limitations of the molecular multi-objective optimization method based on language models and evolutionary algorithms mentioned above, which will not be repeated here. Each module in the aforementioned molecular multi-objective optimization device based on language models and evolutionary algorithms can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.
[0136] The molecular multi-objective optimization device based on language models and evolutionary algorithms provided in this embodiment integrates a semantic-guided generation mechanism based on molecular language models and a multi-objective optimization algorithm for evolutionary iterative optimization. This approach can balance the broad-area exploration of chemical space with goal-oriented convergence, achieving balanced optimization of molecular multi-objectives and generating high-quality protein-targeting molecules that meet multiple key performance indicators.
[0137] Furthermore, embodiments of this application also provide an electronic device, such as... Figure 5 As shown, it illustrates a structural schematic diagram of the electronic device involved in the embodiments of this application, specifically: The electronic device may include components such as a processor 301 with one or more processing cores, a memory 302 with one or more computer-readable storage media, a power supply 303, and an input unit 304. Those skilled in the art will understand that... Figure 5 The electronic device structure shown does not constitute a limitation on the electronic device and may include more or fewer components than shown, or combine certain components, or have different component arrangements. Wherein: The processor 301 is the control center of the electronic device. It connects various parts of the electronic device via various interfaces and lines, and performs various functions and processes data by running or executing software programs and / or modules stored in the memory 302, and by calling data stored in the memory 302, thereby providing overall monitoring of the electronic device. Optionally, the processor 301 may include one or more processing cores; preferably, the processor 301 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 301.
[0138] The memory 302 can be used to store software programs and modules. The processor 301 executes various functional applications and molecular multi-objective optimization methods based on language models and evolutionary algorithms by running the software programs and modules stored in the memory 302. The memory 302 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the electronic device, etc. In addition, the memory 302 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 302 may also include a memory controller to provide the processor 301 with access to the memory 302.
[0139] The electronic device also includes a power supply 303 that supplies power to various components. Preferably, the power supply 303 can be logically connected to the processor 301 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply 303 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.
[0140] The electronic device may also include an input unit 304, which can be used to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.
[0141] Although not shown, the electronic device may also include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 301 in the electronic device loads the executable files corresponding to the processes of one or more applications into the memory 302 according to the following instructions, and the processor 301 runs the applications stored in the memory 302 to realize various functions, as follows: An initial molecular population is constructed, and each molecule in the initial molecular population is evaluated for multiple objectives to obtain a multi-objective evaluation vector for each molecule. Based on the current molecular population and its corresponding multi-objective evaluation vector, an evolutionary iterative optimization is performed by integrating a semantically guided generation mechanism based on a molecular language model and a multi-objective optimization algorithm to update and obtain a new generation of molecular population. The evolutionary iterative optimization process is repeated until a preset termination condition is met, and the generated final molecular population is taken as the optimization result.
[0142] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.
[0143] This application embodiment integrates a semantically guided generation mechanism based on a molecular language model and a multi-objective optimization algorithm for evolutionary iterative optimization. This approach balances the broad-area exploration of chemical space with goal-oriented convergence, achieving balanced optimization of multiple molecular objectives and generating high-quality protein-targeting molecules that meet multiple key performance indicators.
[0144] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.
[0145] Therefore, embodiments of this application provide a storage medium storing multiple instructions that can be loaded by a processor to execute steps in any of the molecular multi-objective optimization methods based on language models and evolutionary algorithms provided in embodiments of this application. For example, the instructions can execute the following steps: An initial molecular population is constructed, and each molecule in the initial molecular population is evaluated for multiple objectives to obtain a multi-objective evaluation vector for each molecule. Based on the current molecular population and its corresponding multi-objective evaluation vector, an evolutionary iterative optimization is performed by integrating a semantically guided generation mechanism based on a molecular language model and a multi-objective optimization algorithm to update and obtain a new generation of molecular population. The evolutionary iterative optimization process is repeated until a preset termination condition is met, and the generated final molecular population is taken as the optimization result.
[0146] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.
[0147] The storage medium may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0148] Since the instructions stored in the storage medium can execute the steps of any of the molecular multi-objective optimization methods based on language models and evolutionary algorithms provided in the embodiments of this application, the beneficial effects that any of the molecular multi-objective optimization methods based on language models and evolutionary algorithms provided in the embodiments of this application can achieve can be realized, as detailed in the preceding embodiments, and will not be repeated here.
[0149] The foregoing has provided a detailed description of a molecular multi-objective optimization method and apparatus based on language models and evolutionary algorithms provided in the embodiments of this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are only for the purpose of helping to understand the method and its core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A molecular multi-objective optimization method based on language models and evolutionary algorithms, characterized in that, include: An initial molecular population is constructed, and a multi-objective evaluation is performed on each molecule in the initial molecular population to obtain the multi-objective evaluation vector corresponding to each molecule; Based on the current molecular population and its corresponding multi-objective evaluation vector, an evolutionary iteration optimization is performed by integrating a semantically guided generation mechanism based on a molecular language model and a multi-objective optimization algorithm to obtain a new generation of molecular population; The evolutionary iteration optimization process is repeated until the preset termination condition is met, and the resulting final molecular population is taken as the optimization result.
2. The molecular multi-objective optimization method based on language model and evolutionary algorithm according to claim 1, characterized in that, The process of constructing an initial molecular population and performing multi-objective evaluation on each molecule in the initial molecular population to obtain a multi-objective evaluation vector for each molecule includes: According to the preset sampling strategy, multiple molecules are extracted from the molecular database to form an initial molecular set; Each molecule in the initial molecule set is subjected to a chemical validity check to filter out invalid structures, thereby obtaining the initial molecule population; Perform a multi-objective evaluation on each molecule in the initial molecular population to obtain the corresponding multi-objective evaluation vector.
3. The molecular multi-objective optimization method based on language model and evolutionary algorithm according to claim 2, characterized in that, The preset sampling strategy is a stratified sampling strategy. The step of extracting multiple molecules from the molecular database according to the preset sampling strategy to form an initial molecule set includes: Molecules in a molecular database are stratified based on at least one attribute, such as molecular weight, structural type, or source database. Molecules are extracted from each layer according to a preset ratio to form the initial molecular set.
4. The molecular multi-objective optimization method based on language model and evolutionary algorithm according to claim 1, characterized in that, The process involves using the current molecular population and its corresponding multi-objective evaluation vectors to perform evolutionary iterative optimization by integrating a semantically guided generation mechanism based on a molecular language model and a multi-objective optimization algorithm, thereby updating and obtaining a new generation of molecular populations, including: A batch of new candidate molecules is generated based on the current molecular population using at least one generation mechanism, said generation mechanism including a semantically guided generation mechanism based on a pre-trained molecular language model; The new candidate molecules are subjected to multi-objective evaluation to obtain the corresponding multi-objective evaluation vector, and the multi-objective evaluation vector is merged with the current molecule population to form a candidate set. Based on the multi-objective evaluation vectors of all molecules in the candidate set, a multi-objective optimization algorithm is used to select the next-generation molecular population.
5. The molecular multi-objective optimization method based on language model and evolutionary algorithm according to claim 4, characterized in that, The method of generating a batch of new candidate molecules based on the current molecular population using at least one generation mechanism, wherein the generation mechanism includes a semantically guided generation mechanism based on a pre-trained molecular language model, comprising: The molecules in the current molecular population are segmented based on chemical reaction rules to obtain a sequence of segments composed of multiple structural fragments according to their connection relationships; A masking operation is performed on selected segments in the segment sequence to generate an input sequence containing at least one mask position; The input sequence is input into the pre-trained molecular language model, which predicts and fills in the new segments corresponding to the masked positions based on the chemical context provided by the unmasked segments, thus forming a predicted segment sequence. According to predetermined connection rules, the fragments in the predicted fragment sequence are chemically bonded together to reconstruct new candidate molecules.
6. The molecular multi-objective optimization method based on language model and evolutionary algorithm according to claim 5, characterized in that, The masking operation employs a dynamic masking strategy. The step of performing a masking operation on selected segments in the segment sequence to generate an input sequence containing at least one mask position includes: Based on the current evolutionary iteration generation, and according to the preset dynamic scheduling rules, calculate the number or proportion of mask segments required for this iteration. Based on the number or proportion of segments to be masked in this iteration, a corresponding number of segments are selected from the segment sequence for masking to generate the input sequence.
7. The molecular multi-objective optimization method based on language model and evolutionary algorithm according to claim 4, characterized in that, The generation mechanism also includes rule-based mutation based on chemical reaction rules and cross-recombination based on the greatest common substructure between molecules. The step of performing multi-objective evaluation on the new candidate molecules to obtain the corresponding multi-objective evaluation vector, and merging the multi-objective evaluation vector with the current molecular population to form a candidate set, further includes: All generated candidate molecules are checked for atomic valence state, bonding mode and ring structure rationality, and molecules that do not conform to chemical rules are filtered out; Candidate molecules that pass the chemical filtration are compared based on normalized molecular representations to remove duplicate molecular structures.
8. The molecular multi-objective optimization method based on language model and evolutionary algorithm according to claim 4, characterized in that, The step of performing a multi-objective evaluation on each molecule in the initial molecular population to obtain a corresponding multi-objective evaluation vector includes: Based on the multi-objective evaluation vectors of all molecules in the candidate set, a non-dominated sort is performed to divide the molecules into several Pareto front levels. For molecules within the same Pareto front level, calculate the crowding distance in the multi-object space; According to the Pareto front order from high to low, and within the same level according to the crowding distance from large to small, a predetermined number of molecules are selected to form the new generation of molecular population.
9. The molecular multi-objective optimization method based on language model and evolutionary algorithm according to claim 1, characterized in that, The repeated evolutionary iterative optimization process continues until a preset termination condition is met, and the resulting final molecular population is taken as the optimization result, including: After each iteration, it is determined whether a preset termination condition is met. The termination condition includes reaching the maximum number of iterations or population performance convergence. When the preset termination condition is met, the iteration stops, and the molecular population obtained from the last iteration is output as the optimization result. The optimization result includes at least the molecules located at the Pareto front and their corresponding multi-objective evaluation vectors.
10. A molecular multi-objective optimization device based on language models and evolutionary algorithms, characterized in that, include: An initial processing module is used to construct an initial molecular population and perform multi-objective evaluation on each molecule in the initial molecular population to obtain a multi-objective evaluation vector corresponding to each molecule. The iterative optimization module is used to perform evolutionary iterative optimization based on the current molecular population and its corresponding multi-objective evaluation vector by integrating a semantic-guided generation mechanism based on molecular language model and a multi-objective optimization algorithm to obtain a new generation of molecular population; The loop control module is used to repeat the evolutionary iterative optimization process until the preset termination condition is met, and the resulting final molecular population is taken as the optimization result.
Citation Information
Cited By
Explanatable molecular optimization method based on large language model
CN121983172A