Method, device and equipment for generating aggregated induced emission molecules based on reinforcement learning

Through a reinforcement learning-based method, the SMILES expression of each chemical molecule in the molecular training set is preprocessed to generate a Token sequence. The pre-trained molecular generation model and AIE activity prediction model are used to construct a multi-objective reward function and optimize the molecular generation model parameters. This solves the problems of low design efficiency and limited structural diversity in existing technologies, and achieves efficient and accurate aggregation-induced emission molecule generation.

CN120544707BActive Publication Date: 2025-10-10SHENZHEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511037910.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-28
Publication Date
2025-10-10
Estimated Expiration
2045-07-28

AI Technical Summary

Technical Problem

Existing technologies are unable to efficiently and accurately design new aggregation-induced emission molecules, and have problems such as low design efficiency, limited structural diversity, and difficulty in systematically covering the chemical space.

Method used

A reinforcement learning-based method is used to preprocess the SMILES expression of each chemical molecule in the molecular training set to generate a token sequence. The pre-trained molecular generation model is used to generate a set of candidate molecules, which are scored using the AIE activity prediction model. A multi-objective reward function is constructed, and adaptive weighted reinforcement learning is performed to optimize the parameters of the molecular generation model to ultimately generate aggregation-induced emission molecules.

Benefits of technology

The generation efficiency and quality of AIE molecules were improved, dynamic trade-offs and collaborative optimization among multiple objectives were achieved, and the generated molecules performed excellently in multiple indicators.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120544707B_ABST
    Figure CN120544707B_ABST
Patent Text Reader

Abstract

The application is suitable for the field of material science and technology, and provides an aggregation-induced emission molecule generation method, device and equipment based on reinforcement learning, which comprises the following steps: preprocessing SMILES expressions of chemical molecules in a molecular training set to generate Token sequences represented by SELFIES; generating a candidate molecule set by using a molecule generation model according to the Token sequences; predicting AIE activity of candidate molecules in the candidate molecule set by using an AIE activity prediction model; constructing a multi-objective reward function based on AIE activity scores obtained by prediction; taking the multi-objective reward function as an optimization target, iteratively optimizing parameters of the molecule generation model by using a reinforcement learning strategy with adaptive weights; and generating AIE molecules by using the optimized molecule generation model, so that a multi-objective guidance mechanism is constructed by using reinforcement learning, the generated molecules dynamically balance and synergistically optimize among multiple targets, and the generation efficiency and quality of the AIE molecules are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of material science and technology, and in particular relates to a method, device and equipment for generating aggregation-induced luminescence molecules based on reinforcement learning. Background Art

[0002] With the rapid development of high-end applications such as smart materials, bioimaging, optoelectronic devices, and sensors, the demand for functional molecules with excellent luminescence properties, aggregation stability, and structural tunability is increasing. Traditional fluorescent molecules often suffer from the aggregation-induced quenching (ACQ) effect in their aggregated state, resulting in a significant decrease in their performance in the solid state, concentrated phase, or biological environments. To address this issue, the concept of aggregation-induced emission (AIE) materials has emerged. AIE molecules emit no or weak fluorescence in dilute solutions but emit strong fluorescence in their aggregated state, overcoming the ACQ drawbacks of traditional molecules and showing broad potential in solid-state lighting, bioprobes, drug delivery, high-contrast imaging, and near-infrared (NIR) optical technologies. Despite the significant advantages of AIE materials, their molecular design still faces significant challenges. Current research and development of AIE molecules still primarily relies on fragment assembly and empirical rules: Using typical AIE motifs such as TPE (tetraphenylethylene), TPA (triphenylamine), and TPP (triphenylphosphine) as core elements, novel molecules are constructed by artificially combining functional groups, and their AIE activity is then verified experimentally. This trial-and-error strategy has several limitations. First, design efficiency is low: relying on expert experience and repeated experiments results in long screening cycles and high costs. Second, structural diversity is limited: excessive focus on derivatives of known templates (such as TPE / TPA) makes it difficult to break through structural inertia and explore truly novel molecular space. Third, the design space is narrow: artificial combinations struggle to systematically cover the complex chemical space, and potentially high-performance structures are easily missed.

[0003] At the same time, the rise of advanced technologies such as artificial intelligence, machine learning, and molecular generative models has also provided new solutions for the de novo design and property prediction of functional molecules. For example, in recent years, some studies have attempted to introduce machine learning (such as random forests and support vector machines) to predict the AIE activity of molecules. While this has improved screening efficiency to some extent, it still has some drawbacks. First, it relies on passive screening rather than active design: the models can only evaluate the activity of manually pre-set candidate molecules and lack the ability to automatically generate and optimize molecules. Second, property prediction is one-sided: they focus on a single metric (such as luminescence intensity or energy gap) and ignore the coordinated optimization of multiple objectives such as synthetic feasibility, structural novelty, and biocompatibility. Third, the reliability of molecular representation is insufficient: traditional generative models based on SMILES (Simplified Molecular Input Line Entry System) expressions are prone to outputting illegal or invalid structures, hindering model convergence and practical application. More importantly, because molecular aggregation behavior is influenced by multiple factors such as conformational changes, intermolecular interactions, and electronic structure coupling, AIE properties are highly nonlinear and structure-sensitive. This makes it difficult to accurately predict AIE activity directly from molecular structure, further exacerbating the difficulty of AIE molecule design. Therefore, there is an urgent need to develop a new method that can efficiently and accurately design new AIE molecules to promote the further development of the AIE materials field and meet the growing demand for high-end applications. Summary of the Invention

[0004] The purpose of the present invention is to provide a method, device and equipment for generating aggregation-induced emission molecules based on reinforcement learning, aiming to solve the problem of low quality of AIE molecules generated due to the inability of existing technologies.

[0005] In a first aspect, the present invention provides a method for generating aggregation-induced luminescence molecules based on reinforcement learning, the method comprising the following steps:

[0006] Preprocess the SMILES expression of each chemical molecule in the molecular training set to generate a token sequence represented by SELFIES for each chemical molecule;

[0007] According to each of the token sequences, a set of candidate molecules is generated using a pre-trained molecule generation model;

[0008] Using a pre-trained AIE activity prediction model to predict the AIE activity of each candidate molecule in the candidate molecule set to obtain a corresponding AIE activity score;

[0009] constructing a multi-objective reward function based on the AIE activity score;

[0010] Taking the multi-objective reward function as the optimization target, iteratively optimizing the parameters of the molecular generation model through a reinforcement learning strategy with adaptive weights;

[0011] Aggregation-induced emission molecules were generated using the optimized molecular generation model.

[0012] In some embodiments, the step of preprocessing the SMILES expression of each chemical molecule in the molecular training set includes:

[0013] Standardizing the SMILES expression of each chemical molecule in the molecular training set;

[0014] The standardized SMILES expressions are converted into SELFIES representations, and the SELFIES representations obtained after the conversion are segmented to generate the Token sequence.

[0015] In some embodiments, the step of constructing a multi-objective reward function based on the AIE activity score includes:

[0016] Performing a synthetic feasibility assessment, a structural novelty assessment, and an energy gap assessment on each candidate molecule in the set of candidate molecules for which AIE activity prediction has been completed, and obtaining a corresponding synthetic feasibility score, structural novelty score, and energy gap score;

[0017] The multi-objective reward function is constructed based on the AIE activity score, the synthetic feasibility score, the structural novelty score and the energy gap score. Expressed as ,in, represents the AIE activity score, represents the synthesis feasibility score, represents the structural novelty score, represents the energy gap score, 、 、 、 Respectively 、 、 and The weight of .

[0018] In some embodiments, the molecular generation model adopts a Transformer architecture, including an embedding layer, a decoder layer composed of multiple layers of Transformer decoders, and an output layer. The step of generating a candidate molecule set based on each token sequence using a pre-trained molecular generation model includes:

[0019] Perform vector mapping on each of the token sequences through the embedding layer to generate an embedding vector containing chemical semantics and position information;

[0020] Performing multi-layer feature extraction on the embedding vector through the decoder layer to generate a hidden state sequence that integrates global and local structural information;

[0021] The hidden state sequence is converted into the probability distribution of the next Token through the output layer, and a complete molecular sequence is constructed through a sampling strategy to generate the candidate molecular set.

[0022] In some embodiments, the step of converting the hidden state sequence into a probability distribution of the next token through the output layer and constructing a complete molecule sequence through a sampling strategy to generate the candidate molecule set includes:

[0023] Mapping the hidden state of each time step in the hidden state sequence to a vector space of the vocabulary size through linear transformation to obtain a corresponding vocabulary space vector;

[0024] Normalizing the word space vector through the Softmax layer to obtain the probability distribution of the next Token;

[0025] A Top-k sampling strategy is used to randomly select tokens from the probability distribution, and the molecular sequence represented by SELFIES is gradually constructed;

[0026] The molecular sequence is inversely decoded into a SMILES expression, and the SMILES expression that has passed the molecular structure legality verification is used as a candidate molecule, and all candidate molecules constitute the candidate molecule set.

[0027] In some embodiments, the AIE activity prediction model includes an input layer, a decision layer composed of a random forest classifier integrating multiple decision trees, and an output layer, wherein the input layer is used to calculate the ECFP4 fingerprint of each candidate molecule in the candidate molecule set, the decision layer is used to enable each decision tree to independently judge the AIE activity of the corresponding candidate molecule based on the ECFP4 fingerprint, and use a majority voting mechanism to output a classification result of whether the candidate molecule has AIE activity, and the output layer is used to generate the AIE activity score based on the classification result.

[0028] In a second aspect, the present invention provides an aggregation-induced luminescence molecule generation device based on reinforcement learning, the device comprising:

[0029] The token sequence generation unit is used to pre-process the SMILES expression of each chemical molecule in the molecular training set and generate a token sequence represented by SELFIES for each chemical molecule;

[0030] A candidate molecule generation unit is used to generate a candidate molecule set based on each of the Token sequences using a pre-trained molecule generation model;

[0031] An AIE activity prediction unit is used to predict the AIE activity of each candidate molecule in the candidate molecule set using a pre-trained AIE activity prediction model to obtain a corresponding AIE activity score;

[0032] a reward function construction unit, configured to construct a multi-objective reward function based on the AIE activity score;

[0033] A model parameter optimization unit, configured to iteratively optimize the parameters of the molecular generation model using a reinforcement learning strategy with adaptive weights, taking the multi-objective reward function as an optimization target;

[0034] The AIE molecule generation unit is used to generate aggregation-induced emission molecules using the optimized molecule generation model.

[0035] In some embodiments, the Token sequence generation unit includes:

[0036] SMILES represents a standardization unit, which is used to standardize the SMILES expression of each chemical molecule in the molecular training set;

[0037] The SELFIES representation segmentation unit is used to perform SELFIES representation conversion on each of the standardized SMILES expressions, and segment each SELFIES representation obtained after the conversion to generate the Token sequence.

[0038] In a third aspect, the present invention further provides a computing device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above-described method when executing the computer program.

[0039] In a fourth aspect, the present invention further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method described above are implemented.

[0040] An embodiment of the present invention preprocesses the SMILES expression of each chemical molecule in a molecular training set to generate a Token sequence represented by SELFIES for each chemical molecule. Based on each Token sequence, a pre-trained molecular generation model is used to generate a set of candidate molecules. A pre-trained AIE activity prediction model is used to predict the AIE activity of each candidate molecule in the candidate molecule set. A multi-objective reward function is constructed based on the predicted AIE activity score. With the multi-objective reward function as the optimization target, the parameters of the molecular generation model are iteratively optimized through a reinforcement learning strategy with adaptive weights. Aggregation-induced emission molecules are generated using the optimized molecular generation model, thereby constructing a multi-objective guidance mechanism through reinforcement learning, so that the generated molecules can dynamically balance and coordinately optimize among multiple objectives, thereby improving the generation efficiency and quality of AIE molecules. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 1 is a flow chart of a method for generating aggregation-induced luminescence molecules based on reinforcement learning provided in Example 1 of the present invention;

[0042] Figure 2 Schematic diagram of the structure of the device for generating aggregation-induced luminescence molecules based on reinforcement learning provided in the second embodiment of the present invention;

[0043] Figure 3 It is a structural diagram of the computing device provided in Example 3 of the present invention. DETAILED DESCRIPTION

[0044] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0045] It should be understood that when used in this specification and the appended claims, the term "comprising" indicates the presence of the described features, integers, steps, operations, elements, and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or combinations thereof. Furthermore, the terminology used in this specification is for the purpose of describing specific embodiments only and is not intended to limit the invention. As used in this specification and the appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise. The terms "first," "second," and similar terms do not denote any order, quantity, or importance, but are simply used to distinguish one component from another. Terms such as "connected" or "connected" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are used only to indicate relative positions; changes in the absolute position of the described objects may also change the relative positions of the objects. The term "plurality" refers to two or more, and other quantifiers are used similarly.

[0046] In order to keep the following description of the embodiments of the present invention clear and concise, detailed descriptions of some known functions and components are omitted in this specification.

[0047] The following describes the specific implementation of the present invention in detail with reference to specific embodiments:

[0048] Example 1:

[0049] Figure 1 The following is an implementation flow of the method for generating aggregation-induced luminescence molecules based on reinforcement learning provided in the first embodiment of the present invention. For ease of illustration, only the part related to the embodiment of the present invention is shown, which is detailed as follows:

[0050] In step S101, the SMILES expression of each chemical molecule in the molecular training set is preprocessed to generate a token sequence represented by SELFIES for each chemical molecule.

[0051] The embodiments of the present invention are applicable to computing devices, such as personal computers and servers. In the embodiments of the present invention, before preprocessing the SMILES expression of each chemical molecule in the molecular training set, the SMILES expression of the chemical molecule is first obtained from a public molecular database (such as ChEMBL, ZINC, and PubChem), and molecules with legal structures are screened according to the following conditions:

[0052] (1) The molecule is an organic compound;

[0053] (2) The molecule has no structural defects (such as broken bonds, ring errors, charge imbalance, etc.);

[0054] (3) The atoms contained in the molecule are limited to common elements (such as C, H, O, N, S, F, Cl, Br, I);

[0055] Finally, a preset number (e.g., approximately 220,000) of SMILES expressions with legal and syntactically correct chemical structures are retained (the number can be determined based on chemical space coverage and computational resource requirements) to form a molecular training set. The following steps are then used to preprocess the SMILES expression of each chemical molecule in the molecular training set:

[0056] (S101.1) Normalize the SMILES expression of each chemical molecule in the molecular training set;

[0057] In an embodiment of the present invention, the RDKit tool is used to standardize the SMILES expression of each chemical molecule in the molecular training set, including hydrogen atom processing, aromaticity unification, topological regularization, and removal of stereo information.

[0058] (S101.2) Convert each standardized SMILES expression into a SELFIES representation, and segment each SELFIES representation obtained after the conversion to generate a token sequence.

[0059] In an embodiment of the present invention, SELFIES (Self-Referencing Embedded Strings) is a more robust molecular encoding method than SMILES. First, a molecular expression conversion tool (such as the SELFIES library) is used to convert each standardized SMILES expression into a SELFIES expression, thereby ensuring the grammatical legitimacy of the generated molecules and avoiding invalid structures, thereby improving the grammatical and chemical structure legitimacy of the subsequently generated molecules. Then, based on the grammatical rules of SELFIES, each chemical unit is divided into an independent token, and a token sequence corresponding to each chemical molecule is generated. The corresponding SELFIES vocabulary is composed of the chemical units after deduplication in the token sequence. At the same time, [START], , and [PAD] are introduced into the token sequence. These special markers do not correspond to any chemical units, but are used to assist the molecular generation model in understanding the structure and boundaries of the sequence. Each token in the Token sequence corresponds to a chemical unit in SELFIES (such as the atomic symbol [C], chemical bond [=], ring structure marker [Ring1], and branch point [Branch1], etc.). [START] is used to mark the beginning of the sequence. For example, when generating molecules, the model starts from the [START] token and gradually generates subsequent chemical units; is used to mark the end of the sequence, indicating the termination of a complete molecular structure (such as [START][C][O] ), to prevent the model from generating infinitely long sequences; [PAD] is used to fill in the sequence to make the sequence length consistent. When the lengths of the Token sequences of different molecules are inconsistent (for example, some molecules contain 5 tokens and some contain 10), [PAD] is used to fill in the shorter sequence to make the length of all sequences uniform and meet the model's requirements for input dimensions.

[0060] As an example, suppose there is a SELFIES string: '[C][O][C]', the generated Token sequence is {[START],[C], [O], [C],}, and the SELFIES vocabulary includes '[C]' and '[O]'.

[0061] The generation of token sequences is achieved through the above steps (S101.1) to (S101.2), thereby utilizing the token combination mechanism to achieve innovative design of asymmetric structures, expand the structural diversity of AIE molecules, and precisely control the generation of branch and ring structures to ensure the accuracy of molecular skeleton construction, effectively avoid the generation of invalid structures, and enhance the chemical legitimacy of AIE molecules.

[0062] In step S102, a candidate molecule set is generated based on each Token sequence using a pre-trained molecule generation model.

[0063] In an embodiment of the present invention, each Token sequence is input into a pre-trained molecular generation model to generate a new molecular structure and obtain new candidate molecules, which constitute a candidate molecule set.

[0064] In a feasible embodiment, the molecular generation model adopts a Transformer architecture, including an embedding layer, a decoder layer composed of multiple layers of Transformer decoders stacked together, and an output layer. The embedding layer is used to perform vector mapping on each Token sequence to generate an embedding vector containing chemical semantics and position information. The decoder layer is used to perform multi-layer feature extraction on the embedding vector output by the embedding layer through a multi-head self-attention mechanism and feedforward transformation to generate a hidden state sequence that integrates global and local structural information. The output layer is used to convert the hidden state sequence output by the decoder layer into the probability distribution of the next Token, and construct a complete molecular sequence through a sampling strategy to achieve the generation of candidate molecules.

[0065] In a feasible embodiment, the generation of the candidate molecule set is achieved through the following steps:

[0066] (S102.1) Mapping each token sequence to a vector using an embedding layer to generate an embedding vector containing chemical semantics and position information;

[0067] In this embodiment of the present invention, the discrete tokens in the input token sequence are mapped into continuous dense vector representations through the embedding layer. Specifically, the input token sequence is mapped into a fixed-dimensional vector through word vector embedding (Token Embedding) to capture the semantic features of the molecular structure, and positional encoding is used to add position information to each token in the token sequence to ensure that the molecular generation model can learn the "relative position relationship between the i-th token and the j-th token". Finally, an embedding vector containing chemical semantic features and position features is generated. The embedding vector is represented as ,in, Represents a Token in a Token sequence, Indicates the position of the token in the sequence order.

[0068] (S102.2) performing multi-layer feature extraction on the embedding vector through a decoder layer to generate a hidden state sequence that incorporates global and local structural information;

[0069] In an embodiment of the present invention, the decoder layer is composed of N layers of stacked Transformer decoders, and each decoder layer is composed of a multi-head self-attention mechanism (Multi-head Self-Attention), a feedforward fully connected network (Feed Forward Network), a residual connection (Residual) and a layer normalization (LayerNorm) operation. The core goal of each decoder layer is to transform the input hidden state (encoding the semantic information of the prefix sequence) into a more abstract context representation through the self-attention mechanism and feedforward transformation. Here, the output of each decoder layer is the input of the next decoder layer. Specifically, each decoder layer processes the input hidden state sequence, and after multi-head self-attention, feedforward neural network, residual connection and layer normalization, it outputs the processed hidden state sequence as the input of the next layer. Finally, after processing by N layers of decoders, the final hidden state sequence for generating the next Token is output. This hidden state sequence will be passed to the output layer for generating the probability distribution of the next Token. Output of layer decoder Expressed as: ,in, It is used to split the input vector into multiple "heads", independently calculate the attention matrix, and capture the dependencies between different subspaces. It uses a feedforward fully connected network to perform nonlinear transformation on the attention output and extract high-order abstract features. It is a layer normalization operation used to stabilize the input distribution of each layer. Indicates the The hidden state sequence output by the layer decoder (or the embedding vector output by the embedding layer, when hour).

[0070] (S102.3) The hidden state sequence is converted into the probability distribution of the next token through the output layer, and the complete molecular sequence is constructed through the sampling strategy to generate a set of candidate molecules.

[0071] In this embodiment of the present invention, the following steps are used to convert the hidden state sequence into the probability distribution of the next token through the output layer, and to construct a complete molecular sequence through a sampling strategy to generate a candidate molecular set:

[0072] (S102.3.1) Map the hidden state at each time step in the hidden state sequence to a vector space of the vocabulary size using a linear transformation to obtain a corresponding vocabulary space vector;

[0073] In this embodiment of the present invention, the hidden state output by the decoder layer at each time step is mapped to a vector space matching the vocabulary size through a linear transformation. The calculation formula of the linear transformation is: ,in, is the linear mapping weight matrix, is the bias vector, The decoder layer The hidden state of time steps, is the space vector of the vocabulary after mapping.

[0074] (S102.3.2) Normalize the vocabulary space vector using the Softmax layer to obtain the probability distribution of the next token;

[0075] In the embodiment of the present invention, the Softmax layer is used to Perform normalization to generate the probability distribution of the next token. The calculation formula is: ,in, Indicates the The probability distribution of the next Token predicted in time steps.

[0076] (S102.3.3) Use a top-k sampling strategy to randomly select tokens from the probability distribution and gradually construct a molecular sequence represented by SELFIES;

[0077] In the embodiment of the present invention, a Top-k sampling strategy is adopted to obtain the probability distribution of each time step. In the process, only the top k tokens with the highest probability are retained (e.g. k = 50), and then a token is randomly sampled from them to avoid generating exactly the same molecules each time, increase structural diversity, and gradually splice the tokens sampled at all time steps to form a complete SELFIES string, forming a molecular sequence represented by SELFIES.

[0078] (S102.3.4) The molecular sequence is deconstructed into a SMILES expression, and the SMILES expressions that have passed the molecular structure validity verification are used as candidate molecules. All candidate molecules constitute a candidate molecule set.

[0079] In an embodiment of the present invention, the molecular sequence represented by SELFIES is reversed into a SMILES expression, and the legality of its molecular structure is verified using RDKit, and only chemically valid molecules are retained and entered into the candidate molecule set.

[0080] The above steps (S102.1) to (S102.3) are based on the autoregressive generation mechanism of Transformer, which realizes the end-to-end generation from token to legal molecules. Different from the traditional molecular design methods based on templates or fragment splicing, it improves the long sequence (complex molecules) modeling and molecular generation capabilities of the molecular generation model.

[0081] In one feasible embodiment, before each token sequence is input into a pre-trained molecular generation model, the molecular generation model is trained by combining a teacher forcing technique with a masked multi-head self-attention mechanism. At each time step, the actual token sequence (rather than the model prediction) is input, and the teacher forcing technique is used to guide the model to learn the conditional probability distribution between tokens (i.e., predicting the next token given the previous token). This reduces the accumulated error in the early stages of training and accelerates model convergence. Simultaneously, in the self-attention calculation of each decoder layer, a mask value (Mask) is set in the lower triangular attention weight matrix to strictly restrict the current time step to only access the token at the current position and its previous position, preventing the model from "cheating" by using future token information. This training strategy can not only improve the stability and convergence speed of model training, but also ensure that the model masters the ability to generate self-sequences token by token, ultimately meeting the requirements of autoregressive generation. That is, the model can start from the starting token and generate subsequent tokens in sequence according to the time step, gradually constructing a complete molecular sequence.

[0082] In step S103, the pre-trained AIE activity prediction model is used to predict the AIE activity of each candidate molecule in the candidate molecule set to obtain a corresponding AIE activity score.

[0083] In an embodiment of the present invention, the candidate molecule set is input into a pre-trained AIE activity prediction model to obtain the AIE activity score of each candidate molecule in the candidate molecule set. The AIE activity score is used to evaluate the luminescence efficiency of the candidate molecule in the aggregated state.

[0084] In a feasible embodiment, the AIE activity prediction model includes an input layer, a decision layer composed of a random forest classifier integrating multiple decision trees, and an output layer, wherein the input layer is used to calculate the ECFP4 (Extended-Connectivity Fingerprint) fingerprint of each candidate molecule in the candidate molecule set, the decision layer is used to enable each decision tree to independently judge the AIE activity of the corresponding candidate molecule based on the ECFP4 fingerprint, and use a majority voting mechanism to output the classification result of whether the candidate molecule has AIE activity, and the output layer is used to generate an AIE activity score based on the classification result.

[0085] In a feasible embodiment, the following steps are used to implement AIE activity prediction for each candidate molecule in the candidate molecule set using a pre-trained AIE activity prediction model:

[0086] ① The input layer converts the SMILES expression of each candidate molecule into a molecular object and calculates the corresponding ECFP4 fingerprint of the molecular object. The ECFP4 fingerprint is a 2048-dimensional sparse vector, where each bit of the sparse vector corresponds to a substructure fragment (such as a benzene ring, a donor, etc.), which is used to characterize the molecular substructure features. This representation method has good local structure representation capabilities and helps the classifier identify key AIE structural features;

[0087] ② The ECFP4 fingerprint of each candidate molecule is input into the decision layer. Each decision tree independently predicts whether the corresponding candidate molecule has AIE activity based on the ECFP4 fingerprint. After all decision trees have predicted the same candidate molecule, the prediction results of all decision trees are summarized using a majority voting mechanism. If more than a preset proportion (e.g., 50%) of the trees predict that the molecule has AIE activity, indicating that its AIE activity is strong, the candidate molecule is determined to have AIE activity. Otherwise, the candidate molecule is determined to have no AIE activity.

[0088] ③ The classification result of the decision layer (i.e., the binary judgment of "having AIE activity" or "not having AIE activity") is converted into a probability value between 0 and 1 through the output layer. This probability value is calculated by the proportion of decision trees predicted to have AIE activity in the decision layer, indicating the predicted probability that the molecule has AIE activity. This probability value is used as the AIE activity score of the candidate molecule and is used in the reinforcement learning process of the molecular generation model as a positive reward feedback signal to guide the generation of candidate molecules with higher AIE activity.

[0089] In a feasible embodiment, before using a pre-trained AIE activity prediction model to predict the AIE activity of each candidate molecule in the candidate molecule set, the AIE activity prediction model is trained. Specifically, 4108 manually annotated SMILES samples are obtained, wherein the ratio of positive samples to negative samples is 1:1, the positive samples are molecules experimentally verified to have AIE activity, and the negative samples are molecules without AIE activity; each SMILES sample is encoded as an ECFP4 fingerprint as an input feature, and the training set and the validation set are divided into a ratio of 8:2; the AIE activity prediction model is trained on the training set and verified on the validation set. The training goal is to maximize the classification accuracy. Finally, the AIE activity prediction model achieves a classification accuracy of 97.69% on the validation set, and is highly robust.

[0090] In step S104, a multi-objective reward function is constructed based on the AIE activity score.

[0091] In an embodiment of the present invention, a multi-objective reward function is constructed based on the AIE activity score to guide the molecular generation model to generate molecules that meet specific requirements through a reinforcement learning (RL) reward mechanism. The multi-objective reward function combines the AIE activity score with other molecular property scores.

[0092] In one feasible embodiment, the following steps are performed to construct a multi-objective reward function based on the AIE activity score:

[0093] (S104.1) performing a synthetic feasibility assessment, a structural novelty assessment, and an energy gap assessment on each candidate molecule in the set of candidate molecules for which AIE activity prediction has been completed, and obtaining a corresponding synthetic feasibility score, structural novelty score, and energy gap score;

[0094] In an embodiment of the present invention, each candidate molecule in the set of candidate molecules for which AIE activity prediction has been completed is subjected to a synthetic feasibility evaluation, a structural novelty evaluation, and an energy gap evaluation, respectively, to obtain a corresponding synthetic feasibility score, a structural novelty score, and an energy gap score, wherein the synthetic feasibility score is used to evaluate whether the molecule can be synthesized by existing chemical methods, the structural novelty score is used to measure whether the molecule is a new structure, and the energy gap score is used to guide the evolution of the molecule to the target energy gap range.

[0095] In a feasible embodiment, when evaluating the synthetic feasibility of each candidate molecule, specifically, first, the candidate molecule is decomposed into common structural fragments, then, based on the statistical frequency of known molecules in public molecular databases (such as PubChem and ChEMBL), a synthesizability score is assigned to each structural fragment obtained by decomposition, and finally, the synthesizability score is normalized to the interval [0,1] to obtain a synthetic feasibility score, wherein a higher synthetic feasibility score indicates that it is easier to synthesize.

[0096] In a feasible embodiment, when evaluating the structural novelty of each candidate molecule, specifically, the maximum Tanimoto similarity between each candidate molecule and the molecular training set is calculated according to the formula Calculate the structural novelty score, where Represents the structural novelty score. The higher the score, the greater the difference between the molecular structure of the corresponding candidate molecule and the molecular structure in the molecular training set. Indicates the maximum Tanimoto similarity.

[0097] In a feasible embodiment, when evaluating the energy gap of each candidate molecule, the energy gap values ​​of the monomer and aggregated state of the candidate molecule are first calculated based on the aggregated conformation of the candidate molecule, and then the energy gap score of the corresponding candidate molecule is determined based on the calculated energy gap value. Specifically, the aggregated conformation of the input candidate molecule is searched through the aggregated conformation search (CREST), and the energy gap values ​​(specifically, HOMO-LUMO gap) of the monomer and aggregated state of the candidate molecule are calculated using xTB based on the aggregated conformation obtained from the search. The energy gap value is normalized to obtain an energy gap score that is closer to the actual luminescence behavior. The energy gap score is used to guide the molecule to evolve towards the target energy gap range (e.g., less than 1.0 eV) during the reinforcement learning process. Among them, the aggregated state energy gap can reflect the emission wavelength characteristics of the molecule, which is helpful for the design of near-infrared II (NIR-II) molecules to make it more consistent with the real physical process of AIE molecules.

[0098] (S104.2) Construct a multi-objective reward function based on AIE activity score, synthetic feasibility score, structural novelty score and energy gap score. Expressed as ,in, represents the AIE activity score, represents the synthetic feasibility score, represents the structural novelty score, represents the energy gap score, 、 、 、 Respectively 、 、 and The weight of .

[0099] In the embodiment of the present invention, a combined multi-objective reward function is constructed based on the AIE activity score, the synthesis feasibility score, the structural novelty score and the energy gap score. , to guide reinforcement learning to optimize the molecular generation model, where the multi-objective reward function supports the weight adaptive dynamic adjustment mechanism, represents the AIE activity score, represents the synthetic feasibility score, represents the structural novelty score, represents the energy gap score, 、 、 、 Respectively 、 、 and The weight of .

[0100] In a specific embodiment, 、 、 and The default weight is set to 、 、 、 .

[0101] Through the above steps (S104.1) to (S104.2), a multi-index evaluation system including AIE activity, synthetic feasibility, novelty and HOMO-LUMO gap is constructed, thereby realizing multi-angle screening of candidate molecules and avoiding performance bias caused by the optimality of a single indicator.

[0102] In step S105 , the parameters of the molecular generation model are iteratively optimized through a reinforcement learning strategy with adaptive weights, taking the multi-objective reward function as the optimization target.

[0103] In an embodiment of the present invention, a multi-objective reward function with dynamically adjusted weights is used as the optimization target, and the reward contribution of various objectives such as AIE activity and synthetic feasibility is automatically balanced through an adaptive weight mechanism to prevent a certain indicator from dominating the training direction; at the same time, the REINFORCE policy gradient algorithm is adopted to iteratively update the parameters of the molecular generation model with the multi-objective total reward as the feedback signal to achieve targeted optimization of the model. Through the above mechanism, the molecular generation model can gradually learn to generate molecules with better structures that are simultaneously optimized in multiple dimensions, ultimately improving the quality of molecules generated by the model.

[0104] In step S106 , aggregation-induced emission molecules are generated using the optimized molecule generation model.

[0105] In an embodiment of the present invention, the molecular generation model optimized by final reinforcement learning is used to batch generate AIE molecules. The molecular structures not only have high AIE activity scores but also take into account synthetic feasibility and novelty, and can directly enter the TD-DFT (Time-Dependent Density Functional Theory) or synthesis stage.

[0106] In an embodiment of the present invention, the SMILES expression of each chemical molecule in the molecular training set is preprocessed to generate a Token sequence represented by SELFIES for each chemical molecule. According to each Token sequence, a pre-trained molecular generation model is used to generate a set of candidate molecules. The AIE activity of each candidate molecule in the candidate molecule set is predicted using a pre-trained AIE activity prediction model. A multi-objective reward function is constructed based on the predicted AIE activity score. With the multi-objective reward function as the optimization target, the parameters of the molecular generation model are iteratively optimized through a reinforcement learning strategy with adaptive weights. Aggregation-induced emission molecules are generated using the optimized molecular generation model, thereby constructing a multi-objective guidance mechanism through reinforcement learning, so that the generated molecules can dynamically balance and coordinately optimize among multiple objectives, thereby improving the generation efficiency and quality of AIE molecules.

[0107] Example 2:

[0108] Figure 2 The structure of the aggregation-induced luminescence molecule generation device based on reinforcement learning provided by the second embodiment of the present invention is shown. For ease of explanation, only the parts related to the embodiment of the present invention are shown, including:

[0109] A token sequence generation unit 21 is used to pre-process the SMILES expression of each chemical molecule in the molecular training set and generate a token sequence represented by SELFIES for each chemical molecule;

[0110] The candidate molecule generation unit 22 is used to generate a candidate molecule set based on each Token sequence using a pre-trained molecule generation model;

[0111] An AIE activity prediction unit 23 is used to predict the AIE activity of each candidate molecule in the candidate molecule set using a pre-trained AIE activity prediction model to obtain a corresponding AIE activity score;

[0112] a reward function construction unit 24 for constructing a multi-objective reward function based on the AIE activity score;

[0113] A model parameter optimization unit 25 is used to iteratively optimize the parameters of the molecular generation model through a reinforcement learning strategy with adaptive weights, taking the multi-objective reward function as the optimization target;

[0114] The AIE molecule generation unit 26 is used to generate aggregation-induced emission molecules using the optimized molecule generation model.

[0115] Preferably, the Token sequence generation unit 21 includes:

[0116] SMILES stands for Standardization Unit, which is used to standardize the SMILES expression of each chemical molecule in the molecular training set;

[0117] SELFIES represents a segmentation unit, used for converting each SMILES expression after normalization into a SELFIES representation, and segmenting each SELFIES representation after conversion to generate a Token sequence.

[0118] In the embodiments of the present application, for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is exemplified, and in actual application, the above-mentioned functions can be implemented by different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to realize all or part of the functions described above. Each unit and module of the device can be realized by a corresponding hardware or software unit, and each unit and module can be an independent software and hardware unit, or can be integrated into a software and hardware unit, which is not used to limit the present application. In addition, the specific names of each functional unit and module are only for easy distinction, and do not limit the protection scope of the present application. The specific working process of the units and modules in the device can be referred to the corresponding description in the foregoing method embodiments, which will not be repeated here.

[0119] Embodiment three:

[0120] Figure 3 The structure of the computing device provided in the third embodiment of the present application is shown, and only the parts related to the embodiments of the present application are shown for convenience of illustration.

[0121] The computing device 3 of the embodiment of the present application includes a processor 30, a memory 31, and a computer program 32 stored in the memory 31 and executable on the processor 30. The processor 30 implements the steps in the above-mentioned embodiment of the method for generating a collective induction light molecule based on reinforcement learning when executing the computer program 32, for example Figure 1 The steps S101 to S106 shown. Alternatively, the processor 30 implements the functions of each unit in the above-mentioned device embodiments when executing the computer program 32, for example Figure 2 The functions of the units shown.

[0122] In an embodiment of the present invention, the SMILES expression of each chemical molecule in the molecular training set is preprocessed to generate a Token sequence represented by SELFIES for each chemical molecule. According to each Token sequence, a pre-trained molecular generation model is used to generate a set of candidate molecules. The AIE activity of each candidate molecule in the candidate molecule set is predicted using a pre-trained AIE activity prediction model. A multi-objective reward function is constructed based on the predicted AIE activity score. With the multi-objective reward function as the optimization target, the parameters of the molecular generation model are iteratively optimized through a reinforcement learning strategy with adaptive weights. Aggregation-induced emission molecules are generated using the optimized molecular generation model, thereby constructing a multi-objective guidance mechanism through reinforcement learning, so that the generated molecules can dynamically balance and coordinately optimize among multiple objectives, thereby improving the generation efficiency and quality of AIE molecules.

[0123] The computing device of the embodiment of the present invention may be a personal computer. The steps implemented when the processor 30 in the computing device 3 executes the computer program 32 to implement the method for generating aggregation-induced emission molecules based on reinforcement learning can be referred to the description of the aforementioned method embodiment, which will not be repeated here.

[0124] Example 4:

[0125] In an embodiment of the present invention, a computer-readable storage medium is provided, which stores a computer program. When the computer program is executed by a processor, the steps in the embodiment of the above-mentioned method for generating aggregation-induced luminescence molecules based on reinforcement learning are implemented, for example, Figure 1 Alternatively, when the computer program is executed by a processor, the functions of each unit in the above-mentioned device embodiments are realized, for example Figure 2 Function of the unit shown.

[0126] In an embodiment of the present invention, the SMILES expression of each chemical molecule in the molecular training set is preprocessed to generate a Token sequence represented by SELFIES for each chemical molecule. According to each Token sequence, a pre-trained molecular generation model is used to generate a set of candidate molecules. The AIE activity of each candidate molecule in the candidate molecule set is predicted using a pre-trained AIE activity prediction model. A multi-objective reward function is constructed based on the predicted AIE activity score. With the multi-objective reward function as the optimization target, the parameters of the molecular generation model are iteratively optimized through a reinforcement learning strategy with adaptive weights. Aggregation-induced emission molecules are generated using the optimized molecular generation model, thereby constructing a multi-objective guidance mechanism through reinforcement learning, so that the generated molecules can dynamically balance and coordinately optimize among multiple objectives, thereby improving the generation efficiency and quality of AIE molecules.

[0127] The computer-readable storage medium of the embodiments of the present invention may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EEPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the embodiments of the present invention, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0128] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that the scope of disclosure involved in the above embodiments is not limited to the technical solutions formed by the specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosed concepts. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

[0129] In addition, although adopting specific order to describe each operation, this should not be interpreted as requiring these operations to be executed in the specific order shown or in sequential order.Under certain environment, multitasking and parallel processing may be advantageous.Similarly, although comprising some specific implementation details in the above discussion, these should not be interpreted as limiting the scope of the present invention.Some features described in the context of independent embodiment can also be implemented in single embodiment in combination.On the contrary, the various features described in the context of independent embodiment also can be implemented in multiple embodiments individually or in the mode of any suitable subcombination.

Claims

1. A method for generating aggregation-induced luminescence molecules based on reinforcement learning, characterized in that: The method comprises the following steps: Preprocess the SMILES expression of each chemical molecule in the molecular training set to generate a token sequence represented by SELFIES for each chemical molecule; According to each of the token sequences, a set of candidate molecules is generated using a pre-trained molecule generation model; Using a pre-trained AIE activity prediction model to predict the AIE activity of each candidate molecule in the candidate molecule set to obtain a corresponding AIE activity score; constructing a multi-objective reward function based on the AIE activity score; Taking the multi-objective reward function as the optimization target, iteratively optimizing the parameters of the molecular generation model through a reinforcement learning strategy with adaptive weights; Generate aggregation-induced emission molecules using the optimized molecular generation model; The step of constructing a multi-objective reward function based on the AIE activity score includes: Performing a synthetic feasibility assessment, a structural novelty assessment, and an energy gap assessment on each candidate molecule in the set of candidate molecules for which AIE activity prediction has been completed, and obtaining a corresponding synthetic feasibility score, structural novelty score, and energy gap score; The multi-objective reward function is constructed based on the AIE activity score, the synthetic feasibility score, the structural novelty score and the energy gap score. Expressed as ,in, represents the AIE activity score, represents the synthesis feasibility score, represents the structural novelty score, represents the energy gap score, 、 、 、 Respectively 、 、 and The weight of .

2. The method according to claim 1, wherein The steps for preprocessing the SMILES expression of each chemical molecule in the molecular training set include: Standardizing the SMILES expression of each chemical molecule in the molecular training set; The standardized SMILES expressions are converted into SELFIES representations, and the SELFIES representations obtained after the conversion are segmented to generate the Token sequence.

3. The method according to claim 1, wherein The molecular generation model adopts a Transformer architecture, including an embedding layer, a decoder layer composed of multiple layers of Transformer decoders, and an output layer. The steps of generating a set of candidate molecules based on each token sequence using a pre-trained molecular generation model include: Perform vector mapping on each of the token sequences through the embedding layer to generate an embedding vector containing chemical semantics and position information; Performing multi-layer feature extraction on the embedding vector through the decoder layer to generate a hidden state sequence that integrates global and local structural information; The hidden state sequence is converted into the probability distribution of the next Token through the output layer, and a complete molecular sequence is constructed through a sampling strategy to generate the candidate molecular set.

4. The method according to claim 3, wherein The step of converting the hidden state sequence into the probability distribution of the next token through the output layer, and constructing a complete molecular sequence through a sampling strategy to generate the candidate molecular set includes: Mapping the hidden state of each time step in the hidden state sequence to a vector space of the vocabulary size through linear transformation to obtain a corresponding vocabulary space vector; Normalizing the word space vector through the Softmax layer to obtain the probability distribution of the next Token; A Top-k sampling strategy is used to randomly select tokens from the probability distribution, and the molecular sequence represented by SELFIES is gradually constructed; The molecular sequence is inversely decoded into a SMILES expression, and the SMILES expression that has passed the molecular structure legality verification is used as a candidate molecule, and all candidate molecules constitute the candidate molecule set.

5. The method according to claim 1, wherein The AIE activity prediction model includes an input layer, a decision layer consisting of a random forest classifier integrating multiple decision trees, and an output layer, wherein the input layer is used to calculate the ECFP4 fingerprint of each candidate molecule in the candidate molecule set, the decision layer is used to enable each decision tree to independently judge the AIE activity of the corresponding candidate molecule based on the ECFP4 fingerprint, and use a majority voting mechanism to output a classification result of whether the candidate molecule has AIE activity, and the output layer is used to generate the AIE activity score based on the classification result.

6. A device for generating aggregation-induced luminescence molecules based on reinforcement learning, characterized in that: The device comprises: The token sequence generation unit is used to pre-process the SMILES expression of each chemical molecule in the molecular training set and generate a token sequence represented by SELFIES for each chemical molecule; A candidate molecule generation unit is used to generate a candidate molecule set based on each of the Token sequences using a pre-trained molecule generation model; An AIE activity prediction unit is used to predict the AIE activity of each candidate molecule in the candidate molecule set using a pre-trained AIE activity prediction model to obtain a corresponding AIE activity score; A reward function construction unit is used to construct a multi-objective reward function based on the AIE activity score, including performing a synthesis feasibility assessment, a structural novelty assessment, and an energy gap assessment on each candidate molecule in the candidate molecule set for which AIE activity prediction has been completed, to obtain a corresponding synthesis feasibility score, a structural novelty score, and an energy gap score; constructing the multi-objective reward function based on the AIE activity score, the synthesis feasibility score, the structural novelty score, and the energy gap score, wherein the multi-objective reward function Expressed as ,in, represents the AIE activity score, represents the synthesis feasibility score, represents the structural novelty score, represents the energy gap score, 、 、 、 Respectively 、 、 and The weight of ; A model parameter optimization unit, configured to iteratively optimize the parameters of the molecular generation model using a reinforcement learning strategy with adaptive weights, taking the multi-objective reward function as an optimization target; The AIE molecule generation unit is used to generate aggregation-induced emission molecules using the optimized molecule generation model.

7. The device according to claim 6, characterized in that The Token sequence generation unit includes: SMILES represents a standardization unit, which is used to standardize the SMILES expression of each chemical molecule in the molecular training set; The SELFIES representation segmentation unit is used to perform SELFIES representation conversion on each of the standardized SMILES expressions, and segment each SELFIES representation obtained after the conversion to generate the Token sequence.

8. A computing device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 5 are implemented.

9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Multi-target attribute molecule generation method and system based on strategy learning

    CN114974461A

  • Drug molecule optimization design method based on deep reinforcement learning and skeleton constraint

    CN119400294A