A Multimodal Chemical Reaction Yield Prediction Method Based on Adaptive Data Screening

CN120783883BActive Publication Date: 2026-09-01ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510833377.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2026-09-01
Estimated Expiration
2045-06-20

AI Technical Summary

Technical Problem

[0004]本发明旨在提供一种基于自适应数据筛选的多模态化学反应产率预测方法,以解决现有技术中化学反应产率预测模型在处理高噪声数据和多模态信息融合方面的不足

Benefits of technology

[0029]本发明提出了一种基于自适应数据筛选的多模态化学反应产率预测方法,能够有效解决现有模型在面对高噪声样本和多模态数据融合时精度不足、鲁棒性差的问题。通过引入弱监督学习机制构建的自适应数据筛选算法,模型可在训练过程中动态调整样本筛选策略,优先使用高质量样本进行拟合,有效抵御低质量数据的干扰,提升训练稳定性。此外,本发明通过构建微观多模态编码器,实现了对一维SMILES序列、二维分子图和三维分子构象等多种化学反应信息的深度融合,能够从不同层次和角度提取有效特征,增强模型的表达能力和对复杂结构关系的建模能力。结合跨模态对比学习策略,本发明进一步提升了多模态特征在潜在空间中的对齐程度,加速了模型的收敛效率。同时,本发明集成的不确定性量化模块可对预测结果进行可信度估计,为实际应用提供决策支持,增强模型的解释性和可控性。经实验验证,本发明方法在多个标准反应数据集(包括Buchwald HTE、Suzuki HTE和USPTO)上均取得优于现有模型的预测性能,表现出更高的准确性、鲁棒性与泛化能力,适用于药物合成、高通量实验筛选等多种化学反应预测场景。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120783883B_ABST
    Figure CN120783883B_ABST
Patent Text Reader

Abstract

This invention discloses a multimodal chemical reaction yield prediction method based on adaptive data screening, belonging to the interdisciplinary field of chemical synthesis and machine learning. The method constructs a multimodal input system including one-dimensional chemical attribute data, two-dimensional molecular structure maps, and three-dimensional spatial configurations to achieve multi-dimensional analysis of molecular interactions. Simultaneously, it employs a multi-stage training process, simulating human cognitive patterns, allowing the model to learn progressively from basic reaction characteristics to complex reaction mechanisms. This invention effectively suppresses noise interference and improves adaptability to long-tailed data by dynamically adjusting the screening threshold. Compared to traditional single-modal prediction models, this invention improves prediction accuracy on noisy datasets, significantly reduces the trial-and-error costs of chemical synthesis experiments, and provides reliable technical support for efficient screening of reaction conditions and accelerated development of new compounds.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of interdisciplinary technology involving chemistry and artificial intelligence, and in particular relates to a method for predicting the yield of multimodal chemical reactions based on adaptive data screening. Background Technology

[0002] Chemical reaction yield prediction is a key technology in the chemical industry, crucial for optimizing production processes, reducing production costs, and improving product quality. Traditional methods typically rely on experimental data and empirical models, which are not only inefficient but also often fail to accurately predict yields when faced with complex chemical reactions. With the development of cheminformatics and machine learning technologies, data-driven yield prediction models have gained increasing attention. However, existing yield prediction models still face numerous challenges in handling noisy data and fusing multimodal information.

[0003] First, chemical reaction data often contains a large amount of noise and outliers, which can severely interfere with model training, causing the model to struggle to converge or get trapped in local optima. Second, chemical reactions involve multimodal data, such as one-dimensional smils sequences, two-dimensional molecular diagrams, and three-dimensional molecular conformations. These modal information have complex relationships, and effectively fusing this information to improve the model's predictive ability is a pressing issue. Furthermore, existing models often lack alignment and fusion mechanisms for different modal features when processing multimodal data, preventing them from fully utilizing the advantages of multimodal information. Therefore, there is an urgent need for a chemical reaction yield prediction model that can effectively resist data noise interference and deeply fuse multimodal information to improve the accuracy and reliability of predictions. Summary of the Invention

[0004] This invention aims to provide a multimodal chemical reaction yield prediction method based on adaptive data filtering, addressing the shortcomings of existing chemical reaction yield prediction models in handling high-noise data and fusing multimodal information. By introducing an adaptive data filtering mechanism and a cross-modal comparative learning strategy, this invention can effectively resist the interference of data noise and deeply fuse multimodal information, thereby improving the accuracy and reliability of chemical reaction yield prediction.

[0005] To achieve the above-mentioned objectives, the present invention specifically adopts the following technical solution:

[0006] In a first aspect, the present invention provides a method for predicting the yield of multimodal chemical reactions based on adaptive data screening, comprising the following steps:

[0007] S1. Obtain one-dimensional SMILES sequences and their corresponding two-dimensional molecular maps, and form an unlabeled multimodal dataset;

[0008] S2. In the pre-training phase, the network parameters of the multilayer perceptron predictor are fixed, and the micro-multimodal encoder is trained on an unlabeled multimodal dataset using a cross-modal contrastive learning training strategy.

[0009] S3. After pre-training, the unlabeled multimodal dataset is labeled to obtain the labeled yield dataset. The network parameters of the multilayer perceptron predictor are unfrozen. The micro-multimodal chemical reaction yield prediction model is constructed by the micro-multimodal encoder and the multilayer perceptron predictor. The micro-multimodal chemical reaction yield prediction model is jointly trained end-to-end based on the labeled yield dataset. The weights of the micro-multimodal chemical reaction yield prediction model are updated synchronously through backpropagation to obtain the trained micro-multimodal chemical reaction yield prediction model.

[0010] During the joint training process, an adaptive data filtering algorithm dynamically removes abnormal training data based on the prediction confidence and feature space distribution, iteratively updates the labeled yield dataset, and calculates the variance estimate of the predicted chemical reaction yield through the uncertainty quantification module, and generates the confidence interval of the predicted chemical reaction yield based on Monte Carlo sampling.

[0011] S4. Input the one-dimensional SMILES sequence to be predicted and its corresponding two-dimensional molecular map into the trained microscopic multimodal chemical reaction yield prediction model. The microscopic multimodal encoder extracts high-dimensional fusion features, and the multilayer perceptron predictor maps the extracted high-dimensional fusion features into chemical reaction yield prediction results. The variance estimate of the chemical reaction yield prediction results is calculated through the uncertainty quantification module. The confidence interval of the chemical reaction yield prediction results is generated based on Monte Carlo sampling, realizing multimodal chemical reaction yield prediction based on adaptive data screening.

[0012] Based on the above scheme, each step can be implemented in the following preferred manner.

[0013] As a preferred embodiment of the first aspect, in the micro-multimodal encoder of step S2, the SMILES sequence encoder encodes the one-dimensional SMILES sequence, maps the one-dimensional SMILES sequence to a vector space to obtain the SMILES feature vector, the molecular graph encoder encodes the two-dimensional molecular graph to generate the global graph representation of the two-dimensional molecular graph, the three-dimensional molecular conformation features are generated based on the two-dimensional molecular graph, and the SMILES feature vector, the global graph representation of the two-dimensional molecular graph and the three-dimensional molecular conformation features are then combined through feature splicing operation to generate a comprehensive feature vector.

[0014] As a preferred embodiment of the first aspect above, the specific method for generating the three-dimensional molecular conformation features is as follows: First, a three-dimensional molecular initial conformation is obtained from a two-dimensional molecular diagram using a conformation prediction-related technique. Then, the energy of the three-dimensional molecular initial conformation is minimized using the MMFF molecular mechanical force field to obtain the corresponding stable conformation and form a set of three-dimensional molecular stable conformations that participate in chemical reactions. Next, the cheminformatics and machine learning toolkit RDKit is used to calculate the stable conformation set to obtain three-dimensional molecular conformation features containing geometric diameter, moment of inertia, and mass distribution. Alternatively, the three-dimensional molecular conformation encoder can encode the stable conformation set to generate three-dimensional molecular conformation features.

[0015] As a preferred embodiment of the first aspect mentioned above, in the cross-modal contrastive learning training strategy, the SMILES feature vectors, global graph representations of two-dimensional molecular graphs, and three-dimensional molecular conformation features belonging to the same chemical reaction are used as positive samples, while the SMILES feature vectors, global graph representations of two-dimensional molecular graphs, and three-dimensional molecular conformation features of different chemical reactions are used as negative samples. The contrast loss between the two different modalities is calculated separately, and the calculated contrast loss is weighted and summed to obtain the total contrast loss. The network parameters of the micro-multimodal encoder are updated based on minimizing the total contrast loss.

[0016] As a preferred embodiment of the first aspect mentioned above, the adaptive data filtering algorithm operates through the following steps:

[0017] S31. First, the micro-multimodal chemical reaction yield prediction model is initialized and trained for several rounds using a labeled yield dataset to enable the micro-multimodal chemical reaction yield prediction model to have preliminary prediction capabilities. During the initial training process, the weights of all training data are set to 1, and a yield data filtering vector is constructed to record the weights of each training data.

[0018] S32. After the initial training is completed, the reliable data screening stage begins. The micro-multimodal chemical reaction yield prediction model evaluates the loss function value of each training data, selects training data with a loss function value less than the preset screening threshold and uses them as reliable data. The labeled yield dataset is updated by the reliable data, and the weight of the reliable data is set to 1, while the weight of the remaining training data is set to 0.

[0019] S33. After the reliable data has been screened, the microscopic multimodal chemical reaction yield prediction model is retrained on the updated labeled yield dataset to increase the screening threshold. New reliable data is generated and the labeled yield dataset is updated according to step S32.

[0020] S34. Iterate the training continuously according to step S33 until the microscopic multimodal chemical reaction yield prediction model converges, and obtain the trained microscopic multimodal chemical reaction yield prediction model.

[0021] As a preferred embodiment of the first aspect above, the SMILES sequence encoder employs the BERT model.

[0022] As a preferred embodiment of the first aspect above, the molecular graph encoder employs a message-passing neural network.

[0023] In a second aspect, the present invention provides a computer program product, including a computer program / instruction, which, when executed by a processor, enables the multimodal chemical reaction yield prediction method based on adaptive data screening as described in any of the solutions of the first aspect above.

[0024] Thirdly, the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the multimodal chemical reaction yield prediction method based on adaptive data screening as described in any of the solutions of the first aspect above.

[0025] Fourthly, the present invention provides a computer electronic device, which includes a memory and a processor;

[0026] The memory is used to store computer programs;

[0027] The processor is configured to, when executing the computer program, implement the multimodal chemical reaction yield prediction method based on adaptive data screening as described in any of the embodiments of the first aspect above.

[0028] Compared with the prior art, the present invention has the following advantages:

[0029] This invention proposes a multimodal chemical reaction yield prediction method based on adaptive data selection, which effectively addresses the problems of insufficient accuracy and poor robustness of existing models when facing high-noise samples and multimodal data fusion. By introducing an adaptive data selection algorithm built with a weakly supervised learning mechanism, the model can dynamically adjust its sample selection strategy during training, prioritizing the use of high-quality samples for fitting, effectively resisting interference from low-quality data and improving training stability. Furthermore, this invention constructs a microscopic multimodal encoder to achieve deep fusion of various chemical reaction information, such as one-dimensional SMILES sequences, two-dimensional molecular diagrams, and three-dimensional molecular conformations. This enables the extraction of effective features from different levels and perspectives, enhancing the model's expressive power and its ability to model complex structural relationships. Combined with a cross-modal contrastive learning strategy, this invention further improves the alignment of multimodal features in the latent space, accelerating the model's convergence efficiency. Simultaneously, the uncertainty quantification module integrated in this invention can estimate the credibility of the prediction results, providing decision support for practical applications and enhancing the model's interpretability and controllability. Experimental results show that the method of this invention outperforms existing models in prediction performance on multiple standard reaction datasets (including Buchwald HTE, Suzuki HTE and USPTO), demonstrating higher accuracy, robustness and generalization ability, and is applicable to various chemical reaction prediction scenarios such as drug synthesis and high-throughput experimental screening. Attached Figure Description

[0030] Figure 1 This is a flowchart of the steps of the present invention;

[0031] Figure 2 This is the overall architecture diagram of the microscopic multimodal chemical reaction yield prediction model provided in this embodiment;

[0032] Figure 3 This is a schematic diagram of the adaptive data filtering algorithm provided in this embodiment;

[0033] Figure 4 This is a schematic diagram of the pre-training process of the microscopic multimodal encoder in this embodiment. Detailed Implementation

[0034] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Many specific details are set forth in the following description to provide a thorough understanding of the present invention. However, the present invention can be practiced in many other ways different from those described herein, and those skilled in the art can make similar modifications without departing from the spirit of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below. Technical features in the various embodiments of the present invention can be combined accordingly without mutual conflict.

[0035] Before providing a further detailed description of the embodiments of the present invention, the nouns and terms used in the embodiments of the present invention are explained, and the nouns and terms used in the embodiments of the present invention are subject to the following interpretations:

[0036] 1) Chemical reaction yield

[0037] Chemical reaction yield refers to the efficiency with which reactants are converted into products in a chemical reaction. It is usually represented by a yield value y, which ranges from 0 to 1. When y = 0, it means that the reaction cannot occur; when y = 1, it means that the conversion rate reaches 100%.

[0038] 2) One-dimensional SMILES sequences

[0039] SMILES (Simplified Molecular Input Line Entry System) is a string format for representing chemical structures. Using SMILES, chemical structures can be represented in concise text form and can be parsed and processed by computer programs. In this invention, SMILES sequences are used to describe reactants and products in chemical reactions.

[0040] 3) Two-dimensional molecular diagram

[0041] Two-dimensional molecular diagrams are a graphical representation of molecules that uses the relationships between vertices and edges to show the connections and local structures of molecules. In this graph structure, each node represents an atom, and the edges between nodes represent the chemical bonds between atoms. Two-dimensional molecular diagrams can provide information about the overall structure and local environment of a molecule.

[0042] 4) Three-dimensional molecular conformation

[0043] Three-dimensional molecular conformation refers to the positional coordinates of atoms in a molecule in three-dimensional space and their three-dimensional relationships. This three-dimensional structural information provides more detailed spatial arrangement and structural details than two-dimensional molecular diagrams, and is crucial for understanding the physical and chemical properties of molecules and simulating intermolecular interactions.

[0044] 5) Adaptive data filtering

[0045] Adaptive data filtering is a mechanism that dynamically adjusts the filtering threshold to prioritize high-quality labeled samples and resist interference from data noise. In this invention, the adaptive data filter simulates human learning strategies, namely, processing information gradually from shallow to deep. In the early stages of training, the model focuses on samples with strong regularity and easy fitting, and as the model's capabilities improve, it gradually transitions to processing more complex samples.

[0046] 6) Multimodal feature extraction

[0047] Multimodal feature extraction refers to the process of extracting feature information from data of multiple modalities. In this invention, the multimodal feature extractor can integrate one-dimensional SMILES sequences, two-dimensional molecular diagrams, and three-dimensional molecular conformation information, enabling the model to gain a deeper understanding of chemical information from multiple perspectives and levels.

[0048] 7) Cross-modal contrastive learning

[0049] Cross-modal contrastive learning is a method that promotes model learning by comparing the similarities and differences between feature vectors from different modalities. In this invention, a cross-modal contrastive learning strategy is used to align and fuse feature vectors from different modalities, thereby accelerating the model's learning process.

[0050] 8) Quantification of uncertainty

[0051] Uncertainty quantification refers to the process of evaluating the reliability of model predictions. In this invention, the uncertainty quantification module is used to evaluate the reliability of model predictions, providing an uncertainty estimate for the model's prediction results and helping users better understand the model's prediction quality and reliability.

[0052] 9) Weakly supervised learning

[0053] Weakly supervised learning is a machine learning method in which the model's training data contains only partial label information or the label information is not completely accurate. In this invention, weakly supervised learning is used to develop an adaptive filtering algorithm for yield data, aiming to prioritize training the model to fit high-quality labeled samples, thereby effectively resisting the interference of data noise.

[0054] 10) Microscopic multimodal features

[0055] Microscopic multimodal features refer to various modal feature information extracted from the microscopic level, including one-dimensional sequence information, two-dimensional topological graph information, and three-dimensional conformational information. In this invention, microscopic multimodal features are integrated through a multimodal feature extractor to facilitate a deeper understanding of chemical information from multiple perspectives and levels.

[0056] Based on the above explanation of the nouns and terms used in the embodiments of this application, the implementation scenario of the present invention will be described first. Addressing the technical shortcomings of existing yield prediction models, which are limited by data label noise and uneven distribution, this invention innovatively proposes a solution combining adaptive progressive learning and multimodal feature fusion, namely, a multimodal chemical reaction yield prediction method based on adaptive data screening. For example... Figure 1 As shown, in a preferred embodiment of the present invention, the above-mentioned multimodal chemical reaction yield prediction method based on adaptive data screening includes the following steps S1 to S4. The specific implementation process of each step will be described in detail below.

[0057] S1. Obtain one-dimensional SMILES sequences and their corresponding two-dimensional molecular maps to form an unlabeled multimodal dataset.

[0058] S2. In the pre-training phase, the network parameters of the multilayer perceptron predictor are fixed, and the micro-multimodal encoder is trained on an unlabeled multimodal dataset using a cross-modal contrastive learning training strategy.

[0059] It should be noted that in the micro-multimodal encoder of step S2 of the present invention, the SMILES sequence encoder encodes the one-dimensional SMILES sequence, maps the one-dimensional SMILES sequence to the vector space to obtain the SMILES feature vector, the molecular graph encoder encodes the two-dimensional molecular graph to generate the global graph representation of the two-dimensional molecular graph, the three-dimensional molecular conformation features are generated based on the two-dimensional molecular graph, and then the SMILES feature vector, the global graph representation of the two-dimensional molecular graph and the three-dimensional molecular conformation features are generated by feature splicing operation to generate a comprehensive feature vector.

[0060] In this embodiment, the SMILES sequence encoder adopts the BERT model, which utilizes its bidirectional attention mechanism to simultaneously consider the contextual information on the left and right sides of each element in the SMILES sequence, thereby effectively capturing the complex relationships of molecular structures.

[0061] In this embodiment, the molecular graph encoder employs a message-passing neural network (MPNN), which aggregates local graph structure information by iteratively passing messages between nodes to ultimately generate a global graph representation. MPNN can capture the complex relationships between atoms in a molecule, including both long-range and short-range interactions.

[0062] It should be noted that in this invention, the three-dimensional molecular conformation features can be encoded by an encoder or extracted directly through computation. In this embodiment, the three-dimensional molecular conformation features are extracted by calculating a series of three-dimensional molecular fingerprints. Specifically, the generation method of the above-mentioned three-dimensional molecular conformation features is as follows: First, the initial three-dimensional molecular conformation is obtained from the two-dimensional molecular diagram using conformation prediction-related techniques, and the energy of the initial three-dimensional molecular conformation is minimized using the MMFF molecular mechanical force field to obtain the corresponding stable conformation and form a set of three-dimensional molecular stable conformations participating in chemical reactions; then, the cheminformatics and machine learning toolkit RDKit is used to calculate the stable conformation set to obtain three-dimensional molecular conformation features including geometric diameter, moment of inertia, and mass distribution, or the three-dimensional molecular conformation encoder encodes the stable conformation set to generate three-dimensional molecular conformation features to enrich the feature representation of molecules.

[0063] In this embodiment, as Figure 4As shown, this invention constructs a microscopic multimodal encoder capable of integrating one-dimensional SMILES sequences, two-dimensional molecular maps, and three-dimensional molecular conformations. Specifically, the SMILES sequence is generated by a SMILES sequence encoder (…). Figure 4 The two-dimensional molecular graph is encoded by the Transformer encoder in the model, and the molecular graph is encoded by the molecular graph encoder. Figure 4 The features are encoded using an MPNN encoder, while the 3D molecular conformation features are obtained by encoding the stable 3D molecular conformation using a SchNet encoder. Finally, these features are concatenated to generate a comprehensive feature vector. Figure 4 Multimodal fusion features are used as input to the multilayer perceptron predictor to help the model understand chemical information more deeply from multiple angles and levels.

[0064] It should be noted that in the cross-modal contrastive learning training strategy, the SMILES feature vectors, global graph representations of two-dimensional molecular graphs, and three-dimensional molecular conformation features belonging to the same chemical reaction are used as positive samples, while the SMILES feature vectors, global graph representations of two-dimensional molecular graphs, and three-dimensional molecular conformation features of different chemical reactions are used as negative samples. The contrastive loss between the two different modalities is calculated separately, and the calculated contrastive loss is weighted and summed to obtain the total contrastive loss. The network parameters of the micro-multimodal encoder are updated based on minimizing the total contrastive loss.

[0065] In this embodiment, a cross-modal contrastive learning training strategy is introduced to promote the fusion of model knowledge. This strategy is based on the idea that the three feature vectors encoded by the same chemical reaction should be brought closer together as positive samples, while the three feature vectors encoded by different chemical reactions should be spaced further apart as negative samples. Through this training strategy, the model can align feature vectors from different modalities in the latent space, thereby accelerating the learning process. Specifically, a nonlinear projection function g(·) is obtained, which maps the multimodal encoded vectors to a fixed-dimensional vector used for contrastive learning. Then, this function is used to train the SMILES feature vectors. Global representation of two-dimensional molecular graphs and three-dimensional molecular conformation features The mapping is performed to obtain the encoding vector derived from the SMILES feature vector. Encoding vector obtained from the global graph representation of the two-dimensional molecular graph and the encoding vector obtained from the three-dimensional molecular conformation features Next, h S h G with h C Each is considered as a mode. First, the contrast loss between any two modes is calculated.

[0066]

[0067] Where h1 and h2 represent h S h G with h C The encoding vectors for any two different modes; N represents the number of chemical reactions; The encoding vector represents the first mode corresponding to the j-th chemical reaction; The encoding vector representing the second mode corresponding to the j-th chemical reaction; The encoding vector represents the first mode corresponding to the k-th chemical reaction; represents the encoding vector for the second mode corresponding to the k-th chemical reaction; j and k are chemical reaction indices; <·,·> are vector inner products used to represent the similarity between two encoding vectors; τ is a temperature parameter, used as entropy to control the distribution.

[0068] Therefore, the total contrast loss for the three modes It can be represented as:

[0069]

[0070] S3. After pre-training, the unlabeled multimodal dataset is labeled to obtain the labeled yield dataset. The network parameters of the multilayer perceptron predictor are unfrozen. The micro-multimodal chemical reaction yield prediction model is constructed by the micro-multimodal encoder and the multilayer perceptron predictor. The micro-multimodal chemical reaction yield prediction model is jointly trained end-to-end based on the labeled yield dataset. The weights of the micro-multimodal chemical reaction yield prediction model are updated synchronously through backpropagation to obtain the trained micro-multimodal chemical reaction yield prediction model.

[0071] During the joint training process, an adaptive data filtering algorithm dynamically removes abnormal training data based on the prediction confidence and feature space distribution, iteratively updates the labeled yield dataset, and calculates the variance estimate of the predicted chemical reaction yield through the uncertainty quantification module. Based on Monte Carlo sampling, the confidence interval of the predicted chemical reaction yield is generated.

[0072] It should be noted that, as Figure 2The diagram shows the overall architecture of a microscopic multimodal chemical reaction yield prediction model employing an adaptive data filtering mechanism, as provided in this embodiment of the invention. This model combines an adaptive data filtering algorithm based on weakly supervised learning and comprises two parts: a microscopic multimodal encoder and a multilayer perceptron predictor. The adaptive data filtering algorithm dynamically adjusts the filtering threshold to prioritize high-quality labeled training data, thus mitigating the interference of data noise. The microscopic multimodal encoder integrates one-dimensional SMILES sequences, two-dimensional molecular diagrams, and three-dimensional molecular conformations, promoting a deeper understanding of chemical information from multiple perspectives and levels. The multilayer perceptron predictor generates the final predicted chemical reaction yield values.

[0073] The adaptive data filtering algorithm designed in this invention will be briefly described below. Figure 3 As shown, this algorithm dynamically adjusts the screening threshold to prioritize high-quality labeled samples. This allows the model to focus on samples with strong regularity and ease of fitting in the early stages of training, gradually transitioning to handling more complex samples as the model's capabilities improve. This process is similar to how humans learn complex knowledge, improving learning efficiency by gradually increasing the difficulty.

[0074] In the specific implementation, the one-dimensional SMILES sequences and their corresponding two-dimensional molecular diagrams are used as training data. The adaptive data filtering algorithm operates through the following steps:

[0075] S31. Model Initialization: First, the microscopic multimodal chemical reaction yield prediction model is iterated and trained several times using a labeled yield dataset to enable it to have preliminary predictive capabilities (i.e., the model can initially learn some basic reaction characterization and yield prediction abilities). During the initial training process, the weights of all training data are set to 1, and a yield data filtering vector is constructed to record the weights of each training data point.

[0076] S32. Reliable Data Screening: After the initial training is completed, the reliable data screening stage begins. The micro-multimodal chemical reaction yield prediction model evaluates the loss function value of each training data, screens out the training data whose loss function value is less than the preset screening threshold, and uses them as reliable data. The labeled yield dataset is updated by the reliable data, and the weight of the reliable data is set to 1, while the weight of the remaining training data is set to 0.

[0077] In this embodiment, this reliable data updates the original labeled yield dataset and is used for subsequent model training to update model parameters and improve the model's predictive performance.

[0078] S33. Dynamically adjust the threshold: After the reliable data has been screened, retrain the microscopic multimodal chemical reaction yield prediction model on the updated labeled yield dataset, increase the screening threshold, generate new reliable data and update the labeled yield dataset according to step S32.

[0079] In this embodiment, as training progresses, the capability of the microscopic multimodal chemical reaction yield prediction model gradually improves, and the screening threshold is dynamically adjusted. In the initial stage, the screening threshold is low to ensure that the model can select reliable data with sufficiently high confidence. As the model performance improves, the screening threshold is gradually increased, and more challenging training data is added to improve the model's generalization ability until the model converges.

[0080] It should be noted that this invention also integrates an uncertainty quantification module for evaluating the reliability of model predictions. This module provides uncertainty estimates for the model's prediction results, helping users better understand the model's prediction quality and reliability, thereby enabling more informed decision-making in practical applications. The implementation of the uncertainty quantification module is prior art; to facilitate understanding of this invention by those skilled in the art, a brief description of the module is provided below.

[0081] In this embodiment, specifically, the uncertainty quantification module operates through the following steps:

[0082] 1. Evaluation of prediction results: For each prediction result, the uncertainty quantification module calculates the variance of its predicted value as an indicator of uncertainty estimation.

[0083] 2. Uncertainty Threshold Setting: Set an uncertainty threshold according to the application scenario requirements. When the uncertainty of the prediction result exceeds the threshold, the model will prompt the user that the reliability of the prediction result is low and further verification is needed.

[0084] 3. Results Feedback and Optimization: Users can further verify and optimize the prediction results based on the prompts from the uncertainty quantification module, thereby improving the model's prediction performance.

[0085] S4. Input the one-dimensional SMILES sequence to be predicted and its corresponding two-dimensional molecular map into the trained microscopic multimodal chemical reaction yield prediction model. The microscopic multimodal encoder extracts high-dimensional fusion features, and the multilayer perceptron predictor maps the extracted high-dimensional fusion features into chemical reaction yield prediction results. The variance estimate of the chemical reaction yield prediction results is calculated through the uncertainty quantification module. The confidence interval of the chemical reaction yield prediction results is generated based on Monte Carlo sampling, realizing multimodal chemical reaction yield prediction based on adaptive data screening.

[0086] In this embodiment, the microscopic multimodal encoder integrates multi-dimensional molecular features into a comprehensive feature vector (i.e., multi-dimensional fusion feature) through a multimodal feature fusion strategy. Specifically, for the SMILES feature vector... Global representation of two-dimensional molecular graphs and three-dimensional molecular conformation features This invention uses a feature fusion function v(·) to... and By combining these features, we can achieve the goal of fusing feature information from different modalities and generating high-dimensional fused features.

[0087]

[0088] in, This indicates a feature concatenation operation. The resulting high-dimensional fused features... The input is fed into a multilayer perceptron predictor to help the model understand chemical information more deeply from multiple perspectives and levels. Finally, the multilayer perceptron predictor processes and outputs the predicted chemical reaction yield.

[0089] Through the above technical solution, the present invention can effectively resist the interference of data noise and deeply integrate multimodal information, thereby improving the accuracy and reliability of chemical reaction yield prediction.

[0090] The present invention will now demonstrate the application effect of the multimodal chemical reaction yield prediction method based on adaptive data screening described in S1 to S4 of the above embodiments on a specific dataset through a specific example, so as to facilitate understanding of the essence of the present invention.

[0091] Example

[0092] The specific implementation process of the multimodal chemical reaction yield prediction method based on adaptive data screening used in this embodiment is as described above and will not be repeated here.

[0093] To verify the effectiveness of the weakly supervised adaptive training strategy and multimodal fusion mechanism proposed in this invention for chemical reaction yield prediction tasks, comparative experiments were conducted on the Buchwald HTE, Suzuki HTE, and USPTO datasets. The datasets used cover representative chemical reaction scenarios with different reaction types, experimental scales, and noise levels. The comparative models include the currently mainstream sequence modeling method YieldBERT, the fingerprint feature method DRFP based on molecular structure changes, the graph neural network model AGNN with an attention mechanism, the MPNN model based on a message passing mechanism, and the Egret method, which uses a conditional contrastive learning strategy to improve model robustness. These models are representative in their structural design, widely used in yield prediction tasks, and have strong reference value. The experimental results are shown in Tables 1-3.

[0094] Table 1 Performance evaluation of the present invention on the Buchwald HTE dataset.

[0095] MAE 3.91 3.90 3.89 2.92 4.01 2.90 RMSE 6.02 6.03 6.01 4.43 6.19 4.35 <![CDATA[R 2 ]]> 0.951 0.950 0.953 0.974 0.940 0.977

[0096] Table 1 shows the yield prediction performance of different models on the Buchwald HTE dataset. This dataset contains 3955 palladium-catalyzed Buchwald-Hartwig C–N coupling reactions, involving 15 aryl halides, 4 ligands, 3 bases, and 23 additive combinations, representing typical high-throughput experimental (HTE) data. The reaction types are relatively uniform, and the molecular structures show high similarity. The results in Table 1 show that the method of this invention achieves high yield predictions on R... 2 It outperforms existing baseline methods in all three metrics: MAE and RMSE, especially in the coefficient of determination R. 2 The value reached 0.977, indicating that it has stronger feature extraction capabilities and fitting accuracy in such low heterogeneous data.

[0097] Table 2 Performance evaluation of the present invention on the Suzuki HTE dataset.

[0098] MAE 8.13 6.77 6.96 6.02 6.76 6.03 RMSE 12.07 10.59 11.01 9.41 10.59 9.22 <![CDATA[R 2 ]]> 0.810 0.850 0.840 0.886 0.850 0.889

[0099] Table 2 shows the prediction results of different models on the Suzuki HTE dataset. This dataset contains 5760 Suzuki-Miyaura C–C coupling reactions, consisting of 15 electronegative / nucleophile combinations, 12 ligands, 8 bases, and 4 solvents. The reaction conditions are more complex and diverse, resulting in greater data heterogeneity and higher noise levels. The results in Table 2 demonstrate that the method of this invention still achieves the best R-value on this dataset. 2 The 0.889 and RMSE values ​​indicate that the proposed model has good generalization ability and noise resistance, and remains stable under complex conditions.

[0100] Table 3 Performance evaluation of the present invention on the USPTO

[0101]

[0102] Table 3 shows the yield prediction results for two subsets of the USPTO real-world patent dataset: USPTO-gram (gram-scale reactions) and USPTO-subgram (milligram-scale reactions). This dataset, derived from publicly available US patent literature, covers various types of real-world reactions, such as C–C coupling, C–N coupling, and redox reactions, and is characterized by significant scale heterogeneity, high noise, and diverse reaction types. Table 3 illustrates the prediction performance of the model in this invention on the two subsets. The results demonstrate that the method of this invention maintains good stability on real-world large-scale complex reaction data, outperforming existing models on multiple metrics, and exhibiting strong generalization ability and practical applicability.

[0103] In summary, as shown in Tables 1 to 3, the method of the present invention exhibits significant advantages on datasets with different reaction types, scales, and distribution characteristics, demonstrating stronger prediction accuracy, robustness, and practical application potential.

[0104] It is understood that the multimodal chemical reaction yield prediction method based on adaptive data screening described in S1 to S4 above can essentially be implemented by a computer program. Therefore, based on the same inventive concept, another preferred embodiment of the present invention also provides a computer program product corresponding to the multimodal chemical reaction yield prediction method based on adaptive data screening provided in the above embodiments. This computer program / instructions, when executed by a processor, can implement the multimodal chemical reaction yield prediction method based on adaptive data screening as described in the above embodiments.

[0105] Similarly, based on the same inventive concept, another preferred embodiment of the present invention also provides a computer electronic device corresponding to the multimodal chemical reaction yield prediction method based on adaptive data screening provided in the above embodiments, which includes a memory and a processor;

[0106] The memory is used to store computer programs;

[0107] The processor is configured to implement the multimodal chemical reaction yield prediction method based on adaptive data screening in the above embodiments when executing the computer program.

[0108] Furthermore, the logical instructions in the aforementioned memory can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.

[0109] Therefore, based on the same inventive concept, another preferred embodiment of the present invention also provides a computer-readable storage medium corresponding to the multimodal chemical reaction yield prediction method based on adaptive data screening provided in the above embodiments. The storage medium stores a computer program, which, when executed by a processor, can realize the multimodal chemical reaction yield prediction method based on adaptive data screening in the above embodiments.

[0110] It is understood that the aforementioned storage media may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device. Furthermore, the storage media may also be various media capable of storing program code, such as USB flash drives, external hard drives, magnetic disks, or optical discs.

[0111] It is understood that the processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0112] It should also be noted that those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process of the system described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here. In the embodiments provided in this application, the division of steps or modules in the system and method is merely a logical functional division, and there may be other division methods in actual implementation. For example, multiple modules or steps may be combined or integrated together, and a module or step may also be split.

[0113] The embodiments described above are merely preferred embodiments of the present invention and are not intended to limit the invention. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the invention. Therefore, all technical solutions obtained through equivalent substitution or transformation fall within the protection scope of the present invention.

Claims

1. A method for predicting the yield of multimodal chemical reactions based on adaptive data screening, characterized in that, Includes the following steps: S1. Obtain one-dimensional SMILES sequences and their corresponding two-dimensional molecular maps, and form an unlabeled multimodal dataset; S2. In the pre-training phase, the network parameters of the multilayer perceptron predictor are fixed, and the micro-multimodal encoder is trained on an unlabeled multimodal dataset using a cross-modal contrastive learning training strategy. S3. After pre-training, the unlabeled multimodal dataset is labeled to obtain the labeled yield dataset. The network parameters of the multilayer perceptron predictor are unfrozen. The micro-multimodal chemical reaction yield prediction model is constructed by the micro-multimodal encoder and the multilayer perceptron predictor. The micro-multimodal chemical reaction yield prediction model is jointly trained end-to-end based on the labeled yield dataset. The weights of the micro-multimodal chemical reaction yield prediction model are updated synchronously through backpropagation to obtain the trained micro-multimodal chemical reaction yield prediction model. During the joint training process, an adaptive data filtering algorithm dynamically removes abnormal training data based on the prediction confidence and feature space distribution, iteratively updates the labeled yield dataset, and calculates the variance estimate of the predicted chemical reaction yield through the uncertainty quantification module, and generates the confidence interval of the predicted chemical reaction yield based on Monte Carlo sampling. S4. Input the one-dimensional SMILES sequence to be predicted and its corresponding two-dimensional molecular map into the trained microscopic multimodal chemical reaction yield prediction model. The microscopic multimodal encoder extracts high-dimensional fusion features, and the multilayer perceptron predictor maps the extracted high-dimensional fusion features into chemical reaction yield prediction results. The variance estimate of the chemical reaction yield prediction results is calculated through the uncertainty quantification module. The confidence interval of the chemical reaction yield prediction results is generated based on Monte Carlo sampling, realizing multimodal chemical reaction yield prediction based on adaptive data screening. In the micro-multimodal encoder of step S2, the SMILES sequence encoder encodes the one-dimensional SMILES sequence and maps the one-dimensional SMILES sequence to the vector space to obtain the SMILES feature vector. The molecular graph encoder encodes the two-dimensional molecular graph to generate the global graph representation of the two-dimensional molecular graph. Based on the two-dimensional molecular graph, the three-dimensional molecular conformation features are generated. Then, the SMILES feature vector, the global graph representation of the two-dimensional molecular graph and the three-dimensional molecular conformation features are combined through feature splicing operation to generate a comprehensive feature vector. The adaptive data filtering algorithm operates through the following steps: S31. First, the micro-multimodal chemical reaction yield prediction model is initialized and trained for several rounds using a labeled yield dataset to enable the micro-multimodal chemical reaction yield prediction model to have preliminary prediction capabilities. During the initial training process, the weights of all training data are set to 1, and a yield data filtering vector is constructed to record the weights of each training data. S32. After the initial training is completed, the reliable data screening stage begins. The micro-multimodal chemical reaction yield prediction model evaluates the loss function value of each training data, selects training data with a loss function value less than the preset screening threshold and uses them as reliable data. The labeled yield dataset is updated by the reliable data, and the weight of the reliable data is set to 1, while the weight of the remaining training data is set to 0. S33. After the reliable data has been screened, the microscopic multimodal chemical reaction yield prediction model is retrained on the updated labeled yield dataset to increase the screening threshold. New reliable data is generated and the labeled yield dataset is updated according to step S32. S34. Iterate the training continuously according to step S33 until the microscopic multimodal chemical reaction yield prediction model converges, and obtain the trained microscopic multimodal chemical reaction yield prediction model.

2. The method for predicting the yield of a multimodal chemical reaction based on adaptive data screening as described in claim 1, characterized in that, The specific generation method of the three-dimensional molecular conformation features is as follows: First, the initial three-dimensional molecular conformation is obtained from the two-dimensional molecular diagram using a conformation prediction-related technology. Then, the energy of the initial three-dimensional molecular conformation is minimized by the MMFF molecular mechanical force field to obtain the corresponding stable conformation and form a set of three-dimensional molecular stable conformations that participate in chemical reactions. Next, the cheminformatics and machine learning toolkit RDKit is used to calculate the stable conformation set to obtain three-dimensional molecular conformation features containing geometric diameter, moment of inertia, and mass distribution. Alternatively, the three-dimensional molecular conformation encoder can encode the stable conformation set to generate three-dimensional molecular conformation features.

3. The method for predicting the yield of a multimodal chemical reaction based on adaptive data screening as described in claim 1, characterized in that, In the cross-modal contrastive learning training strategy, SMILES feature vectors, global graph representations of two-dimensional molecular graphs, and three-dimensional molecular conformation features belonging to the same chemical reaction are used as positive samples, while SMILES feature vectors, global graph representations of two-dimensional molecular graphs, and three-dimensional molecular conformation features belonging to different chemical reactions are used as negative samples. The contrastive loss between the two different modalities is calculated separately, and the calculated contrastive losses are weighted and summed to obtain the total contrastive loss. The network parameters of the micro-multimodal encoder are updated based on minimizing the total contrastive loss.

4. The method for predicting the yield of a multimodal chemical reaction based on adaptive data screening as described in claim 1, characterized in that, The SMILES sequence encoder uses the BERT model.

5. The method for predicting the yield of a multimodal chemical reaction based on adaptive data screening as described in claim 1, characterized in that, The molecular graph encoder employs a message-passing neural network.

6. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instruction is executed by the processor, it can implement the multimodal chemical reaction yield prediction method based on adaptive data screening as described in any one of claims 1 to 5.

7. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, which, when executed by a processor, implements the multimodal chemical reaction yield prediction method based on adaptive data screening as described in any one of claims 1 to 5.

8. A computer electronic device, characterized in that, Including memory and processor; The memory is used to store computer programs; The processor is configured to, when executing the computer program, implement the multimodal chemical reaction yield prediction method based on adaptive data screening as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Molecular large model based on multi-dimensional molecular information, construction method and application

    CN117524353A

  • Hybrid integrated circuit assembly defect detection method based on semi-supervised learning

    CN117670889A