Similar organic reaction retrieval method and storage medium
By constructing an organic chemical reaction formula recall vector database and searching using a similarity algorithm, the problem of insufficient accuracy and speed of organic reaction similarity search in the prior art is solved, and more efficient reaction retrieval and data storage are achieved.
Patent Information
- Application Number
- CN202510296719.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-06-24
AI Technical Summary
The prior art lacks accuracy and speed when searching for organic reaction similarity, especially in large-scale reaction databases, which are difficult for traditional methods to effectively process and retrieve structural conversion information of reaction components.
By constructing an organic chemical reaction formula recall vector database, it includes reaction formula identification, reaction type, reaction center characteristics and reaction formula component molecular characteristics, and searching matching reaction formulas in the database using a preset similarity algorithm to improve the search accuracy and speed.
It significantly improves the accuracy and speed of organic reaction retrieval, can maximize the chemistry cognitive needs of chemists, and improves the efficiency of data storage and retrieval.
Smart Images

Figure CN120199349A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of chemistry, and particularly to a method for retrieving similar organic reactions for calculating and retrieving similarities between organic chemical reactions, and a storage medium. Background Art
[0002] Before designing an organic reaction and conducting experimental operations, in addition to relying on their accumulated experience, organic synthesis chemists generally query reaction databases for reaction formulas that are the same as or similar to the target reaction formula in the literature, and refer to various specific elements such as reaction conditions and experimental procedures, and adjust them in combination with the actual situation to determine the final experimental plan. The basis for the success of this experimental plan design lies in the accuracy and time consumption of the similarity retrieval of the target reaction formula. Based on the current existing technical level, accuracy requires more detailed information on the structure and structure conversion of the reaction components (including reactants, reagents, and products) molecules, but the processing, storage, and retrieval of this information will consume more time. Therefore, it is necessary to achieve a relatively ideal balance state, especially in the case of similarity retrieval on a reaction database containing tens of millions of entries, which is even more important.
[0003] When constructing traditional reaction databases, they are generally split into reactant molecules and product molecules, and the chemical fingerprints (Chemical Fingerprint / CFP) of the molecules are calculated separately. When performing similarity retrieval, a certain method (such as the Tanimoto coefficient method) is used to calculate the similarity distance between compounds. At this time, the result may be very different from the reaction type expected by chemists because the structural conversion process from reactants to products is not considered.
[0004] Some existing technologies also consider the influence of reaction sites and their attached microenvironments and important functional groups in addition to simple compound similarities, and make specific quantifications. However, using traditional relational databases to store chemical fingerprints and performing similarity retrieval based on chemical fingerprints will have a very large time cost in the case of a large amount of organic reaction data. There is no clear and consistent method and result for the definition and recognition of important functional groups, which strongly rely on the experience of chemists. In addition, given that the current various technologies cannot accurately identify reaction centers completely, auxiliary methods are needed to improve this accuracy, or additional classification marks are needed to facilitate screening during retrieval to filter out irrelevant reaction types and improve the accuracy of the retrieved results; finally, these chemical fingerprint data cannot be directly used in machine learning such as deep neural networks, and more efficient expression forms, storage, and retrieval methods are needed. Summary of the Invention
[0005] A series of simplified concepts are introduced in the Summary of the Invention section. These simplified concepts are all simplified from the prior art in this field and will be further described in detail in the Detailed Description section. The Summary of the Invention section of the present invention does not mean to attempt to define the key features and essential technical features of the claimed technical solution, nor does it mean to attempt to determine the protection scope of the claimed technical solution.
[0006] The technical problem to be solved by the present invention is to provide a method for retrieving similar organic reactions that can improve the retrieval accuracy and speed, including the following steps: S1. Construct a vectorized database for recalling organic chemical reaction formulas, which includes: reaction formula identifier UID, reaction type marker ID of the reaction formula, chemical fingerprints of reaction centers at all levels, and molecular characteristics of reaction formula components; the molecular components of the reaction formula are divided into a pre-reaction part and a post-reaction part; the pre-reaction part of the molecular components of the reaction formula is "reactants + reagents", and the post-reaction part of the molecular components of the reaction formula is "products"; Construct an auxiliary database for recalling organic chemical reaction formulas, which includes: additional data items other than the reaction formula associated with the reaction formula identifier UID; A reaction center refers to a specific position or area where a reaction occurs in a chemical reaction or physical process. At this central position, atoms, ions, or molecules undergo material transformation or combination to produce new chemical substances; A chemical fingerprint is to convert the structural characteristics of a chemical molecule into a series of binary codes (bit vectors) or numerical values through specific algorithms or rules to represent information related to the local structure, physicochemical properties, or biological activity of the molecule; S2. For the target reaction formula, obtain the reaction type marker of the target reaction formula, chemical fingerprints of reaction centers at all levels, and molecular characteristics of reaction formula components; S3. Based on the reaction type marker, reaction center characteristics at all levels, and molecular characteristics of reaction formula components of the target reaction formula, retrieve in the vectorized database for recalling organic chemical reaction formulas through a preset similarity algorithm to obtain the reaction formula identifiers of the matching reaction formulas; S4. According to the reaction formula identifiers of the matching reaction formulas, extract the associated additional data items from the auxiliary database for recalling organic chemical reaction formulas; S5. Sort the matching reaction formulas according to preset rules and output a specified number of similar reaction formulas.
[0007] Preferably, further improve the above-mentioned method for retrieving similar organic reactions by using the chemical fingerprint of the reaction center as the reaction center characteristic; Use the chemical fingerprint of the molecular components of the reaction formula as the molecular characteristics of the reaction formula components.
[0008] Preferably, further improve the above-mentioned similar organic reaction retrieval method. Constructing a vectorized database for recalling organic chemical reaction formulas includes the following sub-steps: S1.1. Convert the reaction formula data into the RSMILES format, standardize it to STD-RSMILES, and assign a unique reaction formula identifier UID to each reaction formula. Converting the reaction formula data into the RSMILES format is to standardize the chemical structure of the reaction formula to eliminate the problem of inconsistent chemical structure expressions caused by different forms of metal salts, tautomeric forms, etc. The structural formula form of each component molecule is the standardized SMILES (Simplified Molecular Input Line Entry Specification). Correspondingly, the standardized SMILES is a preferred implementation method, and other standardized formats or custom general formats can also be selected. S1.2. Calculate the reaction type of each reaction formula, and extract the reaction type label ID of the reaction formula as a scalar data item for screening; the reaction type label can be calculated based on organic reaction mechanisms and conversion rules. S1.3. Calculate the chemical fingerprints of at least three levels of reaction centers in each reaction formula, and vectorize them according to requirements. The reaction centers can be identified according to the reaction formula. That is, the chemical fingerprints of the reaction center and its 1 to multiple covalently bonded atoms in its extension; the more the extension, the greater the impact on the similarity of the reaction center and its surrounding environment, and it can meet the different expectations of chemists for the similarity retrieval results of reactions or conversions; multiple levels of reaction fingerprints can be respectively fixed as determined scalar screening conditions corresponding to the previously calculated reaction types, or can be used as vectors for similarity retrieval when chemists need changes in reaction type and reaction center similarity. S1.4. Calculate the chemical fingerprints of the pre-reaction part and the post-reaction part respectively, vectorize them respectively, and form a vectorized database for recalling organic chemical reaction formulas. Calculate the chemical fingerprints of the pre-reaction part and the post-reaction part. The longer the bit length of the chemical fingerprint, the more molecular features it contains, which is beneficial to the accuracy of retrieval, but the more storage space it occupies, affecting the retrieval time. It is necessary to achieve a balance as much as possible according to actual needs.
[0009] Preferably, further improve the above-mentioned similar organic reaction retrieval method. Constructing an auxiliary database for recalling organic chemical reaction formulas includes the following sub-step: associating additional data items other than the reaction formula through the reaction formula identifier.
[0010] Preferably, further improve the above-mentioned similar organic reaction retrieval method. Implementing step S1.3 to vectorize according to requirements includes: If the requirement is to retrieve the similarity of the reaction center, then vectorize the chemical fingerprint of the reaction center. If the requirement is that the reactive formula to be queried must have exactly the same reaction center and corresponding reaction type, the chemical fingerprint of the reaction center is scalarized.
[0011] Preferably, further improving the above-mentioned similar organic reaction retrieval method, the implementation steps of S2 include: S2.1, draw the target reaction formula through a chemical structure drawing tool, including: reactants, reagents and products; S2.2, convert the reaction formula data into the RSMILES format and standardize it to STD-RSMILES; S2.3, calculate the reaction type for STD-RSMILES to obtain its reaction type label ID; S2.4, calculate the chemical fingerprints of at least three levels of reaction centers in the reaction formula and vectorize them according to the requirements; S2.5, calculate the chemical fingerprints of the component molecules of the reaction formula for STD-RSMILES respectively.
[0012] Preferably, further improving the above-mentioned similar organic reaction retrieval method, the preset parameters include the number of returns and / or the similarity threshold; The specified parameter is the product similarity.
[0013] Preferably, further improving the above-mentioned similar organic reaction retrieval method, calculate the similarity through the ChemBert model; the ChemBert model is based on the RoBERTa model, and this model is pre-trained on the task of masked language modeling (MLM); Obtain a specified number of similar reactions through the Reranking model; The Reranking model is a technology used to optimize the sorting of search results in the fields of information retrieval and natural language processing.
[0014] The present invention provides a computer-readable storage medium, which stores a computer program internally, and the computer program is used to implement the steps of the above-mentioned similar organic reaction retrieval method when executed.
[0015] The present invention calculates the chemical fingerprints of the component molecules in the organic reaction formula, calculates the reaction fingerprints at all levels, and adds the accurately calculated reaction type as a screening condition to improve the accuracy of the retrieval and make the retrieval results conform to the chemical cognition of chemists to the greatest extent. The present invention constructs a vectorized database for the component molecules of the reaction formula, the chemical fingerprints of reaction centers at all levels and the reaction type, and uses a deep learning model to perform similarity retrieval through the quantitative database, significantly improving the efficiency of data storage and retrieval. Description of the Drawings
[0016] The accompanying drawings of the present invention are intended to illustrate the general characteristics of the methods, structures, and / or materials used in specific exemplary embodiments of the present invention, and to supplement the descriptions in the specification. However, the accompanying drawings of the present invention are schematic diagrams not drawn to scale, and thus may not accurately reflect the precise structures or performance characteristics of any given embodiment. The accompanying drawings of the present invention should not be construed as limiting or restricting the scope of the numerical values or properties covered by the exemplary embodiments according to the present invention. The present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments: Figure 1 It is a standardized example diagram of the chemical structure of organic reaction component molecules.
[0017] Figure 2 It is a schematic flowchart of the method for retrieving similar organic reactions of the present invention. Specific Embodiments
[0018] The following specific embodiments illustrate the implementation manners of the present invention. Those skilled in the art can fully understand other advantages and technical effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through different specific implementation manners. The details in this specification can also be applied based on different viewpoints, and various modifications or changes can be made without departing from the overall design concept of the invention. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other. The following exemplary embodiments of the present invention can be implemented in many different forms and should not be construed as being limited only to the specific embodiments described herein. It should be understood that these embodiments are provided to make the disclosure of the present invention thorough and complete, and to fully convey the technical solutions of these exemplary specific embodiments to those skilled in the art.
[0019] The First Embodiment; Referring to Figure 2 as shown, the present invention provides a method for retrieving similar organic reactions, including the following steps: Implement step S1. On the premise that a huge reaction data set has been obtained, the scale of the reaction data set affects the retrieval accuracy. Extract the reaction formulas in the reaction data set. These reaction formulas can be in formats commonly used in the chemical industry, such as RDF, RXN, RSMILES, etc. S1. Construct a vectorized database for recalling organic chemical reaction formulas, which includes: reaction formula identifier UID, reaction type marker ID of the reaction formula, characteristics of reaction centers at all levels, and characteristics of reaction formula component molecules; in this embodiment, chemical fingerprints are used as the characteristics of reaction centers, and chemical fingerprint reaction formula component molecule characteristics are used; reaction formula component molecules are divided into the front part of the reaction; Convert the reactive data to the RSMILES format, standardize each reactive RSMILES, and obtain the standardized structure STD-RSMILES. Assign a unique reactive identifier UID to each reaction (subsequent calculations are based on this, unless otherwise stated). Eliminate the differences in the expressions of certain structural fragments or features in the same molecule from different data sources, such as different forms of metal salts, tautomers, etc. as Figure 1 shown, to avoid potential problems they may bring during database construction and retrieval; Calculate the reaction type of each reaction formula based on organic reaction mechanisms and conversion rules, and extract the reaction type label ID of the reaction formula as a scalar data item for screening; this processing method has relatively high accuracy, but it cannot cover all types of organic reactions and is affected by chemists' habits or ways of inputting reaction formulas. The definition and recognition of the following multi-level reaction centers need to be mutually assisted; Calculate the chemical fingerprints of at least three levels of reaction centers in each reaction formula and vectorize them according to requirements; that is, the chemical fingerprints of the reaction center and its 1 to multiple covalently bonded atoms extending outward; the more outward extensions, the greater the impact on the similarity of the reaction center and its surrounding environment, which can meet chemists' different expectations for the similarity retrieval results of reactions or conversions; the reaction fingerprints at multiple levels can be fixed as definite scalar screening conditions corresponding to the reaction types calculated above, or used as vectors for similarity retrieval when chemists need changes in reaction type and reaction center similarity; Exemplarily, the chemical fingerprints can be calculated through the following algorithms: Morgan fingerprints (Circular Fingerprints), MACCS fingerprints (Molecular ACCess System), tomPair fingerprints (Atom Pairs), Topological Torsion fingerprints (Topological Torsions), PubChem fingerprints (PC Fingerprints), Mini FingerPrint (MFP), Barnard Chemistry Information (BCI) fingerprints, SMIlesFingerPrint (SMIFP), Extended Connectivity Fingerprints (ECFPs), Functional-class Fingerprints (FCFPs), Molprint2D, Molprint3D, RDKFingerprint; these algorithms are widely used in chemoinformatics and drug research and development, and different algorithms are suitable for different application scenarios; It should be further noted that there are two processing methods for quantification according to requirements: 1). If there is a need to retrieve the similarity of the reaction center (i.e., the reaction center and the corresponding reaction types may be different, with a similarity less than 1.0), then vectorization is performed; 2). If there is a need to have exactly the same reaction center and the corresponding reaction types as the query reaction formula, then the chemical fingerprint is scalarized, that is, a definite encoding is performed as the scalar data item during screening; The chemical fingerprints of the pre-reaction part and the post-reaction part are calculated and vectorized respectively to form a vectorized database for organic chemical reaction formula recall; Exemplarily, it can be adjusted between 512 - 2048 bits according to the actual situation, such as 1024 bits; They reflect the influence of the whole molecule and atoms or groups far from the reaction center on the similarity of the reaction formula, and are adapted to the attention degree of chemists to the structure of reaction components; Exemplarily, the chemical fingerprint of the reaction formula component molecule includes more than 8 parts (sections); 1: Reactant + reagent molecule 2: Product molecule 3: Reactant-side atoms of the reaction center 4: Product-side atoms of the reaction center 5: Reactant-side atoms of the reaction center expanded by 1 bond 6: Product-side atoms of the reaction center expanded by 1 bond 7: Reactant-side atoms of the reaction center expanded by 2 bonds 8: Product-side atoms of the reaction center expanded by 2 bonds 9, 10,...: Reactant-side atoms and product-side atoms of the reaction center expanded by more bonds; Construct an auxiliary database for organic chemical reaction formula recall, which includes: additional data items other than the reaction formula associated with the reaction formula identifier UID. Exemplary additional data items include reaction conditions, operations, yields, safety matters, detection data, patents, literature citations, etc.; S2. For the target reaction formula, obtain the reaction type marker ID, the chemical fingerprints of each level of the reaction center, and the chemical fingerprints of the reaction formula component molecules; For example, draw a reaction formula from a chemical structure drawing tool, including reactants, reagents, and products, and convert it to the RSMILES format using built-in tools or external tools; Standardize the RSMILES format to obtain STD-RSMILES; Calculate the reaction type for STD-RSMILES to obtain its reaction type marker ID; Calculate the chemical fingerprint of the tertiary or higher-level reaction center for STD-RSMILES, and perform vectorization or scalarization encoding according to the preset requirements; Calculate the chemical fingerprints of the reactants + reagents and products for STD-RSMILES respectively, and perform vectorization; S3. Based on the reaction type label of the target reaction formula, the characteristics of each level of reaction center, and the molecular characteristics of the components of the reaction formula, retrieve from the vectorized database of organic chemical reaction formulas through a preset similarity algorithm, and obtain the reaction formula identifier of the matching reaction formula; Exemplarily, the similarity can be calculated by the following algorithms: text similarity algorithms (Levenshtein Distance algorithm, Damerau-Levenshtein Distance algorithm), vector similarity algorithms (Cosine Similarity algorithm, Dot Product algorithm, Manhattan Distance algorithm, Euclidean Distance algorithm), similarity algorithms in deep learning (Embedding Models algorithm, Hybrid Search algorithm); obtain the required results by specifying the return number and / or similarity threshold; S4. Retrieve the additional data items associated with the reaction formula identifier UID of the matching reaction formula from the auxiliary database of organic chemical reaction formulas according to the reaction formula identifier UID of the matching reaction formula; S5. Sort the matching reaction formulas according to the preset rules, and output a specified number of similar reaction formulas. The preset rules can be specified, for example, specifying the similarity from high to low as the preset rule.
[0020] Using chemical fingerprints can quickly recall known reactions with the same reaction principle as the target reaction, but the number of reactions recalled by this method alone will be very large. In order to sort the recalled results, a rough sorting algorithm based on product similarity is first used.
[0021] Second embodiment; The second embodiment of the present invention is a further improvement based on the first embodiment, and the same parts will not be described in detail; S6. Use a self-supervised pre-training model, such as the ChemBert model, to recalculate the similarity; S7. Use a re-ranking model, such as the Reranking model, to obtain a specified number of similar reactions as the output result.
[0022] It should be noted that in order to further improve the recommendation effect and make full use of user behavior data to optimize the user adoption rate index, a deep learning model based on BERT is used to measure the similarity between two reactions. First, the currently existing reaction data is used as the training data for the basic model, and the data set is divided into clusters.
[0023] Construct a positive sample data set with exactly the same conditions. For the data constructed from two reactions within a reaction cluster with exactly the same conditions, the predicted regression result is 1.
[0024] Then, construct a similarity set through data of the same reaction type. If the relationship between two reactions is of the same type, then in the prediction regression stage, the predicted value is 0.5.
[0025] Finally, construct a negative sample set for the data that is not in these two data sets, and the predicted regression value is -1.
[0026] Construct 1B data pairs through a data set of tens of millions to train our backbone ChemBert Model. As a basic data model. Then, use the user feedback data, click data, and favorite data as fine-tuning data to further optimize the model effect.
[0027] BERT (Bidirectional Encoder Representations from Transformers) is a pre-trained language representation model that can be used for various natural language processing tasks, including classification and regression tasks. For regression tasks, the model structure of BERT can be adjusted. The following is an introduction to the model structure of using BERT for regression tasks: Input layer: The input is a pair of SMILES. After tokenization and encoding, it is converted into token IDs, segment IDs, and position IDs.
[0028] BERT encoder: The input passes through the multi-layer Transformer encoder of BERT to generate the context representation of each token. The encoder of BERT consists of multiple layers of Transformers, and each layer contains a self-attention mechanism and a feed-forward neural network.
[0029] [CLS] token representation: In the output of BERT, the representation of the first token ([CLS]) is used as the aggregated representation of the entire input sequence. This representation contains the context information of the entire pair of SMILES.
[0030] Regression layer: Add a regression layer (usually a fully connected layer) to the representation of the [CLS] token, and output a real value, which is the regression value constructed above. The output of this fully connected layer can represent the target value (similarity) of the regression task.
[0031] Loss function: For regression tasks, the mean squared error (MSE) is usually used as the loss function to measure the difference between the predicted value and the true value.
[0032] Third embodiment; A computer-readable storage medium of the present invention stores a computer program therein, and when the computer program is executed, it is used to implement the steps of the similar organic reaction retrieval method described in any one of the first embodiment or the second embodiment.
[0033] The computer-readable medium includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.
[0034] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. It will also be understood that terms such as those defined in a general dictionary should be interpreted as having a meaning consistent with their meaning in the context of the relevant art, and not in an ideal or overly formal sense, unless expressly defined herein.
[0035] The present invention has been described in detail above through specific embodiments and examples, but these do not constitute a limitation to the present invention. Without departing from the principles of the present invention, those skilled in the art can also make many modifications and improvements, which should also be regarded as the protection scope of the present invention.
Claims
1. A similar organic reaction retrieval method, characterized in that: The following steps are involved: S1, constructing a quantitative database for recalling organic chemical reaction formulas, which includes: reaction formula identification, reaction type label of the reaction formula, characteristics of reaction centers at all levels and characteristics of reaction formula component molecules; the reaction formula component molecules are divided into pre-reaction part and post-reaction part; Constructing an organic chemical reaction formula recall auxiliary database, which includes: additional data items other than the reaction formula associated with the reaction formula identifier; S2, target reaction formula, obtaining the reaction type label of the target reaction formula, the chemical fingerprints of the reaction centers at each level and the molecular characteristics of the reaction formula components; S3, based on the reaction type label of the target reaction formula, the characteristics of the reaction centers at each level and the molecular characteristics of the reaction formula components, a search is performed in the organic chemical reaction formula recall vector database through a preset similarity algorithm to obtain a reaction formula identifier that matches the reaction formula; S4, extracting associated additional data items from an organic chemical reaction formula recall auxiliary database according to the reaction formula identifier of the matching reaction formula; S5, sorting the matching reaction formulas according to preset rules, and outputting a specified number of similar reaction formulas.
2. The similar organic reaction retrieval method according to claim 1, characterized in that: Also includes: S6, re-ranking and calculating similarity using the self-supervised pre-trained model; S7, using the re-ranking model, obtains a specified number of similar reactions as output results.
3. The similar organic reaction retrieval method according to claim 1, characterized in that: The chemical fingerprint of the reaction center is used as the reaction center feature; The chemical fingerprints of the reaction formula component molecules are used as the molecular characteristics of the reaction formula component molecules.
4. The similar organic reaction retrieval method according to claim 1, characterized in that: The construction of organic chemical reaction recall vector database includes the following sub-steps: S1.1, convert the reaction data into RSMILES format and standardize it into STD-RSMILES, and assign a unique reaction identifier to each reaction; S1.2, calculating the reaction type of each reaction formula, and extracting the reaction type label of the reaction formula as a scalar data item for screening; S1.3, calculate the chemical fingerprints of at least three levels of reaction centers in each reaction formula and vectorize them according to the requirements; S1.4, respectively calculate the chemical fingerprints of the pre-reaction part and the post-reaction part and vectorize them to form an organic chemical reaction formula recall vectorized database.
5. The similar organic reaction retrieval method according to claim 4, characterized in that: The construction of an organic chemical reaction formula recall auxiliary database includes the following sub-steps: associating additional data items other than the reaction formula through the reaction formula identifier.
6. The similar organic reaction retrieval method according to claim 4, characterized in that ,Implementation step S1.3 is quantified according to the needs including: If the requirement is to retrieve the similarity of reaction centers, the chemical fingerprint of the reaction center is vectorized; If the requirement is that the reaction center and the corresponding reaction type must be exactly the same as those in the query reaction formula, the chemical fingerprint of the reaction center is scalarized.
7. The similar organic reaction retrieval method according to claim 1, characterized in that , the implementation step S2 includes: S2.1, draw the target reaction formula using the chemical structure drawing tool, including reactants, reagents and products; S2.2, convert the reactive data into RSMILES format and standardize it into STD-RSMILES; S2.3, calculating the reaction type of STD-RSMILES and obtaining its reaction type label; S2.4, calculate the chemical fingerprints of reaction centers at each level in the reaction formula and vectorize them according to the requirements; S2.5, calculate the chemical fingerprints of the component molecules of the reaction formula for STD-RSMILES respectively.
8. The similar organic reaction retrieval method according to claim 1, characterized in that: Preset parameters include return count and / or similarity threshold; The specified parameter is product similarity.
9. The similar organic reaction retrieval method according to claim 2, characterized in that: Similarity was calculated by rearrangement of the ChemBert model; Get the specified number of similar responses through the Reranking model.
10. A computer-readable storage medium, characterized in that: A computer program is stored therein, and when the computer program is executed, it is used to implement the steps of the similar organic reaction retrieval method according to any one of claims 1 to 9.