Compound vector database construction method, compound similarity search method and device
By converting the compound structure into characteristic vectors and constructing a compound vector database, the problem of low compound similarity search efficiency in traditional methods is solved, and the rapid and accurate retrieval of large-scale compound databases is achieved, and applications in the fields of drug discovery and material design are supported.
Patent Information
- Application Number
- CN202510024254.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-07
- Publication Date
- 2025-05-16
AI Technical Summary
Traditional compound similarity search methods have high computational complexity and low search efficiency when processing large-scale data, and lack the ability to quickly retrieve similar compounds.
By converting compound structures into eigenvectors and using vector databases and approximate nearest neighbor search algorithms, a compound vector database is constructed to achieve fast and accurate similarity search.
It has achieved rapid and accurate retrieval of compound structures in large-scale compound databases, significantly improved search efficiency, and provided powerful tools to support the fields of drug discovery and material design.
Smart Images

Figure CN120015181A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of chemical informatics, and in particular to a compound vector database construction method, a compound similarity search method and a device. Background Art
[0002] In the fields of drug discovery, material design, and chemical research, similarity search of compound structures is a key step in finding new compounds and optimizing the performance of existing compounds. However, traditional similarity search methods based on text descriptions or graph matching face the problems of high computational complexity and low search efficiency when processing large-scale data. The rise of vector database technology provides a new solution to this challenge. It converts compound structures into vector representations and uses efficient vector indexing and similarity matching algorithms to achieve fast search. Summary of the invention
[0003] In view of this, the embodiments of the present invention provide a method for constructing a compound vector database, a method and device for compound similarity search, so as to eliminate or improve one or more defects existing in the prior art and solve the problem that the prior art lacks the ability to search for similar compounds.
[0004] One aspect of the present invention provides a method for constructing a compound vector database, the method comprising the following steps:
[0005] Acquire compound structure data based on multiple preset data sources, and perform data cleaning on the compound structure data; the compound structure data includes SMILES string format, JSON structure character format string and image format;
[0006] Using a preset chemical information processing tool to perform structural encoding on the cleaned compound structure data and convert it into characteristic data of the corresponding compound; the characteristic data includes one or more molecular descriptors for describing the basic properties and structural attributes of the molecule;
[0007] According to a preset rule, each molecular descriptor in the characteristic data is vectorized and combined to obtain a characteristic vector of the corresponding compound;
[0008] The characteristic vector of each compound is stored in a preset vector database, and an index is constructed to obtain a compound vector database.
[0009] In some embodiments, data cleaning of the compound structure data includes:
[0010] Delete missing value records or fill in missing values;
[0011] and / or, removing, scaling, or replacing outliers;
[0012] And / or, remove duplicate data and keep only the latest version of inconsistent data.
[0013] In some embodiments, in the method, the molecular descriptors used to describe the basic properties of the molecule are one-dimensional descriptors, including: molecular weight, molecular formula, atomic technology, bond count, and element ratio;
[0014] The molecular descriptors in the method also include one-dimensional descriptors for describing chemical properties, topological properties, electronic properties, statistical properties, and medicinal chemical properties;
[0015] Among them, the one-dimensional descriptors used to describe chemical properties include: polar surface area, total charge, hydrogen bond donor count, hydrogen bond acceptor count, and partition coefficient;
[0016] One-dimensional descriptors used to describe topological properties include: the number of rotatable bonds, the number of rings, the number of aromatic rings, and the molecular topological index;
[0017] One-dimensional descriptors used to describe electronic properties include: dipole moment, highest occupied molecular orbital energy, lowest unoccupied molecular orbital energy, and ionization energy;
[0018] One-dimensional descriptors used to describe statistical properties include: molecular complexity and Lipinski rule counts;
[0019] One-dimensional descriptors used to describe the chemical properties of drugs include: pharmacophore counts and toxicity risk indicators.
[0020] In some embodiments, the molecular descriptor used to describe the structural attributes in the method is a two-dimensional descriptor using Morgan fingerprint.
[0021] In some embodiments, each molecular descriptor in the feature data is vectorized according to a preset rule and combined to obtain a feature vector of the corresponding compound, including: using a preset chemical informatics tool to vectorize each molecular descriptor, the preset chemical informatics tool is RDKit, Open Babel or ChemoPy;
[0022] Alternatively, a pre-trained graph neural network model is used to perform vectorized feature extraction on each molecular descriptor.
[0023] In some embodiments, the feature vector of each compound is stored in a preset vector database and an index is constructed, including:
[0024] A relational database or a vector database is used to store the vectors of the compounds, wherein the relational database is an SQLite database, and the vector database is a FAISS or Milvus database; and an HNSW index is established for query.
[0025] On the other hand, the present invention also provides a chemical structure similarity search method based on a compound vector library, the method comprising:
[0026] For the target compound to be searched, a preset chemical information processing tool is used to perform structural encoding on the compound structure data, and convert it into target characteristic data of the corresponding compound; the target characteristic data includes one or more molecular descriptors for describing the basic properties and structural attributes of the molecule;
[0027] According to a preset rule, each molecular descriptor in the target feature data is vectorized and combined to obtain a target feature vector of the corresponding compound;
[0028] The compound vector database in the above-mentioned compound vector database construction method is searched based on the target feature vector, and the target result is found based on the approximate nearest neighbor search algorithm.
[0029] In some embodiments, the method includes: taking a predetermined number of items with the highest similarity obtained through searching as the target results.
[0030] On the other hand, the present invention further provides a computer-readable storage medium having a computer program / instruction stored thereon, which implements the steps of the above method when the computer program / instruction is executed by a processor.
[0031] On the other hand, the present invention also provides a computer program product, comprising a computer program / instruction, which implements the steps of the above method when executed by a processor.
[0032] The beneficial effects of the present invention are at least:
[0033] The compound vector database construction method, compound similarity search method and device of the present invention clean, structure encode and vectorize the compound structure data of multiple data sources, and form a uniform feature vector for each compound according to its basic molecular properties and structural attributes. The feature vector is stored and indexed through a vector database to obtain a compound vector database. In the application process, for the structural data of the compound to be retrieved, the structure is firstly encoded and vectorized, and then the compound vector database is searched and the most similar multiple compounds are output as the search results. With the help of the index, the compound structure in the large-scale compound database can be quickly and accurately retrieved, which has broad application prospects in the fields of drug discovery, material design, etc., and provides powerful tool support for scientific researchers.
[0034] Additional advantages, purposes, and features of the present invention will be described in part in the following description, and will become apparent to those skilled in the art after studying the following, or may be learned from the practice of the present invention. The purposes and other advantages of the present invention may be achieved and obtained by the structures specifically indicated in the specification and the accompanying drawings.
[0035] Those skilled in the art will appreciate that the objectives and advantages that can be achieved with the present invention are not limited to the above specific description, and the above and other objectives that can be achieved by the present invention will be more clearly understood from the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] The drawings described herein are used to provide a further understanding of the present invention, constitute a part of the present application, and do not constitute a limitation of the present invention. In the drawings:
[0037] Figure 1 The figure is a flow chart of a method for constructing a compound vector database according to an embodiment of the present invention.
[0038] Figure 2 The figure is a flow chart of a chemical structure similarity search method based on a compound vector library according to an embodiment of the present invention.
[0039] Figure 3 The present invention is a schematic diagram of the process of storing compound structure data in a method for searching compound structure similarity based on a vector database according to an embodiment of the present invention.
[0040] Figure 4 The figure is a schematic diagram of the data cleaning process in the compound structure similarity search method based on the vector database according to an embodiment of the present invention.
[0041] Figure 5 The figure is a schematic diagram of the compound structure search process in the compound structure similarity search method based on a vector database according to an embodiment of the present invention. DETAILED DESCRIPTION
[0042] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the embodiments and the accompanying drawings. Here, the illustrative embodiments of the present invention and their descriptions are used to explain the present invention, but are not intended to limit the present invention.
[0043] It should also be noted that, in order to avoid obscuring the present invention due to unnecessary details, only structures and / or processing steps closely related to the solutions according to the present invention are shown in the accompanying drawings, while other details that are not closely related to the present invention are omitted.
[0044] It should be emphasized that the term “include / comprises” when used herein refers to the presence of features, elements, steps or components, but does not exclude the presence or addition of one or more other features, elements, steps or components.
[0045] In the fields of drug discovery, material design and chemical research, similarity search of compound structures is an important means of innovation-driven, and is widely used in tasks such as the discovery of new compounds, the optimization of existing compound performance, and molecular library screening. However, traditional similarity search methods mainly rely on string matching based on text descriptions (such as SMILES or InChI) or substructure matching based on molecular graphs. Although these methods can accurately reflect the structural characteristics of compounds, they face significant computational bottlenecks when processing large-scale molecular databases. On the one hand, methods based on text descriptions usually rely on exact matching or simple character similarity metrics, which are difficult to capture the complex relationships between molecules. On the other hand, methods based on graph matching need to compare the topological structures of molecular graphs one by one, and the computational complexity is usually at the polynomial level, especially when facing a compound library of millions or even billions of scales, the computational efficiency is extremely low. In addition, these traditional methods are sensitive to noise and inconsistency in molecular representation, and show obvious deficiencies in scenarios such as high-throughput screening that require rapid processing. Therefore, with the rapid growth of the scale of molecular databases, compound similarity search is in urgent need of more efficient methods, and the present invention is based on the approximate nearest neighbor search technology of vectorized embedding. By converting molecular structures into high-dimensional vectors and using efficient indexing structures (such as HNSW and FAISS), candidate molecules similar to target compounds can be quickly located in large-scale databases, significantly improving search efficiency and applicability, and providing powerful tool support for modern chemical research and molecular design.
[0046] Specifically, one aspect of the present invention provides a method for constructing a compound vector database, such as Figure 1 As shown, the method includes the following steps S101 to S104:
[0047] Step S101: Acquire compound structure data based on multiple preset data sources and perform data cleaning on the compound structure data; the compound structure data includes SMILES string format, JSON structure character format string and image format.
[0048] Step S102: Use a preset chemical information processing tool to perform structural encoding on the cleaned compound structure data and convert it into characteristic data of the corresponding compound; the characteristic data includes one or more molecular descriptors for describing the basic properties and structural attributes of the molecule.
[0049] Step S103: vectorizing each molecular descriptor in the feature data according to a preset rule and combining them to obtain a feature vector of the corresponding compound.
[0050] Step S104: storing the characteristic vector of each compound into a preset vector database, and constructing an index to obtain a compound vector database.
[0051] In step S101, the sources for obtaining compound structure data are very extensive, including public databases, proprietary databases, and experimentally generated data sets, which can all be used as preset data sources. Among public databases, PubChem is one of the most well-known compound databases, which is maintained by the National Center for Biotechnology Information (NCBI) of the United States and contains the structure, properties, and biological activity information of millions of chemical substances. ChEMBL is another important data source, focusing on medicinal chemistry, providing drug targets, activity values, and other data related to drug design. The ZINC database focuses on virtual screening and contains a collection of purchasable compounds and their structural information. For crystal structures and chemical reactions, Cambridge Structural Database (CSD) and Reaxys provide high-quality data. In addition, DrugBank integrates information on drug molecules and their targets, and is a key resource in the field of drug discovery. In terms of proprietary data, many pharmaceutical companies and scientific research institutions maintain internal compound libraries for proprietary research. Experimentally generated data sets are usually obtained through high-throughput screening (HTS) or chemical synthesis experiments, supplemented by characterization techniques such as nuclear magnetic resonance (NMR) and mass spectrometry (MS) to record structural information. In addition, compound structure data can also be obtained from chemical papers and patents, and SMILES or InChI representations can be extracted through text mining technology. Combining these data sources can provide comprehensive chemical information support for drug development and material design.
[0052] There are various ways to record compound structure data. Common formats include SMILES strings, JSON structured strings, and image formats. SMILES (Simplified Molecular Input Line Entry System) is a simple linear symbolic representation that records the topological structure of a molecule through a string. For example, the SMILE representation of ethanol is CCO. It is intuitive and compact and is widely used in the field of cheminformatics. JSON structured strings are a hierarchical data format that usually stores detailed structural information of compounds, such as atoms, bond types, coordinates, etc., in the form of key-value pairs, which is suitable for integration with modern programming languages. For example, the JSON of ethanol can contain the symbols of atoms ["C","C","O"] and bond information [[0,1],[1,2]], which is convenient for program parsing and operation. Image formats display molecular structures in a visual way and are often used in chemistry textbooks, research papers, or user interaction interfaces. For example, ethanol can be presented through a two-dimensional chemical structure diagram showing carbon and oxygen atoms and their bond connections. Image formats are usually saved in the form of vector graphics (such as SVG) or raster graphics (such as PNG). Although they are intuitive and easy to read, they are not suitable for direct computational analysis. These three formats each have their own advantages. SMILES is simple and compact, suitable for storage and fast retrieval, JSON is suitable for data interaction and complex analysis, and images are easy for humans to observe and understand.
[0053] In some embodiments, step S101 performs data cleaning on the compound structure data, including: deleting missing value records or filling missing values; and / or, deleting, scaling or replacing abnormal values; and / or, deleting duplicate data and retaining only the latest version of inconsistent data.
[0054] In the processing of compound structure data, data cleaning is a key step to improve data quality and ensure model accuracy. First, for the processing of missing values, records containing missing values (such as missing SMILES strings, because they cannot parse molecular structures) can be deleted, or the mean, median or interpolation method can be used to fill in the missing numerical data (such as molecular weight) to ensure data integrity. Secondly, for outliers (such as extreme molecular weight or unreasonable chemical property values), statistical methods (such as 3 times the standard deviation or IQR method) can be used to detect anomalies and then delete them, or the impact of abnormal data can be adjusted by scaling (such as normalization) and replacement (such as replacement with the median). Cleaning up duplicate data is another important step. Completely duplicated data can be deleted by checking duplicate SMILES representations, or the latest version can be retained based on metadata such as timestamps to ensure that the data is not redundant and updated. Finally, in response to the format inconsistency problem caused by the diversity of data sources, the SMILES representation can be normalized (for example, using the RDKit tool) to ensure that all records are represented in a unified format. In addition, the logical consistency of the compound data needs to be checked, such as whether the key-value relationship is correct, to avoid subsequent analysis deviations due to incorrect structures. By deleting and filling missing values, processing outliers, deduplication and normalization, data cleaning can reduce noise while providing reliable data support for compound similarity search, model training and chemical property analysis. It is an indispensable basic link in chemical informatics research.
[0055] In step S102, the preset chemical information processing tool may be RDKit. In addition to RDKit, there are other open source tools that can also implement these functions for analyzing and extracting compound structures, such as Open Babel: OpenBabel is an open source chemical information toolbox that supports multiple programming languages, including C++, Python, and Perl. CDK (Chemistry Development Kit): CDK is an open source Java library for chemical informatics and drug design. ChemoPy: ChemoPy is an open source Python library for calculating chemical and physical properties and performing molecular fingerprint analysis.
[0056] Molecular descriptors are a set of numerical features extracted from molecular structures, which can be used to represent various chemical and physical properties of molecules.
[0057] In some embodiments, in the method, the molecular descriptors used to describe the basic properties of molecules are one-dimensional descriptors, including: molecular weight, molecular formula, atomic technology, bond count and element ratio.
[0058] The molecular descriptors in the method also include one-dimensional descriptors for describing chemical properties, topological properties, electronic properties, statistical properties and medicinal chemistry properties.
[0059] Among them, the one-dimensional descriptors used to describe chemical properties include: polar surface area, total charge, hydrogen bond donor count, hydrogen bond acceptor count and partition coefficient.
[0060] One-dimensional descriptors used to describe topological properties include: the number of rotatable bonds, the number of rings, the number of aromatic rings, and the molecular topological index.
[0061] One-dimensional descriptors used to describe electronic properties include: dipole moment, highest occupied molecular orbital energy, lowest unoccupied molecular orbital energy, and ionization energy.
[0062] One-dimensional descriptors used to describe statistical properties include: molecular complexity and Lipinski rule count.
[0063] One-dimensional descriptors used to describe the chemical properties of drugs include: pharmacophore counts and toxicity risk indicators.
[0064] In some embodiments, the molecular descriptor used to describe the structural attributes in the method is a two-dimensional descriptor using Morgan fingerprint.
[0065] Morgan fingerprint is a molecular fingerprint method used in chemical informatics. It is based on molecular graph theory and represents the structural information of molecules by encoding the environment of each atom in the molecule. Morgan fingerprint is determined by radius and length, where: the radius determines the range of structural information contained in the Morgan fingerprint. A larger radius can capture the interactions between atoms farther away, but it will also increase the calculation time and complexity of the fingerprint. In the present invention, the radius is 3. The length refers to the total number of bits in the fingerprint. Longer fingerprints can provide higher discrimination, but they will also increase storage requirements and computing costs. In the present invention, the length is 4096 bits.
[0066] Furthermore, ECFP may be used as a two-dimensional descriptor. ECFP is a special form of Morgan fingerprint, and generally refers to a Morgan fingerprint with a specific number of iterations.
[0067] In other embodiments, in order to record more complex molecular structures of compounds, three-dimensional descriptors may be used to describe the three-dimensional structure of the molecule, such as molecular volume, surface area, various geometric parameters, and the like.
[0068] In step S103, each molecular descriptor in the feature data is vectorized and combined according to a preset rule to obtain a feature vector of the corresponding compound, including: using a preset chemical informatics tool to vectorize each molecular descriptor, the preset chemical informatics tool is RDKit, Open Babel or ChemoPy. Or, using a pre-trained graph neural network model to extract features of each molecular descriptor for vectorization.
[0069] A pre-trained graph neural network (GNN) model is used to perform vectorized feature extraction on molecular descriptors. The implementation process includes the following steps S201 to S207:
[0070] Step S201: Data preparation requires collecting or creating molecular datasets and ensuring that they are represented in an appropriate format. Molecules can usually be represented in the form of graphs, where atoms are nodes and chemical bonds are edges. These graphs can be two-dimensional planar structures or three-dimensional structures. Each atom and bond may also have attributes, such as atom type, charge, hybridization state, etc.
[0071] Step S202: Select a pre-trained model, and select a pre-trained GNN model suitable for molecular graph data. Common GNN models include graph convolutional networks (GCN), graph attention networks (GAT), message passing neural networks (MPNN), etc. You can choose a pre-trained model provided by the open source community.
[0072] Step S203: Model adaptation: If the selected pre-trained model does not fully match your dataset, you may need to make some adjustments to the model. This may include modifying the input layer to adapt to the structure of your molecular data, or adjusting the model parameters to suit a specific task.
[0073] Step S204: Feature extraction: Use the pre-trained model to extract features from the molecule. This usually means inputting the molecular graph into the model and obtaining node-level feature representations through the first few layers of the model. These features can be used for downstream tasks such as molecular property prediction, activity prediction in drug discovery, etc.
[0074] Step S205: Post-processing: Post-process the extracted features as needed. For example, node features may be aggregated into a feature vector of the entire molecule, or feature dimensions may be reduced through a pooling operation.
[0075] Step S206: Evaluation and optimization. In practical applications, the effect of feature extraction needs to be evaluated. This can be done through cross-validation, comparison with other baseline methods, etc. Based on the evaluation results, the model may need to be adjusted and optimized.
[0076] Step S207: Deployment. Once the performance of the model reaches a satisfactory level, it can be deployed in a production environment for actual molecular descriptor vectorization feature extraction tasks.
[0077] In step S104, before performing compound vector storage, it is first necessary to select a suitable vector database. According to the needs and scenarios, open source solutions (such as FAISS, Milvus) or cloud services (such as Pinecone, Weaviate) can be selected. Taking Milvus as an example, after installing the vector database, the database is connected through the Python client (pymilvus). Next, the structure of the vector table is designed, including the primary key field (such as compound ID) and the field for storing feature vectors, and the dimension and data type of the vector (such as a 2048-dimensional floating-point vector) are specified. After the table structure design is completed, the table is created and initialized to ensure that the database can receive subsequent vector data operations. After the vector table structure is created, the compound ID and the corresponding feature vector need to be inserted into the database in batches. Usually, the compound ID and feature vector are stored in the form of a list or array, and the insertion interface of the database is used for batch submission. For example, in Milvus, the collection.insert() method can insert a batch of IDs and vectors into the table in a grouped form. After the insertion is completed, the success of the data storage can be checked through the API, such as calling collection.num_entities to verify whether the total number of vectors is consistent with the input. To improve efficiency, it is recommended to perform deduplication and cleaning before data insertion to ensure the uniqueness and normalization of the vector. After storage is completed, in order to speed up the retrieval of large-scale vectors, it is necessary to create an index for the feature vector field. Commonly used vector indexing methods include HNSW (graph-based approximate nearest neighbor search) and IVF_FLAT (inverted file combined with exact matching). According to the retrieval requirements and database scale, select appropriate algorithms and parameters. For example, in HNSW, you can adjust the graph connectivity and construction efficiency. After building the index, you can verify its successful creation by querying the status of the index and check the correctness of the configuration. The existence of the index significantly improves the efficiency of vector retrieval while ensuring the accuracy and scalability of the results.
[0078] In some embodiments, the characteristic vector of each compound is stored in a preset vector database, and an index is constructed, including: using a relational database or a vector database to store the vector of the compound, the relational database is a SQLite database, and the vector database is a FAISS or Milvus database; and establishing an HNSW index for query.
[0079] On the other hand, the present invention also provides a chemical structure similarity search method based on a compound vector library, such as Figure 2 As shown, the method includes steps S301 to S303:
[0080] Step S301: using a preset chemical information processing tool to perform structural encoding on the compound structure data of the target compound to be retrieved, and converting it into target feature data of the corresponding compound; the target feature data includes one or more molecular descriptors for describing the basic properties and structural attributes of the molecule.
[0081] Step S302: vectorizing each molecular descriptor in the target feature data according to a preset rule and combining them to obtain a target feature vector of the corresponding compound.
[0082] Step S303: searching the compound vector database in the compound vector database construction method described in steps S101 to S104 based on the target feature vector, and searching for the target result based on an approximate nearest neighbor search algorithm.
[0083] The process of structural encoding and vectorization of the target compound in steps S301 and S302 can refer to the above steps S101 to S103. In step S303, the approximate nearest neighbor search algorithm is a search algorithm based on L2 distance, cosine similarity, Tanimoto coefficient and other metrics.
[0084] In some embodiments, the method includes: taking a first set number of items with the highest similarity obtained through searching as target results.
[0085] On the other hand, the present invention further provides a computer-readable storage medium having a computer program / instruction stored thereon, which implements the steps of the above method when the computer program / instruction is executed by a processor.
[0086] On the other hand, the present invention also provides a computer program product, comprising a computer program / instruction, which implements the steps of the above method when executed by a processor.
[0087] The present invention is described below in conjunction with a specific embodiment:
[0088] This embodiment provides a method for searching compound structure similarity based on a vector database, such as Figure 3 As shown, the following steps are included:
[0089] 1. System construction
[0090] Install the relevant dependent environment on the server, then deploy the Faiss vector database system on the server and configure the database parameters.
[0091] 2. Data Preprocessing
[0092] Reference Figure 4, import the compound structure data source file, preprocess the file, and perform noise removal operations. The specific operations are as follows:
[0093] a. Handling missing values: There may be missing values in compound structure data, for example, some attribute information of some compounds may be missing. Methods for handling missing values include deleting missing value records, filling missing values (such as using mean, median, mode, or interpolation) and filling missing values based on prediction models;
[0094] b. Outlier processing: Outliers are observations that deviate significantly from most data points. For compound structure data, outliers may indicate data recording errors or abnormal molecular structures. Processing methods include deleting outliers, scaling or replacing outliers. The interquartile range (IQR) is usually used for outlier detection.
[0095] c. Remove duplicate records: Remove duplicate data records in compound structure data.
[0096] 3. Data formatting
[0097] a. File format conversion: Unify the format of compound structure data and convert it into a format suitable for subsequent analysis or calculation according to computer processing requirements.
[0098] b. Chemical diagram formatting: For compound structure data stored in text form, a chemical diagram formatting step may be required, that is, the text elements are formatted and split to obtain the chemical diagram information of the compound.
[0099] 4. Structural Coding
[0100] a. Select encoding method: Select a suitable structural encoding method according to the analysis or calculation requirements. Common encoding methods include molecular fingerprints (such as ECFP, Morgan fingerprint, etc.), molecular descriptors (such as topological descriptors, geometric descriptors, electronic descriptors, etc.), etc. The present invention selects Morgan fingerprints combined with one-dimensional molecular descriptor vectors.
[0101] b. Generate code: Use the selected code to convert the compound's structural information into vector form. Calculate the feature vectors of the one-dimensional molecular descriptor and the two-dimensional molecular descriptor respectively.
[0102] 5. Vectorization
[0103] The features of the two molecular descriptors of the compound are calculated separately, and then the two types of feature vectors are concatenated to form the vector of the compound.
[0104] 6. Vectors are stored in the vector database
[0105] The structural information vectors and other metadata information data of the compounds obtained from the data source are imported into the vector database.
[0106] 7. Build index
[0107] Build appropriate indexes for the vector database according to application requirements. Indexes can speed up similarity search operations and improve query efficiency. Common index structures include KD trees, ball trees, and LSH (local sensitive hashing). The present invention uses HNSW indexes.
[0108] HNSW: A graph-based indexing algorithm that builds a multi-layer navigation structure for images according to certain rules. In this structure, the upper layers are sparser and the distance between nodes is farther; the lower layers are denser and the distance between nodes is closer. The search starts from the top layer, finds the node closest to the target in that layer, and then enters the next layer to start another search. After multiple iterations, it can quickly approach the target position
[0109] 8. Query
[0110] Reference Figure 5 , receiving the compounds to be queried submitted by the user, converting the compounds into vectors, and then submitting the vectors to the vector database, and then using the query function provided by the vector database to perform similarity search or other types of queries. Get the final query result, which is the vector in the TopN hits.
[0111] 9. Interface development
[0112] Design and implement the user interface, including query condition input, query result display, filtering and sorting functions.
[0113] 10. Deployment and Application
[0114] The system is deployed in actual application environments to provide researchers with compound structure similarity search services.
[0115] The efficient compound structure similarity search system and method based on vector database proposed in the present invention realizes the rapid and accurate retrieval of compound structures in large-scale compound databases through advanced vectorization technology, efficient vector database and similarity search algorithm. The system has broad application prospects in the fields of drug discovery, material design, etc., and provides powerful tool support for scientific researchers.
[0116] Corresponding to the above method, the present invention also provides an apparatus / system, which includes a computer device, the computer device includes a processor and a memory, the memory stores computer instructions, the processor is used to execute the computer instructions stored in the memory, and when the computer instructions are executed by the processor, the apparatus / system implements the steps of the method described above.
[0117] The embodiment of the present invention also provides a computer-readable storage medium on which a computer program is stored, and when the computer program is executed by a processor, the steps of the aforementioned edge computing server deployment method are implemented. The computer-readable storage medium can be a tangible storage medium, such as a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a floppy disk, a hard disk, a removable storage disk, a CD-ROM, or any other form of storage medium known in the technical field.
[0118] The embodiment of the present invention further provides a computer device, which may include a processor, a memory and an image acquisition device, wherein the processor and the memory may be connected via a bus or other means, and this embodiment takes the bus connection as an example. The image acquisition device may be connected to the processor and the memory via a wired or wireless manner.
[0119] The processor may be a central processing unit (CPU). The processor may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, or a combination of the above chips.
[0120] The memory is a non-transitory computer-readable storage medium that can be used to store non-transitory software programs, non-transitory computer executable programs and modules. The processor executes various functional applications and data processing of the processor by running the non-transitory software programs, instructions and modules stored in the memory, that is, the image color correction method in the above method embodiment is implemented.
[0121] The memory may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required by at least one function; the data storage area may store data created by the processor, etc. In addition, the memory may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some embodiments, the storage may optionally include a memory remotely disposed relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0122] The one or more modules are stored in the memory and when executed by the processor, perform the following steps: Figure 1-Figure 5 Method in the embodiment shown.
[0123] As an implementation method, the functions of the receiver and the transmitter in the present invention may be implemented by a transceiver circuit or a dedicated chip for transceiver, and the processor may be implemented by a dedicated processing chip, a processing circuit or a general-purpose chip.
[0124] As another implementation, it is possible to use a general-purpose computer to implement the authentication device and authentication server provided in the embodiment of the present invention, that is, the program code for implementing the functions of the processor, receiver, and transmitter is stored in a memory, and the general-purpose processor implements the functions of the processor, receiver, and transmitter by executing the code in the memory.
[0125] In summary, the compound vector database construction method, compound similarity search method and device of the present invention clean, structure encode and vectorize the compound structure data of multiple data sources, and form a unified feature vector for each compound according to its basic molecular properties and structural attributes. The feature vector is stored and indexed through a vector database to obtain a compound vector database. In the application process, for the structural data of the compound to be retrieved, the structure is first encoded and vectorized, and then the compound vector database is searched and the most similar multiple compounds are output as the search results. With the help of the index, the compound structure in the large-scale compound database can be quickly and accurately retrieved, which has broad application prospects in the fields of drug discovery, material design, etc., and provides powerful tool support for scientific researchers.
[0126] It should be understood by those skilled in the art that the exemplary components, systems and methods described in conjunction with the embodiments disclosed herein can be implemented in hardware, software or a combination of the two. Whether it is performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention. When implemented in hardware, it can be, for example, an electronic circuit, an application specific integrated circuit (ASIC), appropriate firmware, a plug-in, a function card, etc. When implemented in software, the elements of the present invention are programs or code segments used to perform the required tasks. The program or code segment can be stored in a machine-readable medium, or transmitted on a transmission medium or a communication link via a data signal carried in a carrier.
[0127] It should be clear that the present invention is not limited to the specific configuration and processing described above and shown in the figures. For the sake of simplicity, a detailed description of the known method is omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present invention is not limited to the specific steps described and shown, and those skilled in the art can make various changes, modifications and additions, or change the order between the steps after understanding the spirit of the present invention.
[0128] In the present invention, features described and / or illustrated for one embodiment may be used in the same or similar manner in one or more other embodiments, and / or combined with features of other embodiments or replace features of other embodiments.
[0129] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the embodiments of the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for constructing a compound vector database, characterized in that: The method comprises the following steps: Acquire compound structure data based on multiple preset data sources, and perform data cleaning on the compound structure data; the compound structure data includes SMILES string format, JSON structure character format string and image format; Using a preset chemical information processing tool to perform structural encoding on the cleaned compound structure data and convert it into characteristic data of the corresponding compound; the characteristic data includes one or more molecular descriptors for describing the basic properties and structural attributes of the molecule; According to a preset rule, each molecular descriptor in the characteristic data is vectorized and combined to obtain a characteristic vector of the corresponding compound; The characteristic vector of each compound is stored in a preset vector database, and an index is constructed to obtain a compound vector database.
2. The method for constructing a compound vector database according to claim 1, characterized in that: The compound structure data is cleaned, including: Delete missing value records or fill in missing values; and / or, removing, scaling, or replacing outliers; And / or, remove duplicate data and keep only the latest version of inconsistent data.
3. The method for constructing a compound vector database according to claim 1, characterized in that: In the method, the molecular descriptors used to describe the basic properties of the molecule are one-dimensional descriptors, including: molecular weight, molecular formula, atomic technology, bond count and element ratio; The molecular descriptors in the method also include one-dimensional descriptors for describing chemical properties, topological properties, electronic properties, statistical properties, and medicinal chemical properties; Among them, the one-dimensional descriptors used to describe chemical properties include: polar surface area, total charge, hydrogen bond donor count, hydrogen bond acceptor count, and partition coefficient; One-dimensional descriptors used to describe topological properties include: the number of rotatable bonds, the number of rings, the number of aromatic rings, and the molecular topological index; One-dimensional descriptors used to describe electronic properties include: dipole moment, highest occupied molecular orbital energy, lowest unoccupied molecular orbital energy, and ionization energy; One-dimensional descriptors used to describe statistical properties include: molecular complexity and Lipinski rule counts; One-dimensional descriptors used to describe the chemical properties of drugs include: pharmacophore counts and toxicity risk indicators.
4. The method for constructing a compound vector database according to claim 1, characterized in that: The molecular descriptor used to describe the structural attributes in the method is a two-dimensional descriptor using Morgan fingerprint.
5. The method for constructing a compound vector database according to claim 1, characterized in that: Vectorizing each molecular descriptor in the feature data according to a preset rule and combining them to obtain a feature vector of the corresponding compound, including: vectorizing each molecular descriptor using a preset chemical informatics tool, wherein the preset chemical informatics tool is RDKit, OpenBabel or ChemoPy; Alternatively, a pre-trained graph neural network model is used to perform vectorized feature extraction on each molecular descriptor.
6. The method for constructing a compound vector database according to claim 1, characterized in that: The characteristic vector of each compound is stored in a preset vector database and an index is constructed, including: A relational database or a vector database is used to store the vectors of the compounds, wherein the relational database is an SQLite database, and the vector database is a FAISS or Milvus database; and an HNSW index is established for query.
7. A chemical structure similarity search method based on a compound vector library, characterized in that: The method comprises: For the target compound to be searched, a preset chemical information processing tool is used to perform structural encoding on the compound structure data, and convert it into target characteristic data of the corresponding compound; the target characteristic data includes one or more molecular descriptors for describing the basic properties and structural attributes of the molecule; According to a preset rule, each molecular descriptor in the target feature data is vectorized and combined to obtain a target feature vector of the corresponding compound; Based on the target feature vector, the compound vector database in the compound vector database construction method according to any one of claims 1 to 6 is searched, and the target result is found based on an approximate nearest neighbor search algorithm.
8. The chemical structure similarity search method based on the compound vector library according to claim 7, characterized in that: The method comprises: taking a preset number of items with the highest similarity obtained through searching as the target results.
9. A computer-readable storage medium having a computer program / instruction stored thereon, characterized in that: When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.
10. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
Chemical industry-based search method
CN110569420A
Method for predicting retention time of liquid chromatogram under different chromatographic conditions
CN118243842A
Pharmaceutical industry drug discovery and function prediction method and system based on large model and vector database
CN118800368A
Chemical functional homologue search using inconsistent canonicalized query
WO2024238699A2
Cited By
Drug intermediate database construction and AI intelligent retrieval method
CN120199374A