Drug intermediate database construction and AI intelligent retrieval method
By constructing a structured intermediate database and using AI intelligent search methods, the dilemma of drug intermediate data management and retrieval is solved, and efficient and accurate drug research and development support is achieved.
Patent Information
- Application Number
- CN202510679543.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-26
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-05-26
AI Technical Summary
The data management and retrieval of drug intermediate data are scattered, the data format is incompatible, the lack of efficient means of extracting key information, and the database structure does not meet the needs of fast and accurate retrieval, resulting in low R&D efficiency and frequent errors.
By obtaining chemical structure data and synthetic path data, using feature extraction models to identify and extract molecular descriptors, a structured intermediate database is constructed, and an AI search model is used to perform intelligent search and multi-dimensional sorting based on graph neural network algorithm.
It realizes the unified acquisition and processing of drug intermediate data from multiple data sources, improves the storage and retrieval efficiency of data, reduces artificial errors, and provides more accurate and efficient drug development support.
Smart Images

Figure CN120199374A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of drug research and development, specifically to the construction of a drug intermediate database and an AI intelligent retrieval method. Background Art
[0002] In the field of drug research and development, drug intermediates, as the key raw materials for synthesizing the final drugs, the management and utilization of their related data are crucial for the R & D process. However, the current management and retrieval of drug intermediate data face many difficulties.
[0003] From the perspective of data acquisition, the sources of chemical structure data and synthesis path data are scattered. Chemical structure data are widely distributed in various chemical databases, such as SciFinder, Reaxys, etc. The data formats and storage standards of different databases vary greatly; the data generated by experimental recording devices are also difficult to collect uniformly due to different device models and manufacturers. For example, the chemical structure data format recorded by a nuclear magnetic resonance instrument used in a certain laboratory is incompatible with that of another laboratory, which makes it extremely difficult for researchers to integrate multi-source data, and they spend a lot of time and effort on data format conversion and cleaning, seriously affecting the R & D efficiency.
[0004] In terms of data processing, traditional methods are difficult to effectively extract and utilize key information. The chemical structures of drug intermediates are complex, containing important features such as atomic connectivity, electron cloud distribution, and functional group topological relationships, but the existing technologies lack means to efficiently and accurately extract these molecular descriptors. Researchers often rely on manual analysis of chemical structure data, which is not only inefficient but also prone to errors due to human factors, resulting in misjudgments of the properties and reaction activities of drug intermediates.
[0005] Regarding database construction, most of the existing databases lack structured and multi-dimensional indexing. The data storage is scattered and the correlation is poor, unable to meet the requirements of rapid and accurate retrieval. When R & D personnel need to search for drug intermediates with specific structures and meeting certain synthesis conditions, it is difficult to quickly locate relevant information in the existing databases, hindering the R & D process.
[0006] The retrieval technology also has obvious deficiencies. The existing retrieval methods are mostly based on simple text matching or structure similarity search, unable to fully consider various factors in drug research and development, such as reaction conditions, yield, synthesis difficulty, and cost. In actual R & D, R & D personnel hope to find drug intermediates that can not only meet the structural requirements but also have suitable synthesis conditions and lower costs, but traditional retrieval methods cannot effectively provide such comprehensive information.
[0007] As competition in drug research and development becomes increasingly fierce, the requirements for R&D efficiency and innovation capabilities continue to increase. There is an urgent need for a method that can integrate multi-source data, efficiently extract key information, build a structured database and realize intelligent retrieval to solve the difficulties in existing technologies, accelerate the drug research and development process, and reduce R&D costs. Summary of the invention
[0008] The purpose of the present invention is to provide a pharmaceutical intermediate database construction and an AI intelligent retrieval method to solve the problems raised in the above background technology.
[0009] To achieve the above object, the present invention provides the following technical solution: a pharmaceutical intermediate database construction and AI intelligent retrieval method, the method comprising: Obtaining chemical structure data and synthesis route data of a target drug intermediate, wherein the chemical structure data is derived from at least one chemical database or experimental recording device, and the synthesis route data includes reaction conditions, catalysts, and yield parameters; Inputting the chemical structure data into a preset feature extraction model to identify and extract molecular descriptors in the chemical structure data, wherein the molecular descriptors include atomic connectivity, electron cloud distribution, and functional group topological relationships; Constructing a structured intermediate database based on the molecular descriptors, wherein the database associates chemical structures, synthesis pathways, and multi-dimensional index tags of physicochemical properties; Inputting the structured intermediate database into a preset AI retrieval model to generate dynamic retrieval results, wherein the AI retrieval model matches the similarity between user queries and database records based on a graph neural network algorithm; The dynamic search results are sorted in multiple dimensions using a preset optimization algorithm, and a standardized search list is output.
[0010] Preferably, the step of training the feature extraction model includes: obtaining an initial training sample set, wherein each sample in the initial training sample set is labeled with a corresponding molecular descriptor label; iteratively training an initial feature extraction network using the initial training sample set until the matching rate between the feature prediction result output by the initial feature extraction network and the molecular descriptor label is greater than or equal to a preset threshold; during each iteration, adjusting the weight parameters of the initial feature extraction network according to the prediction error of the previous iteration, and eliminating samples in the initial training sample set with errors higher than the set threshold, to form a reduced training subset for the next iterative training.
[0011] Preferably, obtaining the chemical structure data of the target drug intermediate includes obtaining chemical data from one of the data sources in the following manner: Connect to the target chemical database server and / or experimental equipment real-time data interface; Extract raw chemical data from at least one of the target chemical database server and the real-time data interface of the experimental equipment according to a preset data acquisition protocol, where the preset data acquisition protocol includes data format, storage precision, and metadata association rules.
[0012] Preferably, before inputting the chemical structure data into a preset feature extraction model, the method further includes: Filter redundant information from the raw chemical data to eliminate format differences between different data sources; Use a multi-scale noise reduction algorithm to remove outliers from the raw chemical data and retain key information of the molecular structure.
[0013] Preferably, constructing a structured intermediate database based on the molecular descriptors includes: Segment the chemical structure map according to the atomic connectivity in the molecular descriptors; Generate hierarchical storage units based on the segmented map, and associate the corresponding synthesis paths and physicochemical property parameters within each unit.
[0014] Preferably, inputting the structured intermediate database into a preset AI retrieval model includes: Set query constraint conditions in the AI retrieval model to simulate the retrieval requirements of users according to structural similarity, reaction conditions, or yield range; Match the multi-dimensional similarity scores between database records and query conditions by iteratively calculating the graph node embedding vectors.
[0015] Preferably, using a preset optimization algorithm to perform multi-dimensional sorting on the dynamic retrieval results includes: Assign weights to the dynamic retrieval results to generate a derivative sorted list according to structural similarity, synthesis difficulty, or cost priority; Inject optimization parameters that conform to chemical logic into the derivative sorted list based on an adaptive adjustment algorithm.
[0016] Preferably, the method further includes training the optimization algorithm in the following manner: Construct an optimization training set containing real chemical synthesis cases, where each case is labeled with the corresponding synthesis path truth value; Train the optimization algorithm using a reinforcement learning framework so that the distribution difference between the sorted list generated by the optimization algorithm and the real chemical synthesis cases is less than a set tolerance range.
[0017] Preferably, the method further includes: Input the standardized retrieval list into a preset verification model and calculate the error matrix between the retrieval results and the experimental record data; Adjust the graph network depth and feature embedding dimension of the AI retrieval model in reverse according to the error matrix; The verification model is constructed through the following steps: Collect the synthetic route data recorded in the laboratory as the verification benchmark; Establish an error evaluation network based on the attention mechanism, and output an error weight vector by comparing the multi-dimensional feature maps of the retrieval result and the verification benchmark.
[0018] Preferably, the method further includes a data security control step: When receiving an access request from an external system to the structured intermediate database, verify the matching of the identity token provided by the requestor with the preset permission list; When the matching is successful, transmit a data subset that conforms to the access permission to the external system through an encrypted link.
[0019] Compared with the prior art, the beneficial effects of the present invention are: The method for constructing a pharmaceutical intermediate database and AI intelligent retrieval proposed by the present invention brings many significant beneficial effects to the field of pharmaceutical research and development. In terms of data acquisition and processing, this method can obtain chemical structure data and synthetic route data from various chemical databases and experimental recording devices. Through the preset data acquisition protocol, the data formats, storage precisions, and metadata association rules of different data sources can be unified, greatly improving the efficiency and accuracy of data acquisition. At the same time, before the data is input into the feature extraction model, redundant information filtering and multi-scale noise reduction processing will be performed to effectively eliminate format differences and outliers, retain key information, provide a high-quality data basis for subsequent analysis, and avoid research deviations caused by data quality problems.
[0020] The feature extraction model is carefully trained to accurately identify and extract molecular descriptors in chemical structure data, such as atomic connectivity, electron cloud distribution, and functional group topological relationships. This process not only improves the extraction efficiency of key information, but also provides an important basis for subsequent database construction and intelligent retrieval. Through iterative training and sample screening of a large number of samples, the accuracy and stability of the model are continuously improved. Compared with traditional manual analysis, the human error is greatly reduced, enabling researchers to understand the chemical properties of pharmaceutical intermediates more deeply and accurately.
[0021] The construction of a structured intermediate database is a highlight. The chemical structure map is segmented based on molecular descriptors, and hierarchical storage units are generated, while multi-dimensional index labels of chemical structure, synthesis path and physicochemical properties are associated. This structured design significantly improves the efficiency of data storage and retrieval. R&D personnel can quickly locate the required data from multiple dimensions, such as screening drug intermediates based on specific atomic connection methods, reaction conditions or physicochemical properties, which greatly saves the time of searching for data and improves the overall efficiency of R&D work.
[0022] The AI retrieval model is based on a graph neural network algorithm, which can more accurately match the similarity between user queries and database records and generate dynamic retrieval results. Compared with traditional retrieval methods, this model can comprehensively consider multiple factors, such as structural similarity, reaction conditions, yield, etc., to provide users with retrieval results that are more in line with actual needs. In addition, the preset optimization algorithm is used to sort the retrieval results in multiple dimensions and output a standardized retrieval list. Through weight allocation and adaptive adjustment algorithms, R&D personnel can obtain retrieval results of corresponding priorities according to their own needs, such as paying more attention to structural similarity, synthesis difficulty or cost, and further improve the practicality and pertinence of the retrieval.
[0023] In addition, the present invention also sets a data security control step to verify the matching of the requester's identity token with the preset permission list, and transmits data through an encrypted link. This effectively ensures the security of sensitive data in the database, prevents data leakage, meets the strict requirements of the drug research and development industry for data security, and provides reliable protection for the data assets of enterprises and research institutions. In general, the invention comprehensively improves the management and utilization level of drug intermediate data, strongly promotes the progress of drug research and development, reduces research and development costs, and has significant economic and social benefits. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 This is a working principle diagram of the pharmaceutical intermediate database construction and AI intelligent retrieval method of the present invention; Figure 2 Flowchart for feature extraction model training; Figure 3 Flowchart for chemical structure data preprocessing; Figure 4 Flowchart for multi-dimensional sorting of search results. DETAILED DESCRIPTION
[0025] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0026] Please refer to Figures 1-4 , the present invention provides a technical solution: The present invention relates to a method for constructing a drug intermediate database and AI intelligent retrieval, and the specific implementation steps are as follows: Obtain data: Obtain the chemical structure data and synthesis path data of the target drug intermediate. The chemical structure data has a wide range of sources and is taken from at least one chemical database, such as the commonly used ChemSpider, PubChem, etc. These databases store a large amount of chemical structure information of known compounds; it can also be sourced from experimental recording devices. For example, when conducting drug intermediate synthesis experiments in the laboratory, the experimental data recorded by high-resolution spectrometers, chromatographs, and other devices. The synthesis path data includes reaction conditions, such as reaction temperature, reaction time, reaction pressure, etc.; catalysts, that is, substances that participate in chemical reactions but whose chemical properties remain unchanged before and after the reaction; and yield parameters, which are used to measure the efficiency of the synthesis reaction to produce the target drug intermediate.
[0027] Extract molecular descriptors: Input the obtained chemical structure data into a preset feature extraction model. This model can identify and extract the molecular descriptors in the chemical structure data. The molecular descriptors cover atomic connectivity, which is used to describe the connection mode and topological structure between each atom in the molecule; electron cloud distribution, which reflects the distribution state of electrons in the molecule and has an important impact on the chemical activity and reaction performance of the molecule; and functional group topological relationship, which clarifies the spatial position and interaction relationship between different functional groups in the molecule.
[0028] Construct a structured intermediate database: Construct a structured intermediate database based on the extracted molecular descriptors. This database associates multi-dimensional index tags of chemical structures, synthesis paths, and physical and chemical properties, enabling the data in the database to be quickly retrieved and analyzed from multiple perspectives, facilitating researchers to obtain the required information.
[0029] Generate dynamic retrieval results: Input the constructed structured intermediate database into a preset AI retrieval model. This AI retrieval model is based on the graph neural network algorithm, which matches the similarity between the user's query and the database records through this algorithm, thereby generating dynamic retrieval results and providing the user with information on drug intermediates related to the query.
[0030] Multi-dimensional sorting and output: Use a preset optimization algorithm to perform multi-dimensional sorting on the dynamic retrieval results. Considering various factors such as structural similarity, synthesis difficulty, cost, etc., the sorted results are output in the form of a standardized retrieval list for the convenience of users to view and use.
[0031] The present invention will be further described below in conjunction with Embodiments 1 to 5:
[0032] Embodiment 1: In this embodiment, the specific process of training the feature extraction model is mainly described.
[0033] Obtain an initial training sample set, and each sample in this sample set is labeled with a corresponding molecular descriptor label. These samples can be collected from public chemical data sets or a large amount of experimental data. For example, select data containing the chemical structure of drug intermediates and corresponding molecular descriptors from the compound data sets released by some professional chemical research institutions as the initial samples.
[0034] Use the initial training sample set to iteratively train the initial feature extraction network. During the training process, continuously adjust the weight parameters of the network to make the feature prediction results output by the network match the molecular descriptor labels as much as possible. Here, the concept of matching rate is introduced. After each training, calculate the matching rate between the prediction result and the label. Let the prediction result be , and the molecular descriptor label be , the calculation formula of the matching rate is:
[0035] where is the number of samples, is the judgment function. When matches , , otherwise it is . When the matching rate is greater than or equal to the preset threshold, the training ends.
[0036] In each iteration process, adjust the weight parameters of the initial feature extraction network according to the prediction error of the previous iteration. The prediction error can be calculated by the mean square error (MSE). Let the predicted value be , and the true value be , then the mean square error
[0037] where is the number of samples in the current batch of training. At the same time, eliminate the samples in the initial training sample set with errors higher than the set threshold, and form a reduced training subset for the next iteration training. In this way, gradually optimize the network and improve the accuracy of feature extraction.
[0038] Example 2: This example details the specific method for obtaining the chemical structure data of the target drug intermediate.
[0039] When obtaining chemical structure data, it is possible to connect to the target chemical database server and / or the real-time data interface of experimental equipment. Taking the connection to the commonly used chemical database ChemSpider as an example, through the API interface it provides, connection code is written using a specific programming language (such as Python) to achieve communication with the database server. For the real-time data interface of experimental equipment, assume that a certain model of nuclear magnetic resonance spectrometer is used in the laboratory, and this instrument is equipped with a dedicated data transmission interface. Through the corresponding data transmission line and driver, the data collected by the instrument is transmitted to the computer in real time.
[0040] Extract the original chemical data from at least one of the target chemical database server and the real-time data interface of experimental equipment according to the preset data acquisition protocol. The preset data acquisition protocol includes data format, storage precision, and metadata association rules. In terms of data format, for example, the data extracted from the database may be in JSON format, and the data obtained from the experimental equipment may be in CSV format. The storage precision is set according to actual needs. For example, for chemical shift data, the storage precision can be set to three decimal places. The metadata association rules are used to clarify the relationship between the data and other relevant information, such as the association between the name and source of the compound and the chemical structure data.
[0041] Before inputting the chemical structure data into the preset feature extraction model, the original chemical data needs to be processed. First, redundant information is filtered. The data from different data sources may contain duplicate or irrelevant information. By writing specific program code, based on the characteristics and laws of the data, this redundant information is removed, and at the same time, the format differences between different data sources are eliminated, and all data is unified into a format suitable for subsequent processing. Then, a multi-scale noise reduction algorithm is used to remove the outliers in the original chemical data. The multi-scale noise reduction algorithm is based on the principle of wavelet transform, decomposes the data into different scale spaces, estimates and removes the noise at each scale, retains the key information of the molecular structure, improves the data quality, and provides a reliable data basis for subsequent feature extraction.
[0042] Example 3: When constructing a structured intermediate database, a specific drug intermediate - ibuprofen intermediate is used as an example for illustration. The chemical structure of the ibuprofen intermediate contains multiple complex atomic connection modes and functional groups, and these atoms and functional groups are interconnected to form a specific chemical structure map.
[0043] The chemical structure map of ibuprofen intermediates is segmented according to the atomic connectivity in molecular descriptors. First, analyze the atomic connectivity data to determine the bond connection relationships and topological structures between atoms in the map. For example, through the study of its chemical structure, it is found that the carbon atoms on the benzene ring form stable covalent bonds with other atoms, and there are specific connection sequences and spatial position relationships with the atoms on the side chain. Use relevant algorithms in graph theory, such as the minimum spanning tree algorithm (MST), which can construct a tree-like structure with the minimum sum of connection edge weights while ensuring that all atoms are connected, to divide different parts of the map. Starting from the chemical structure map of ibuprofen intermediates, taking a certain key atom as the starting point, gradually add connection edges according to atomic connectivity to construct the minimum spanning tree. In this process, the map is segmented into several sub-maps, and each sub-map represents a relatively independent and functionally specific part of the chemical structure. For example, the benzene ring part is taken as a sub-map, and the carboxyl group and related connected atoms on the side chain are taken as another sub-map.
[0044] Generate hierarchical storage units based on the segmented map. Take these sub-maps as nodes in the hierarchical structure to construct a tree-shaped storage system. Taking the node corresponding to the benzene ring sub-map as the parent node, establish a hierarchical relationship with the nodes of the side chain sub-maps that are closely related to it as child nodes. Associate the corresponding synthesis path and physical and chemical property parameters in each storage unit. For the storage unit of the benzene ring part of ibuprofen intermediates, its synthesis path parameters may include: in a certain reaction, using benzene as the starting material, under specific temperature (such as 50 - 60 °C) and pressure (1 - 2 atm) conditions, using aluminum trichloride as a catalyst, and undergoing a Friedel - Crafts alkylation reaction with halogenated hydrocarbons for 3 - 4 hours. Its physical and chemical property parameters include: melting point of [X] °C, solubility in common organic solvents such as ethanol of [X] g / 100 mL, etc. For the storage unit related to the side chain carboxyl group, the synthesis path may involve using a specific oxidant (such as potassium permanganate) in a specific oxidation reaction, reacting under alkaline conditions, controlling the reaction temperature at [X] °C, and the reaction time is [X] hours. In terms of physical and chemical properties, information such as its solubility in water and acid dissociation constant is also stored in this unit. In this way, the chemical structure, synthesis path, and physical and chemical property information of complex drug intermediates are stored in a structured manner, facilitating subsequent data query, analysis, and management.
[0045] Example 4: In the process of inputting the structured intermediate database into a preset AI retrieval model, also take the retrieval related to ibuprofen intermediates as an example. Suppose the user wants to find drug intermediates that are structurally similar to ibuprofen intermediates and have a high yield, and set query constraint conditions in the AI retrieval model.
[0046] For the retrieval of structural similarity, the model adopts a similarity calculation method based on molecular fingerprints. Molecular fingerprints are a way to encode molecular structures, which convert the structural information of molecules into a series of binary digits. In this embodiment, Extended Connectivity Fingerprint (ECFP) is selected to describe the molecular structure. ECFP generates fingerprints by performing substructure searches with increasing radii on the molecular structure and hashing these substructures. The structural similarity threshold is set to 0.7, that is, when the ECFP fingerprint similarity between the target molecule and the molecule in the database is greater than 0.7, the two are considered structurally similar. The similarity calculation uses the Tanimoto coefficient, and its calculation formula is:
[0047] where, is the number of bits that are both 1 in the two molecular fingerprints, and are the total number of 1s in the two molecular fingerprints respectively.
[0048] For the retrieval of the yield range, assume that the user wants to find drug intermediates with a yield greater than 75%. The query condition for the yield is set to greater than 75% in the model.
[0049] By iteratively calculating the graph node embedding vectors, the multi-dimensional similarity scores between the database records and the query conditions are matched. The Graph Attention Network (GAT) algorithm is used to calculate the graph node embedding vectors. In GAT, each node updates its own feature representation based on the features and relative importance of its neighbor nodes. For each drug intermediate molecule in the database, its chemical structure map is converted into graph structure data and input into the GAT model. Let the feature vector of node at the -th iteration be , then the update formula is:
[0050] where, is the activation function, such as the ReLU function; is the weight matrix at the -th iteration; is the set of neighbor nodes of node ; is the attention coefficient between node and node , which is calculated by the following formula:
[0051] Here is a learnable parameter vector, Represents the vector concatenation operation. Through multiple iterations, stable node embedding vectors are obtained. Then, combining structural similarity and yield conditions, a multi-dimensional similarity score between database records and query conditions is calculated. For each drug intermediate in the database, a multi-dimensional similarity score is comprehensively calculated based on its structural similarity score (calculated based on the Tanimoto coefficient) and whether the yield meets the conditions. For example, the structural similarity score accounts for 0.6, and the yield score accounts for 0.4 (weights are set according to user needs). Finally, the comprehensive score of each record is obtained, sorted according to the score, and the retrieval results that meet the user's query requirements are output, providing valuable reference information for drug R & D personnel.
[0052] Example 5: This example focuses on introducing the multi-dimensional sorting of dynamic retrieval results using a preset optimization algorithm and the related training process, and also involves content related to the verification model. Weight assignment is performed on the dynamic retrieval results to generate a derivative sorted list according to the priorities of structural similarity, synthesis difficulty, or cost. For example, for structural similarity, a higher weight is given according to the previously calculated similarity score; for synthesis difficulty, it can be quantitatively evaluated according to factors such as the number of steps required for the synthesis reaction and the harshness of the reaction conditions, and the corresponding weight is assigned; for cost, consider raw material cost, catalyst cost, equipment cost, etc., and assign weights after comprehensive evaluation. Assume the weight of structural similarity is , the weight of synthesis difficulty is , the weight of cost is , and . Through weighted calculation, the comprehensive score of each retrieval result is obtained, thereby generating a derivative sorted list.
[0053] Based on the adaptive adjustment algorithm, optimization parameters that conform to chemical logic are injected into the derivative sorted list. The adaptive adjustment algorithm can dynamically adjust the weights according to the feedback information of the retrieval results. For example, if it is found through multiple retrievals that the user pays more attention to results with high structural similarity and low synthesis difficulty, the algorithm can appropriately increase the weight of structural similarity , and reduce the weight of synthesis difficulty .
[0054] When training the optimization algorithm, an optimization training set containing real chemical synthesis cases is constructed, and each case is labeled with the corresponding synthesis path truth value. A large number of real chemical synthesis cases are collected from actual chemical synthesis literature and experimental records. For example, 1000 synthesis cases of different drug intermediates are collected. The optimization algorithm is trained using a reinforcement learning framework so that the distribution difference between the sorted list generated by the optimization algorithm and the distribution of real chemical synthesis cases is less than the set tolerance range. The distribution difference can be measured by calculating the Kullback-Leibler divergence (KL divergence). Let the distribution of the sorted list generated by the optimization algorithm be , the distribution of real chemical synthesis cases is , then the KL divergence ,
[0055] where is the length of the sorted list. By continuously adjusting the parameters of the optimization algorithm, the KL divergence is made less than the set tolerance range to improve the performance of the optimization algorithm.
[0056] During the entire retrieval process, a verification model is also involved. The standardized retrieval list is input into a preset verification model to calculate the error matrix between the retrieval results and the experimental record data. The steps for constructing the verification model are as follows: Collect the synthetic route data recorded in the laboratory as the verification benchmark, and extract a large amount of reliable synthetic route data from the experimental record database in the laboratory. Establish an error evaluation network based on the attention mechanism, and output the error weight vector by comparing the multi-dimensional feature maps of the retrieval results and the verification benchmark. The attention mechanism can make the network pay more attention to the differences in important features between the retrieval results and the verification benchmark, improve the accuracy of error evaluation, and then inversely adjust the graph network depth and feature embedding dimension of the AI retrieval model according to the error matrix to continuously optimize the performance of the retrieval model. At the same time, the present invention also includes a data security control step. When receiving an access request from an external system to the structured intermediate database, verify the matching of the identity token provided by the requestor with the preset permission list. If the matching is successful, transmit a data subset that conforms to the access permission to the external system through an encrypted link to ensure the security and confidentiality of the data.
[0057] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device.
[0058] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for constructing a pharmaceutical intermediate database and AI intelligent retrieval, characterized in that, include: Obtaining chemical structure data and synthesis route data of a target drug intermediate, wherein the chemical structure data is derived from at least one chemical database or experimental recording device, and the synthesis route data includes reaction conditions, catalysts, and yield parameters; Inputting the chemical structure data into a preset feature extraction model to identify and extract molecular descriptors in the chemical structure data, wherein the molecular descriptors include atomic connectivity, electron cloud distribution, and functional group topological relationships; Constructing a structured intermediate database based on the molecular descriptors, wherein the database associates chemical structures, synthesis pathways, and multi-dimensional index tags of physicochemical properties; Inputting the structured intermediate database into a preset AI retrieval model to generate dynamic retrieval results, wherein the AI retrieval model matches the similarity between user queries and database records based on a graph neural network algorithm; The dynamic search results are sorted in multiple dimensions using a preset optimization algorithm, and a standardized search list is output.
2. The method for constructing a drug intermediate database and AI intelligent retrieval according to claim 1, wherein The step of training the feature extraction model includes: obtaining an initial training sample set, wherein each sample in the initial training sample set is labeled with a corresponding molecular descriptor label; using the initial training sample set to iteratively train an initial feature extraction network until the matching rate between the feature prediction result output by the initial feature extraction network and the molecular descriptor label is greater than or equal to a preset threshold; during each iteration, adjusting the weight parameter of the initial feature extraction network according to the prediction error of the previous iteration, and eliminating samples in the initial training sample set with errors higher than the set threshold, so as to form a reduced training subset for the next iterative training.
3. The method for constructing a pharmaceutical intermediate database and AI intelligent retrieval according to claim 1, wherein Obtaining the chemical structure data of the target drug intermediate includes obtaining chemical data from one of the data sources in the following manner: Connect to the target chemical database server and / or experimental equipment real-time data interface; Extracting raw chemical data from at least one of the target chemical database server and the experimental equipment real-time data interface according to a preset data acquisition protocol, wherein the preset data acquisition protocol includes data format, storage accuracy and metadata association rules.
4. The method for constructing a drug intermediate database and AI intelligent retrieval according to claim 3, wherein, Before inputting the chemical structure data into a preset feature extraction model, the method further comprises: filtering redundant information on the raw chemical data to eliminate format differences between different data sources; A multi-scale noise reduction algorithm is used to remove outliers in the raw chemical data and retain key information of the molecular structure.
5. The method for constructing a pharmaceutical intermediate database and AI intelligent retrieval according to claim 1, wherein Constructing a structured intermediate database based on the molecular descriptors comprises: segmenting the chemical structure map according to atomic connectivity in the molecular descriptors; A hierarchical storage unit is generated based on the segmented graph, and the corresponding synthesis path and physicochemical property parameters are associated in each unit.
6. The method for constructing a pharmaceutical intermediate database and AI intelligent retrieval according to claim 5, wherein Inputting the structured intermediate database into a preset AI retrieval model comprises: Setting query constraints in the AI retrieval model to simulate the user's retrieval requirements based on structural similarity, reaction conditions or yield range; By iteratively computing graph node embedding vectors, we match database records with multi-dimensional similarity scores of query conditions.
7. The method for constructing a drug intermediate database and AI intelligent retrieval according to claim 1, wherein, Performing multi-dimensional sorting on the dynamic retrieval results using a preset optimization algorithm includes: Assigning weights to the dynamic retrieval results to generate a derivative sorted list according to structural similarity, synthesis difficulty, or cost priority; Injecting optimization parameters that conform to chemical logic into the derivative sorted list based on an adaptive adjustment algorithm.
8. The method for constructing a drug intermediate database and AI intelligent retrieval according to claim 7, wherein The method further includes training the optimization algorithm in the following manner: Constructing an optimization training set containing real chemical synthesis cases, where each case is labeled with the corresponding synthesis path truth value; Training the optimization algorithm using a reinforcement learning framework such that the distribution difference between the sorted list generated by the optimization algorithm and the real chemical synthesis cases is less than a set tolerance range.
9. The method for constructing a pharmaceutical intermediate database and AI intelligent retrieval according to claim 1, wherein The method further includes: Inputting the standardized retrieval list into a preset verification model and calculating the error matrix between the retrieval results and the experimental record data; Back-adjusting the graph network depth and feature embedding dimension of the AI retrieval model according to the error matrix; The verification model is constructed through the following steps: Collecting the synthesis path data recorded in the laboratory as a verification benchmark; Establishing an error evaluation network based on the attention mechanism and outputting an error weight vector by comparing the multi-dimensional feature maps of the retrieval results and the verification benchmark.
10. The method for constructing a drug intermediate database and AI intelligent retrieval according to claim 1, wherein The method further includes a data security control step: When receiving an access request from an external system to the structured intermediate database, verifying the matching of the identity token provided by the requestor with a preset permission list; When the matching is successful, transmitting a data subset that conforms to the access permission to the external system through an encrypted link.
Citation Information
Patent Citations
Construction method and system of drug-induced target organ toxicity prediction model
CN116343907A
Drug activity screening method and device
CN118506911A
Medical intermediate quality control method based on machine learning driving
CN118711714A
Drug molecule generation method capable of explaining federal large model
CN119920356A
Compound vector database construction method, compound similarity search method and device
CN120015181A