Molecular data storage method and apparatus, application method and apparatus
By verifying and classifying molecular data in drug development, the problem of low data storage and analysis efficiency is solved, and efficient management and retrieval of molecular data are achieved, making it suitable for complex scenarios such as drug development.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-30
- Publication Date
- 2026-03-20
AI Technical Summary
In existing technologies, the large amount of molecular data generated during drug development cannot be efficiently stored and uniformly analyzed, leading to difficulties in the screening process, low data access efficiency, and a large workload for analysis. It also cannot support the storage of semi-structured and unstructured data, thus affecting the efficiency of drug development.
A molecular data storage method and apparatus are provided. By verifying the molecular data to be processed, the incremental auxiliary data type is determined, and structured data is saved to a database, while semi-structured and unstructured data are saved to files. The association between the files and the database is established to achieve effective data management and query.
It enables convenient and effective management of molecular data, supports various computing needs, enhances data computing capabilities, simplifies user query and usage processes, and is suitable for complex scenarios such as drug development.
Smart Images

Figure CN114300064B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of data processing, in particular to a molecular data storage method and device, and further relates to a molecular data application method and device. BACKGROUND
[0002] In early screening in drug research and development, there are a large number of molecular data generated from artificial intelligence (AI) and virtual screening of large compound libraries. These molecular data generally include the physical and chemical properties of the molecules themselves, such as molecular mass, molecular smiles (simplified molecular input line entry system) formula, and other basic properties, as well as three-dimensional coordinate structure files of the molecules. In the screening process, various computational chemistry, quantum chemistry, and AI model prediction methods such as free energy perturbation (FEP) calculation, quantum chemistry (QM) calculation, molecular dynamics (MD) simulation, ADMET (Absorption, Distribution, Metabolism, Excretion, and Toxicity) prediction, etc. are used to obtain energy information, absorption decomposition, and toxicity prediction information of the molecules; and some molecules are synthesized and subjected to various biological activity experiments to obtain activity data of the molecules.
[0003] In the entire molecular screening process, a large amount of structured molecular data such as molecular properties, protein binding energy, and activity data is generated, and a large amount of semi-structured and unstructured molecular data such as molecular structure, molecular synthesis report, and molecular activity experiment report is also generated. In addition, there are also molecular related metadata such as docking protein information and molecular scaffold patent files. In the screening process, a large number of simple or complex algorithms are run on the molecular data. The algorithms are independent of each molecule or are batch screening of molecules.
[0004] Due to the existence of a large amount of data information, the construction of the screening process in the related art is very difficult, the data generated in the screening steps cannot be efficiently accessed, and the series connection of the process relies on various non-standardized methods to establish, resulting in that data comprehensive analysis and review cannot be unified, the data analysis workload is huge, and the efficiency is low. SUMMARY
[0005] To solve or partially solve the problems in the related art, the present application provides a molecular data storage method and device, an application method and device, which can conveniently and effectively manage all related data of molecules, and facilitate user query and use of the data.
[0006] The first aspect of the present application provides a molecular data storage method, which comprises: receiving molecular data to be processed; checking the molecular data to obtain data passing the check; determining the incremental accessory data of the data passing the check, the incremental accessory data comprising any one or more of the following: structured incremental accessory data, semi-structured incremental accessory data, and unstructured incremental accessory data; saving the data passing the check and the structured incremental accessory data into a database; saving the semi-structured incremental accessory data and / or the unstructured incremental accessory data into a file, and adding the directory index of the file to the database to establish the association between the file and the database.
[0007] The second aspect of the present application provides a molecular data storage device, which comprises: a data receiving module, an analysis module, a data summarizing module, and a data management module. The data receiving module is configured to receive molecular data to be processed. The analysis module is configured to check the molecular data to obtain data passing the check. The data summarizing module is configured to determine the incremental accessory data of the data passing the check, the incremental accessory data comprising any one or more of the following: structured incremental accessory data, semi-structured incremental accessory data, and unstructured incremental accessory data. The data management module is configured to save the data passing the check and the structured incremental accessory data into a database, save the semi-structured incremental accessory data and / or the unstructured incremental accessory data into a file, and add the directory index of the file to the database to establish the association between the file and the database.
[0008] The third aspect of the present application provides a molecular data application method, which comprises: receiving a calculation method submitted by a user through an API; obtaining calculation data related to the calculation method, the calculation data comprising calculation data obtained from a database and / or a file; the database storing molecular data and structured incremental accessory data thereof, the file storing semi-structured incremental accessory data and / or unstructured incremental accessory data of the molecular data, and the database containing the directory index of the file; calculating the calculation data using the calculation method to obtain a calculation result; and saving the calculation result into the database and / or the file.
[0009] The fourth aspect of the present application provides a molecular data application device, which comprises: an application interface module configured to receive a calculation method submitted by a user; a calculation data acquisition module configured to acquire calculation data related to the calculation method, the calculation data comprising calculation data acquired from a database and / or a file; the database stores molecular data and structured incremental accessory data thereof, the file stores semi-structured incremental accessory data and / or unstructured incremental accessory data of the molecular data, and the database comprises a directory index of the file; and a calculation processing module configured to perform calculation on the calculation data by using the calculation method to obtain a calculation result, and save the calculation result to the database and / or the file.
[0010] The fifth aspect of the present application provides an electronic device, comprising: a processor; a memory having executable code stored thereon, when the executable code is executed by the processor, the processor performs the above method.
[0011] The sixth aspect of the present application further provides a computer-readable storage medium having executable code stored thereon, when the executable code is executed by the processor of the electronic device, the processor performs the above method.
[0012] The seventh aspect of the present application further provides a computer program product comprising executable code, when the executable code is executed by the processor, the above method is implemented.
[0013] The molecular data storage method and device provided by the embodiments of the present application perform verification on the molecular data to be processed, determine the incremental accessory data of the data that passes the verification, and according to the feature that the incremental accessory data can include one or more different types of data, different storage methods are used according to the different types, the data that passes the verification and the structured molecular data related thereto are saved to the database, the semi-structured molecular data and the unstructured molecular data related thereto are saved to the file, and the directory index of the file is added to the database, thereby establishing the association between the file and the database, so that the management of all related data of the molecule is facilitated and effective, and the subsequent user query and use of the data are facilitated.
[0014] The molecular data application method and device provided by the embodiments of the present application receive the calculation method submitted by the user through the application programming interface (API), such as quantum chemistry algorithm, computational chemistry algorithm, AI model algorithm, etc., based on the effective storage and association of the molecular data and the different types of incremental accessory data thereof, acquire the calculation data related to the calculation method, perform corresponding calculation to obtain the calculation result, and save the calculation result to the database or the file, so that the calculation result is also effectively stored.
[0015] Further, the technical solution of the present application not only supports local calculation, but also can submit part or all data to a remote cluster server for relevant calculation, greatly improving the calculation capacity of data, thereby meeting various different calculation requirements of users.
[0016] Further, the technical solution of the present application can also receive query information submitted by a user through an API, can query according to different query information such as molecular substructure, molecular similarity, molecular attribute parameters, read data from a database, and display the read data, thereby facilitating the user to query and use the molecular data.
[0017] It should be understood that the foregoing general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF DRAWINGS
[0018] The above and other objects, features and advantages of the present application will become more apparent from the following detailed description when taken in conjunction with the accompanying drawings in which like reference characters refer to like parts throughout the different views of the drawings, and in which:
[0019] Figure 1 An exemplary system architecture to which the molecular data storage method and device, and the application method and device according to embodiments of the present application can be applied is schematically shown;
[0020] Figure 2 A flowchart of a molecular data storage method according to embodiments of the present application is schematically shown;
[0021] Figure 3 A molecular dimension model diagram in embodiments of the present application is schematically shown;
[0022] Figure 4 A flowchart of a molecular data application method according to embodiments of the present application is schematically shown;
[0023] Figure 5 A structural block diagram of a molecular data storage device according to embodiments of the present application is schematically shown;
[0024] Figure 6 A structural block diagram of a molecular data application device according to embodiments of the present application is schematically shown;
[0025] Figure 7 Another structural block diagram of a molecular data application device according to embodiments of the present application is schematically shown;
[0026] Figure 8 A block diagram of an electronic device to realize embodiments of the present application is schematically shown. DETAILED DESCRIPTION
[0027] Embodiments of the present application will be described in more detail with reference to the drawings. Although the embodiments of the present application are shown in the drawings, it is understood that the present application can be embodied in various forms and should not be limited by the embodiments set forth herein. Rather, these embodiments are provided so that this application will be thorough and complete, and fully convey the scope of the application to those skilled in the art.
[0028] The terminology used in the present application is for the purpose of describing particular embodiments only and is not intended to be limiting of the present application. As used herein, the terms "comprises", "comprising", "includes", "including" and the like are specifically intended to be open-ended and to mean that the features, steps, operations and / or components listed are present, but not excluding the presence or addition of one or more other features, steps, operations, components, and so forth.
[0029] All terms used herein, including technical and scientific terms, have the meanings commonly understood by one of ordinary skill in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted as having a meaning that is consistent with the context of the specification, and should not be interpreted in an idealized or overly formal way.
[0030] It should be understood that although the terms "first", "second", "third", etc. can be employed in this application to describe various information, these information should not be limited by these terms. These terms are only used to distinguish one piece of information from another piece of information. For example, without departing from the scope of the present application, the first information can also be referred to as the second information, and similarly, the second information can also be referred to as the first information. Therefore, the features defined with "first", "second" can explicitly or implicitly include one or more of the features. In the description of the present application, the meaning of "a plurality of" is two or more, unless otherwise specifically limited.
[0031] Before the technical solutions of the present application are described, some technical terms related to the field of the present application are explained.
[0032] Data warehouse is a storage structure of structured data, which is a core component for supporting report generation, building data mart, and business intelligence.
[0033] JS object score (JavaScript Object Notation, referred to as JSON) is a lightweight data exchange format. JSON is based on a subset of js specification (ECMAScript) formulated by European Computer Association, and uses a text format completely independent of programming language to store and represent data. JSON is easy for people to read and write, and is easy for machines to parse and generate, and effectively improves network transmission efficiency.
[0034] Simplified Molecular Input Line Entry System (SMILES) is a specification for unambiguously describing the structure of chemical molecules using ASCII strings. SMILES strings can be imported into a variety of molecular editing software and converted into two-dimensional drawings or three-dimensional models of molecules.
[0035] Due to the large amount of molecular data generated from AI-generated early screening and virtual screening of large compound libraries in drug research and development, the molecular data is not only large in quantity, but also various in type. In the absence of a molecular data storage system, it is difficult to construct a screening process, the data generated in the screening step cannot be efficiently accessed, and data analysis and review are difficult to unify, resulting in extremely low data analysis efficiency. The relational database, data warehouse and file storage (object storage) system in the related art cannot support complete storage of molecular data. The relational database is mostly applied to OLTP (online transaction processing) and supports transaction processing; the data warehouse is applied to OLAP (online analytical processing) and can support data analysis, but cannot support semi-structured and unstructured data storage and cannot be used for molecular screening algorithm running.
[0036] Therefore, an embodiment of the present application provides a molecular data storage method and device. For the molecular data to be processed, first, the molecular data is verified. For the data that passes the verification, not only the molecular data itself is stored, but also the incremental affiliated data is determined. Considering that the incremental affiliated data can include one or more different types of data, different storage methods are used according to the different types of data. The data that passes the verification and the related structured data are saved to the database, the related semi-structured molecular data and unstructured molecular data are saved to the file, the directory index of the file is added to the database, and the association between the file and the database is established, so that all types of data of the molecule can be effectively stored, which is especially suitable for complex and long-period specific use scenarios such as drug research and development.
[0037] The following will be described in detail Figures 1 to 8 A molecular data storage method and device, application method and device of an embodiment of the present application will be described in detail.
[0038] Figure 1 An exemplary system architecture to which the molecular data storage method and device, application method and device according to an embodiment of the present application can be applied is schematically shown. It should be noted that Figure 1 The system architecture shown is only an example of a system architecture to which the embodiment of the present application can be applied, to help those skilled in the art understand the technical content of the present application, but does not mean that the embodiment of the present application cannot be used in other devices, systems, environments or scenarios.
[0039] Referring to Figure 1 According to the system architecture 100 of this embodiment, it can include terminal devices 101, 102, 103, a network 104 and a server 105. The network 104 is a medium for providing communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 can include various connection types, such as wired, wireless communication links or optical fiber cables, etc.
[0040] The user can use the terminal devices 101, 102, 103 to interact with other terminal devices and the server 105 through the network 104 to receive or send information, etc., such as sending calculation methods, calculation data, etc. The terminal devices 101, 102, 103 can be installed with various applications, such as drug development applications, material design applications, web browser applications, database applications, search applications, instant messaging tools, mailbox clients, social platform software, etc.
[0041] The terminal devices 101, 102, 103 include, but are not limited to, smart desktop computers, tablet computers, laptop computers, etc., which can support modeling, analysis calculation, design, web browsing, etc.
[0042] The server 105 can receive calculation methods, calculation data, etc., and also can send calculation results to the terminal devices 101, 102, 103. For example, the server 105 can be a background management server, a server cluster, etc.
[0043] It should be noted that the number of terminal devices, networks and servers is only illustrative. According to the implementation needs, there can be any number of terminal devices, networks and clouds.
[0044] The embodiments of the present application provide a molecular data application method and device, which utilizes the above database and file to facilitate the use and query of molecular data by users, and can save the calculation results to the database and file.
[0045] As Figure 2 shown, Figure 2 a flowchart of a molecular data storage method according to an embodiment of the present application is shown.
[0046] The molecular data storage method of this embodiment includes operations S210-S240.
[0047] In operation 210, the molecular data to be processed is received.
[0048] It should be noted that the molecular data to be processed can be any molecular data involved in drug development, such as molecular design, synthesis, experimental testing, and evaluation data related to molecular MIMI, molecular mass, and molecular activity.
[0049] In operation 220, the molecular data is verified, and the data that passes the verification is obtained.
[0050] In this embodiment, validation can be performed based on the dimensions of the molecular data. Specifically, the dimensions of the molecular data can be determined according to a molecular dimension model, and the molecular data can be validated based on the determined dimensions to obtain validated data.
[0051] Molecular dimensional models can be determined based on the type of molecular data. A molecular dimensional model can include: a table of basic molecular properties and one or more tables of molecular attributes. The molecular attribute tables include: a table of computational attributes and a table of experimental attributes. The table of basic molecular properties contains an identifier field for each molecular attribute table.
[0052] like Figure 3 The diagram illustrates a molecular dimension model in an embodiment of this application. In this example, *Molecule* represents the molecular basic property dimension table, where the basic molecular attributes are the two-dimensional molecular structural formulas (smiles), *inchi-key* is the tag obtained based on the smiles hash operation, and the other *id* are association identifiers, i.e., the identifier fields of each molecular attribute dimension table. Corresponding molecular attribute dimension tables include, for example, *FEPResults*, *Structure*, *Activity*, *ADMETResults*, and *Synthesis Results*. In the *FEPResults* dimension table, *pdb* refers to the target protein used in drug research, *dg* is the binding free energy between the molecule and the protein, and *fep_method* refers to the calculation method used. In the *ADMETResults* dimension table, *clogp* refers to the molecule's hydrophobicity constant, *caco-2* refers to the absorption capacity of the *caco-2* cell line, and *water_solubility* refers to the molecule's water solubility. In the *Structure* dimension table, *name* refers to the name of the molecular three-dimensional structure file, and *file_path* is the file storage path. In the *SynthesisResults* dimension table, *weight* and *purity* refer to the quality and purity of the molecule synthesis, respectively, and *report* records the synthesis report information. The Activity dimension table records the values of different activity properties tested using different methods. Among them, name is used to identify the record with a custom name, property refers to the activity property being tested, and method records the method used for the activity experiment.
[0053] It should be noted that the above Figure 3 The above is only a simple example of the molecular dimension model in the embodiments of the present application, and is not intended to limit the specific structure of the molecular dimension model.
[0054] In the verification of the molecular data, the fields and attribute information thereof contained can be determined according to the dimension of the molecular data, the fields are verified according to the attribute information of the fields, and the data passing the verification is obtained.
[0055] Specifically, the molecular data can be verified according to the fields defined in the data detail table, such as whether it is a field that must be contained, the field type (such as integer, floating point number or string), and the like.
[0056] In operation 230, the incremental accessory data of the data passing the verification is determined.
[0057] The incremental accessory data refers to the data related to the molecular data other than the basic data of the molecule itself, such as molecular state data (such as whether the molecule contains FEP data, whether it contains experimental data, molecular fingerprint), specific information data corresponding to the molecular state, and the like, which can be set as needed.
[0058] The incremental accessory data can generally be used for screening and filtering of batches of molecules, such as screening molecules that have undergone a certain type of calculation or experiment.
[0059] The types of the incremental accessory data of the molecular data can include any one or more of the following types: structured molecular data, semi-structured molecular data, unstructured molecular data.
[0060] Through the determination of the incremental accessory data of the molecular data, all related data information of the molecular data can be more richly and comprehensively obtained.
[0061] In operation 240, the data passing the verification and the structured incremental accessory data are saved into the database; the semi-structured incremental accessory data and / or the unstructured incremental accessory data are saved into a file, and the directory index of the file is added to the database to establish the association between the file and the database.
[0062] Further, in the verification process of the molecular data, if the molecular data fails the verification, an error can also be reported, such as by displaying error information, so that the operator can promptly know whether the molecular data is correct, facilitating the processing of the molecular data, such as correction, deletion and the like.
[0063] The molecular data storage method and device provided in the embodiments of the present application check the molecular data to be processed, determine the incremental accessory data of the data that passes the check, and according to the feature that the incremental accessory data can include one or more different types of data, save the data that passes the check and the structured molecular data related to the data in a database, save the semi-structured molecular data and unstructured molecular data related to the data in a file, add the directory index of the file to the database, and establish the association between the file and the database, so that the molecular data is conveniently and effectively managed, and the subsequent user query and use of the data are facilitated.
[0064] Correspondingly, based on the storage of the molecular data and the incremental accessory data thereof in the database and the file, the embodiments of the present application further provide a molecular data application method, which provides an effective solution for the use of the molecular data by the user.
[0065] As shown in FIG. 5, a flowchart of the molecular data application method according to the embodiments of the present application is schematically shown. Figure 4
[0066] The molecular data application method of the embodiments includes the following operations:
[0067] In operation 410, the calculation method submitted by the user through the API is received.
[0068] The calculation method can be, but is not limited to, any one of the following: quantum chemistry algorithm, computational chemistry algorithm, AI model algorithm, etc.
[0069] In specific applications, RESTFUL or graphql style API can be used, and the embodiments of the present application do not limit this.
[0070] RESTFUL (Representational State Transfer) is a design style and development method of a network application, which is based on the Hyper Text Transfer Protocol (HTTP) and can be defined in XML format or JSON format. RESTFUL is suitable for the scene of mobile Internet vendors as a business interface, and realizes the function of calling mobile network resources by a third-party OTT, and the action type is to add, change or delete the called resources.
[0071] GraphQL is a Query Language especially good for querying Graph (graph data), so it is called GraphQL. It shares the QL suffix with SQL. GraphQL can choose NoSQL type database, SQL type database or other storage.
[0072] At operation 420, computing data related to the computing method is acquired, the computing data including computing data acquired from a database and / or a file.
[0073] It should be noted that the database and the file here refer to the database and the file described above that store molecular data and structured incremental data thereof, and the database contains a directory index of the file.
[0074] Further, the computing data related to the computing method can also include user input computing data.
[0075] That is, the computing data required for the computing method submitted by the user to participate in the calculation can be partially from the already stored data, such as the molecular three-dimensional structure file stored in the file storage system, or the molecular material property data stored in the database, etc., and partially from the user input computing data, such as parameter configuration information for indicating which batch of molecular data to use for calculation. Of course, it can also be all from the above-mentioned file and database, and the embodiments of the present application do not limit this.
[0076] At operation 430, the computing data is calculated by using the computing method to obtain a calculation result.
[0077] It should be noted that in specific applications, part or all of the calculation can be completed locally, or part or all of the calculation can be submitted to a remote cluster server for calculation and the calculation result returned by the remote cluster server can be received. For example, for light calculation, the calculation can be completed locally and the calculation result can be returned in real time; for complex and long-consuming calculation tasks, asynchronous calculation can be used to continuously monitor the calculation progress and the calculation result.
[0078] At operation 440, the calculation result is saved to a database and / or a file.
[0079] Further, error information of a calculation error can also be stored, such as saved to a log file.
[0080] The molecular data application method and device provided in the embodiments of the present application can receive a user-submitted calculation method, such as a quantum chemistry algorithm, a computational chemistry algorithm, an AI model algorithm, and the like, obtain calculation data related to the calculation method based on the effective storage and correlation of the molecular data and the different types of incremental attached data of the molecular data, perform corresponding calculation to obtain a calculation result, and save the calculation result in a database or a file, so that the calculation result is also effectively stored.
[0081] In another embodiment of the molecular data application method of the present application, the above database and file can also be used to facilitate user query of molecular data information. Specifically, query information submitted by a user through an API is received, the query information can include but is not limited to any one or more of the following: molecular substructure, molecular similarity, molecular attribute parameters, and the like; data is read from the database according to the query information; and the read data is displayed.
[0082] When performing the query, a query strategy can be generated according to the query information, and data is read from the database according to the query strategy; and the read data is format-converted.
[0083] For example, there are already 1 million pieces of molecular data in the database, and each piece of data has a smiles attribute, which describes the two-dimensional structure of the molecule and can be used for substructure matching search. When the query function is used, a smiles describing substructure information can be used as query information, and an API (application program interface) is called to submit the query information to the system. The system will run the substructure matching search algorithm, perform 1 million data traversal filtering, and return the finally matched molecular data. If the submitted query information contains the query information of the substructure and “molecular mass greater than 100”, the system will automatically identify and generate a query optimization strategy, that is, the molecules with a molecular mass greater than 100 are found first, and then the substructure search algorithm is run on these molecules to perform filtering query. The structure of a molecule can be described by smiles, or can use the sdf file format or the mol file format to generate a three-dimensional structure with coordinate description. According to the data format requirement specified in the query information, the system will automatically convert smiles into a specific format.
[0084] It should be noted that the query process is a general capability, which is processed according to specific query fields. For example, when a molecular substructure is queried, a molecular substructure fragment can be filled into the query information and submitted through an API. The system will automatically perform data retrieval on all molecules in the database based on the query strategy of the molecular substructure search algorithm. According to the display configuration, the system will support format conversion of smiles according to different rendering, and also support conversion into other molecular data formats such as sdf.
[0085] Correspondingly, the embodiment of the present application also provides a molecular data storage device, such as Figure 5 As shown in a non-limiting embodiment, the molecular data storage device 500 includes a data receiving module 510, an analysis module 520, a data summarizing module 530, and a data management module 540. Among them:
[0086] The data receiving module 510 is configured to receive molecular data to be processed.
[0087] The analysis module 520 is configured to verify the molecular data to obtain data that passes the verification.
[0088] The data summarizing module 530 is configured to determine the incremental affiliated data of the data that passes the verification, and the incremental affiliated data includes any one or more of the following: structured incremental affiliated data, semi-structured incremental affiliated data, and unstructured incremental affiliated data.
[0089] The data management module 540 is configured to save the data that passes the verification and the structured incremental affiliated data into a database, save the semi-structured incremental affiliated data and / or the unstructured incremental affiliated data into a file, and add a directory index of the file to the database to establish an association between the file and the database.
[0090] In a non-limiting embodiment, the analysis module 520 described above can include a data parsing unit and a data verification unit.
[0091] The data parsing unit is configured to determine the dimension of the molecular data according to a molecular dimension model.
[0092] The data verification unit is configured to verify the molecular data according to the determined dimension of the molecular data to obtain data that passes the verification.
[0093] The data verification unit described above can include a data detail determination unit and a field verification unit.
[0094] The data detail determination unit is configured to determine each field and its attribute information contained in the molecular data according to the dimension of the molecular data.
[0095] The field verification unit is configured to verify each field according to its attribute information to obtain data that passes the verification.
[0096] In specific applications, the molecular dimension model can include a molecular basic property dimension table and one or more molecular attribute dimension tables.
[0097] The molecular attribute dimension table can include a calculation attribute dimension table and an experimental attribute dimension table, and the molecular basic property dimension table can contain an identification field of each molecular attribute dimension table.
[0098] By using the molecular data management apparatus provided in the embodiments of the present application, the molecular data and various types of incremental accessory data thereof can be conveniently and effectively managed, and subsequent users can conveniently query and use the data, which is especially suitable for complex and long-period specific use scenarios such as drug research and development.
[0099] Correspondingly, the present application further provides a molecular data application apparatus, which can use the database and files storing the molecular data and the incremental accessory data thereof to provide calculation and query functions for users.
[0100] As shown in Figure 6 in a non-limiting embodiment, the molecular data application apparatus 600 includes the following modules: an application interface module 610, a calculation data acquisition module 620, and a calculation processing module 630.
[0101] The application interface module 610 is configured to receive a calculation method submitted by a user, which includes but is not limited to any one of the following: quantum chemistry algorithm, computational chemistry algorithm, AI model algorithm, etc.
[0102] The calculation data acquisition module 620 is configured to acquire calculation data related to the calculation method.
[0103] The calculation processing module 630 is configured to calculate the calculation data by using the calculation method to obtain a calculation result, and save the calculation result into a database and / or a file.
[0104] It should be noted that the calculation data can specifically include calculation data acquired from the database and / or the file. The database and the file are the database and the file established by the above-mentioned molecular data storage apparatus. The database stores the molecular data and the structured incremental accessory data thereof, the file stores the semi-structured incremental accessory data and / or the unstructured incremental accessory data of the molecular data, and the database contains a directory index of the file.
[0105] Further, the calculation data can further include calculation data input by a user, such as parameter configuration information for indicating which batch of molecular data is used for calculation, etc.
[0106] In specific applications, the above-mentioned calculation processing module 630 can include a local calculation unit and / or a remote calculation unit.
[0107] The local calculation unit is configured to perform local calculation on part or all of the data by using the calculation method to obtain a calculation result.
[0108] The remote calculation unit is configured to submit the calculation method and part or all of the calculation data to a remote cluster server for calculation, and receive a calculation result returned by the remote cluster server.
[0109] For example, for lightweight computing, the computing can be completed locally and the computing result is returned in real time. For complex and time-consuming computing tasks, asynchronous computing can be used to continuously monitor the computing progress and the computing result.
[0110] As shown in FIG. 6, in another non-limiting embodiment, different from the embodiment shown in FIG. 5, the molecular data application apparatus 600 can further include a query and export module 640 and a display module 650. Figure 7 Figure 6 In this embodiment, the application interface module 610 is further configured to receive query information submitted by a user.
[0111] Correspondingly, the query and export module 640 is configured to read data from the database and / or the file according to the query information. The display module 650 is configured to present the data read by the query and export module 640.
[0112] It should be noted that the query information can include, but is not limited to, any one or more of the following: a molecular substructure, a molecular similarity, a molecular attribute parameter, etc.
[0113] The query and export module 640 can specifically include a filtering and sorting unit and a format conversion unit.
[0114] The filtering and sorting unit is configured to generate a query strategy according to the query information, and read data from the database and / or the file according to the query strategy.
[0115] The format conversion unit is configured to perform format conversion on the data read by the filtering and sorting unit.
[0116] The format conversion unit is configured to perform format conversion on the data read by the filtering and sorting unit.
[0117] As for the molecular data storage apparatus 500 and the molecular data application apparatus 600 in the above embodiments, the specific manners in which the modules and units perform operations have been described in detail in the embodiments of the method, and will not be described in detail here.
[0118] It should be noted that in specific applications, the molecular data storage apparatus 500 and the molecular data application apparatus 600 can be integrated in a system, and the system can be divided into a storage layer, a business layer, and a display layer. The modules and units in the molecular data storage apparatus 500 and the molecular data application apparatus 600 can be arranged in the business layer and the display layer, and the database and the file can be arranged in the storage layer. Different permissions can be set for the creation and use of data to ensure the security of the data.
[0119] Another aspect of the present application also provides an electronic device, which can implement the molecular data storage method provided by the embodiments of the present application or implement the molecular data application method provided by the embodiments of the present application. Another aspect of the present application also provides an electronic device, which can implement the molecular data storage method provided by the embodiments of the present application or implement the molecular data application method provided by the embodiments of the present application.
[0120] As Figure 8 shown, Figure 8 a block diagram of an electronic device for implementing an embodiment of the present application is schematically shown.
[0121] Referring to Figure 8 , the electronic device 800 includes a memory 810 and a processor 820.
[0122] The processor 810 can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gates or transistor logic, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can be any conventional processor.
[0123] The memory 820 can include various types of storage units, such as a system memory, a read-only memory (ROM), and a permanent storage device. Among them, the ROM can store static data or instructions required by the processor 820 or other modules of the computer. The permanent storage device can be a read-write storage device. The permanent storage device can be a non-volatile storage device that does not lose stored instructions and data even after the computer is powered off. In some embodiments, the permanent storage device uses a mass storage device (such as a magnetic or optical disk, a flash memory) as a permanent storage device. In some other embodiments, the permanent storage device can be a removable storage device (such as a floppy disk, an optical drive). The system memory can be a read-write storage device or a volatile read-write storage device, such as a dynamic random access memory. The system memory can store some or all of the instructions and data required by the processor during runtime. In addition, the memory 810 can include a combination of any computer readable storage media, including various types of semiconductor storage chips (such as DRAM, SRAM, SDRAM, flash memory, programmable read-only memory), magnetic disks and / or optical disks. In some embodiments, the memory 810 can include a read and / or write removable storage device, such as a compact disc (CD), a read-only digital versatile disc (such as DVD-ROM, double-layer DVD-ROM), a read-only Blu-ray disc, an ultra-density optical disc, a flash memory card (such as an SD card, a min SD card, a Micro-SD card, etc.), a magnetic floppy disk, etc. The computer readable storage medium does not include a carrier wave and a transient electronic signal transmitted through a wireless or wired transmission.
[0124] The executable code stored in the memory 810 can cause the processor 820 to perform part or all of the methods described in the above embodiments when the executable code is processed by the processor 820.
[0125] In addition, the method according to the present application can also be implemented as a computer program or computer program product, which includes computer program code instructions for performing part or all of the operations in the above method of the present application.
[0126] Alternatively, the present application can also be implemented as a computer readable storage medium (or non-transitory machine readable storage medium or machine readable storage medium) having stored executable code (or computer program or computer instruction code) which, when executed by a processor of an electronic device (or a server, etc.), causes the processor to perform part or all of the operations of the above method according to the present application.
[0127] The above has described the embodiments of the present application, and the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and changes are obvious to those skilled in the art without departing from the scope and spirit of the described embodiments. The selection of the terms used herein is intended to best explain the principles of the embodiments, practical application, or improvement to the technology in the market, or to enable other ordinary skilled in the art to understand the embodiments disclosed herein.
Claims
1. A method for storing molecular data, characterized in that, The method includes: Receive molecular data to be processed; The molecular data is validated to obtain validated data; this includes: determining the dimensions of the molecular data according to a molecular dimension model; and validating the molecular data according to the dimensions to obtain validated data. The incremental auxiliary data of the data that passed the verification is determined. The incremental auxiliary data includes any one or more of the following: structured incremental auxiliary data, semi-structured incremental auxiliary data, and unstructured incremental auxiliary data. The incremental auxiliary data includes data other than the basic data of the molecule itself related to the molecular data, including molecular state data and specific information data corresponding to the molecular state. The verified data and the structured incremental supplementary data are saved to the database; the semi-structured incremental supplementary data and / or the unstructured incremental supplementary data are saved to a file, and the directory index of the file is added to the database to establish the association between the file and the database.
2. The method according to claim 1, characterized in that, The step of validating the molecular data according to the dimension to obtain data that passes the validation includes: The molecular data and its attribute information are determined based on the dimensions described above. The fields are validated based on their attribute information to obtain data that passes the validation.
3. The method according to claim 1, characterized in that, The molecular dimension model includes: a molecular basic property dimension table and one or more molecular attribute dimension tables. The molecular attribute dimension tables include: a computational attribute dimension table and an experimental attribute dimension table. The molecular basic property dimension table contains an identifier field for each molecular attribute dimension table.
4. A method for applying molecular data, characterized in that, The method includes: Receive calculation methods submitted by users via API; The computational data related to the computational method is acquired, including computational data obtained from a database and / or files; the database stores molecular data and its structured incremental ancillary data, the files store semi-structured incremental ancillary data and / or unstructured incremental ancillary data of the molecular data, and the database contains a directory index of the files; wherein, any one of the structured incremental ancillary data, the semi-structured incremental ancillary data, and the unstructured incremental ancillary data includes data other than the basic data of the molecule itself related to the molecular data; the computational data includes molecular state data and specific information data corresponding to the molecular state; The calculation data is calculated using the aforementioned calculation method to obtain the calculation result; The calculation results are saved to the database and / or the file.
5. The method according to claim 4, characterized in that, The calculation method includes any one or more of the following: quantum chemistry algorithm, computational chemistry algorithm, and AI model algorithm.
6. The method according to claim 4, characterized in that, The calculation data also includes: calculation data input by the user.
7. The method according to claim 4, characterized in that, The calculation of the data using the aforementioned calculation method to obtain the calculation result includes: The calculation method described above is used to perform local calculations on part or all of the data to obtain the calculation results; and / or The calculation method and some or all of the data are submitted to a remote cluster server for calculation, and the calculation results returned by the remote cluster server are received.
8. The method according to any one of claims 4 to 7, characterized in that, The method further includes: Receive query information submitted by users via API; Data is read from the database based on the query information; Display the read data.
9. The method according to claim 8, characterized in that, The query information includes any one or more of the following: molecular substructure, molecular similarity, and molecular attribute parameters.
10. The method according to claim 8, characterized in that, The step of reading data from the database based on the query information includes: A query strategy is generated based on the query information, and data is read from the database according to the query strategy; The read data is formatted.
11. A molecular data storage device, characterized in that, The device includes: a data receiving module, an analysis module, a data aggregation module, and a data management module; The data receiving module is used to receive molecular data to be processed; The analysis module is used to verify the molecular data and obtain data that passes verification; including: determining the dimensions of the molecular data according to a molecular dimension model; and verifying the molecular data according to the dimensions to obtain data that passes verification. The data aggregation module is used to determine the incremental auxiliary data of the data that has passed the verification. The incremental auxiliary data includes any one or more of the following: structured incremental auxiliary data, semi-structured incremental auxiliary data, and unstructured incremental auxiliary data. The incremental auxiliary data includes data other than the basic data of the molecule itself related to the molecular data, including molecular state data and specific information data corresponding to the molecular state. The data management module is used to save the verified data and the structured incremental auxiliary data to the database; save the semi-structured incremental auxiliary data and / or the unstructured incremental auxiliary data to a file, and add the directory index of the file to the database to establish the association between the file and the database.
12. A molecular data application device, characterized in that, The device includes: The application interface module is used to receive calculation methods submitted by users; A computational data acquisition module is used to acquire computational data related to the computational method. The computational data includes data acquired from a database and / or files. The database stores molecular data and its structured incremental ancillary data, and the files store semi-structured incremental ancillary data and / or unstructured incremental ancillary data of the molecular data. The database also includes a directory index for the files. Any one of the structured incremental ancillary data, the semi-structured incremental ancillary data, and the unstructured incremental ancillary data includes data other than the basic data of the molecule itself related to the molecular data. The computational data includes molecular state data and specific information data corresponding to the molecular states. The calculation processing module is used to perform calculations on the calculation data using the calculation method, obtain calculation results, and save the calculation results to the database and / or the file.
13. An electronic device, characterized in that, include: processor; as well as A memory having executable code stored thereon, which, when executed by the processor, causes the processor to perform the method as described in any one of claims 1-3, or to perform the method as described in any one of claims 4-10.
14. A computer-readable storage medium, characterized in that, It stores executable code that, when executed by a processor of an electronic device, causes the processor to perform the method as described in any one of claims 1-3, or the method as described in any one of claims 4-10.
15. A computer program product, characterized in that, Includes executable code, which, when executed by a processor, implements the method according to any one of claims 1-3, or implements the method according to any one of claims 4-10.
Citation Information
Patent Citations
Interactive database system for whole drug virtual screening process
CN110827927A
Systems, methods, and apparatus for processing documents to identify structures
US20110276589A1