Database construction method and device, computer program product and electronic equipment
Through the database construction method of splitting and correlation analysis, the problems of low efficiency and insufficient accuracy of database construction in the existing technology are solved, and efficient and accurate material database construction is achieved to ensure that the target molecules match the material attributes.
Patent Information
- Application Number
- CN202510397105.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-04
AI Technical Summary
The prior art consumes a lot of manpower and time when building material-related databases, resulting in inefficient database construction and insufficient accuracy, and easy introduction of unrelated molecules, affecting the reliability of subsequent material attribute prediction and molecular design.
By constructing the initial database, using preset splitting strategies to split the chemical structure, and combining reference attribute information for correlation analysis, the target molecules are generated to build the target database, ensuring that the database information is highly correlated with the target material attributes, and avoiding the inefficiency and subjectivity of manual splitting.
It improves the efficiency and accuracy of database construction, ensures that the generated target molecules have clear attribute characteristics, avoids the disconnection between molecular design and actual needs, and realizes the rapid screening of substructures related to the attributes of the target material.
Smart Images

Figure CN120260738A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technologies, and more particularly, to a method for constructing a database, a device for constructing a database, a computer program product, and an electronic device. Background Art
[0002] In scenarios such as material property prediction, reverse molecular design, and synthesis prediction, accurately constructing a required material-related database is of great significance for tasks in the above scenarios. Currently, it is necessary to manually analyze target materials to construct a database, which not only consumes a great deal of manpower and time, but also introduces a large number of irrelevant molecules into the constructed database, affecting the accuracy and effectiveness of database construction.
[0003] It should be noted that the information disclosed in the above background art is only used to enhance the understanding of the background of the present disclosure, and thus may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0004] An object of the present disclosure is to provide a method and device for constructing a database, a computer program product, and an electronic device, so as to at least to a certain extent improve the construction efficiency and accuracy of a material database.
[0005] Other features and advantages of the present disclosure will become apparent through the following detailed description, or will be partially learned through the practice of the present disclosure.
[0006] According to one aspect of the present disclosure, there is provided a method for constructing a database, including: constructing an initial database according to the acquisition data corresponding to a target material, where the initial database includes reference attribute information and a chemical structure corresponding to the reference attribute information, and the reference attribute information is an associated attribute of the target material; splitting the chemical structure in the initial database through a preset splitting strategy to obtain a plurality of sub-structures, and performing a correlation analysis according to the reference attribute information corresponding to the chemical structure and the plurality of sub-structures to obtain the corresponding relationship between the sub-structures and the material properties; according to the target material and the corresponding relationship, generating a target molecule by combining the sub-structures and determining the molecular properties of the target molecule, so as to construct a target database according to the target molecule and the molecular properties.
[0007] In an exemplary embodiment of the present disclosure, constructing an initial database according to the acquisition data corresponding to a target material includes: performing data preprocessing on the acquisition data to obtain a preprocessing result; performing entity recognition on the preprocessing result based on the target material attributes to obtain attribute entities, and determining reference attribute information according to the attribute entities; performing chemical structure detection on the preprocessing result to obtain a chemical structure; and constructing an initial database according to the reference attribute information and the chemical structure based on the matching relationship between the reference attribute information and the chemical structure.
[0008] In an exemplary embodiment of the present disclosure, entity recognition is performed on the preprocessing result based on the target material attributes to obtain attribute entities, and reference attribute information is determined according to the attribute entities, including: collecting a keyword set corresponding to the target material attributes, and performing data augmentation on the keyword set to obtain a material attribute dictionary; extracting attribute entities from the preprocessing text based on the material attribute dictionary; extracting the entity relationship of the attribute entities according to the target material, and semantically understanding the original text information and entity relationship corresponding to the attribute entities respectively to obtain a first understanding result and a second understanding result; determining reference attribute information based on the first understanding result and the second understanding result according to the entity relationship and the attribute entities.
[0009] In an exemplary embodiment of the present disclosure, based on the matching relationship between the reference attribute information and the chemical structure, an initial database is constructed according to the reference attribute information and the chemical structure, including: performing a structured representation on the chemical structure to obtain structured information; performing relationship recognition on the structured information and the reference attribute information based on a pre-trained relationship recognition model to obtain a relationship recognition result, and the pre-trained relationship recognition model is trained based on sample data of chemical structures and attribute descriptions; in response to the relationship recognition result indicating a match between the structured information and the reference attribute information, taking the reference attribute information and the chemical structure as a data pair to construct an initial database according to the obtained data pair.
[0010] In an exemplary embodiment of the present disclosure, the chemical structure in the initial database is split through a preset splitting strategy to obtain a plurality of substructures, including: performing chemical bond recognition on the chemical structure, and if a target chemical bond is recognized, splitting the chemical structure at the position of the target chemical bond to obtain a first splitting result; determining the substructure corresponding to the chemical structure according to the first splitting result.
[0011] In an exemplary embodiment of the present disclosure, determining the substructure corresponding to the chemical structure according to the first splitting result further includes: representing the chemical structure as a structure diagram with atoms as nodes and chemical bonds as edges, performing subgraph mining on the structure diagram to obtain a specific functional structure, and splitting the specific functional structure from the chemical structure to obtain a second splitting result; determining the substructure of the chemical structure according to the first splitting result and the second splitting result.
[0012] In an exemplary embodiment of the present disclosure, performing subgraph mining on the structure diagram to obtain a specific functional structure includes: adding chemical attribute information to the nodes and edges in the structure diagram to obtain a functional structure diagram; obtaining functional requirement information, and performing subgraph mining from the functional structure diagram based on the functional requirement information to obtain a specific functional structure.
[0013] In an exemplary embodiment of the present disclosure, a correlation analysis is performed based on the reference attribute information corresponding to the chemical structure and multiple substructures to obtain the correspondence between the substructures and the material properties, including: performing a statistical analysis based on the reference attribute information corresponding to the chemical structure and the substructures to obtain a first correlation result between the substructures and the material properties; determining the correspondence between the substructures and the material properties according to the first correlation result.
[0014] In an exemplary embodiment of the present disclosure, determining the correspondence between the substructures and the material properties according to the first correlation result further includes: performing a correlation analysis on the reference attribute information corresponding to the chemical structure and the substructures through a pre-trained correlation analysis model to obtain a second correlation result; performing a consistency check on the first correlation result and the second correlation result, and determining the correspondence from the first correlation result and the second correlation result according to the consistency check result; or, performing a weighted fusion of the first correlation result and the second correlation result according to a preset reliability weight to obtain a fusion result, and determining the correspondence according to the fusion result.
[0015] In an exemplary embodiment of the present disclosure, performing a statistical analysis based on the reference attribute information corresponding to the chemical structure and the substructures to obtain a first correlation result between the substructures and the material properties includes: for each molecular structure, statistically analyzing the distribution of at least one material property of the chemical structures containing the substructure; determining the first correlation result between the substructures and the material properties based on the distribution of at least one material property.
[0016] In an exemplary embodiment of the present disclosure, according to the target material and the correspondence, using the substructures to combine and generate target molecules and determining the molecular properties of the target molecules, so as to construct a target database according to the target molecules and the molecular properties, includes: based on the correspondence, screening candidate substructures related to the properties of the target material from the substructures; combining the candidate substructures so that the obtained candidate molecules meet the preset combination rules; predicting the properties of the candidate molecules using a pre-trained property prediction model to obtain the molecular properties of the candidate molecules; determining the target molecules from the candidate molecules according to the properties of the target material and the molecular properties of the candidate molecules, and constructing a target database according to the target molecules and the molecular properties corresponding to the target molecules.
[0017] In an exemplary embodiment of the present disclosure, the preset combination rules include at least one of a chemical feasibility rule, a connection mode rule, and a reaction mechanism rule.
[0018] According to an aspect of the present disclosure, there is provided a database construction device, including: a first construction module for constructing an initial database according to the acquisition data corresponding to the target material, the initial database including reference attribute information and the chemical structure corresponding to the reference attribute information, and the reference attribute information being the associated attribute of the target material; a structure splitting module for splitting the chemical structure in the initial database by a preset splitting strategy to obtain a plurality of sub-structures, and performing a correlation analysis based on the reference attribute information corresponding to the chemical structure and the plurality of sub-structures to obtain the correspondence between the sub-structures and the material attributes; a second construction module for generating a target molecule by combining the sub-structures according to the target material and the correspondence, and determining the molecular attributes of the target molecule, so as to construct a target database according to the target molecule and the molecular attributes.
[0019] According to an aspect of the present disclosure, there is provided a computer program product including a computer program, which when executed by a processor implements the method of any one of the above.
[0020] According to an aspect of the present disclosure, there is provided an electronic device, including: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to execute the method of any one of the above by executing the executable instructions.
[0021] The database construction method in the exemplary embodiments of the present disclosure, on the one hand, constructs an initial database including reference attribute information and the chemical structure corresponding to the reference attribute information based on the acquisition data corresponding to the target material, where the reference attribute information is the associated attribute of the target material. When constructing the initial database, the attributes of the target material are directly associated, ensuring that the information in the database is highly relevant to the actual attributes of the target material, avoiding the problem of possible mismatch between attributes and materials in the database, and improving the accuracy and practicality of the data; on the other hand, by presetting a splitting strategy to split the chemical structure in the initial database and combining the correlation analysis with the reference attribute information corresponding to the chemical structure, the internal relationship between the sub-structure and the material attributes can be accurately revealed, avoiding the problem of only relying on the overall chemical structure and ignoring local features, thereby improving the accuracy of the database; on the further hand, according to the target material and the corresponding relationship, sub-structures are used to combine to generate target molecules and the molecular attributes of the target molecules are determined, so as to construct a target database according to the target molecules and the molecular attributes, ensuring that the generated target molecules have clear attribute characteristics, avoiding the problem of disconnection between molecular design and actual needs, and further improving the accuracy of the database. In addition, by presetting a splitting strategy to automatically split the chemical structure, the inefficiency and subjectivity of manual splitting are avoided, significantly improving the efficiency of chemical structure analysis, laying a foundation for subsequent correlation analysis and target molecule generation. At the same time, the correlation analysis is used to quickly establish the mapping relationship between the sub-structure and the material attributes, avoiding the inefficient process of a large number of experimental verifications, that is, through a data-driven method, the sub-structures related to the attributes of the target material can be quickly screened out, improving the efficiency of database construction.
[0022] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. Brief Description of the Drawings
[0023] By reading the following detailed description with reference to the accompanying drawings, the above and other objects, features, and advantages of the exemplary embodiments of the present disclosure will become readily understood. In the drawings, several embodiments of the present disclosure are shown in an exemplary but not restrictive manner.
[0024] Figure 1 Shows an application environment according to an exemplary embodiment of the present disclosure.
[0025] Figure 2 Shows a flowchart of a database construction method according to an exemplary embodiment of the present disclosure.
[0026] Figure 3 Shows a schematic diagram of associating a chemical structure with reference attribute information according to an exemplary embodiment of the present disclosure.
[0027] Figure 4The flowchart of an implementation manner for constructing an initial database according to an exemplary embodiment of the present disclosure is shown.
[0028] Figure 5 The flowchart of determining reference attribute information according to an exemplary embodiment of the present disclosure is shown.
[0029] Figure 6 The flowchart of an implementation manner for splitting a chemical structure according to an exemplary embodiment of the present disclosure is shown.
[0030] Figure 7 The flowchart of determining a sub-structure corresponding to a chemical structure according to a first splitting result according to an exemplary embodiment of the present disclosure is shown.
[0031] Figure 8 The flowchart of an implementation manner for performing a correlation analysis on reference attribute information and a sub-structure according to an exemplary embodiment of the present disclosure is shown.
[0032] Figure 9 The flowchart of an implementation manner for constructing a target database according to a generated target molecule according to an exemplary embodiment of the present disclosure is shown.
[0033] Figure 10 The schematic diagram of the composition of a database construction device according to an exemplary embodiment of the present disclosure is shown.
[0034] Figure 11 The block diagram of an electronic device according to an exemplary embodiment of the present disclosure is shown.
[0035] In the drawings, the same or corresponding reference numerals indicate the same or corresponding parts. Detailed implementation manners
[0036] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, the exemplary embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be more complete and comprehensive, and will fully convey the concept of the exemplary embodiments to those skilled in the art. The same reference numerals in the figures denote the same or similar structures, and thus their detailed descriptions will be omitted.
[0037] In addition, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present disclosure. However, those skilled in the art will realize that the technical solutions of the present disclosure can be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. can be adopted. In other cases, well-known structures, methods, devices, implementations, or operations are not shown or described in detail to avoid obscuring aspects of the present disclosure.
[0038] The block diagrams shown in the drawings are only functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or implemented in one or more software-hardened modules, or implemented in different networks and / or processor devices and / or microcontroller devices.
[0039] Currently, the analysis of target materials mainly relies on manual work, including chemical structure analysis, property annotation, data collation, etc. However, in the fields of materials science, drug design, etc., the types and quantities of target materials may be very large. Manual analysis is difficult to meet the needs of large-scale data processing, has high labor costs, and the manual analysis process takes a long time, and the database construction cycle is too long to quickly respond to the urgent needs in scientific research or industrial applications. Moreover, during the manual analysis process, due to subjective judgment or knowledge limitations, a large amount of molecular data irrelevant to the properties of the target material may be introduced. The introduction of irrelevant molecules will reduce the accuracy and effectiveness of the database, affecting the reliability of subsequent tasks such as material property prediction and molecular design. For example, in reverse molecular design or synthesis prediction, irrelevant molecules may lead to incorrect design schemes or synthesis paths, increasing the trial-and-error cost. In addition, it is currently difficult to quickly build a customized database according to the properties of specific target materials, or a large amount of manual intervention is required, further increasing the cost and time consumption.
[0040] Based on this, the exemplary embodiments of the present disclosure provide a database construction method. By constructing an initial database, splitting the chemical structures in the initial database according to a preset splitting strategy, and performing correlation analysis of combined data, etc., target molecules are generated and molecular properties are determined to construct a target database. By using the properties of the target material as a guide throughout the process of constructing the database (structure splitting, molecular combination), a database related to the properties of the target material can be constructed, improving the construction efficiency and accuracy of the material database.
[0041] It should be noted that the database construction method of the exemplary embodiments of the present disclosure can be applied to fields such as material property prediction, reverse molecular design, synthesis prediction, etc., such as high-performance (such as high strength, high thermal conductivity) material design, drug design, etc., and there is no limitation thereto.
[0042] The database construction method provided by the exemplary embodiments of the present disclosure can be applied to an application environment as Figure 1 shown. Among them, the terminal 101 communicates with the server 102 through the network. The data storage system can store the data that the server 102 needs to process. The data storage system can be integrated on the server 102, or placed in the cloud or other network servers.
[0043] In an exemplary embodiment, the database construction method provided by the exemplary embodiment of the present disclosure can be executed by the server 102, and the corresponding database construction device is set in the server 102. Correspondingly, in this manner executed by the server 102, the server 102 can start executing the technical solution in the exemplary embodiment of the present disclosure in response to a trigger command, where the trigger command can be sent by the terminal used by the user, or can be locally triggered by the server in response to some automated events.
[0044] Among them, the server 102 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server 102 can execute background tasks.
[0045] Furthermore, in another exemplary embodiment, the terminal 101 can also have a similar function to the server 102, so as to execute the database construction method provided by the exemplary embodiment of the present disclosure.
[0046] Among them, the terminal 101 can be a smart phone, a tablet computer, a laptop computer, or a desktop computer. The terminal 101 can also be referred to as a mobile terminal, a terminal device, a mobile device, etc. The exemplary embodiment of the present disclosure does not limit the type of the terminal 101.
[0047] In addition, the technical solution of the exemplary embodiment of the present disclosure can also be executed jointly by the terminal 101 and the server 102. In this manner of joint execution by the terminal 101 and the server 102, some steps in the technical solution provided by the exemplary embodiment of the present disclosure are executed by the terminal 101, while some other steps are executed by the server 102. It should be noted that in this manner of joint execution by the terminal 101 and the server 102, the steps respectively executed by the terminal 101 and the server 102 can be dynamically adjusted according to the actual situation, and no special limitation is imposed thereon.
[0048] Among them, the terminal 101 and the server 102 can be directly or indirectly connected through a wireless communication method, and the exemplary embodiment of the present disclosure does not impose special limitations thereon.
[0049] As Figure 2 shown is a flowchart of a database construction method according to an exemplary embodiment of the present disclosure. The database construction method includes steps S210 to S230, which are specifically as follows:
[0050] In step S210, an initial database is constructed based on the acquisition data corresponding to the target material. The initial database includes reference attribute information and the chemical structure corresponding to the reference attribute information, and the reference attribute information is the associated attribute of the target material.
[0051] In an exemplary embodiment of the present disclosure, before starting to collect data, it is necessary to clarify the material scope to be analyzed, which is the basis for subsequent data collection and analysis. Among them, the target material is a specific material that needs to be analyzed, designed, and optimized, with clear chemical composition, structural characteristics, and functional attributes, and can meet the requirements of specific application scenarios. The exemplary embodiment of the present disclosure indicates the material scope to be analyzed through the target material.
[0052] For example, the target material to be analyzed can be determined according to application requirements, such as functional materials (such as blue phosphorescent materials), structural materials (such as high-strength alloys), and biological materials (such as biocompatible materials), etc. The exemplary embodiment of the present disclosure is described with the target material being a blue phosphorescent material. For example, a blue phosphorescent material is composed of organometallic complexes or organic small molecules, and its core structure contains heavy metal atoms (such as iridium, platinum) and organic ligands, has a specific molecular structure, can achieve blue phosphorescence emission through energy transfer in the excited state (such as ligand structure, metal center), has a luminescence wavelength in the range of 450 - 490 nm, exhibits blue phosphorescence, a thermal analysis temperature higher than 300 °C, and other properties.
[0053] Among them, after determining the target material, relevant data related to the target material can be collected through various tools and methods to obtain the acquisition data. Specifically, a crawler tool or API (Application Programming Interface) can be used to retrieve literature related to the target material from scientific literature databases (such as ScienceDirect, etc.), or patent literature related to the target material can be retrieved from a patent retrieval platform. Information such as the chemical structure, molecular formula, and molecular weight of the target material can also be obtained from chemical databases (such as ChemSpider, etc.). Of course, if there is existing experimental data, it can also be included in the acquisition scope. The exemplary embodiment of the present disclosure includes, but is not limited to, the above ways of obtaining the acquisition data corresponding to the target material.
[0054] After obtaining the collected data, since the formats of the collected data are diverse, including web files, PDF files, etc., directly extracting the text may cause partial information loss or format chaos. Moreover, the collected data may contain special fonts, symbols, or layouts, and directly extracting the text may lead to a decrease in readability. Therefore, the collected data can also be saved as an image file. For example, chemical structure diagrams, spectrograms, experimental data tables, etc. are usually presented in image form. Saving as an image can ensure the integrity of these key information, and saving as an image can support the integration and analysis of multi-modal data, providing accurate and complete data support for subsequent processing. Of course, for other text-based collected data, the form of the text file can also be retained, or the collected data can be preprocessed, including but not limited to deduplication, format unification, missing value processing, etc., without specific restrictions on this.
[0055] Furthermore, after preprocessing the collected data, the associated attributes of the target material, that is, the reference attribute information, can be extracted from the collected data, and the chemical structure corresponding to the reference attribute information can be extracted to construct an initial database. The reference attribute information refers to the material attributes related to the attributes of the target material in the collected data. For example, OCR (Optical Character Recognition) is used to recognize the text information of the image file to obtain the reference attribute information. For example, the CTPN (Connectionist Text Proposal Network, a deep model for natural scene text detection) model is used for information detection, and then the CRNN (Convolutional Recurrent Neural Network) model is used to recognize based on the detected information to obtain the reference attribute information. Key information can also be extracted through a named entity recognition model to obtain the reference attribute information, and the chemical structure can be obtained from the collected data through a molecular structure diagram detection model. For example, a pre-trained YOLO (such as YoloV8) model can be fine-tuned with sample data labeled with molecular structures to obtain a model to identify the molecular structure diagram and obtain the chemical structure. Furthermore, the obtained chemical structure can be associated with its corresponding reference attribute information to construct an initial database, providing accurate and comprehensive data support for constructing the target database. Exemplarily, as Figure 3 shown in the collected data, the chemical structure in Figure 1-1 is associated with its corresponding reference attribute information (i.e., the property description in the text information).
[0056] In step S220, the chemical structure in the initial database is split through a preset splitting strategy to obtain multiple sub-structures, and a correlation analysis is performed based on the reference attribute information corresponding to the chemical structure and the multiple sub-structures to obtain the corresponding relationship between the sub-structures and the material attributes.
[0057] In an exemplary embodiment of the present disclosure, the purpose of the preset splitting strategy is to decompose a complex chemical structure into multiple sub-structures with specific functions (such as functional groups, molecular fragments, etc.). It may include splitting based on chemical bonds, splitting based on functional groups, splitting based on molecular fragments, etc. Among them, the chemical structure can be split based on graph theory algorithms and machine learning models. For example, the chemical structure of a blue phosphorescent material is split into a metal center (such as iridium, platinum) and an organic ligand (such as phenylpyridine, phenylquinoline), and a polymer material is split into repeating units and end groups, and so on.
[0058] Among them, the material property refers to one or some properties in the reference property information (such as thermal stability, optical properties). Statistical analysis methods, machine learning methods or graph analysis methods can be used for correlation analysis according to the reference property information corresponding to the chemical structure and multiple sub-structures, so as to associate the sub-structures with the material properties (such as optical properties, thermal stability, chemical stability, etc.) in the reference property information, and obtain the corresponding relationship between the sub-structures and the material properties, providing a basis for subsequent material design and optimization. For example, it is analyzed that a certain sub-structure (such as phenylpyridine ligand) is positively correlated with high quantum efficiency, and a certain sub-structure (such as platinum metal center) is positively correlated with high thermal stability, etc.
[0059] Optionally, before splitting the chemical structures in the initial database, the chemical structures in the initial database can be standardized to ensure that the representation methods of the chemical structures are consistent (such as SMILES (Simplified molecular input line entry system) format, molecular graph format).
[0060] Through the preset splitting strategy and correlation analysis, the relationship between the sub-structures and the material properties can be accurately mined, providing data support for quickly designing new materials that meet the target material properties.
[0061] In step S230, according to the target material and the corresponding relationship, the sub-structures are used for combination to generate a target molecule and the molecular properties of the target molecule are determined, so as to construct a target database according to the target molecule and the molecular properties.
[0062] In an exemplary embodiment of the present disclosure, the generation of the target molecule is a process of designing a molecule that meets the properties of the target material by combining sub-structures based on the corresponding relationship between the sub-structures and the material properties. Among them, when generating the target molecule, it is necessary to make it meet the molecular design rules, that is, the preset combination rules, including at least one of chemical feasibility rules, connection mode rules and reaction mechanism rules. That is to say, the target molecule meets both the properties of the target material and the requirements of chemical rationality.
[0063] Among them, through deep learning models, genetic algorithms, etc., substructures can be used for combination to generate target molecules and determine the molecular properties of the target molecules, so as to construct a target database. Among them, after generating the target molecules, by determining their molecular properties, it is also possible to verify whether they meet the requirements of the target material properties.
[0064] For the database construction method in the exemplary embodiments of the present disclosure, on the one hand, an initial database including reference attribute information and the chemical structure corresponding to the reference attribute information is constructed according to the acquisition data corresponding to the target material, where the reference attribute information is the associated attribute of the target material. When constructing the initial database, the attributes of the target material are directly associated, ensuring that the information in the database is highly relevant to the actual attributes of the target material, avoiding the problem of possible mismatch between the attributes and the material in the database, and improving the accuracy and practicality of the data; on the other hand, by presetting a splitting strategy to split the chemical structures in the initial database and combining the reference attribute information corresponding to the chemical structures for correlation analysis, the internal relationship between the substructures and the material properties can be accurately revealed, avoiding the problem of only relying on the overall chemical structure and ignoring local features, thereby improving the accuracy of the database; on the other hand, according to the target material and the corresponding relationship, substructures are used for combination to generate target molecules and determine the molecular properties of the target molecules, so as to construct a target database according to the target molecules and the molecular properties, ensuring that the generated target molecules have clear attribute characteristics, avoiding the problem of disconnection between molecular design and actual needs, and further improving the accuracy of the database. In addition, by presetting a splitting strategy to automatically split the chemical structures, the inefficiency and subjectivity of manual splitting are avoided, significantly improving the efficiency of chemical structure analysis, laying a foundation for subsequent correlation analysis and target molecule generation. At the same time, the correlation analysis is used to quickly establish the mapping relationship between the substructures and the material properties, avoiding the inefficient process of a large number of experimental verifications, that is, through a data-driven method, the substructures related to the properties of the target material can be quickly screened out, improving the efficiency of database construction.
[0065] The above steps will be described in detail below.
[0066] In an exemplary embodiment, an implementation manner for constructing an initial database is provided. As Figure 4 shown, constructing an initial database according to the acquisition data corresponding to the target material may include:
[0067] Step S410: Perform data preprocessing on the acquisition data to obtain a preprocessing result.
[0068] As can be seen above, performing data preprocessing on the acquisition data may include operations such as cleaning and formatting the acquisition data, and no excessive limitations are imposed on this.
[0069] Step S420: Perform entity recognition on the preprocessing result based on the attributes of the target material to obtain attribute entities, and determine reference attribute information according to the attribute entities.
[0070] Entity recognition is a process of extracting key information related to the attributes of the target material from the preprocessing result and determining reference attribute information. Among them, named entity recognition technology, predefined rules (such as regular expressions), pre-trained machine learning models (such as BERT (Bidirectional Encoder Representations from Transformers, a language representation model)), pre-trained large language models, etc. can be used to perform entity recognition on the preprocessing result.
[0071] Step S430: Perform chemical structure detection on the preprocessing result to obtain a chemical structure.
[0072] The purpose of performing chemical structure detection on the preprocessing result is to identify and extract chemical structure information in the preprocessing result. Optionally, for text data, a chemical structure dictionary can be used to match chemical terms or molecular formulas in the text to extract the chemical structure. Optionally, regular expressions can be used to match chemical formulas or SMILES strings to extract the chemical structure. Optionally, a pre-trained large language model can also be used to identify chemical structure information in the text. Optionally, a pre-trained molecular structure diagram detection model can also be used to locate the position of the two-dimensional molecular structure diagram and perform structure extraction to obtain the chemical structure. It should be noted that the training method of the pre-trained model involved in the exemplary embodiments of the present disclosure is the same as the conventional model training method, and the present disclosure does not specifically limit the training process and method of the pre-trained model, and will not be repeated below.
[0073] Step S440: Based on the matching relationship between the reference attribute information and the chemical structure, construct an initial database according to the reference attribute information and the chemical structure.
[0074] Based on the matching relationship between the reference attribute information and the chemical structure, the correlation between the chemical structure and the reference attribute information can be reflected, and unnecessary chemical structures can be avoided from being introduced.
[0075] In some optional embodiments, the chemical structure can be structurally represented to obtain structural information, and then a pre-trained relationship recognition model is used to perform relationship recognition on the structural information and the reference attribute information to obtain a relationship recognition result. In response to the relationship recognition result indicating that the structural information and the reference attribute information match, the reference attribute information and the chemical structure are used as a data pair to construct an initial database according to the obtained data pair. Among them, the pre-trained relationship recognition model is trained based on sample data of chemical structure and attribute description.
[0076] Among them, by structuring the chemical structure, the chemical structure can be converted into a SMILES string to accurately represent the chemical structure. The pre-trained relationship recognition model is trained based on a large number of sample data of chemical structures and property descriptions, and can accurately identify the association relationship between the chemical structure and the property information, ensuring the accuracy of the matching.
[0077] When it is recognized that the structured information and the reference property information match, it indicates that the reference property information is a description of the corresponding chemical structure, thereby establishing a connection between the reference property information and the chemical structure.
[0078] Through the structure representation and the pre-trained model, the matching of the chemical structure and the property information can be automatically completed, reducing manual intervention, improving the efficiency of database construction, and the data pairs generated by the structure representation and the relationship recognition model have clear structures and semantics, which can improve the interpretability and reusability of the data.
[0079] In an exemplary embodiment, as Figure 5 shown is a flowchart for determining reference property information, and the process includes:
[0080] Step S510: Collect a keyword set corresponding to the properties of the target material, and perform data augmentation on the keyword set to obtain a material property dictionary.
[0081] For the target material properties, a keyword set related to them can be collected. For example, it can be extracted from existing material databases, literature, patent documents, or relevant technical materials, so that the keyword set can cover all aspects of the target material properties, including but not limited to physical properties, chemical properties, mechanical properties, etc.
[0082] Among them, data augmentation refers to expanding and optimizing the initially collected keyword set through a series of technical means to improve its coverage, accuracy, and applicability. The methods of data augmentation include but are not limited to synonym expansion, inflection expansion, domain term expansion, etc. For example, through a synonym dictionary or natural language processing tools, synonyms can be added to each keyword to expand the coverage of the keyword set. For example, for "hardness", "strength", "high strength", "strength value", "strength", etc. can be further expanded. Through data augmentation, the finally obtained material property dictionary can better match the diverse expression forms in the text, thereby improving the effect of subsequent entity recognition and property extraction.
[0083] Step S520: Extract property entities from the preprocessed text based on the material property dictionary.
[0084] A "property entity" refers to a specific information unit extracted from the preprocessed text that is related to the properties of the target material, that is, a property entity is a noun, phrase, or term that is clearly mentioned in the preprocessed text and can characterize the material properties.
[0085] Specifically, the vocabulary in the preprocessed text can be first matched with the keywords in the material property dictionary to identify entities related to the properties of the target material. Among them, exact matching or fuzzy matching can be used during the matching process to ensure that different forms of expressions can be identified. At the same time, during the matching process, the context information can also be combined to determine whether the vocabulary actually represents a material property. For example, some vocabulary may have different meanings in a specific context, and through context analysis, mis-matching situations can be excluded.
[0086] As an example, the material property dictionary can be loaded into memory and an efficient lookup structure (such as a hash table) can be formed. Then, the preprocessed text can be segmented and part-of-speech tagged to generate a sequence of vocabulary. For example, if the original preprocessed text is "This material has high hardness and excellent tensile strength", the segmentation result is "This", "material", "has", "high hardness", "and", "excellent", "tensile strength", ".". Next, each vocabulary in the preprocessed text can be traversed and exactly matched with the keywords in the material property dictionary. If the vocabulary exactly matches a certain keyword in the dictionary, the vocabulary is marked as an attribute entity. Of course, for the vocabulary that fails to pass the exact matching, partial matching or fuzzy matching can be attempted. For example, match the prefix, suffix, or root of the vocabulary, or perform a synonym match (such as "rigidity" matches "hardness"), then "hardness" is identified as an attribute entity.
[0087] Step S530: According to the target material, extract the entity relationships of the attribute entities, and perform semantic understanding on the original text information and entity relationships corresponding to the attribute entities respectively to obtain a first understanding result and a second understanding result.
[0088] The entity relationship refers to the relevance between attribute entities, which is manifested as the interaction or dependence relationship between attribute entities in the text. For example, the relationship between the hardness and strength of a certain material, such as a causal relationship (such as high hardness leads to enhanced wear resistance), a comparison relationship (such as the strength of a material is higher than that of another material), a composition relationship (an alloy is composed of iron, carbon, and chromium), etc. Among them, based on the target material, a pre-trained relationship extraction model can be used to identify the entity relationships of the attribute entities, or the syntactic dependency relationships between the vocabulary in the preprocessed text can also be analyzed (such as using a syntactic analysis tool) to determine the relevance between the attribute entities. The exemplary embodiments of the present disclosure do not limit the manner of extracting the entity relationships of the attribute entities.
[0089] The first understanding result is the result of deeply understanding the specific meaning and context information of the original text where the attribute entity is located. A pre-trained language model (such as BERT) can be used to process the original text information to capture its semantic information, so as to obtain the second understanding result. The second understanding result is the result of deeply understanding the specific meaning and relevance of the relationships between attribute entities. A pre-trained relationship classification model can be used to classify the entity relationships to determine the relationship classification result, and the specific meaning of the entity relationships can be supplemented by combining the context information of the attribute entities (such as pre-processed text) to obtain the second understanding result.
[0090] Step S540: Based on the first understanding result and the second understanding result, determine the reference attribute information according to the entity relationships and attribute entities.
[0091] The accuracy of the extracted entity relationships and attribute entities can be verified by comparing the first understanding result and the second understanding result. Among them, it can be checked whether the first understanding result and the second understanding result are semantically consistent, and whether the entity relationships conform to logic or domain knowledge. For example, the first understanding result is "High hardness means that the hardness value of the material is significantly higher than the conventional level", and the second understanding result is "High hardness leads to a significant increase in wear resistance", and the two are semantically consistent. Another example is that through the second understanding result, it is determined that there is a decisive relationship between "tensile strength" and "heat treatment process", and according to material science knowledge, the heat treatment process can indeed affect the tensile strength, and the logic is consistent. Based on this, the accuracy of the extracted entity relationships and attribute entities can be ensured. Further, the attribute entities and entity relationships can be integrated into reference attribute information.
[0092] The exemplary embodiments of the present disclosure extract attribute entities and entity relationships through a material attribute dictionary, and ensure the accuracy of the extracted attribute entities and entity relationships by semantically understanding the original text information and entity relationships respectively, thereby improving the accuracy of the reference attribute information.
[0093] In an exemplary embodiment, an implementation manner of splitting a chemical structure is provided, such as Figure 6 shown, splitting the chemical structure in the initial database through a preset splitting strategy to obtain multiple sub-structures may include:
[0094] Step S610: Identify chemical bonds in the chemical structure. If a target chemical bond is identified, split the chemical structure at the position of the target chemical bond to obtain a first splitting result.
[0095] Among them, the target chemical bond can be a pre-specified chemical bond type, that is, the chemical bond to be split, such as an easily breakable chemical bond, a chemical bond of a specific functional group, etc. For chemical structure chemical bond recognition, chemical structure analysis tools or pre-trained models can be used to identify the chemical bond types in the chemical structure (such as single bonds, double bonds, triple bonds, aromatic bonds, etc.). The chemical structure can be broken at the position of the target chemical bond to generate two or more sub-structures to obtain the first splitting result. Similarly, the broken sub-structures can be represented in a structured form, such as SMILES, molecular graphs, etc., and there is no limitation on this.
[0096] Step S620: Determine the sub-structures corresponding to the chemical structure according to the first splitting result.
[0097] After obtaining the first splitting result, the sub-structures corresponding to the chemical structure can be determined according to the first splitting structure, and the association relationship between the sub-structures and the original chemical structure can also be recorded. For example, if the original chemical structure is "C=O", the sub-structure 1 obtained after splitting is "C" and the sub-structure 2 is "O", then the sub-structure 1, sub-structure 2 and their original chemical structure are stored in an associated manner for subsequent analysis and use.
[0098] Through steps such as chemical bond recognition, chemical structure splitting, and sub-structure determination, the splitting of the chemical structure and the extraction of sub-structures are realized, which can provide data support for flexibly analyzing the relationship between its sub-structures and material properties, so as to support more complex query and analysis requirements.
[0099] Based on the foregoing exemplary embodiments, as Figure 7 shown, determining the sub-structures corresponding to the chemical structure according to the first splitting result may further include:
[0100] Step S710: Represent the chemical structure as a structure diagram with atoms as nodes and chemical bonds as edges, then perform sub-graph mining according to the structure diagram to obtain a specific functional structure, and split the specific functional structure from the chemical structure to obtain a second splitting result.
[0101] Represent the chemical structure as a structure diagram with atoms as nodes and chemical bonds as edges. This process is to represent the chemical structure as a graph structure for sub-graph mining. Among them, each atom in the chemical structure can be represented as a node in the graph. For example, for the chemical structure "C=O", the corresponding nodes are C and O. The chemical bonds in the chemical structure are represented as edges in the graph. For example, the double bond in "C=O" is the edge between nodes C and O (edge C-O), thus representing the chemical structure as a structure diagram.
[0102] Among them, the specific functional structure refers to the sub-graph pattern to be mined, such as specific functional groups, ring structures, chain structures, etc. The sub-graph matching algorithm can be used to match the structure with specific functions in the structure diagram. For example, the Ullmann algorithm (enumeration algorithm) can be used, and there is no limitation in this regard. Furthermore, the specific functional structure can be split from the chemical structure to obtain the second splitting result.
[0103] Furthermore, sub-graph mining based on the structure diagram to obtain the specific functional structure further includes:
[0104] First, chemical attribute information is added to the nodes and edges in the structure diagram to obtain the functional structure diagram; then, the functional requirement information is obtained, and sub-graph mining is performed from the functional structure diagram based on the functional requirement information to obtain the specific functional structure.
[0105] Among them, adding chemical attribute information to the nodes and edges in the structure diagram to obtain the functional structure can add chemical attribute information to each atomic node in the structure diagram, such as atomic type, charge, hybridization state, etc. Chemical attribute information can be added to each chemical bond edge in the structure diagram, such as bond type, bond length, bond energy, etc. The functional requirement information refers to the functional requirement information related to the target material attributes that guides sub-graph mining, expressing specific requirements for the material function. For example, it is necessary to obtain sub-structures with specific pharmacophore groups, catalytic sites or reaction activities. As an example, specific functional structures to be mined can be defined according to the target material attributes, such as functional groups, ring structures, chain structures, etc.
[0106] The functional requirement information is represented as a graph structure pattern, so that the specific functional structure can be mined from the functional structure diagram based on the functional requirement information. Among them, the sub-graph matching algorithm can be used to match the functional requirement graph corresponding to the functional requirement information in the functional structure diagram, so as to extract the matched specific functional structure from the functional structure diagram to obtain the specific functional structure.
[0107] By adding chemical attribute information, the specific functional structure can be identified more accurately. Guiding sub-graph mining through functional requirement information can meet the personalized needs of users for specific functional structures, thus supporting the material task requirements based on functional structures.
[0108] Step S720: Determine the sub-structure of the chemical structure according to the first splitting result and the second splitting result.
[0109] After obtaining the first splitting result and the second splitting result, the sub-structures in the first splitting result and the second splitting result can be integrated, removing duplicate sub-structures while enriching the types of sub-structures of the chemical structure.
[0110] In an exemplary embodiment, an implementation manner for performing correlation analysis on the reference attribute information and the sub-structure is provided. Such asFigure 8 As shown, perform a correlation analysis based on the reference attribute information corresponding to the chemical structure and the sub-structure to obtain the correspondence between the sub-structure and the material attribute, including:
[0111] Step S810: Perform a statistical analysis based on the reference attribute information corresponding to the chemical structure and the sub-structure to obtain the first correlation result between the sub-structure and the material attribute.
[0112] Step S820: Determine the correspondence between the sub-structure and the material attribute according to the first correlation result.
[0113] Among them, for each molecular structure, the distribution of at least one material attribute of the chemical structure containing the sub-structure can be statistically analyzed, and then based on the distribution of at least one material attribute, the first correlation result between the sub-structure and the material attribute can be determined. The relationship between the sub-structure and the material attribute can be quantified by statistical methods, such as correlation coefficient analysis, regression analysis, etc.
[0114] For example, the reference attribute information corresponding to the chemical structure corresponds to 100 sub-structures obtained by splitting. The attribute values of the molecules containing each sub-structure can be statistically analyzed. For example, for 200 molecules containing sub-structure 1, the mean value of the Tg (glass transition temperature) is 400 degrees, and the standard deviation is 10 degrees; for 150 molecules containing sub-structure 2, the mean value of the Tg is 200 degrees, and the standard deviation is 20 degrees. Then relatively speaking, sub-structure 1 has a higher correlation with the Tg attribute (because the variance is small, which proves more stable), and it is more relevant to the Tg material with a higher temperature; sub-structure 2 has a lower correlation with the Tg attribute and is more relevant to the Tg with a lower temperature.
[0115] Based on the foregoing exemplary embodiments, determining the correspondence between the sub-structure and the material attribute according to the first correlation result may further include:
[0116] First, perform a correlation analysis on the reference attribute information corresponding to the chemical structure and the sub-structure through a pre-trained correlation analysis model to obtain a second correlation result; then perform a consistency check on the first correlation result and the second correlation result, and determine the correspondence from the first correlation result and the second correlation result according to the consistency check result.
[0117] Among them, the pre-trained correlation analysis model is trained according to the sample data of the attribute information and its sub-structure. The exemplary embodiments of the present disclosure do not limit the network structure of the pre-trained correlation analysis model, etc.
[0118] Consistency checking refers to comparing the correlation results obtained by two methods to check whether there is consistency between them. If both methods indicate a strong correlation between certain molecular structures and material properties, then the correlation results may have a high degree of credibility. That is to say, if there is consistency between the first correlation result and the second correlation result, the corresponding relationship between the substructure and the material property can be determined based on the two results. For example, a certain molecular structure (such as a platinum metal center) is positively correlated with high thermal stability.
[0119] Among them, it is possible to judge whether they are consistent by comparing the correlation scores in the first correlation result and the second correlation result. A consistency threshold can be set according to actual needs to be used to judge whether the correlation results are consistent. If the first correlation result and the second correlation result are consistent, then the result with the higher correlation score is adopted as the final corresponding relationship. Exemplarily, if the first correlation result is "the correlation score between 'C=O' and 'high hardness' is 0.80", and the second correlation result is "the correlation score between 'C=O' and 'high hardness' is 0.85", then it is determined that the correlation score between 'C=O' and 'high hardness' is 0.85. On the contrary, if the first correlation result and the second correlation result are inconsistent, the correlation analysis can be performed again or manual review can be carried out. Finally, the corresponding relationship is determined according to the reliability score.
[0120] Optionally, the first correlation result and the second correlation result can also be weighted and fused according to a preset reliability weight to obtain a fusion result, so as to determine the corresponding relationship based on the fusion result.
[0121] The preset reliability weights can be assigned according to the reliability and accuracy of each method. If the statistical analysis method has passed strict verification and has wide applications in the field, a higher weight can be given. Similarly, if the machine learning model performs well on similar problems and has been fully trained and verified, a higher weight can also be given.
[0122] Among them, the first correlation result and the second correlation result can be weighted and fused. For example, the final correlation score = the correlation of the first correlation result × reliability weight 1 + the correlation of the second correlation result × reliability weight 2. Thus, based on the final fusion result (correlation score), according to the level of correlation between the substructure and the material property, the corresponding relationship between the substructure and the material property is determined, such as obtaining the result with a correlation score higher than the preset threshold as the corresponding relationship between the substructure and the material property.
[0123] Through the second correlation analysis and consistency checking / reliability weighted fusion, the corresponding relationship between the substructure and the material property can be determined more accurately.
[0124] In an exemplary embodiment, an implementation method for constructing a target database based on the generated target molecules is also provided. As Figure 9 shown, according to the target material and the corresponding relationship, using the substructures to generate target molecules and determine the molecular properties of the target molecules, so as to construct a target database based on the target molecules and the molecular properties may include:
[0125] Step S910: Based on the corresponding relationship, screen candidate substructures related to the target material properties from the substructures.
[0126] First, the key properties of the target material (such as mechanical strength, conductivity, thermal stability, etc.) need to be clarified, so as to set the target range of the properties. For example, the conductivity > 100 S / cm. Furthermore, based on the corresponding relationship, candidate substructures related to the properties of the target material can be screened from the substructures. For example, if the property of the target material is high conductivity, substructures positively correlated with conductivity (such as conjugated π systems) are screened out.
[0127] Step S920: Combine the candidate substructures so that the resulting candidate molecules satisfy the preset combination rules.
[0128] Combine the candidate substructures so that the resulting candidate molecules satisfy the preset combination rules. The preset combination rules include but are not limited to chemical feasibility rules, connection mode rules, and reaction mechanism rules. Among them, the chemical feasibility rules are used to ensure that the combination of substructures conforms to chemical rules (such as valence bond rules, steric hindrance, etc.), the connection mode rules are used to ensure the connection mode between substructures (such as single bonds, double bonds, cyclization reactions, etc.), and the reaction mechanism rules can combine known chemical reaction mechanisms to ensure that the combined molecules can be synthesized through experiments. When combining molecular structures, a combination algorithm can be used to generate, such as genetic algorithms, Monte Carlo tree search and other algorithms.
[0129] Taking the genetic algorithm as an example for illustration, a group of molecules can be randomly generated according to the candidate substructures first, then calculate the target properties of each molecule, and select the molecule with the optimal properties as the parent generation. Then, combine the substructures of the parent generation molecules to generate offspring molecules. Finally, randomly mutate the offspring molecules (such as replacing substructures). By repeating evaluation, selection, crossover, and mutation until the termination condition is met (such as reaching the number of repetitions), the determined offspring molecules are determined as candidate molecules. Of course, the present disclosure does not specifically limit the way of combining molecular structures, and other algorithms can also be selected according to actual needs.
[0130] Step S930: Use a pre-trained property prediction model to predict the properties of the candidate molecules to obtain the molecular properties of the candidate molecules.
[0131] Step S940: Determine a target molecule from the candidate molecules according to the properties of the target material and the molecular properties of the candidate molecules, so as to construct a target database based on the target molecule and the molecular properties corresponding to the target molecule.
[0132] Among them, the pre-trained property prediction model can be a random forest, a graph neural network, etc., or the Gaussian program can be used to predict the properties of the candidate molecules, and there is no limitation on this. According to the property prediction results, the molecular properties of the candidate molecules are obtained. Furthermore, target molecules related to the properties of the target material can be obtained from the candidate molecules, so as to construct a target database based on the target molecules and their molecular properties. For example, target molecules with a conductivity > 100 S / cm are screened out from the candidate molecules.
[0133] By combining the substructure-property correspondence relationship, combinatorial algorithms, machine learning models, etc., new molecular structures with the properties of the target material can be efficiently generated, enhancing the usability and scalability of the database.
[0134] In addition, the exemplary embodiments of the present disclosure can also perform operations such as adding, deleting, modifying, and querying the molecular structures or related properties in the constructed target database, so as to facilitate and efficiently operate the target database and utilize the data in the target database.
[0135] The database construction method in the exemplary embodiments of the present disclosure, on the one hand, constructs an initial database including reference attribute information and the chemical structure corresponding to the reference attribute information based on the acquisition data corresponding to the target material, where the reference attribute information is the associated attribute of the target material. When constructing the initial database, the attributes of the target material are directly associated, ensuring that the information in the database is highly relevant to the actual attributes of the target material, avoiding the problem of possible mismatch between attributes and materials in the database, and improving the accuracy and practicality of the data. On the other hand, by presetting a splitting strategy to split the chemical structure in the initial database and combining the reference attribute information corresponding to the chemical structure for correlation analysis, the internal relationship between the substructure and the material attributes can be accurately revealed, avoiding the problem of only relying on the overall chemical structure and ignoring local features, thereby improving the accuracy of the database. On the further hand, according to the target material and the corresponding relationship, the substructures are used to combine to generate target molecules and the molecular attributes of the target molecules are determined, so as to construct a target database based on the target molecules and the molecular attributes, ensuring that the generated target molecules have clear attribute characteristics, avoiding the problem of disconnection between molecular design and actual needs, and further improving the accuracy of the database. In addition, by presetting a splitting strategy to automatically split the chemical structure, the inefficiency and subjectivity of manual splitting are avoided, significantly improving the efficiency of chemical structure analysis, laying a foundation for subsequent correlation analysis and target molecule generation. At the same time, the correlation analysis is used to quickly establish the mapping relationship between the substructure and the material attributes, avoiding the inefficient process of a large number of experimental verifications, that is, through a data-driven method, the substructures related to the attributes of the target material can be quickly screened out, improving the efficiency of database construction.
[0136] In the exemplary embodiments of the present disclosure, a database construction device is also provided. Refer to Figure 10 As shown, the database construction device 1000 may include a first construction module 1010, a structure splitting module 1020, and a second construction module 1030. Specifically:
[0137] The first construction module 1010 is configured to construct an initial database according to the acquisition data corresponding to the target material. The initial database includes reference attribute information and the chemical structure corresponding to the reference attribute information, and the reference attribute information is the associated attribute of the target material. The structure splitting module 1020 is configured to split the chemical structure in the initial database by a preset splitting strategy to obtain a plurality of substructures, and perform correlation analysis according to the reference attribute information corresponding to the chemical structure and the plurality of substructures to obtain the corresponding relationship between the substructure and the material attributes. The second construction module 1030 is configured to generate target molecules by combining the substructures according to the target material and the corresponding relationship and determine the molecular attributes of the target molecules, so as to construct a target database according to the target molecules and the molecular attributes.
[0138] In an exemplary embodiment of the present disclosure, the first construction module 1010 is configured to perform: performing data preprocessing on the collected data to obtain a preprocessing result; performing entity recognition on the preprocessing result based on the attributes of the target material to obtain attribute entities, and determining reference attribute information according to the attribute entities; performing chemical structure detection on the preprocessing result to obtain a chemical structure; and constructing an initial database according to the reference attribute information and the chemical structure based on the matching relationship between the reference attribute information and the chemical structure.
[0139] In an exemplary embodiment of the present disclosure, the first construction module 1010 is configured to perform: collecting a keyword set corresponding to the attributes of the target material, and performing data augmentation on the keyword set to obtain a material attribute dictionary; extracting attribute entities from the preprocessing text based on the material attribute dictionary; extracting the entity relationships of the attribute entities according to the target material, and respectively performing semantic understanding on the original text information and the entity relationships corresponding to the attribute entities to obtain a first understanding result and a second understanding result; and determining reference attribute information according to the entity relationships and the attribute entities based on the first understanding result and the second understanding result.
[0140] In an exemplary embodiment of the present disclosure, the first construction module 1010 is configured to perform: performing a structured representation on the chemical structure to obtain structured information; performing relationship recognition on the structured information and the reference attribute information based on a pre-trained relationship recognition model to obtain a relationship recognition result, where the pre-trained relationship recognition model is trained based on sample data of chemical structures and attribute descriptions; and in response to the relationship recognition result indicating a match between the structured information and the reference attribute information, using the reference attribute information and the chemical structure as a data pair to construct an initial database according to the obtained data pair.
[0141] In an exemplary embodiment of the present disclosure, the structure splitting module 1020 is configured to perform: performing chemical bond recognition on the chemical structure, and if a target chemical bond is recognized, splitting the chemical structure at the position of the target chemical bond to obtain a first splitting result; and determining the sub-structure corresponding to the chemical structure according to the first splitting result.
[0142] In an exemplary embodiment of the present disclosure, the structure splitting module 1020 is configured to perform: representing the chemical structure as a structure diagram with atoms as nodes and chemical bonds as edges, performing sub-graph mining on the structure diagram to obtain a specific functional structure, and splitting the specific functional structure from the chemical structure to obtain a second splitting result; and determining the sub-structure of the chemical structure according to the first splitting result and the second splitting result.
[0143] In an exemplary embodiment of the present disclosure, the structure splitting module 1020 is configured to perform: adding chemical attribute information to the nodes and edges in the structure diagram to obtain a functional structure diagram; obtaining functional requirement information, and performing subgraph mining from the functional structure diagram based on the functional requirement information to obtain a specific functional structure.
[0144] In an exemplary embodiment of the present disclosure, the structure splitting module 1020 is configured to perform: performing statistical analysis according to the reference attribute information corresponding to the chemical structure and the sub-structure to obtain a first correlation result between the sub-structure and the material attribute; determining the corresponding relationship between the sub-structure and the material attribute according to the first correlation result.
[0145] In an exemplary embodiment of the present disclosure, the structure splitting module 1020 is configured to perform: performing correlation analysis on the reference attribute information corresponding to the chemical structure and the sub-structure through a pre-trained correlation analysis model to obtain a second correlation result; performing consistency check on the first correlation result and the second correlation result, and determining the corresponding relationship from the first correlation result and the second correlation result according to the consistency check result; or, performing weighted fusion on the first correlation result and the second correlation result according to the preset reliability weight to obtain a fusion result, so as to determine the corresponding relationship according to the fusion result.
[0146] In an exemplary embodiment of the present disclosure, the structure splitting module 1020 is configured to perform: for each molecular structure, statistically analyze the distribution of at least one material attribute of the chemical structure containing the sub-structure; determining a first correlation result between the sub-structure and the material attribute based on the distribution of at least one material attribute.
[0147] In an exemplary embodiment of the present disclosure, the second construction module 1030 is configured to perform: based on the corresponding relationship, screening candidate sub-structures related to the attributes of the target material from the sub-structures; combining the candidate sub-structures so that the obtained candidate molecules satisfy the preset combination rules; using a pre-trained attribute prediction model to perform attribute prediction on the candidate molecules to obtain the molecular attributes of the candidate molecules; determining the target molecule from the candidate molecules according to the target material attributes and the molecular attributes of the candidate molecules, so as to construct a target database according to the target molecule and the molecular attributes corresponding to the target molecule.
[0148] In an exemplary embodiment of the present disclosure, the preset combination rules include at least one of a chemical feasibility rule, a connection mode rule, and a reaction mechanism rule.
[0149] Since the detailed content of each functional module of the database construction device in the exemplary embodiment of the present disclosure has been described in the exemplary embodiment of the above database construction method, it will not be repeated here.
[0150] It should be noted that although several modules or units of the database construction device are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0151] An exemplary embodiment of the present disclosure also provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, it implements the above-mentioned database construction method.
[0152] In one embodiment, the computer program product may be a tangible product containing the computer program, such as a computer-readable storage medium storing the computer program. The readable storage medium may be a storage medium based on signals such as electricity, magnetism, light, electromagnetic, infrared, etc., including but not limited to: random access memory (RAM), read-only memory (ROM), magnetic tape, floppy disk, flash memory (Flash), mechanical hard disk (HDD), solid-state drive (SSD), etc. Exemplarily, the computer program product can be implemented as a non-volatile storage medium storing the computer program, such as read-only memory, Nand Flash, etc.
[0153] In one embodiment, the computer program product may be an intangible product containing the computer program. Exemplarily, the computer program product can be implemented as a virtual digital product, such as an executable file storing the computer program, an installation package and other digital files.
[0154] The code of the computer program can be written in one or more programming languages. Programming languages such as C language, Java, C++, etc. The program code can be executed entirely on the user's computing device, or partially on the user's computing device, or executed as an independent software package, or partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user's computing device through any type of network, such as a local area network (LAN), a wide area network (WAN), etc., or can be connected to an external computing device (for example, through the Internet connection provided by an operator).
[0155] A computer program can be carried or transmitted by signals such as electricity, magnetism, light, electromagnetic, infrared, etc. An electronic device can convert the signal carrying the computer program into a digital signal and then run the computer program. When the computer program runs on the electronic device, its code is used to cause the electronic device to execute (more specifically, to cause the processor of the electronic device to execute) the method steps of various exemplary embodiments of the present disclosure, such as the database construction method described above.
[0156] In addition, in the exemplary embodiments of the present disclosure, an electronic device capable of implementing the above method is also provided. Those skilled in the art can understand that various aspects of the present disclosure can be implemented as a system, a method, or a program product. Therefore, various aspects of the present disclosure can be specifically implemented in the following forms, namely: a complete hardware embodiment, a complete software embodiment (including firmware, microcode, etc.), or an embodiment combining hardware and software aspects, which can be collectively referred to as "circuit", "module", or "system" here.
[0157] Refer to the following Figure 11 to describe the electronic device 1100 according to this embodiment of the present disclosure. Figure 11 The electronic device 1100 shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present disclosure.
[0158] As Figure 11 shown, the electronic device 1100 is presented in the form of a general-purpose computing device. The components of the electronic device 1100 may include but are not limited to: at least one of the above-mentioned processing units 1110, at least one of the above-mentioned storage units 1120, a bus 1130 connecting different system components (including the storage unit 1120 and the processing unit 1110), and a display unit 1140.
[0159] Among them, the storage unit stores program code, and the program code can be executed by the processing unit 1110, so that the processing unit 1110 executes the steps according to various exemplary embodiments of the present disclosure described in the above "Exemplary Method" section of this specification.
[0160] The storage unit 1120 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 1121 and / or a cache storage unit 1122, and may further include a read-only storage unit (ROM) 1123.
[0161] The storage unit 1120 may further include a program / utilities 1124 having a set (at least one) of program modules 1125. Such program modules 1125 include but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment.
[0162] The bus 1130 can represent one or more of several types of bus architectures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of the various bus architectures.
[0163] The electronic device 1100 can also communicate with one or more external devices 1200 (such as a keyboard, a pointing device, a Bluetooth device, etc.), and can also communicate with one or more devices that enable a user to interact with the electronic device 1100, and / or communicate with any device that enables the electronic device 1100 to communicate with one or more other computing devices (such as a router, a modem, etc.). Such communication can be carried out through the input / output (I / O) interface 1150. Moreover, the electronic device 1100 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through the network adapter 1160. As shown in the figure, the network adapter 1160 communicates with other modules of the electronic device 1100 through the bus 1130. It should be understood that although not shown in the figure, other hardware and / or software modules can be used in combination with the electronic device 1100, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0164] Through the description of the above embodiments, those skilled in the art can easily understand that the exemplary embodiments described herein can be implemented by software, or can be implemented by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.
[0165] In addition, the above drawings are only schematic illustrations of the processes included in the method according to the exemplary embodiments of the present disclosure, rather than for limiting purposes. It is easy to understand that the processes shown in the above drawings do not indicate or limit the time sequence of these processes. Additionally, it is also easy to understand that these processes can be executed synchronously or asynchronously, for example, in multiple modules.
[0166] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known common general knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and the embodiments are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the claims.
Claims
1. A method for constructing a database, characterized in that, Including: Construct an initial database based on the acquisition data corresponding to the target material, where the initial database includes reference attribute information and the chemical structure corresponding to the reference attribute information, and the reference attribute information is the associated attribute of the target material; Split the chemical structure in the initial database through a preset splitting strategy to obtain multiple sub-structures, and perform a correlation analysis based on the reference attribute information corresponding to the chemical structure and the multiple sub-structures to obtain the corresponding relationship between the sub-structures and the material attributes; According to the target material and the corresponding relationship, use the sub-structures to combine to generate a target molecule and determine the molecular attributes of the target molecule, so as to construct a target database based on the target molecule and the molecular attributes.
2. The method according to claim 1, characterized in that, The constructing of the initial database based on the acquisition data corresponding to the target material includes: Perform data preprocessing on the acquisition data to obtain a preprocessing result; Perform entity recognition on the preprocessing result based on the attributes of the target material to obtain attribute entities, and determine the reference attribute information according to the attribute entities; Perform chemical structure detection on the preprocessing result to obtain the chemical structure; Based on the matching relationship between the reference attribute information and the chemical structure, construct the initial database according to the reference attribute information and the chemical structure.
3. The method according to claim 2, wherein The performing of entity recognition on the preprocessing result based on the attributes of the target material to obtain attribute entities, and determining the reference attribute information according to the attribute entities includes: Collect a keyword set corresponding to the attributes of the target material, and perform data augmentation on the keyword set to obtain a material attribute dictionary; Extract the attribute entities from the preprocessing text based on the material attribute dictionary; According to the target material, extract the entity relationship of the attribute entities, and perform semantic understanding on the original text information corresponding to the attribute entities and the entity relationship respectively to obtain a first understanding result and a second understanding result; Based on the first understanding result and the second understanding result, determine the reference attribute information according to the entity relationship and the attribute entities.
4. The method according to claim 2, characterized in that, The constructing of the initial database according to the reference attribute information and the chemical structure based on the matching relationship between the reference attribute information and the chemical structure includes: Perform a structured representation on the chemical structure to obtain structured information; Perform relationship recognition on the structured information and the reference attribute information based on a pre-trained relationship recognition model to obtain a relationship recognition result, and the pre-trained relationship recognition model is trained based on sample data of chemical structures and attribute descriptions; In response to the relationship recognition result indicating that the structured information and the reference attribute information match, use the reference attribute information and the chemical structure as a data pair to construct the initial database according to the obtained data pair.
5. The method according to claim 1, characterized in that The splitting of the chemical structure in the initial database through a preset splitting strategy to obtain multiple sub-structures includes: Perform chemical bond recognition on the chemical structure. If a target chemical bond is recognized, split the chemical structure at the position of the target chemical bond to obtain a first splitting result; Determine the sub-structure corresponding to the chemical structure according to the first splitting result.
6. The method according to claim 5, characterized in that, The step of determining the sub-structure corresponding to the chemical structure according to the first splitting result further includes: Taking atoms as nodes and chemical bonds as edges, representing the chemical structure as a structure diagram, performing sub-graph mining on the structure diagram to obtain a specific functional structure, and splitting the specific functional structure from the chemical structure to obtain a second splitting result; Determine the sub-structure of the chemical structure according to the first splitting result and the second splitting result.
7. The method according to claim 6, wherein The step of performing sub-graph mining on the structure diagram to obtain a specific functional structure includes: Adding chemical attribute information to the nodes and edges in the structure diagram to obtain a functional structure diagram; Obtain functional requirement information, and perform sub-graph mining from the functional structure diagram based on the functional requirement information to obtain the specific functional structure.
8. The method according to claim 1, wherein The step of performing correlation analysis on the reference attribute information corresponding to the chemical structure and the multiple sub-structures to obtain the corresponding relationship between the sub-structure and the material attribute includes: Performing statistical analysis on the reference attribute information corresponding to the chemical structure and the sub-structure to obtain a first correlation result between the sub-structure and the material attribute; Determine the corresponding relationship between the sub-structure and the material attribute according to the first correlation result.
9. The method according to claim 8, characterized in that The step of determining the corresponding relationship between the sub-structure and the material attribute according to the first correlation result further includes: Performing correlation analysis on the reference attribute information corresponding to the chemical structure and the sub-structure through a pre-trained correlation analysis model to obtain a second correlation result; Performing consistency check on the first correlation result and the second correlation result, and determining the corresponding relationship from the first correlation result and the second correlation result according to the consistency check result; Alternatively, according to a preset reliability weight, perform weighted fusion on the first correlation result and the second correlation result to obtain a fusion result, and determine the corresponding relationship according to the fusion result.
10. The method according to claim 8, characterized in that, The step of performing statistical analysis on the reference attribute information corresponding to the chemical structure and the sub-structure to obtain a first correlation result between the sub-structure and the material attribute includes: For each molecular structure, statistically analyze the distribution of at least one material attribute of the chemical structure containing the sub-structure; Based on the distribution of the at least one material attribute, determine the first correlation result between the sub-structure and the material attribute.
11. The method according to claim 1, characterized in that, The step of generating a target molecule using the sub-structure according to the target material and the corresponding relationship and determining the molecular attribute of the target molecule, and constructing a target database according to the target molecule and the molecular attribute includes: Based on the corresponding relationship, screen candidate sub-structures related to the attributes of the target material from the sub-structures; Combine the candidate sub-structures so that the resulting candidate molecule satisfies a preset combination rule; Perform attribute prediction on the candidate molecule using a pre-trained attribute prediction model to obtain the molecular attribute of the candidate molecule; Determine the target molecule from the candidate molecules according to the properties of the target material and the molecular properties of the candidate molecules, so as to construct the target database according to the target molecule and the molecular properties corresponding to the target molecule.
12. The method according to claim 11, wherein The preset combination rules include at least one of a chemical feasibility rule, a connection mode rule, and a reaction mechanism rule.
13. A database construction device, characterized in that, Comprising: A first construction module, configured to construct an initial database according to the collected data corresponding to the target material, where the initial database includes reference attribute information and a chemical structure corresponding to the reference attribute information, and the reference attribute information is the associated attribute of the target material; A structure splitting module, configured to split the chemical structure in the initial database through a preset splitting strategy to obtain a plurality of sub-structures, and perform a correlation analysis according to the reference attribute information corresponding to the chemical structure and the plurality of sub-structures to obtain the corresponding relationship between the sub-structures and the material properties; A second construction module, configured to generate a target molecule by combining the sub-structures according to the target material and the corresponding relationship, and determine the molecular properties of the target molecule, so as to construct a target database according to the target molecule and the molecular properties.
14. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method according to any one of claims 1 to 12.
15. An electronic device, characterized in that, Comprising: A processor; And A memory, configured to store executable instructions of the processor; Wherein, the processor is configured to execute the method according to any one of claims 1 to 12 by executing the executable instructions.
Citation Information
Cited By
Artificial intelligence enhanced material database system
CN121636649A