Data processing method, training method, determination method, design method and device

CN116682506BActive Publication Date: 2026-08-11SHENZHEN JINGTAI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-11
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0003]然而,申请人发现相关技术得到的屏蔽电荷(σ)的分布密度函数(也叫做表面电荷密度为σ的概率分布,简称σ-profile)的时间成本和消耗的计算资源过高

Benefits of technology

[0016] The data processing method, training method, determination method, design method, and apparatus provided in this application store the three-dimensional structure of the target object and the distribution density function of the shielding charge corresponding to the three-dimensional structure in association, thereby obtaining a training dataset. This establishes a correspondence between the three-dimensional structure of the target object and the distribution density function of the shielding charge, enabling the training of a distribution density function prediction model using this training dataset. This allows for the direct prediction of the σ-profile from the structure of the molecule or crystal, bypassing the complex intermediate calculation of the shielding charge, and effectively reducing the time, computational resources, and computational costs consumed in predicting the σ-profile.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116682506B_ABST
    Figure CN116682506B_ABST
Patent Text Reader

Abstract

This application relates to a data processing method, a training method, a determination method, a design method, and an apparatus. The data processing method includes: obtaining the three-dimensional structure of a target object, and obtaining the distribution density function of the shielding charge corresponding to each three-dimensional structure; storing the three-dimensional structure and the distribution density function of the shielding charge corresponding to the three-dimensional structure in association to obtain a training dataset, so as to train a distribution density function prediction model based on the training dataset, wherein the trained distribution density function prediction model can predict the distribution density function of the shielding charge corresponding to the three-dimensional structure. This application effectively reduces the time cost and computational resources consumed in predicting the σ-profile.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer simulation technology, and in particular to a data processing method, training method, determination method, design method and apparatus. Background Technology

[0002] With the rapid development of computer technology and artificial intelligence technology, computer simulation technology is being applied to more and more scenarios, such as materials design and drug design.

[0003] However, the applicant found that the time cost and computational resources consumed in obtaining the distribution density function of the shielding charge (σ) (also called the probability distribution of the surface charge density σ, or σ-profile for short) obtained by the relevant technology were too high. Summary of the Invention

[0004] To address or partially address the problems existing in related technologies, this application provides a data processing method, training method, determination method, design method, and apparatus that can effectively reduce the time cost and computational resources consumed in predicting σ-profile.

[0005] The first aspect of this application provides a data processing method, comprising: obtaining a three-dimensional structure of a target object, and obtaining a distribution density function of shielding charge corresponding to each of the three-dimensional structures; storing the three-dimensional structures and the distribution density functions of shielding charge corresponding to the three-dimensional structures in association to obtain a training dataset, so as to train a distribution density function prediction model based on the training dataset, wherein the trained distribution density function prediction model is able to predict the distribution density function of shielding charge corresponding to the three-dimensional structure.

[0006] The second aspect of this application provides a method for training a distribution density function prediction model, comprising: inputting training data obtained as described above into a distribution density function prediction model to be trained, adjusting the model parameters to make the loss function converge, and obtaining a trained distribution density function prediction model, wherein the training data includes the three-dimensional structure of the target object and the distribution density function of the shielding charge, and the loss function includes the correlation between the distribution density function prediction result and the distribution density function annotation result.

[0007] A third aspect of this application provides a method for determining the distribution density function of shielding charge, comprising: processing a three-dimensional structure or virtual image of a target object using a distribution density function prediction model trained as described above, to obtain the distribution density function of shielding charge for the target object, wherein the virtual image is generated based on the three-dimensional structure of the target object.

[0008] A fourth aspect of this application provides a design method comprising: determining, according to the method described above, a distribution density function of the shielding charge of a target object, and / or determining, according to the method described above, the chemical potential of the target object; and designing a candidate drug or material based on the distribution density function of the shielding charge and / or the chemical potential of the target object.

[0009] A fifth aspect of this application provides a data processing apparatus, comprising: an information acquisition module and an associated storage module. The information acquisition module is used to acquire the three-dimensional structure of a target object and to acquire the distribution density function of the shielding charge corresponding to each three-dimensional structure. The associated storage module is used to associate and store the three-dimensional structures and the distribution density functions of the shielding charges corresponding to the three-dimensional structures to obtain a training dataset, so as to train a distribution density function prediction model based on the training dataset. The trained distribution density function prediction model is capable of predicting the distribution density function of the shielding charge corresponding to the three-dimensional structure.

[0010] The sixth aspect of this application provides an apparatus for training a distribution density function prediction model, comprising: a model training module for inputting training data obtained based on the apparatus described above into a distribution density function prediction model to be trained, adjusting model parameters to make the loss function converge, and obtaining a trained distribution density function prediction model, wherein the training data includes the three-dimensional structure of the target object and the distribution density function of the shielding charge, and the loss function includes the correlation between the distribution density function prediction result and the distribution density function annotation result.

[0011] The seventh aspect of this application provides an apparatus for determining a distribution density function of shielding charge, comprising: processing a three-dimensional structure or virtual image of a target object using a distribution density function prediction model trained according to the above apparatus to obtain a distribution density function of shielding charge for the target object, wherein the virtual image is generated based on the three-dimensional structure of the target object.

[0012] An eighth aspect of this application provides a design apparatus, comprising: an information determination module for determining a distribution density function of the shielding charge of a target object according to the apparatus described above, and / or determining the chemical potential of the target object; and a design module for designing candidate drugs or materials based on the distribution density function of the shielding charge and / or the chemical potential of the target object.

[0013] A ninth aspect of this application provides an electronic device, including: a processor; and a memory having executable code stored thereon, which, when executed by the processor, causes the processor to perform the method described above.

[0014] The tenth aspect of this application also provides a computer-readable storage medium having executable code stored thereon, which, when executed by a processor of an electronic device, causes the processor to perform the above-described method.

[0015] The eleventh aspect of this application also provides a computer program product, including executable code that implements the above-described method when executed by a processor.

[0016] The data processing method, training method, determination method, design method, and apparatus provided in this application store the three-dimensional structure of the target object and the distribution density function of the shielding charge corresponding to the three-dimensional structure in association, thereby obtaining a training dataset. This establishes a correspondence between the three-dimensional structure of the target object and the distribution density function of the shielding charge, enabling the training of a distribution density function prediction model using this training dataset. This allows for the direct prediction of the σ-profile from the structure of the molecule or crystal, bypassing the complex intermediate calculation of the shielding charge, and effectively reducing the time, computational resources, and computational costs consumed in predicting the σ-profile.

[0017] Furthermore, embodiments of this application use virtual graphs to describe molecular or crystal structures and employ equivariant graph neural network models to achieve the transformation of structure-property relationships, thereby obtaining accurate σ-profiles.

[0018] Furthermore, σ-profile is a fixed-dimensional vector, which is easier to predict than shielding charge, and chemical potential can be calculated conveniently and quickly based on σ-profile.

[0019] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description

[0020] The above and other objects, features and advantages of this application will become more apparent from the more detailed description of exemplary embodiments thereof in conjunction with the accompanying drawings, wherein the same reference numerals generally represent the same components in the exemplary embodiments thereof.

[0021] Figure 1 This illustration schematically shows an exemplary system architecture to which data processing methods, training methods, determination methods, design methods, and apparatus can be applied according to embodiments of this application;

[0022] Figure 2 The flowchart illustrating the relevant techniques for calculating chemical potential is shown in the diagram.

[0023] Figure 3 A flowchart illustrating a data processing method according to an embodiment of this application is shown schematically.

[0024] Figure 4 A schematic diagram of a σ-profile according to an embodiment of this application is shown;

[0025] Figure 5 A flowchart illustrating the training of a distribution density function prediction model according to an embodiment of this application is shown schematically.

[0026] Figure 6 The schematic diagram illustrates the structure of an isovariant graph convolutional network according to an embodiment of this application;

[0027] Figure 7 The diagram illustrates the logic of a training distribution density function prediction model according to an embodiment of this application.

[0028] Figure 8 A schematic diagram of the structure of a convolutional layer according to an embodiment of this application is shown.

[0029] Figure 9 A schematic diagram illustrating the structure of a fully connected network according to an embodiment of this application is shown.

[0030] Figure 10 A flowchart illustrating a method for determining the distribution density function of shielding charge according to an embodiment of this application is shown schematically.

[0031] Figure 11 The illustration schematically shows a logic diagram for determining the distribution density function of the shielding charge according to an embodiment of this application;

[0032] Figure 12 It is a comparison between the σ-profile predicted values ​​and the actual values ​​corresponding to the 9 molecular structures in the test set;

[0033] Figure 13 It is a comparison between the predicted and actual values ​​of the σ-profile corresponding to the nine crystal structures in the test set;

[0034] Figure 14 A flowchart illustrating a design method according to an embodiment of this application is shown schematically.

[0035] Figure 15 A block diagram of a data processing apparatus according to an embodiment of this application is shown schematically;

[0036] Figure 16 A block diagram of an apparatus for training a distribution density function prediction model according to an embodiment of this application is shown schematically;

[0037] Figure 17 A block diagram schematically illustrates an apparatus for determining a distribution density function of shielding charge according to an embodiment of this application;

[0038] Figure 18A block diagram of a design device according to an embodiment of this application is shown schematically;

[0039] Figure 19 A block diagram of an electronic device according to an embodiment of this application is shown schematically. Detailed Implementation

[0040] Embodiments of this application will now be described in more detail with reference to the accompanying drawings. While embodiments of this application are shown in the drawings, it should be understood that this application may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided to make this application more thorough and complete, and to fully convey the scope of this application to those skilled in the art.

[0041] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to limit the application. The terms "comprising," "including," etc., as used herein indicate the presence of features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0042] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.

[0043] It should be understood that although the terms "first," "second," "third," etc., may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this application, "multiple" means two or more, unless otherwise explicitly specified.

[0044] Before describing the technical solution of this application, some of the technical terms in this field will be explained.

[0045] A graph is a data structure describing related objects, consisting of nodes and edges.

[0046] Structure refers to molecular or crystal structure.

[0047] A structure descriptor represents a molecule or crystal as a data structure that a computer program can process.

[0048] Virtual graph: A structural descriptor that represents atoms as nodes and the relationships between atoms as edges. Unlike ordinary graphs, which establish edges based on bonding information between atoms, virtual graphs establish edges based on cutoff radii. For periodic systems (such as crystals), periodic boundary conditions need to be considered when establishing nodes and edges.

[0049] The cutoff radius is crucial because if we were to establish edges between an atom and all other atoms within a molecule, the number of edges would be enormous, leading to excessive computation. Considering that the influence of atoms farther away from the atom decreases, a cutoff radius is chosen, allowing edges to be established only between the atom and atoms within that radius. Atoms outside the cutoff radius have their interactions ignored. Similarly, for an atom within a unit cell of a crystal structure, establishing edges between it and all other atoms in the structure would also result in a large number of edges and excessive computation. Therefore, a cutoff radius can also be used for an atom within a unit cell of a crystal structure to reduce computational complexity.

[0050] Loss function: A loss function maps the values ​​of a random event or its related random variables to non-negative real numbers to represent the "risk" or "loss" of that random event. The distribution density function (σ-profile) of the shielded charge is a vector composed of a set of numbers. Unlike scalars, vector prediction cannot simply use absolute error or root mean square error as the loss function; the correlation between vectors must also be considered. In this invention, spectral information divergence and spectral information similarity are preferably introduced as loss functions.

[0051] A coordinate system is a reference system used to describe the position and orientation of an object; it is also called a reference frame or frame of reference. For example, a coordinate system can be one created during simulation, such as a Cartesian coordinate system.

[0052] Coordinates are used to represent the absolute position of an object in a specific coordinate system. In mathematics, coordinates are essentially ordered logarithms.

[0053] The shielding charge distribution density function describes the distribution density of the shielding charge (σ) on the solvated surface of a molecule or crystal. For most systems, the shielding charge value is between -0.030 and 0.030. For ease of calculation, this range is usually discretized, and the σ-profile is calculated on discrete σ values.

[0054] Solvation is a reaction process driven by the interaction between solute and solvent molecules. It is a key step in drug development processes of interest, such as crystal nucleation, chemical reactions, drug metabolism, drug interactions, and drug-receptor interactions.

[0055] The Conductor-like Screening Model for Real Solvents (COSMO-RS) is a quantum chemistry-based equilibrium thermodynamic model primarily used for calculating chemical potentials.

[0056] By utilizing chemical potential, macroscopic physicochemical properties with significant industrial applications, such as activation coefficient, partition coefficient, solvation free energy, solubility, and viscosity, can be obtained. This enables COSMO-RS to be widely used in industries such as chemical engineering and pharmaceuticals, specifically in areas such as solvent screening, drug screening, crystal nucleation, chemical reactions, and chemical engineering design.

[0057] Calculating chemical potentials in related technologies requires significant computational resources and incurs substantial time and computational costs, hindering large-scale industrial applications due to their long development cycles and high costs. While some technologies address this computational time-consuming issue by establishing empirical structure-property relationship models to predict desired properties, the applicant's extensive research and analysis have revealed two drawbacks: first, the accuracy of this method is not particularly high, insufficient to guarantee that the calculated chemical potential meets application requirements; second, the prediction of shielding charges is difficult due to their complexity, depending on the system size and surface element division.

[0058] In order to achieve fast and accurate prediction of chemical potential, this application provides a data processing method that can generate training data and a machine learning model that can describe the conversion relationship between the three-dimensional structure of the target object and the distribution density function of the shielding charge, so as to improve the accuracy of the predicted distribution density function of the shielding charge.

[0059] The following will be through Figures 1 to 19 This application provides a detailed description of a data processing method, training method, determination method, design method, and apparatus according to embodiments of the present application.

[0060] Figure 1 This illustration schematically depicts an exemplary system architecture to which data processing methods, training methods, determination methods, design methods, and apparatus can be applied according to embodiments of this application. It should be noted that... Figure 1The examples shown are merely examples of system architectures that can be applied to the embodiments of this application, in order to help those skilled in the art understand the technical content of this application, but do not mean that the embodiments of this application cannot be used in other devices, systems, environments or scenarios.

[0061] See Figure 1 The system architecture 100 according to this embodiment may include terminal devices 101, 102, and 103, a network 104, and a server 105. The network 104 serves as a medium for providing a communication link between the terminal devices 101, 102, and 103 and the server 105. The network 104 may include various connection types, such as wired or wireless communication links, or fiber optic cables, etc.

[0062] Users can use terminal devices 101, 102, and 103 to interact with other terminal devices and server 105 via network 104 to receive or send information, such as sending model training requests, requests for prediction of the distribution density function of shielded charge, and receiving model training results and the distribution density function of shielded charge. Terminal devices 101, 102, and 103 can be equipped with various communication client applications, such as drug development applications, materials design applications, web browser applications, database applications, search applications, instant messaging tools, email clients, social media platforms, and other applications.

[0063] Terminal devices 101, 102, and 103 include, but are not limited to, smart desktop computers, tablet computers, laptop computers, and other electronic devices that can support functions such as internet access, modeling, analysis and calculation, and design.

[0064] Server 105 can receive model training requests, shielded charge distribution density function requests, etc., adjust model parameters, store model topology, distribute model parameters, output shielded charge distribution density function, etc., and can also send the shielded charge distribution density function to terminal devices 101, 102, and 103. For example, server 105 can be a backend management server, server cluster, etc.

[0065] It should be noted that the number of terminal devices, networks, and servers is merely illustrative. Depending on the implementation requirements, any number of terminal devices, networks, and cloud components can be included.

[0066] COSMO-RS is an equilibrium thermodynamic model that can calculate the chemical potential of matter using quantum chemical methods. This model divides the solvation surface of a molecule into multiple facets and further calculates the chemical potential by statistically analyzing the distribution of shielding charge σ at these facets.

[0067] Figure 2 The flowchart illustrating the relevant techniques for calculating chemical potential is shown.

[0068] See Figure 2 The COSMO-RS calculation of chemical potential can include the following steps. First, COSMO calculations are performed using quantization software to obtain the shielding charges of all facet elements on the solvated surface of the molecule or crystal (which can be stored as a COSMO file). Quantization software includes, but is not limited to, Gaussian (a quantum chemistry software for semi-empirical and ab initio calculations), TURBOMOLE, and DMol3. Then, the histogram distribution of the shielding charges is statistically analyzed to obtain a σ-profile characterizing the specific numerical distribution of each charge interval. Next, the σ-profile is input into chemical thermodynamics calculation software (such as COSMOtherm) to perform COSMO-RS calculations to obtain the chemical potential.

[0069] Therefore, in order to obtain the chemical potential of a target object for candidate drug development, material development, etc., it is necessary to first calculate the distribution density function of the shielding charge of the target object. How to efficiently and accurately obtain the σ-profile of the target object has become an urgent problem to be solved.

[0070] In recent years, with the rapid development of machine learning technology, an increasing number of machine learning models have been applied to fields such as chemistry, physics, and medicine, significantly enhancing their ability to solve problems in these areas. This application addresses these problems by constructing relevant models and methods to meet the needs of industrial applications.

[0071] Figure 3 A flowchart illustrating a data processing method according to an embodiment of this application is shown.

[0072] like Figure 3 As shown, this embodiment provides a data processing method, which includes operations S310 to S320, as detailed below:

[0073] In operation S310, the three-dimensional structure of the target object is obtained, as well as the distribution density function of the shielding charge corresponding to each three-dimensional structure.

[0074] In this embodiment, the three-dimensional structure of the target object and the distribution density function σ-profile of the shielding charge can be represented using strings or similar methods. The three-dimensional structure can characterize the spatial properties of at least some atoms in the target object. These spatial properties include, but are not limited to, spatial position attributes and conformations. Spatial position attributes can be coordinates in Cartesian coordinates or polar coordinates.

[0075] For example, the three-dimensional structure of a target object can be represented as a string in (x,y,z) format. For instance, the string could include the three-dimensional conformation of a molecule and the (x,y,z) coordinates of each atom within the molecule. For example, reading the (x,y,z) string of a molecule or crystal structure yields the atomic position information of the atoms contained within that molecule or crystal structure. It should be noted that the formats of the three-dimensional molecular and crystal structures shown above are merely examples; other common three-dimensional structure representation formats are also possible.

[0076] Furthermore, the same coordinate system, especially a spatial coordinate system, can be used in calculating the distribution density function of molecular shielding charge and in drug design. For example, the geometric center coordinates of a molecule, the atomic coordinates of a molecule, the geometric center coordinates of a crystal structure, and some or all of the atomic coordinates in a crystal structure are in the same coordinate system. It should be understood that the geometric center coordinates of a molecule, the atomic coordinates in a molecule, the geometric center coordinates of a candidate drug, and some or all of the atomic coordinates in a candidate drug can also be in different coordinate systems, but coordinates in different coordinate systems can be converted to each other.

[0077] Various related techniques can be used to determine the distribution density function of the shielding charge corresponding to each three-dimensional structure. For example, the distribution density function of the shielding charge of the target object can be determined through simulation or testing.

[0078] For example, taking the simulation to obtain the distribution density function of the shielding charge of the target object as an example, the above-mentioned obtaining the distribution density function of the shielding charge corresponding to each three-dimensional structure can include the following operations.

[0079] First, determine the element types of the solvated surface of the target object. Divide the solvated surface into small grids, each of which is called an element type.

[0080] Then, the surface elements of the solvated surface of the target object are processed using a quantization algorithm to obtain the shielding charge of the target object.

[0081] Next, the histogram distribution of the shielding charge of the target object is statistically analyzed to obtain the distribution density function of the shielding charge corresponding to the three-dimensional structure of the target object. This allows us to obtain the three-dimensional structural information and their respective σ-profile information of molecules, or the three-dimensional structural information and their respective σ-profile information of crystal structures.

[0082] In addition, some commercially available simulation software can be used to calculate the distribution density function of the shielding charge corresponding to the three-dimensional structure of the target object.

[0083] Figure 4A schematic diagram of a σ-profile according to an embodiment of this application is shown.

[0084] See Figure 4 This shows the σ-profile of water molecules. For example, the σ-profile can take values ​​in increments of 0.001 between -0.030 and 0.030, for a total of 61 values, in units of charge density (eA). -2 The σ-profile represents the statistical count corresponding to each σ. For ease of calculation, the σ-profile is generally normalized. Simply put, the σ-profile is a normalized probability density function corresponding to 61 σ values. It should be noted that the 61 values ​​are taken at intervals of 0.001 within the range [-0.030, 0.030]. These intervals are examples and can be adjusted according to actual precision requirements.

[0085] In operation S320, the three-dimensional structure and the distribution density function of the shielding charge corresponding to the three-dimensional structure are stored in association to obtain a training dataset, so as to train a distribution density function prediction model based on the training dataset. The trained distribution density function prediction model can predict the distribution density function of the shielding charge corresponding to the three-dimensional structure.

[0086] Specifically, the distribution density function of the shielding charge of the target object can be used to annotate its 3D structure; that is, the distribution density function of the shielding charge is used as the annotation information of the 3D structure of the target object. The 3D structure of the target object and the distribution density function of the shielding charge can be used as a training dataset. The training dataset can include multiple training datasets. This allows the distribution density function prediction model to be trained using this training dataset. After processing the 3D structure of the object to be predicted, the trained distribution density function prediction model can output the distribution density function of the shielding charge of the object to be predicted.

[0087] Compared to related technologies that require the computation and generation of a cosmo file (a process that can take anywhere from several hours to several days), this application constructs a training dataset for the distribution density function prediction model. This allows the distribution density function prediction model trained on this dataset to predict the σ-profile of the target object without needing to generate a cosmo file, effectively reducing computational resource consumption and time costs. This at least partially solves the problems of long cycles and high costs faced in large-scale industrial applications. Furthermore, this application's embodiment, in calculating the distribution density function of the shielding charge, does not depend on the size of the system or the division of surface elements, resulting in a more accurate distribution density function of the shielding charge, thereby improving the accuracy of the calculated chemical potential.

[0088] In some embodiments, to facilitate better machine understanding of the three-dimensional structure of the target object, a virtual graph can be used to represent the three-dimensional structure of the target object. For example, the three-dimensional structure of the target object in string form in the dataset can be converted into a virtual graph. The target object includes the three-dimensional structure of molecules and crystals.

[0089] Specifically, the above method may also include the following operation: converting the three-dimensional structure of the target object into a virtual graph.

[0090] Accordingly, storing the three-dimensional structure and the distribution density function of the shielding charge corresponding to the three-dimensional structure in association may include storing the virtual graph and the distribution density function of the shielding charge corresponding to the virtual graph in association.

[0091] The different properties of molecules and crystals lead to differences in their corresponding virtual diagrams.

[0092] In this system, the molecule is an isolated system, and we only need to consider the interactions between the atoms within the molecule. For example, all the atoms in the molecule form nodes, and edges are established between nodes that meet certain conditions. The maximum number of edges between two atoms is 1, and the minimum is 0.

[0093] Crystals are scalable systems with periodicity, thus requiring consideration of the interactions between atoms within a unit cell and their neighboring units. This application provides virtual graph generation methods for molecules and crystals in different embodiments. For example, when constructing a virtual graph for a crystal structure, nodes represent all atoms in the crystal; however, periodic boundary conditions must be considered when establishing edges. In this case, the number of edges between nodes will be relatively large, theoretically a maximum of 8 and a minimum of 0.

[0094] In some embodiments, the target object includes molecules. The following example illustrates the conversion of the three-dimensional structure of a molecule into a virtual diagram.

[0095] Specifically, converting the three-dimensional structure of a target object into a virtual graph can include the following operations.

[0096] First, obtain the i-th atom A in the three-dimensional structure of the molecule. i The position information in the spatial coordinate system, and the position information of the i-th atom A. i Point scalar feature N s and the i-th atom A i Point vector features N v , where i is an integer greater than or equal to 1. For example, reading a string yields the atom type and position information contained in the structure, corresponding to the generation of the point set {A}. i} and the set of node positions {(x i ,y i ,z i)}, where i = 1, 2, ..., N, N represents the number of atoms in the molecule, and A represents the atom type. For example, by setting the feature dimension F, the set of points is represented as an N×F×1 matrix representing the scalar features of the nodes, while an N×F×3 matrix with all zero elements is initialized to represent the vector features of the nodes.

[0097] Then, according to the preset cutoff radius r cut And position information to obtain the i-th atom A i Adjacent sets of adjacent atoms {N j For a single atom within a molecule, establishing edges between it and all other atoms would result in a large number of edges, consuming excessive computational resources and incurring high time costs. Considering that the influence of other atoms further away from the current atom is smaller, a cutoff radius is chosen. This allows the current atom to establish edges only with neighboring atoms whose distance is less than the cutoff radius. Atoms outside the cutoff radius are approximately ignored in their interactions.

[0098] Next, obtain the i-th atom A. i With adjacent atom set {N j The set of edges {(i,j)} between atoms in} is used to obtain the edge length and position vector of edge (i,j), where j is an integer greater than or equal to 1.

[0099] Then, the edge scalar feature E is determined based on the length of edge (i,j). s Furthermore, the edge vector feature E is determined based on the position vector of edge (i,j). v .

[0100] Next, using the point scalar characteristic N of the i-th atom s Point vector features N v , boundary scalar feature E s and edge vector feature E v To form a virtual graph.

[0101] In one specific embodiment, the target molecule may include N atoms, and multiple nodes in the point set each have F-dimensional features.

[0102] Correspondingly, the point scalar feature N s The dimensions include N×F×1, and the point vector features are N. v The dimensions include N×F×3, and the edge scalar features are E. s The dimensions include E×F×1, and the edge vector features are E v The dimensions include E×F×3, where F is the preset feature dimension.

[0103] Specifically, the feature dimension F is set, such that the value of F is not less than 64. Optionally, the value of F is an integer power of 2. Elements in the point set are embedded and encoded according to the node type (e.g., atom type), representing the point set as an N×F×1 dimensional matrix, which represents the point scalar feature N. s Simultaneously, initialize an N×F×3 matrix consisting entirely of zero elements, which represents the feature N of the point vector. v Embedding encoding can refer to randomly representing a node as an F-dimensional vector, which is then updated during subsequent model training.

[0104] Next, for the point set {A} i In the set {}, each element i can be connected to at least some of the remaining elements j, forming an edge set {(i,j)}, where i = 1, 2, ..., N, j ∈ A. i The number of edges is denoted as E.

[0105] For each edge in the edge set, calculate the distance and position vector between the two nodes corresponding to the edge to obtain the edge distance set, as shown in Equation (1).

[0106]

[0107] The set of edge position vectors is shown in equation (2).

[0108] {r v,ij |r v,ij =(x i -x j ,y i -y j ,z i -z j Equation (2)

[0109] Furthermore, the set of edge distances is represented as an E×F×1 dimensional matrix representing the scalar feature E of the edge. s The set of edge position vectors is represented as an E×F×3 dimensional matrix, representing the feature E of the edge vectors. v .

[0110] In some embodiments, considering the limited range of interatomic forces and in order to reduce the computational resources consumed, the above-mentioned point set {A} can be used. i Select the set of adjacent atoms {N} that satisfy the cutoff radius requirement from the set. j}, then determine the set of adjacent atoms {N} j Let {(i,j)} be the set of edges of}.

[0111] Specifically, the above method may also include the following operations.

[0112] First, determine the cutoff radius r.cut Wherein, the cutoff radius r cut This can be a preset value, which can be a value determined based on expert experience or simulation results, such as... or wait.

[0113] Then, determine from the set of points that the distance between nodes is less than or equal to the cutoff radius r. cut The target node is obtained, resulting in the target point set. This target point set is then used as the adjacent atom set {N}. j}

[0114] Specifically, the cutoff radius r can be set first. cut For each element i in the point set, determine the set of adjacent atoms {N} within the cutoff radius for each element i. j}, as shown in equation (3).

[0115]

[0116] For element i, establish an edge (i,j) with all nodes j in the adjacent atom sets, forming an edge set {(i,j)}, where i = 1, 2, ..., N, j ∈ A. i The number of edges is denoted as E.

[0117] Accordingly, the generation of edge scalar features and edge vector features for the point set based on the coordinate information of each node in the node location set includes:

[0118] Generate edge scalar features E for the target point set based on the coordinate information of the target node in the node location set. s and edge vector feature E v .

[0119] In some embodiments, due to the point set {A} i The set of adjacent atoms {N} was selected from the group. j}, correspondingly, edge scalar features and E s and edge vector feature E v The dimensions will also change.

[0120] Specifically, the target point set includes E nodes, each of which has F-dimensional features. The point scalar features are N. s The dimensions include N×F×1, and the point vector features are N. v The dimensions include N×F×3, and the edge scalar features are E. s The dimensions include E×F×1, and the edge vector features are E v The dimensions include E×F×3.

[0121] Thus, based on the point scalar feature Ns Point vector features N v , boundary scalar feature E s With edge vector feature E v This constitutes a virtual graph. The virtual graph includes scalar and vector information of atoms, as well as scalar and vector information of the relationships between atoms. It is a universal and accurate descriptor that can effectively improve the accuracy of the determined distribution density function of molecular shielding charge.

[0122] In this embodiment, the three-dimensional information of molecules is represented as a virtual graph. In addition to not losing the three-dimensional information of molecules, different molecular conformations can also be strictly distinguished. Compared with the two-dimensional descriptors used by machine learning methods in related technologies, the description of molecules is more accurate.

[0123] In some embodiments, the target object includes a crystal structure. The following example illustrates the conversion of a crystal structure into a virtual diagram.

[0124] Specifically, the above-mentioned conversion of the three-dimensional structure of the target object into a virtual graph may include the following operations.

[0125] First, by treating at least some atoms in the crystal structure as nodes, the position information of these nodes in the spatial coordinate system is obtained, and the point scalar characteristic N of the nodes is also obtained. s The point vector features of the nodes N v .

[0126] Then, based on the preset boundary conditions, construct the edge set {(i,j)} for the node to obtain the edge length and position vector of edge (i,j), where i is an integer greater than or equal to 1 and j is an integer greater than or equal to 1.

[0127] Next, the edge scalar feature E is determined based on the length of edge (i,j). s Furthermore, the edge vector feature E is determined based on the position vector of edge (i,j). v .

[0128] Then, using the point scalar feature N of the node s Point vector features N v , boundary scalar feature E s and edge vector feature E v To form a virtual graph.

[0129] It should be noted that when constructing the virtual graph, the nodes are all the atoms in the crystal, but when establishing the edges, periodic boundary conditions need to be considered. In this case, the number of edges between nodes will be large, theoretically up to 8 and down to 0.

[0130] For example, starting with a crystal molecule (unit cell), expand the cell to form a 3×3×3 supercell (containing 27 identical unit cells). Then, using one atom in the central unit cell as the center, find neighboring atoms within the cutoff radius of that atom and connect them to form edges. Repeat this process to find edges for all atoms within the ()central) unit cell. The difference between targeting a crystal structure and targeting a molecule is that for crystal structures, a cell expansion process is required, and the search for atoms is not limited to a single molecule.

[0131] Specifically, constructing an edge set {(i,j)} for a node based on preset boundary conditions to obtain the edge length and position vector of edge (i,j) can include the following operations.

[0132] First, expand the cell with one unit cell in the crystal structure as the center to obtain a supercell.

[0133] Then, repeat the following operation until the set of adjacent atoms of each atom in the central unit cell within the supercell is obtained: taking the i-th atom A in the central unit cell as an example. i Centered on the upper limit of the number of edges, the i-th atom A is determined from the supercell according to the preset truncation radius. i The set of adjacent atoms {N j}

[0134] Next, the i-th atom A of the central unit cell is obtained. i With adjacent atom set {N j The set of edges {(i,j)} between atoms in} is used to obtain the edge length and position vector of edge (i,j), where j is an integer greater than or equal to 1.

[0135] In one specific embodiment, several σ-profile data points of a molecule or crystal can be calculated using quantization software. For each data point, the 3D structure of the molecule (corresponding to a specific conformation) is stored in the dataset as a string in xyz format, and the 3D structure of the crystal is stored in the dataset as a string in cif format. The corresponding σ-profile is also stored in the dataset as a list of 61 floating-point numbers. The 3D structure of the molecule or crystal is at the conformational level.

[0136] This embodiment addresses the issues of long calculation times and high costs associated with COSMO-RS chemical potential calculations, as well as the low accuracy of structure-property relationship models. It provides a scheme for rapid and accurate prediction of the σ-profile. Specifically, to address the difficulty in predicting the shielding charge, this embodiment directly predicts the σ-profile from the molecular or crystal structure, thus skipping the complex intermediate step of shielding charge calculation. Since the σ-profile is a fixed-dimensional vector, its prediction is easier than that of the shielding charge, and further calculations of the chemical potential can be conveniently and quickly performed based on the σ-profile.

[0137] Furthermore, to address the issue of low accuracy in structure-activity relationship models, this embodiment uses virtual graphs to describe molecular or crystal structures and employs an equivariant graph neural network model to transform structure-activity relationships, thereby obtaining an accurate σ-profile.

[0138] Another aspect of this application provides a method for training a distribution density function prediction model.

[0139] Figure 5 A flowchart illustrating the training distribution density function prediction model according to an embodiment of this application is shown.

[0140] See Figure 5 The training method includes operation S510, which inputs the training data obtained as described above into the distribution density function prediction model to be trained, and adjusts the model parameters to make the loss function converge, thereby obtaining the trained distribution density function prediction model. The training data includes the three-dimensional structure of the target object and the distribution density function of the shielding charge, and the loss function includes the correlation between the distribution density function prediction result and the distribution density function annotation result.

[0141] Distribution density function (PDF) prediction models can employ various machine learning models, such as deep learning models. The following example illustrates a topology using an isovariant graph neural network for PDF prediction. Atoms in molecules are influenced by complex chemical interactions. For molecular data, point scalar features are typically atomic features, and the connectivity between nodes is either provided by chemical bonds or obtained based on distance thresholds. The virtual graph assigns not only a feature to each node but also a geometric vector, which, when input into the isovariant graph neural network, achieves good prediction results.

[0142] Figure 6 A schematic diagram illustrating the structure of an isovariant graph convolutional network according to an embodiment of this application is shown. See also Figure 6 The distribution density function prediction model may include a sequentially connected feature extraction module and a prediction module.

[0143] The feature extraction module includes a convolutional network used to update the point scalar features N based on the virtual graph. s and point vector features N v The scalar feature of the update point New_N is obtained. s and update point vector features New_N v And update the point scalar feature New_N s Update point vector features New_N v , boundary scalar feature E s and edge vector feature E v As structural feature X. The determination of molecular feature X can employ various related techniques, such as methods similar to those used for extracting molecular features X based on molecular diagrams. The virtual diagram is determined based on the three-dimensional structure of the target object, and includes point scalar features N. s Point vector features N v , boundary scalar feature E s and edge vector feature E v For details, please refer to the relevant sections of the virtual graph above.

[0144] The prediction module includes a fully connected network for determining the distribution density function of shielding charge on the solvated surface of the target object based on structural feature X.

[0145] Specifically, the isomorphic graph convolutional network consists of four basic operations: Hadamard (⊙), Cross (×), and... It is combined with Sigma(∑). Among them, Hadamard is the matrix multiplication operation, Cross is the matrix cross product operation, Sum is the matrix summation operation, and Sigma is the summation operation for vector dimensions.

[0146] In addition, the above distribution density function prediction model may also include an encoding module for encoding the three-dimensional structure of the input target object.

[0147] Specifically, the i-th atom A i Together with the remaining N-1 atoms, they form N-1 edge sets {(i,j), j=1,2,...,i-1,i+1,...,N}, where the edge set is related to A. i Atoms A are composed of j atoms whose distance between atoms is less than or equal to a predetermined cutoff radius. i The set of adjacent atoms of an atom {N j}, the corresponding edge set {(i,j),r ij ≤r cut The number of edges in} is E i The number of edges E corresponding to all atoms i The sum of these is denoted as E, which represents the number of edges in the virtual graph corresponding to the target object.

[0148] Accordingly, the distribution density function prediction model also includes: an encoding module connected to the feature extraction module, used to encode the point scalar features N. s The embedding encoding is represented as an N×F×1 dimensional matrix, and the edge scalar features E are... s The embedding encoding is represented as an E×F×1 dimensional matrix, and the point vector features N are... v The embedding encoding is represented as an N×F×3 dimensional matrix, and the edge vector features E are... v The embedding encoding is represented as an E×F×3 dimensional matrix, where F is the preset feature dimension.

[0149] Specifically, the convolutional network is used to take the result of the matrix correspondence summation operation between the first and second update matrices as the scalar feature New_N of the update point. s And / or, the result of the matrix summation operation between the third and fourth update matrices is used as the update point vector feature New_N. v The first update matrix is ​​(E s ⊙N s The second update matrix is ​​∑(E) v ⊙N v The third update matrix is ​​(E). s ⊙N v The fourth update matrix is ​​the result of the matrix cross product operation between the fifth and sixth update matrices. The fifth update matrix is ​​(E... v ⊙N s The sixth update matrix is ​​(E) v ⊙N v ), ⊙ represents the matrix addition operation, and ∑ represents the summation operation.

[0150] Figure 7 The diagram illustrates a logical schematic of a training distribution density function prediction model according to an embodiment of this application.

[0151] See Figure 7 In this embodiment, the backpropagation algorithm is used for model training. Specifically, training data (including the three-dimensional structure of the target object and annotation information: the distribution density function of the shielding charge) can be input into the distribution density function prediction model. The network parameters are adjusted until the model converges. For example, the similarity between the distribution density function output by the distribution density function prediction model and the annotation information is determined based on the loss function to see if the convergence condition is met.

[0152] In this embodiment, the distribution density function prediction model can specifically include three parts: a structure encoding network, an isomorphic graph convolutional network, and a fully connected network. Specifically, the structure encoding network can convert the string-format three-dimensional structures (three-dimensional structures of molecules or crystals) in the dataset into virtual graphs; the isomorphic graph convolutional network can convert the virtual graphs of molecules or crystals into structural features; and the fully connected network can convert the structural features into σ-profiles.

[0153] The following provides illustrative examples of the equivariant graph convolutional network, fully connected network, and loss function for the distribution density function prediction model.

[0154] Figure 8 A schematic diagram of the structure of a convolutional layer according to an embodiment of this application is shown.

[0155] See Figure 8 By combining the four basic operations as shown below, a convolutional layer of the equivariant graph convolutional network is formed, the purpose of which is to transform N... s N v E s and E v The constructed virtual graph is transformed into a new point scalar feature, New N. s and the new point vector feature New N v As shown in equations (4) and (5).

[0156]

[0157]

[0158] Specifically, the number of convolutional layers, num_conv, can be set. For the input (molecule or crystal) virtual graph, E is kept constant. s and E v Remain unchanged, use New N s and New N v Update Nodes and Nodev iteratively num_conv times, and obtain New N on the num_conv-th iteration. s By compressing the last dimension, we transform it into an N×F matrix representing the structural feature X.

[0159] Figure 8 In this context, Linear represents the linearization operation, and the Softplus function can be viewed as a smoothing of the ReLU function.

[0160] In this embodiment, the point scalar and point vector features are updated through interaction with each other, resulting in new point scalar and point vector features. Specifically, edge information is integrated into node information to form new features, thus improving the feature New N.s With New N v The improved structural representation capabilities make it easier for the model to extract information related to the distribution density function of the shielding charge, ultimately leading to more accurate predictions. It should be noted that the scalar feature New N at the update point is updated based on the updated virtual graph. s and update point vector features New N v The logic and Figure 8 The logic shown is the same or similar, and will not be elaborated further here.

[0161] As shown above, a virtual graph can be used as a descriptor to represent molecules. This virtual graph includes more complete three-dimensional features of the molecules, which helps to improve the accuracy of the determined distribution density function of molecular shielding charge.

[0162] In this embodiment, in addition to using scalar features, vector features are also used in the convolution. Compared with the method that only uses scalar features in related technologies, it is easier and more accurate to extract molecular features.

[0163] Figure 9 A schematic diagram of the structure of a fully connected network according to an embodiment of this application is shown.

[0164] See Figure 9 The fully connected network includes: a first linear layer, a first activation function layer, a second linear layer, and a second activation function layer connected in sequence. The activation functions of the first activation function layer and the second activation function layer are different, and the second activation function layer can perform a normalization operation.

[0165] Fully connected networks typically consist of two parts: linear layers and activation layers. Linear layers, also known as fully connected layers, have each neuron connected to all neurons in the previous layer, achieving linear combinations and transformations of the preceding layer. Furthermore, various activation functions can be used, and no specific restrictions are imposed here.

[0166] Specifically, a fully connected network consists of a combination of three operations: Linear, Softplus, and Softmax. Linear is a matrix linear transformation operation, Softplus is an activation operation, and Softmax is a normalization operation. Through methods such as... Figure 6 The method shown combines three operations to form the entire fully connected network, the purpose of which is to transform the structural feature X into a 1×61-dimensional σ-profile. For the input structural feature X, it can be transformed into a σ-profile after passing through the fully connected network.

[0167] The first activation layer used in this embodiment is the softplus activation layer (because the model's actual prediction performance is good). The second is the softmax activation layer, which directly outputs the σ-profile without using a third linear layer. This is because the data output by softmax is normalized, which is completely consistent with the property of σ-profile (the sum of σ-profiles is 1). Therefore, it simplifies the model structure, improves computational efficiency, and reduces computational resource consumption.

[0168] In some embodiments, the loss function includes spectral information divergence or spectral information similarity.

[0169] The formula for calculating the spectral information divergence SID is shown in equation (6). Here, X represents the feature.

[0170]

[0171] The formula for calculating spectral information similarity (SIS) is shown in equation (7).

[0172] SIS = 1 / (1+SID) Equation (7)

[0173] Specifically, the loss function and convergence criterion can be set first. Input the structure string into the model to obtain the predicted value of σ-profile. Combine the predicted value with the true value to calculate the loss function based on equation (6) or equation (7). If the loss function does not reach the convergence criterion, update the network parameters and enter the next iteration. If the convergence criterion is reached, output the network parameters.

[0174] It should be noted that, in practice, only one of the two loss functions provided in this embodiment needs to be selected; it is simply a one-to-one operation between the two vectors. In addition to comparing the numbers in the two vectors, the loss function in this embodiment also compares the correlation between the two vectors (used to evaluate their similarity). Furthermore, this embodiment does not exclude the possibility of using both loss functions in combination; for example, the two loss functions can be weighted and summed, used in segments, or used in multiple steps, etc., which are not limited here.

[0175] In this embodiment, considering that loss functions such as MAE and RMSE can only describe the difference between two numbers, and cannot describe the similarity and correlation between two vectors, a suitable loss function is constructed in this embodiment, such as using spectral information divergence and / or spectral information similarity as the loss function.

[0176] In some embodiments, in order to improve the accuracy of the prediction results of the trained distribution density function prediction model, the trained distribution density function prediction model can be fine-tuned using a test dataset.

[0177] Specifically, the above method may also include the following operation: dividing the training data set into a training data subset and a test data subset.

[0178] Accordingly, training the probability density function (PDF) prediction model may include the following steps: First, train the PDF prediction model using a subset of training data to obtain the trained PDF prediction model. Then, test the PDF prediction model using a subset of test data to adjust the trained PDF prediction model.

[0179] In some embodiments, the above method can also train multiple distribution density function prediction models, and determine the final prediction result based on the prediction sub-results of each of the multiple distribution density function prediction models. For example, the distribution density function of the shielding charge of the target object can be determined by weighted summation, weighted averaging, or other methods.

[0180] Specifically, the above method may also include the following operations.

[0181] First, the training dataset is divided into a specified number of sub-training datasets. For example, the specified number of sub-datasets can be determined based on expert experience or the accuracy of predicting the distribution density function of the molecular shielding charge. The specified number of sub-datasets could be 3, 5, 8, 10, 13, 18, 20, etc.

[0182] Then, construct the same number of distribution density function prediction models as the specified number of parts. This allows training multiple distribution density function prediction models, from which the model with the best accuracy in predicting the distribution density function of the shielding charge can be selected, or the average or weighted average of the outputs of multiple models can be used as the final prediction result.

[0183] Accordingly, inputting training data into the molecular encoding network includes: inputting the training data from each sub-training dataset into the molecular encoding network of different distribution density function prediction models, so as to train the different distribution density function prediction models separately and obtain multiple trained distribution density function prediction models with the same number as the specified number of subsets.

[0184] Specifically, for each distribution density function prediction model, the model is initialized using the network parameters obtained from the training model, and the structural information of the target object to be predicted (such as molecules in xyz format or crystals in cif format) is input into the above distribution density function prediction model to obtain the predicted σ-profile.

[0185] This embodiment provides an isomorphic graph neural network model for σ-profile prediction. The prediction process of this model consists of three steps: structure encoding, isomorphic graph convolution, and σ-profile prediction. The structure encoding step represents the molecule or crystal as a virtual graph with feature encoding. The isomorphic graph convolution step transforms the virtual graph into a matrix-like feature representation. The σ-profile prediction step predicts the σ-profile based on the feature representation of the structure using a fully connected neural network.

[0186] In addition, the dataset can be equally divided into ten parts, and the model can be trained using ten-fold cross-validation. This continues until the loss function on the validation set no longer decreases (i.e., the loss function converges, the difference between two consecutive loss functions is less than a preset value, such as 0.0005, 0.001, 0.002, etc.), resulting in ten equally variable graphical neural network models. It should be noted that five-fold cross-validation or k-fold cross-validation can also be used to improve model accuracy.

[0187] Another aspect of this application provides a method for determining the distribution density function of the shielding charge. Figure 10 A flowchart illustrating a method for determining the distribution density function of shielding charge according to an embodiment of this application is shown schematically.

[0188] See Figure 10 The method for determining the distribution density function of the shielding charge as described above includes: operating S1010, using the distribution density function prediction model trained as described above, processing the three-dimensional structure or virtual image of the target object, and obtaining the distribution density function of the shielding charge for the target object, wherein the virtual image is generated based on the three-dimensional structure of the target object.

[0189] Figure 11 The illustration shows a logical diagram of determining the distribution density function of the shielding charge according to an embodiment of this application.

[0190] See Figure 11 The model parameters (such as network parameters) obtained from training the distribution density function prediction model can be transferred to the constructed model (such as downloading the model topology from the cloud to the local machine) to obtain a trained distribution density function prediction model. This allows the trained distribution density function prediction model to process structures (such as the structural information of molecules and / or crystals) and obtain the corresponding σ-profile.

[0191] In some embodiments, the target object includes molecules and / or crystal structures. Accordingly, Figure 11 The structure can be a molecular structure and / or a crystal structure.

[0192] Specific embodiment one, taking the σ-profile of the predicted molecule as an example, is illustrated by way of example.

[0193] First, 40,000 conformational molecular structures listed in the COSMOtherm database were collected, and the σ-profiles corresponding to these 40,000 structures were calculated using COSMOtherm. The 40,000 structures were stored in the dataset in xyz format, and the corresponding 40,000 σ-profiles were stored in the dataset as floating-point numbers, forming the training set.

[0194] Then, set the feature dimension F to 64 and the cutoff radius rcut to... The number of convolutional layers, num_conv, is set to 4. Initialize the equivariant graph neural network model.

[0195] Next, the loss function is set to spectral information divergence, and the convergence criterion is set to 0.03. The training set is then divided according to... Figure 7 The input is fed into the isotropic graph neural network model as shown, and the predicted value of the σ-profile is calculated. The loss function is then calculated by combining the predicted value with the true value. If the loss function is greater than the convergence criterion of 0.03, the network parameters are updated, a new σ-profile is generated, and the loss function is recalculated. If the loss function is less than or equal to 0.03, the network parameters are output, and training ends. It should be noted that the model training process is not necessary every time; a trained distribution density function prediction model can be obtained through model parameter transfer.

[0196] Next, 8000 conformational molecular structures not included in the training set were collected from the COSMOtherm database, and σ-profiles corresponding to these 8000 structures were calculated using COSMOtherm. The 8000 structures were stored in the dataset in xyz format, and the corresponding 8000 σ-profiles were stored in the dataset as floating-point numbers, forming the test set. The structures in the test set were then compared with the network parameters output in the previous step according to... Figure 11 The model is input as shown, and the predicted value of the σ-profile is calculated. Table 1 shows the spectral information divergence between the true values ​​and the model's predicted values ​​on the training and test sets. It can be seen that the spectral information divergence of the model on the test and training sets is small, indicating that the predicted values ​​are extremely accurate.

[0197] Table 1. Performance of the model on the training and test sets.

[0198] training set 40000 0.02441 test set 8000 0.03275

[0199] To more intuitively illustrate the similarity between the two, Figure 12 The predicted and true σ-profiles corresponding to nine structural information points in the test set are shown. Solid lines represent true values, and dashed lines represent predicted values.

[0200] Specific embodiment two is illustrated by taking the prediction of crystal structure σ-profile as an example.

[0201] Two hundred and ten crystal structures were collected, and the σ-profiles corresponding to these two structures were calculated using COSMOtherm. The two hundred and ten structures were stored in the dataset in CIF format, and the two hundred and ten σ-profiles were stored in the dataset as floating-point numbers, forming the training set.

[0202] Set the feature dimension F to 128 and the cutoff radius rcut to... The number of convolutional layers, num_conv, is set to 4. Initialize the equivariant graph neural network model.

[0203] The loss function is set to spectral information similarity, and the convergence criterion is set to 0.95. The training set is then divided into... Figure 7 The input is fed into the isomorphic graph neural network model as shown, and the predicted value of the σ-profile is calculated. The loss function is then calculated by combining the predicted value with the true value. If the loss function is less than the convergence criterion of 0.95, the network parameters are updated, a new σ-profile is generated, and the loss function is recalculated. If the loss function is greater than or equal to 0.95, the network parameters are output, and training ends.

[0204] Collect 3000 crystal structures not in the training set, and use COSMOtherm to calculate the σ-profiles corresponding to these 3000 structures. Store the 3000 structures in CIF format in the dataset, and the corresponding 3000 σ-profiles in floating-point format in the dataset, forming the test set. Then, combine the structures in the test set with the network parameters output from the previous step according to... Figure 11 The model is input as shown, and the predicted value of the σ-profile is calculated. Table 2 shows the spectral information similarity between the true values ​​and the model's predicted values ​​on the training and test sets. It can be seen that the spectral information similarity between the model and the training set is very high.

[0205] It should be noted that the two distribution density function prediction models in Specific Embodiment 1 and Specific Embodiment 2 can be used simultaneously. There is no difference between using them separately and simultaneously. A large prediction model can include sub-models that predict σ-profiles of molecular and crystal structures. The loss functions in Specific Embodiment 1 and Specific Embodiment 2 can be the same or different, that is, at least two loss functions can be used. In addition, a single loss function can be used, such as spectral information divergence or spectral information similarity. Loss functions can also be selected simultaneously, but only one of them is used as the loss function for simultaneous evaluation and adjustment, while the other is only used as the evaluation loss function.

[0206] For example, spectral information similarity can be used as the loss function in Specific Implementation One, while spectral information divergence can be used as the loss function in Specific Implementation Two. Smaller spectral information divergence is better, generally less than 0.05, but can be adjusted. Larger spectral information similarity is better, generally greater than 0.9, but can also be adjusted. Smaller spectral information divergence and larger spectral information similarity result in a more accurate model, and vice versa.

[0207] Table 2. Model performance on training and test sets.

[0208]

[0209]

[0210] To more intuitively illustrate the similarity between the two, Figure 13 The test set shows the predicted and actual σ-profiles for nine structures. Solid lines represent actual values, and dashed lines represent predicted values.

[0211] In some embodiments, after obtaining the distribution density function of the shielding charge of the target object, the above method may further include the following operation: determining the chemical potential of the target object based on the distribution density function of the shielding charge of the target object.

[0212] As shown above, the distribution density function prediction model can include at least one of the following networks.

[0213] The structure encoding network is configured to transform three-dimensional structures (such as the three-dimensional structures of molecules or crystals) in string form from a dataset into virtual graphs.

[0214] Equivariant graph convolutional networks are configured to transform virtual graphs corresponding to molecular or crystal structures into structural features.

[0215] Fully connected networks are configured to transform structural features into σ-profiles.

[0216] This embodiment represents the three-dimensional information of the structure as a virtual graph in the structure encoding network. Besides preserving the three-dimensional information of the molecule or crystal, it can also strictly distinguish different conformations, providing a more accurate description of the three-dimensional structure compared to the structure-property relationship model which uses sequential descriptors. In addition to accuracy, the isovariant graph convolution model predicts the σ-profile in just a few seconds, a significant improvement over the hours to days required by the original quantization calculation scheme.

[0217] Furthermore, the predicted σ-profile can be directly input into software such as COSMOtherm to calculate various physicochemical properties, which is convenient, fast, and requires no other operations.

[0218] Another aspect of this application provides a design method.

[0219] Figure 14 A flowchart illustrating a design method according to an embodiment of this application is shown schematically.

[0220] See Figure 14 The design method may include operation S1410 and operation S1420.

[0221] In operation S1410, the distribution density function of the shielding charge of the target object is determined according to the method shown above, and / or the chemical potential of the target object is determined.

[0222] In operation S1420, candidate drug or material design is performed based on the distribution density function of the shielding charge and / or chemical potential of the target object.

[0223] Another aspect of this application provides a data processing apparatus.

[0224] Figure 15 A block diagram of a data processing apparatus according to an embodiment of this application is shown schematically.

[0225] See Figure 15 The data processing device 1500 may include an information acquisition module 1510 and an associated storage module 1520.

[0226] The information acquisition module 1510 is used to acquire the three-dimensional structure of the target object, and to acquire the distribution density function of the shielding charge corresponding to each three-dimensional structure.

[0227] The associated storage module 1520 is used to associate and store the three-dimensional structure and the distribution density function of the shielding charge corresponding to the three-dimensional structure to obtain a training dataset, so as to train a distribution density function prediction model based on the training dataset. The trained distribution density function prediction model can predict the distribution density function of the shielding charge corresponding to the three-dimensional structure.

[0228] In some embodiments, the data processing apparatus 1500 may further include a virtual graph conversion module for converting the three-dimensional structure of the target object into a virtual graph.

[0229] Accordingly, the associated storage module 1520 is specifically used to associatedly store the virtual graph and the distribution density function of the shielding charge corresponding to the virtual graph.

[0230] In some embodiments, the target object includes molecules, and the information acquisition module 1510 includes: a first position information acquisition unit, a first adjacent atom set acquisition unit, a first edge set acquisition unit, a first edge feature acquisition unit, and a first virtual graph acquisition unit.

[0231] The first position information acquisition unit is used to obtain the i-th atom A in the three-dimensional structure of the molecule. i The position information in the spatial coordinate system, and the position information of the i-th atom A. i Point scalar feature N s and the i-th atom A i Point vector features N v , where i is an integer greater than or equal to 1.

[0232] The first adjacent atom set acquisition unit is used to obtain the atom set of the i-th atom A according to the preset cutoff radius and position information. i Adjacent sets of adjacent atoms {N j}

[0233] The first set of elements is used to obtain the i-th atom A. i With adjacent atom set {N j The set of edges {(i,j)} between atoms in} is used to obtain the edge length and position vector of edge (i,j), where j is an integer greater than or equal to 1.

[0234] The first edge feature acquisition unit is used to determine the edge scalar feature E based on the length of edge (i,j). s Furthermore, the edge vector feature E is determined based on the position vector of edge (i,j). v .

[0235] The first virtual graph acquisition unit is used to utilize the point scalar feature N of the i-th atom. s Point vector features N v , boundary scalar feature E s and edge vector feature E v To form a virtual graph.

[0236] In some embodiments, the target object includes a crystal structure, and the information acquisition module 1510 includes: a second position information acquisition unit, a second side set acquisition unit, a second side feature acquisition unit, and a second virtual graph acquisition unit.

[0237] The second position information acquisition unit is used to treat at least some atoms in the crystal structure as nodes, obtain the position information of the nodes in the spatial coordinate system, and obtain the point scalar feature N of the nodes. s The point vector features of the nodes N v ;

[0238] The second edge set acquisition unit is used to construct an edge set {(i,j)} for a node based on preset boundary conditions, so as to obtain the edge length and position vector of edge (i,j), where i is an integer greater than or equal to 1 and j is an integer greater than or equal to 1.

[0239] The second edge feature acquisition unit is used to determine the edge scalar feature E based on the length of edge (i,j). s Furthermore, the edge vector feature E is determined based on the position vector of edge (i,j). v ;

[0240] The second virtual graph acquisition unit is used to utilize the scalar feature N of the nodes. s Point vector features N v , boundary scalar feature E s and edge vector feature E v To form a virtual graph.

[0241] In some embodiments, the second edge set acquisition unit includes: a cell expansion subunit, a cycle subunit, and an edge set acquisition subunit.

[0242] The cell expansion subunit is used to expand the cell with one unit cell in the crystal structure as the center to obtain a supercell.

[0243] The cyclic subunit is used to repeat the following operation until the set of adjacent atoms of each atom in the central unit cell within the supercell is obtained: taking the i-th atom A in the central unit cell as an example. i Centered on the upper limit of the number of edges, the i-th atom A is determined from the supercell according to the preset truncation radius. i The set of adjacent atoms {N j}

[0244] The edge set obtains the sub-unit used to obtain the i-th atom A of the central unit cell. i With adjacent atom set {N j The set of edges {(i,j)} between atoms in} is used to obtain the edge length and position vector of edge (i,j), where j is an integer greater than or equal to 1.

[0245] In some embodiments, the target object comprises N atoms; the point scalar feature N s The dimensions include N×F×1, and the point vector features are N. v The dimensions include N×F×3, and the edge scalar features are E. s The dimensions include E×F×1, and the edge vector features are E v The dimensions include E×F×3, where F is the preset feature dimension.

[0246] In some embodiments, the information acquisition module 1510 includes a surface division unit, a shielded charge acquisition unit, and a histogram distribution statistics unit.

[0247] The surface element division unit is used to determine the surface elements of the solvated surface of the target object.

[0248] The shielding charge acquisition unit is used to process the surface elements of the solvated surface of the target object using a quantization algorithm to obtain the shielding charge of the target object.

[0249] The histogram distribution statistical unit is used to statistically analyze the histogram distribution of the shielding charge of the target object, thereby obtaining the distribution density function of the shielding charge corresponding to the three-dimensional structure of the target object.

[0250] Another aspect of this application provides an apparatus for training a distribution density function prediction model.

[0251] Figure 16 A block diagram of an apparatus for training a distribution density function prediction model according to an embodiment of this application is shown schematically.

[0252] See Figure 16 The apparatus 1600 for training the distribution density function prediction model includes a model training module 1610.

[0253] The model training module 1610 is used to input the obtained training data into the distribution density function prediction model to be trained, and to obtain the trained distribution density function prediction model by adjusting the model parameters to make the loss function converge. The training data includes the three-dimensional structure of the target object and the distribution density function of the shielding charge. The loss function includes the correlation between the distribution density function prediction result and the distribution density function annotation result.

[0254] In some embodiments, the distribution density function prediction model includes a sequentially connected feature extraction module and a prediction module.

[0255] The feature extraction module includes a convolutional network used to update the point scalar features N based on the virtual graph. s and point vector features N v The scalar feature of the update point New_N is obtained. s and update point vector features New_N v And update the point scalar feature New_N s Update point vector features New_N v , boundary scalar feature E s and edge vector feature E v As a structural feature X, the virtual graph is determined based on the three-dimensional structure of the target object, and the virtual graph includes point scalar features N. s Point vector features N v , boundary scalar feature E s and edge vector feature E v .

[0256] The prediction module includes a fully connected network for determining the distribution density function of shielding charge on the solvated surface of the target object based on structural feature X.

[0257] In some embodiments, the i-th atom Ai and the remaining N-1 atoms form N-1 edge sets {(i,j),j=1,2,...,i-1,i+1,...,N}. Atoms j in the edge set whose distance to atom Ai is less than or equal to a preset cutoff radius form the neighboring atom set {Nj} of atom Ai. The corresponding edge set {(i,j),r ij ≤r cut The number of edges in} is E i The number of edges E corresponding to all atoms i The sum of these is denoted as E, which represents the number of edges in the virtual graph corresponding to the target object.

[0258] The distribution density function prediction model also includes an encoding module connected to the feature extraction module, used to encode the point scalar features N. s The embedding encoding is represented as an N×F×1 dimensional matrix, and the edge scalar features E are... s The embedding encoding is represented as an E×F×1 dimensional matrix, and the point vector features N are... v The embedding encoding is represented as an N×F×3 dimensional matrix, and the edge vector features E are... v The embedding encoding is represented as an E×F×3 dimensional matrix, where F is the preset feature dimension.

[0259] In some embodiments, the convolutional network is specifically used to take the result of the matrix correspondence summation operation between the first update matrix and the second update matrix as the scalar feature New_N of the update point. s And / or, the result of the matrix summation operation between the third and fourth update matrices is used as the update point vector feature New_N. v The first update matrix is ​​(E s ⊙N s The second update matrix is ​​∑(E) v ⊙N v The third update matrix is ​​(E). s ⊙N v The fourth update matrix is ​​the result of the matrix cross product operation between the fifth and sixth update matrices. The fifth update matrix is ​​(E... v ⊙N s The sixth update matrix is ​​(E) v ⊙N v ), ⊙ represents the matrix addition operation, and ∑ represents the summation operation.

[0260] In some embodiments, the fully connected network includes: a first linear layer, a first activation function layer, a second linear layer, and a second activation function layer connected in sequence, wherein the activation functions of the first activation function layer and the second activation function layer are different, and the second activation function layer can perform a normalization operation.

[0261] In some embodiments, the loss function includes spectral information divergence or spectral information similarity.

[0262] Spectral information divergence

[0263] Spectral information similarity SIS = 1 / (1+SID).

[0264] In some embodiments, the apparatus 1600 may further include a dataset partitioning module. This dataset partitioning module is used to partition the training dataset into a training data subset and a test data subset.

[0265] The model training module 1610 includes a training unit and a testing unit.

[0266] The training unit uses a subset of training data to train a probability density function prediction model, thus obtaining the trained probability density function prediction model.

[0267] The test unit uses a subset of test data to test the distribution density function prediction model in order to adjust the trained distribution density function prediction model.

[0268] Another aspect of this application provides an apparatus for determining the distribution density function of shielding charge.

[0269] Figure 17 A block diagram of an apparatus for determining the distribution density function of shielding charge according to an embodiment of this application is shown schematically.

[0270] See Figure 17 The apparatus 1700 for determining the distribution density function of the shielding charge includes a prediction module 1710.

[0271] The prediction module 1710 uses a trained distribution density function prediction model to process the three-dimensional structure or virtual image of the target object to obtain the distribution density function of the shielding charge for the target object. The virtual image is generated based on the three-dimensional structure of the target object.

[0272] In some embodiments, the target object includes molecules and / or crystal structures.

[0273] In some embodiments, the above-described apparatus 1700 further includes a chemical potential prediction module.

[0274] This chemical potential prediction module is used to determine the chemical potential of a target object based on the distribution density function of the shielding charge on the target object.

[0275] Another aspect of this application provides a design device.

[0276] Figure 18A block diagram of a design device according to an embodiment of this application is shown schematically.

[0277] See Figure 18 The device 1800 may include an information determination module 1810 and a design module 1820.

[0278] The information determination module 1810 is used to determine the chemical potential of the target object based on the distribution density function of the shielding charge of the target object and / or the chemical potential of the target object.

[0279] Design module 1820 is used for candidate drug design or material design based on the distribution density function of the shielding charge and / or chemical potential of the target object.

[0280] Regarding the devices 1500, 1600, 1700, and 1800 in the above embodiments, the specific methods by which each module and unit performs operations have been described in detail in the embodiments related to the method, and will not be elaborated further here.

[0281] Another aspect of this application provides an electronic device.

[0282] Figure 19 A block diagram illustrating an embodiment of the present application is shown schematically.

[0283] See Figure 19 The electronic device 1900 includes a memory 1910 and a processor 1920.

[0284] The processor 1920 can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.

[0285] Memory 1910 may include various types of storage units, such as system memory, read-only memory (ROM), and permanent storage devices. ROM may store static data or instructions required by processor 1920 or other modules of the computer. Permanent storage devices may be read-write storage devices. Permanent storage devices may be non-volatile storage devices that retain stored instructions and data even when the computer is powered off. In some embodiments, permanent storage devices use mass storage devices (e.g., magnetic or optical disks, flash memory) as permanent storage devices. In other embodiments, permanent storage devices may be removable storage devices (e.g., floppy disks, optical drives). System memory may be a read-write storage device or a volatile read-write storage device, such as dynamic random access memory. System memory may store some or all of the instructions and data required by the processor during operation. Furthermore, memory 1910 may include any combination of computer-readable storage media, including various types of semiconductor memory chips (e.g., DRAM, SRAM, SDRAM, flash memory, programmable read-only memory), and disks and / or optical disks may also be used. In some embodiments, the memory 1910 may include a removable storage device that is readable and / or writable, such as a laser disc (CD), a read-only digital versatile optical disc (e.g., DVD-ROM, dual-layer DVD-ROM), a read-only Blu-ray disc, an ultra-high density optical disc, a flash memory card (e.g., SD card, mini SD card, Micro-SD card, etc.), a magnetic floppy disk, etc. Computer-readable storage media do not contain carrier waves or transient electronic signals transmitted wirelessly or via wired connections.

[0286] The memory 1910 stores executable code, which, when processed by the processor 1920, can cause the processor 1920 to execute some or all of the methods described above.

[0287] Furthermore, the method according to this application can also be implemented as a computer program or computer program product, which includes computer program code instructions for performing some or all of the steps in the method described above.

[0288] Alternatively, this application may be implemented as a computer-readable storage medium (or a non-transitory machine-readable storage medium or a machine-readable storage medium) storing executable code (or computer program or computer instruction code) thereon, which, when executed by a processor of an electronic device (or server, etc.), causes the processor to perform part or all of the steps of the methods described above according to this application.

[0289] The various embodiments of this application have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or improvement of the technology in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. A method for obtaining training data for a probability density function prediction model, characterized in that, include: The three-dimensional structure of the target object is obtained, as well as the distribution density function of the shielding charge corresponding to each of the three-dimensional structures. The target object includes molecular and / or crystal structures. The three-dimensional structure and the distribution density function of the shielding charge corresponding to the three-dimensional structure are stored in association to obtain a training dataset, so as to train a distribution density function prediction model based on the training dataset, wherein the trained distribution density function prediction model can predict the distribution density function of the shielding charge corresponding to the three-dimensional structure. The training dataset is obtained by storing the three-dimensional structure and the distribution density function of the shielding charge corresponding to the three-dimensional structure in association, including: The training dataset is obtained by annotating the three-dimensional structure of the molecule and / or crystal structure with the distribution density function of the shielding charge corresponding to the three-dimensional structure. The distribution density function of the shielding charge represents a fixed-dimensional vector that is discretized within a preset numerical range of the shielding charge in the three-dimensional structure. The distribution density function of the shielding charge is used to determine the chemical potential of the target object in order to design candidate drugs or materials.

2. The method for obtaining training data for the distribution density function prediction model according to claim 1, characterized in that, Also includes: The three-dimensional structure of the target object is converted into a virtual graph; The associated storage of the three-dimensional structure and the distribution density function of the shielding charge corresponding to the three-dimensional structure includes: associated storage of the virtual graph and the distribution density function of the shielding charge corresponding to the virtual graph.

3. The method for obtaining training data for the distribution density function prediction model according to claim 2, characterized in that, When the target object is a molecule, the step of converting the three-dimensional structure of the target object into a virtual image includes: Obtain the i-th atom A in the three-dimensional structure of the molecule. i The position information in the spatial coordinate system, and the position information of the i-th atom A are obtained. i Point scalar feature N s and the i-th atom A i Point vector features N v , where i is an integer greater than or equal to 1; According to the preset cutoff radius and the position information, the i-th atom A is obtained. i Adjacent sets of adjacent atoms {N j }; Obtain the i-th atom A i With adjacent atom set {N j The set of edges between atoms in} to obtain the edge The edge length and position vector, where j is an integer greater than or equal to 1; Based on edge The length of the edge scalar feature E is determined s And based on the edge Position vector determines edge vector feature E v ; Using the point scalar feature N of the i-th atom s Point vector features N v , boundary scalar feature E s and edge vector feature E v This constitutes the virtual graph.

4. The method for obtaining training data for the distribution density function prediction model according to claim 2, characterized in that, When the target object is a crystal structure, the step of converting the three-dimensional structure of the target object into a virtual image includes: By using at least a portion of the atoms in the crystal structure as nodes, the position information of these nodes in the spatial coordinate system is obtained, and the point scalar feature N of the nodes is also obtained. s and the point vector features N of the node v ; Construct an edge set for the node based on preset boundary conditions. to obtain the edge The edge length and position vector, where i is an integer greater than or equal to 1, and j is an integer greater than or equal to 1; Based on edge The length of the edge scalar feature E is determined s And based on the edge Position vector determines edge vector feature E v ; Utilizing the point scalar feature N of the node s Point vector features N v , boundary scalar feature E s and edge vector feature E v This constitutes the virtual graph.

5. The method for obtaining training data for the distribution density function prediction model according to claim 4, characterized in that, The set of edges for the node is constructed based on preset boundary conditions. to obtain the edge The side lengths and position vectors include: By expanding the cell around one of the unit cells in the crystal structure, a supercell is obtained; Repeat the following operation until the set of adjacent atoms of each atom in the central unit cell within the supercell is obtained: taking the i-th atom A in the central unit cell as an example. i Centered on the upper limit of the number of edges, the i-th atom A is determined from the supercell according to a preset truncation radius. i The set of adjacent atoms {N j }; Obtain the i-th atom A of the central unit cell i With adjacent atom set {N j The set of edges between atoms in} to obtain the edge The edge length and position vector, where j is an integer greater than or equal to 1.

6. The method for obtaining training data for the distribution density function prediction model according to claim 3 or 4, characterized in that, The target object comprises N atoms; The point scalar feature N s The dimensions include N×F×1 dimensions, and the point vector features N v The dimensions include N×F×3 dimensions, and the edge scalar feature E s The dimensions include E×F×1 dimensions, and the edge vector feature E v The dimensions include E×F×3, where F is the preset feature dimension.

7. The method for obtaining training data for the distribution density function prediction model according to claim 1, characterized in that, The step of obtaining the distribution density function of the shielding charge corresponding to each of the three-dimensional structures includes: Determine the surface elements of the solvated surface of the target object; The surface elements of the solvated surface of the target object are processed using a quantization algorithm to obtain the shielding charge of the target object; The histogram distribution of the shielding charge of the target object is statistically analyzed to obtain the distribution density function of the shielding charge corresponding to the three-dimensional structure of the target object.

8. A method for training a probability density function prediction model, characterized in that: The training data obtained by the method for obtaining training data of the distribution density function prediction model according to any one of claims 1 to 7 is input into the distribution density function prediction model to be trained. The model parameters are adjusted to make the loss function converge, and the trained distribution density function prediction model is obtained. The training data includes the three-dimensional structure of the target object and the distribution density function of the shielding charge. The loss function includes the correlation between the distribution density function prediction result and the distribution density function annotation result.

9. The method for training a distribution density function prediction model according to claim 8, characterized in that, The distribution density function prediction model includes: a feature extraction module and a prediction module connected in sequence; The feature extraction module includes a convolutional network used to update the point scalar features N based on the virtual graph. s and point vector features N v The scalar feature of the update point New_N is obtained. s and update point vector features New_N v And the update point scalar feature New_N s Update point vector features New_N v , boundary scalar feature E s and edge vector feature E v As structural feature X; the virtual graph is determined based on the three-dimensional structure of the target object, and the virtual graph includes point scalar features N. s Point vector features N v , boundary scalar feature E s and edge vector feature E v ; The prediction module includes a fully connected network for determining the distribution density function of shielding charge on the solvated surface of the target object based on the structural feature X.

10. The method for training a distribution density function prediction model according to claim 9, characterized in that, The i-th atom Ai forms an N-1 edge set with the remaining N-1 atoms. In the edge set, atoms j whose distance to atom Ai is less than or equal to the preset cutoff radius constitute the neighboring atom set {Nj} of atom Ai, and the corresponding edge set... The number of edges in the middle is E i The number of edges E corresponding to all atoms i The sum is denoted as E, which represents the number of edges in the virtual graph corresponding to the target object; The distribution density function prediction model further includes: an encoding module connected to the feature extraction module, used to encode the point scalar features N. s The embedding encoding is represented as an N×F×1 dimensional matrix, and the edge scalar feature E is... s The embedding encoding is represented as an E×F×1 dimensional matrix, and the point vector features N are... v The embedding encoding is represented as an N×F×3 dimensional matrix, and the edge vector feature E is... v The embedding encoding is represented as an E×F×3 dimensional matrix, where F is the preset feature dimension.

11. The method for training a distribution density function prediction model according to claim 9, characterized in that, Specifically, the convolutional network is used to take the result of the matrix correspondence summation operation between the first update matrix and the second update matrix as the scalar feature New_N of the update point. s And / or, the result of the matrix summation operation between the third and fourth update matrices is used as the update point vector feature New_N. v Wherein, the first update matrix is ​​(E s ⊙N s The second update matrix is ​​∑(E) v ⊙N v The third update matrix is ​​(E) s ⊙N v The fourth update matrix is ​​the result of the matrix cross product operation between the fifth and sixth update matrices, and the fifth update matrix is ​​(E... v ⊙N s The sixth update matrix is ​​(E) v ⊙N v ), ⊙ represents the matrix addition operation, and ∑ represents the summation operation.

12. The method for training a distribution density function prediction model according to claim 9, characterized in that, The fully connected network includes: a first linear layer, a first activation function layer, a second linear layer, and a second activation function layer connected in sequence, wherein the activation functions of the first activation function layer and the second activation function layer are different, and the second activation function layer can perform a normalization operation.

13. The method for training a distribution density function prediction model according to any one of claims 8 to 12, characterized in that, The loss function includes spectral information divergence or spectral information similarity; The spectral information divergence The similarity of the spectral information .

14. The method for training a distribution density function prediction model according to any one of claims 8 to 12, characterized in that, Also includes: The training dataset is divided into a training data subset and a test data subset; Training the distribution density function prediction model includes: The probability density function prediction model is trained using the subset of training data to obtain the trained probability density function prediction model. The distribution density function prediction model is tested using the subset of test data to adjust the trained distribution density function prediction model.

15. A method for determining the distribution density function of shielding charge, characterized in that, include: The distribution density function prediction model trained using the method according to any one of claims 8 to 14 processes the three-dimensional structure or virtual graph of the target object to obtain the distribution density function of the shielding charge for the target object, wherein the virtual graph is generated based on the three-dimensional structure of the target object.

16. The method for determining the distribution density function of the shielding charge according to claim 15, characterized in that, The target objects include molecules and / or crystal structures.

17. The method for determining the distribution density function of shielding charge according to claim 15, characterized in that, Also includes: The chemical potential of the target object is determined based on the distribution density function of the shielding charge of the target object.

18. A design method, characterized in that, The method includes: The method according to any one of claims 1 to 16 determines the distribution density function of the shielding charge of the target object, and / or, the method according to claim 17 determines the chemical potential of the target object; Candidate drug or material design is performed based on the distribution density function of the shielding charge and / or chemical potential of the target object.

19. A device for obtaining training data for a probability density function prediction model, characterized in that, include: An information acquisition module is used to acquire the three-dimensional structure of a target object and to acquire the distribution density function of the shielding charge corresponding to each of the three-dimensional structures. The target object includes molecular and / or crystal structures. An associated storage module is used to store the three-dimensional structure and the distribution density function of the shielding charge corresponding to the three-dimensional structure in association, to obtain a training dataset, so as to train a distribution density function prediction model based on the training dataset, wherein the trained distribution density function prediction model can predict the distribution density function of the shielding charge corresponding to the three-dimensional structure. The training dataset is obtained by storing the three-dimensional structure and the distribution density function of the shielding charge corresponding to the three-dimensional structure in association, including: The training dataset is obtained by annotating the three-dimensional structure of the molecule and / or crystal structure with the distribution density function of the shielding charge corresponding to the three-dimensional structure. The distribution density function of the shielding charge represents a fixed-dimensional vector that is discretized within a preset numerical range of the shielding charge in the three-dimensional structure. The distribution density function of the shielding charge is used to determine the chemical potential of the target object in order to design candidate drugs or materials.

20. An apparatus for training a distribution density function prediction model, characterized in that, include: The model training module is used to input the training data obtained based on the device of claim 19 into the distribution density function prediction model to be trained, and to obtain the trained distribution density function prediction model by adjusting the model parameters to make the loss function converge. The training data includes the three-dimensional structure of the target object and the distribution density function of the shielding charge, and the loss function includes the correlation between the distribution density function prediction result and the distribution density function annotation result.

21. An apparatus for determining the distribution density function of shielding charge, characterized in that, include: The prediction module is used to process the three-dimensional structure or virtual image of the target object using a distribution density function prediction model trained by the device according to claim 20, to obtain the distribution density function of the shielding charge for the target object, wherein the virtual image is generated based on the three-dimensional structure of the target object.

22. A design device, characterized in that, include: An information determination module is used to determine the distribution density function of the shielding charge of the target object by the apparatus according to claim 21, and / or to determine the chemical potential of the target object; The design module is used to design candidate drugs or materials based on the distribution density function of the shielding charge and / or chemical potential of the target object.

23. An electronic device, characterized in that, include: processor; as well as A memory having executable code stored thereon, which, when executed by the processor, causes the processor to perform the method according to any one of claims 1-18.

24. A computer-readable storage medium, characterized in that, It stores executable code that, when executed by a processor of an electronic device, causes the processor to perform the method according to any one of claims 1-18.

Citation Information

Patent Citations

  • Plane boundary surface charge density extraction method combined with boundary integral equation method and random method

    CN104376135A

  • Data processing method and device, model training method and free energy prediction method

    CN114218869A