A biological data processing method and device, computer equipment and storage medium
By acquiring atomic property information of macromolecules and using a machine translation model to generate reaction intensity values, the problem of inaccurate prediction of macromolecular affinity in existing technologies has been solved, achieving more accurate prediction of material affinity.
Patent Information
- Application Number
- CN202110087985.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-01-22
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2041-01-22
AI Technical Summary
In existing technologies, when predicting the affinity between macromolecules using machine learning, the overall characteristics of the macromolecules are usually integrated for prediction, which leads to inaccurate prediction results.
By acquiring atomic property information of the first and second molecules, a machine translation model is used to generate reaction strength values between atoms, and the affinity of substances is predicted based on these strength values. An attention mechanism is then used to improve the accuracy of the prediction.
It reduces the expenditure of manpower and resources, enables accurate prediction of the affinity between macromolecules, and improves the accuracy of prediction.
Smart Images

Figure CN114822715B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data processing, and in particular to a biological data processing method and device, a computer device, and a storage medium. BACKGROUND
[0002] With the continuous development of computer networks, machine learning is applied in various scenarios of life. For example, machine learning is applied in the medical field.
[0003] In the prior art, when predicting the affinity between two macromolecular substances through machine learning, a set of rules for learning molecular features can be pre-set, so that the model learns the features related to the affinity between different molecular substances according to the rules. Then, when the model is applied, the model can be input with information related to the two macromolecular substances, and the model can predict the affinity between the two macromolecular substances.
[0004] However, the model trained in this way is usually an integrated prediction of the overall features of macromolecular substances. Since macromolecular substances contain complex molecular structures composed of a plurality of atoms, they usually contain many complex dispersed features, which are not easy to integrate. Therefore, the affinity between macromolecular substances predicted by the model trained by the prior art is not accurate. SUMMARY
[0005] The present application provides a biological data processing method, device, computer device, and storage medium, which can improve the accuracy of the obtained substance affinity.
[0006] In one aspect, the present application provides a biological data processing method, comprising:
[0007] Obtaining atomic attribute information of a first atom included in a first molecular substance and atomic attribute information of a second atom included in a second molecular substance;
[0008] Inputting the atomic attribute information of the first atom and the atomic attribute information of the second atom into a machine translation model;
[0009] In the machine translation model, generating a first reaction intensity value between the first atom and each atom in an atomic set, and a second reaction intensity value between the second atom and each atom; the plurality of atoms in the atomic set include the first atom and the second atom;
[0010] According to the first reaction intensity value between the first atom and each atom and the second reaction intensity value between the second atom and each atom, predicting the substance affinity between the first molecular substance and the second molecular substance.
[0011] The application provides a biological data processing device, comprising:
[0012] An attribute acquisition module is configured to acquire atomic attribute information of a first atom included in a first molecular substance and atomic attribute information of a second atom included in a second molecular substance.
[0013] An attribute input module is configured to input the atomic attribute information of the first atom and the atomic attribute information of the second atom into a machine translation model.
[0014] An intensity generation module is configured to generate, in the machine translation model, a first reaction intensity value between the first atom and each atom in an atomic set and a second reaction intensity value between the second atom and each atom in the atomic set; and the atomic set includes the first atom and the second atom.
[0015] An affinity prediction module is configured to predict a substance affinity between the first molecular substance and the second molecular substance according to the first reaction intensity value between the first atom and each atom and the second reaction intensity value between the second atom and each atom.
[0016] The attribute acquisition module includes:
[0017] A kind acquisition unit is configured to acquire atomic kind information of the first atom and atomic kind information of the second atom.
[0018] A position acquisition unit is configured to scan the first molecular substance based on an electromagnetic wave ray to obtain spatial position information of the first atom, and scan the second molecular substance based on the electromagnetic wave ray to obtain spatial position information of the second atom.
[0019] An attribute generation unit is configured to generate the atomic attribute information of the first atom according to the atomic kind information and the spatial position information of the first atom, and generate the atomic attribute information of the second atom according to the atomic kind information and the spatial position information of the second atom.
[0020] The intensity generation module includes:
[0021] An atomic distance determination unit is configured to determine, in the machine translation model, an atomic distance between the first atom and each atom and an atomic distance between the second atom and each atom according to the spatial position information of the first atom and the spatial position information of the second atom.
[0022] An intensity value generation unit is configured to generate the first reaction intensity value between the first atom and each atom and the second reaction intensity value between the second atom and each atom according to the atomic kind information of the first atom, the atomic distance between the first atom and each atom, the atomic kind information of the second atom, and the atomic distance between the second atom and each atom.
[0023] The apparatus further comprises:
[0024] The substance obtaining module is configured to obtain the protein substance and the ligand substance.
[0025] The pocket obtaining module is configured to obtain a docking pocket in the protein substance for the ligand substance; the docking pocket comprises at least two atoms in the protein substance.
[0026] The substance determining module is configured to determine the first molecular substance according to the docking pocket, and determine the ligand substance as the second molecular substance.
[0027] The affinity prediction module comprises:
[0028] The vector generating unit is configured to perform a pooling operation on the first reaction intensity value between each first atom and each atom and the second reaction intensity value between each second atom and each atom, to generate a prediction feature vector.
[0029] The affinity prediction unit is configured to predict the substance affinity between the first molecular substance and the second molecular substance according to the prediction feature vector.
[0030] The affinity determining unit is configured to determine the substance affinity as the affinity between the protein substance and the ligand substance.
[0031] The apparatus further comprises:
[0032] The binding module is configured to bind the protein substance and the ligand substance to obtain a bound substance structure when the substance affinity between the protein substance and the ligand substance is greater than the affinity threshold.
[0033] The auxiliary module is configured to generate auxiliary analysis information for a target drug structure based on the bound substance structure.
[0034] The number of first atoms is at least two; and the number of second atoms is at least two.
[0035] The binding module comprises:
[0036] The reaction pair determining unit is configured to determine an atomic reaction pair between each first atom and each second atom according to the first reaction intensity value between each first atom and each atom and the second reaction intensity value between each second atom and each atom when the substance affinity between the protein substance and the ligand substance is greater than the affinity threshold; one atomic reaction pair comprises one first atom and one second atom.
[0037] The matter binding unit is configured to bind the protein matter and the ligand matter according to an atomic reaction pair to obtain a bound matter structure, and the atomic reaction pair is configured to determine a binding mode of the protein matter and the ligand matter.
[0038] The attribute input module is configured to:
[0039] The atomic attribute information of the first atom and the atomic attribute information of the second atom are input into an encoder of the machine translation model.
[0040] The intensity generation module is configured to:
[0041] The encoder includes an attention mechanism, and the intensity generation module is configured to generate, based on the attention mechanism, the first reaction intensity value between the first atom and each atom and the second reaction intensity value between the second atom and each atom.
[0042] In an aspect, the present application provides a biological data processing device, which includes:
[0043] The sample attribute acquisition module is configured to acquire atomic attribute information of a first sample atom included in a first sample molecular matter and atomic attribute information of a second sample atom included in a second sample molecular matter.
[0044] The sample attribute input module is configured to input the atomic attribute information of the first sample atom and the atomic attribute information of the second sample atom into an initial machine translation model.
[0045] The sample intensity generation module is configured to generate, in the initial machine translation model, a first sample reaction intensity value between the first sample atom and each sample atom in a sample atom set and a second sample reaction intensity value between the second sample atom and each sample atom, wherein the sample atom set includes the first sample atom and the second sample atom.
[0046] The sample affinity prediction module is configured to predict a sample matter affinity between the first sample molecular matter and the second sample molecular matter according to the first sample reaction intensity value between the first sample atom and each sample atom and the second sample reaction intensity value between the second sample atom and each sample atom.
[0047] The parameter correction module is configured to correct model parameters of the initial machine translation model according to an actual matter affinity between the first sample molecular matter and the second sample molecular matter and the sample matter affinity to obtain a machine translation model.
[0048] The parameter correction module includes:
[0049] The parameter correction unit is configured to correct the model parameters of the initial machine translation model according to the actual matter affinity and the sample matter affinity.
[0050] The revised affinity obtaining unit is configured to obtain a revised substance affinity between the first sample molecular substance and the second sample molecular substance based on the initial machine translation model revised based on the model parameter;
[0051] The model determining unit is configured to determine the initial machine translation model revised based on the model parameter as the machine translation model when the affinity gap between the revised affinity and the actual substance affinity is less than the gap threshold.
[0052] In an aspect, a computer device is provided, including a memory and a processor, the memory storing a computer program, the computer program being executed by the processor to cause the processor to perform the method in the aspect.
[0053] In an aspect, a computer readable storage medium is provided, the computer readable storage medium storing a computer program, the computer program including program instructions, the program instructions being executed by a processor to cause the processor to perform the method in the aspect.
[0054] According to an aspect of the present application, a computer program product or computer program is provided, the computer program product or computer program including computer instructions stored in a computer readable storage medium. The processor of the computer device reads the computer instructions from the computer readable storage medium, and the processor executes the computer instructions to cause the computer device to perform the method provided in the aspect and various optional manners.
[0055] The present application obtains atomic attribute information of a first atom included in a first molecular substance and atomic attribute information of a second atom included in a second molecular substance; inputs the atomic attribute information of the first atom and the atomic attribute information of the second atom into a machine translation model; in the machine translation model, generates a first reaction intensity value between the first atom and each atom in an atomic set and a second reaction intensity value between the second atom and each atom; the multiple atoms in the atomic set include the first atom and the second atom; and predicts a substance affinity between the first molecular substance and the second molecular substance according to the first reaction intensity value between the first atom and each atom and the second reaction intensity value between the second atom and each atom. As can be seen, the method proposed in the present application can predict the substance affinity between the first molecular substance and the second molecular substance through the machine translation model, reducing the human and material resources expenditure in detecting the substance affinity between large molecular substances. In addition, since the machine translation model has an attention mechanism, the first atom of the first molecular substance and the second atom of the second molecular substance can be focused on stronger attention through the machine translation model, thereby realizing accurate prediction of the substance affinity between the first molecular substance and the second molecular substance. BRIEF DESCRIPTION OF DRAWINGS
[0056] In order to more clearly illustrate the technical solutions in the application or the prior art, the drawings needed to be used in the embodiments or prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the application, and other drawings can be obtained by those skilled in the art without creative effort on the basis of these drawings.
[0057] Figure 1 is a structural schematic diagram of a network architecture provided by an embodiment of the application;
[0058] Figure 2 is a scene schematic diagram of data prediction provided by the application;
[0059] Figure 3 is a flow schematic diagram of a biological data processing method provided by the application;
[0060] Figure 4 is a scene schematic diagram of obtaining a molecular substance provided by the application;
[0061] Figure 5 is a scene schematic diagram of predicting affinity provided by the application;
[0062] Figure 6 is a scene schematic diagram of predicting affinity provided by the application;
[0063] Figure 7 is a scene schematic diagram of weight visualization provided by the application;
[0064] Figure 8 is a scene schematic diagram of substance combination provided by the application;
[0065] Figure 9 is a flow schematic diagram of a biological data processing method provided by the application;
[0066] Figure 10 is a scene schematic diagram of model training provided by the application;
[0067] Figure 11 is a structural schematic diagram of a biological data processing device provided by the application;
[0068] Figure 12 is a structural schematic diagram of a biological data processing device provided by the application;
[0069] Figure 13 is a structural schematic diagram of a computer device provided by the application;
[0070] Figure 14is a structural schematic diagram of a computer device provided in the present application. DETAILED DESCRIPTION
[0071] The technical solutions in the present application will be clearly and completely described below with reference to the drawings in the present application. Obviously, the described embodiments are only some of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts fall within the scope of the present application.
[0072] The present application relates to artificial intelligence related technologies. Among them, artificial intelligence (Artificial Intelligence, AI) is to use digital computers or digital computer controlled machines to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology of computer science, which tries to understand the essence of intelligence and produce a new intelligent machine that can react in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines, so that machines have the functions of perception, reasoning and decision-making.
[0073] Artificial intelligence technology is a comprehensive discipline, involving a wide range of fields, both hardware and software technologies. Artificial intelligence basic technologies generally include technologies such as sensors, special artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction system, mechatronics, etc. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning, etc.
[0074] Machine learning in artificial intelligence is mainly involved in the present application. Among them, machine learning (Machine Learning, ML) is a multi-field interdisciplinary subject, involving probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory, etc. It is a subject that studies how computers simulate or implement human learning behavior to acquire new knowledge or skills, and reorganize existing knowledge structure to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental approach to making computers intelligent, and its applications are widespread in various fields of artificial intelligence. Machine learning and deep learning usually include artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and inductive learning.
[0075] Machine learning involved in the present application mainly refers to how to train a machine translation model to realize the prediction of the substance affinity between the first molecular substance and the second molecular substance through the machine translation model. For details, please refer to the following Figure 3The method provided in the application can also be applied to the field of AI medical treatment and the field of AI drug research and development, for example, the affinity between macromolecular substances predicted by the method provided in the application can be used to analyze and develop drugs containing the macromolecular substances.
[0076] Please refer to Figure 1 , Figure 1 is a structural schematic diagram of a network architecture provided by an embodiment of the application. As shown in Figure 1 , the network architecture can include a server 200 and a terminal device cluster, which can include one or more terminal devices, and the number of terminal devices will not be limited here. As shown in Figure 1 , the plurality of terminal devices can specifically include a terminal device 100a, a terminal device 101a, a terminal device 102a, …, and a terminal device 103a; as shown in Figure 1 , the terminal device 100a, the terminal device 101a, the terminal device 102a, …, and the terminal device 103a can all be network-connected with the server 200, so that each terminal device can perform data interaction with the server 200 through network connection.
[0077] As shown in Figure 1 , the server 200 can be a stand-alone physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDNs, and basic cloud computing services such as big data and artificial intelligence platforms. The terminal device can be a smart terminal such as a smartphone, a tablet computer, a notebook computer, a desktop computer, a smart television, etc. The following will take the communication between the terminal device 100a and the server 200 as an example to specifically describe the embodiments of the application.
[0078] Please refer to Figure 2 , Figure 2 is a scenario schematic diagram of data prediction provided by the application. As shown in Figure 2 , the terminal device 100a can obtain atomic species information and atomic position information of atoms contained in a protein substance 100b (as shown in block 102b). The atomic species information of the atoms contained in the protein substance 100b represents the species of the atoms contained in the protein substance 100b, for example, the species of the atoms can be hydrogen, oxygen, or helium, etc. The atomic position information of the atoms contained in the protein substance 100b can be three-dimensional coordinate information obtained by the terminal device 100a by scanning the protein substance 100b with x-rays (electromagnetic wave rays).
[0079] In addition, the terminal device 100a can also obtain the atomic species information and the atomic position information of the atoms contained in the ligand substance 101b (as shown in block 103b). Similarly, the atomic species information of the atoms contained in the ligand substance 101b represents the types of the atoms contained in the ligand substance 101b, for example, the types of the atoms can be nitrogen, carbon, or sulfur, etc. The atomic position information of the atoms contained in the ligand substance 101b can also be three-dimensional coordinate information obtained by the terminal device 100a by calling x-ray to scan the ligand substance 101b.
[0080] The terminal device 100a can give the obtained atomic species information and the atomic position information of the atoms contained in the protein substance 100b and the atomic species information and the atomic position information of the atoms contained in the ligand substance 101b to the server 200. The server 200 can input the obtained atomic species information and the atomic position information of the atoms contained in the protein substance 100b and the atomic species information and the atomic position information of the atoms contained in the ligand substance 101b into the machine translation model 104b.
[0081] The machine translation model 104b can be a transformer model, and the machine translation model 104b includes an attention mechanism. The machine translation model 104b can be used to predict the affinity (i.e., chemical affinity) between macromolecular substances, which can represent the strength of the reaction between macromolecular substances. The greater the affinity between two macromolecular substances, the stronger the reaction between the two macromolecular substances, and vice versa. The above-mentioned protein substance 100b can be a macromolecular substance, and the ligand substance 101b can be a macromolecular substance. The specific process of training the machine translation model 104b can be referred to in the description of the corresponding embodiments below. Figure 9
[0082] The atoms contained in the protein substance 100b and the atoms contained in the ligand substance 101b can be collectively referred to as detection atoms. The machine translation model 104b can generate a weight corresponding to each detection atom using the contained attention mechanism. The weight corresponding to each detection atom can include a weight value between the detection atom and itself and each other detection atom, which represents the strength of the reaction between the detection atoms. The greater the weight value, the stronger the reaction between the corresponding detection atoms, and the smaller the weight value, the weaker the reaction between the corresponding detection atoms. The weight corresponding to each detection atom generated by the machine translation model 104b can be embodied by a matrix, that is, the machine translation model 104b can generate a matrix that can include the weight corresponding to each detection atom, which can be referred to as a weight matrix. As shown in the following example:
[0083] For example, all the detected atoms can include atom 1, atom 2 and atom 3, in other words, the protein substance 100b and the ligand substance 101b can include atom 1, atom 2 and atom 3 in total. The weight matrix generated by the machine translation model 104b is matrix 105b. As shown in the matrix 105b, the weight corresponding to atom 1 is the first row of the matrix 105b, which includes the weight value x1 between atom 1 and atom 1, the weight value x2 between atom 1 and atom 2 and the weight value x3 between atom 1 and atom 3. The weight corresponding to atom 2 is the second row of the matrix 105b, which includes the weight value x4 between atom 2 and atom 1, the weight value x5 between atom 2 and atom 2 and the weight value x6 between atom 2 and atom 3. The weight corresponding to atom 3 is the third row of the matrix 105b, which includes the weight value x7 between atom 3 and atom 1, the weight value x8 between atom 3 and atom 2 and the weight value x9 between atom 3 and atom 3.
[0084] Further, the machine translation model 105b can predict the affinity between the protein substance 100b and the ligand substance 101b according to the weights corresponding to each detected atom generated, which can be referred to as the substance affinity 106b. Then, the server 200 can send the substance affinity 106b between the protein substance 100b and the ligand substance 101b predicted by the machine translation model 104b to the terminal device 100a, and the terminal device 100a can output the substance affinity 106b on the terminal page to display the substance affinity 106b to the researcher, and the researcher can further analyze the protein substance 100b and the ligand substance 101b according to the substance affinity 106b.
[0085] By the method provided in the present application, the affinity between macromolecular substances can be predicted by a machine translation model, and the affinity is also predicted by the weights corresponding to each detected atom generated by the machine translation model based on the attention mechanism. Therefore, it can be understood that the affinity is predicted by paying attention to each detected atom involved and according to the association information between each detected atom (such as the weight value between the detected atoms), which makes the predicted affinity between macromolecular substances very accurate.
[0086] Please refer to Figure 3 , Figure 3 is a flowchart of a biological data processing method provided by the present application, as Figure 3 shown, the method can include:
[0087] In step S101, atomic attribute information of first atoms contained in a first molecular substance is acquired, and atomic attribute information of second atoms contained in a second molecular substance is acquired.
[0088] Specifically, the execution subject in the embodiment of the present application can be a computer device or a computer device cluster composed of multiple computer devices. The computer device can be a server or a terminal device. Therefore, the execution subject in the embodiment of the present application can be a server, a terminal device, or a combination of a server and a terminal device. Here, the execution subject is taken as a server as an example to describe the embodiment of the present application.
[0089] The first molecular substance can be any macromolecular substance, and the first molecular substance can include multiple atoms. The atoms contained in the first molecular substance can be referred to as first atoms, and the number of first atoms can be multiple (at least two). Similarly, the second molecular substance can also be any macromolecular substance, and the second molecular substance can also include multiple atoms. The atoms contained in the second molecular substance can be referred to as second atoms, and the number of second atoms can also be multiple (at least two).
[0090] The server can acquire the atomic attribute information of the first atoms, and the atomic attribute information of the first atoms can include atomic species information of the first atoms and spatial position information of the first atoms.
[0091] The atomic attribute information of the first atoms can be the atomic species of the first atoms, such as the species of hydrogen atoms, the species of oxygen atoms, the species of helium atoms, the species of chlorine atoms, and / or the species of carbon atoms. One first atom can correspond to one atomic species information.
[0092] Optionally, the atomic species information of the first atoms can be manually annotated and input, or obtained by identifying related atomic type identification devices. The spatial position information of the first atoms can be three-dimensional coordinate information of the first atoms contained in the first molecular substance obtained by the server by calling x-rays (electromagnetic wave rays) to scan the first molecular substance. One first atom can correspond to one spatial position information.
[0093] Therefore, the server can obtain the atomic species information of the first atoms and the spatial position information of the first atoms as the atomic attribute information of the first atoms.
[0094] Similarly, the atomic attribute information of the second atoms can be the atomic species of the second atoms, such as the species of hydrogen atoms, the species of oxygen atoms, the species of helium atoms, the species of chlorine atoms, and / or the species of carbon atoms. One second atom can correspond to one atomic species information.
[0095] Optionally, the atomic species information of the second atom can be manually annotated input or obtained by identifying the relevant atomic type identification device. The spatial position information of the second atom can be obtained by the server by calling the x-ray (electromagnetic wave ray) to scan the first molecular substance, and the three-dimensional coordinate information of the second atom contained in the first molecular substance. One second atom can correspond to one spatial position information.
[0096] Therefore, the server can obtain the atomic species information of the second atom and the spatial position information of the second atom as the atomic attribute information of the second atom.
[0097] Optionally, the server can also obtain the protein substance and the ligand substance, and the server can obtain the docking pocket of the protein substance for the ligand substance. The docking pocket is the region of the protein substance that wraps the ligand substance, which can be understood as the region of the protein substance that can combine with the ligand substance. The docking pocket can include multiple (at least two) atoms in the protein substance.
[0098] Therefore, the server can obtain the molecule composed of the atoms in the docking pocket as the first molecular substance, and the ligand substance as the second molecular substance. Therefore, in this case, the affinity (which can be called substance affinity, also called chemical affinity) between the first molecular substance and the second molecular substance is predicted, that is, the affinity between the protein substance and the ligand substance is predicted.
[0099] The affinity between macromolecular substances is used to characterize the strength of the reaction between macromolecular substances. The greater the affinity between macromolecular substances, the stronger the reaction between macromolecular substances. Conversely, the smaller the affinity between macromolecular substances, the weaker the reaction between macromolecular substances.
[0100] Please refer to Figure 4 , Figure 4 is a scene diagram provided by the present application for obtaining a molecular substance. As Figure 4 indicated, the server can obtain the docking pocket 101c of the protein substance 100c for the ligand substance 102c, and the docking pocket 101c is also the part of the protein substance 100c that can combine with the ligand substance 102c. In other words, the docking pocket 101c is also the region of the protein substance 100c that wraps the ligand 102c.
[0101] The server can obtain a first molecular substance by obtaining a plurality of atoms included in the interface pocket 101c, and the first molecular substance can actually be the interface pocket 101c. The server can also directly use the ligand substance 102c as a second molecular substance.
[0102] Through the above process, the server can obtain the first molecular substance and the second molecular substance, and further, the server can also obtain atomic attribute information of the first atom included in the first molecular substance and atomic attribute information of the second atom included in the second molecular substance.
[0103] In step S102, the atomic attribute information of the first atom and the atomic attribute information of the second atom are input into a machine translation model.
[0104] Specifically, the server can input the obtained atomic attribute information of the first atom and the atomic attribute information of the second atom into a machine translation model (i.e., a transformer model). The machine translation model can be used to predict the affinity between macromolecular substances, and the specific training process of the machine translation model can be referred to in the process described in the corresponding embodiments. Figure 9
[0105] Optionally, since the conventional machine translation model includes an encoder and a decoder, in the present embodiment, only the encoder in the machine translation model is used when predicting the affinity between macromolecular substances, and the decoder is not used. In addition, since the machine translation model usually includes a Positional Embedding Layer for processing text, in the present embodiment, the Positional Embedding Layer in the machine translation model can be removed, i.e., the Positional Embedding Layer is not used.
[0106] Therefore, the server can input the obtained atomic attribute information of the first atom and the atomic attribute information of the second atom into the encoder of the machine translation model.
[0107] In step S103, the machine translation model generates a first reaction intensity value between the first atom and each atom in an atomic set, and a second reaction intensity value between the second atom and each atom in the atomic set; the plurality of atoms in the atomic set include the first atom and the second atom.
[0108] Specifically, the set formed by all the first atoms and all the second atoms can be referred to as an atomic set, and the atomic set can include a plurality of atoms, and the plurality of atoms include the first atom and the second atom.
[0109] The machine translation model can obtain the distance (which can be referred to as atomic distance) between the first atom and each atom in the atomic set according to the spatial position information of the first atom in the input atomic attribute information of the first atom and the spatial position information of the second atom in the input atomic attribute information of the second atom. Therefore, it can be understood that the machine translation model can obtain the atomic distance between each atom in the atomic set and itself and other atoms except itself.
[0110] Therefore, the server can generate the weight value between the first atom and each atom in the atomic set and the weight value between the second atom and each atom in the atomic set according to the atomic species information of the first atom in the obtained atomic attribute information of the first atom, the atomic species information of the second atom in the atomic attribute information of the second atom, the atomic distance between the first atom and each atom in the atomic set, and the atomic distance between the second atom and each atom in the atomic set.
[0111] The first atom and an atom in the atomic set can correspond to a weight value, which represents the strength of the reaction between the first atom and the corresponding atom in the atomic set. Therefore, the weight value between the first atom and each atom in the atomic set can be referred to as the first reaction strength value between the first atom and each atom in the atomic set.
[0112] Similarly, the second atom and an atom in the atomic set can correspond to a weight value, which represents the strength of the reaction between the second atom and the corresponding atom in the atomic set. Therefore, the weight value between the second atom and each atom in the atomic set can be referred to as the second reaction strength value between the second atom and each atom in the atomic set.
[0113] The first reaction strength value and the second reaction strength value can be represented by a matrix. The server can generate a weight matrix according to the atomic species information of the first atom, the atomic species information of the second atom, the atomic distance between the first atom and each atom in the atomic set, and the atomic distance between the second atom and each atom in the atomic set. The weight matrix includes the first reaction strength value between the first atom and each atom in the atomic set and the second reaction strength value between the second atom and each atom in the atomic set.
[0114] More specifically, the principle of the machine translation model generating the weight matrix F can be referred to as the following formula (1):
[0115] F = Q*Q T-Gamma*D (1)
[0116] wherein Q is a matrix including atom type information of each atom in the atom set, and the atom type information in Q can be represented by a set symbol. In other words, Q includes atom type information of the first atom and atom type information of the second atom. T denotes a transpose of the matrix Q. Gamma is a learning matrix, and Gamma is a model parameter of the machine translation model. D is also a matrix including atom distances between the first atom and each atom in the atom set, and atom distances between the second atom and each atom in the atom set, the atom distances being atom distances in a crystal structure.
[0117] For example, it is assumed that the first atom includes atom a1, atom a2, atom a3, atom a4, and atom a5, and the second atom includes atom b1, atom b2, and atom b3. Therefore, the plurality of atoms in the atom set includes atom a1, atom a2, atom a3, atom a4, atom a5, atom b1, atom b2, and atom b3.
[0118] Therefore, the weight matrix generated by the machine translation model can include 8 rows and 8 columns, and one row in the weight matrix can be regarded as a weight vector. The first 5 rows of the weight matrix can be weight vectors corresponding to each first atom, and the weight vector corresponding to one first atom includes first reaction strength values (i.e., weight values) between the first atom and each atom in the atom set. The last 3 rows of the weight matrix can be weight vectors corresponding to each second atom, and the weight vector corresponding to one second atom includes second reaction strength values (i.e., weight values) between the second atom and each atom in the atom set.
[0119] For example, the first row of the weight matrix can be a weight vector corresponding to atom a1, and the first row can include a first reaction strength value between atom a1 and atom a1, a first reaction strength value between atom a1 and atom a2, a first reaction strength value between atom a1 and atom a3, a first reaction strength value between atom a1 and atom a4, a first reaction strength value between atom a1 and atom a5, a first reaction strength value between atom a1 and atom b1, a first reaction strength value between atom a1 and atom b2, and a first reaction strength value between atom a1 and atom b3.
[0120] The second row of the weight matrix can be a weight vector corresponding to atom a2, and the second row can include a first reactivity strength value between atom a2 and atom a1, a first reactivity strength value between atom a2 and atom a2, a first reactivity strength value between atom a2 and atom a3, a first reactivity strength value between atom a2 and atom a4, a first reactivity strength value between atom a2 and atom a5, a first reactivity strength value between atom a2 and atom b1, a first reactivity strength value between atom a2 and atom b2, and a first reactivity strength value between atom a2 and atom b3.
[0121] The third row of the weight matrix can be a weight vector corresponding to atom a3, and the third row can include a first reactivity strength value between atom a3 and atom a1, a first reactivity strength value between atom a3 and atom a2, a first reactivity strength value between atom a3 and atom a3, a first reactivity strength value between atom a3 and atom a4, a first reactivity strength value between atom a3 and atom a5, a first reactivity strength value between atom a3 and atom b1, a first reactivity strength value between atom a3 and atom b2, and a first reactivity strength value between atom a3 and atom b3.
[0122] The fourth row of the weight matrix can be a weight vector corresponding to atom a4, and the fourth row can include a first reactivity strength value between atom a4 and atom a1, a first reactivity strength value between atom a4 and atom a2, a first reactivity strength value between atom a4 and atom a3, a first reactivity strength value between atom a4 and atom a4, a first reactivity strength value between atom a4 and atom a5, a first reactivity strength value between atom a4 and atom b1, a first reactivity strength value between atom a4 and atom b2, and a first reactivity strength value between atom a4 and atom b3.
[0123] The fifth row of the weight matrix can be a weight vector corresponding to atom a5, and the fifth row can include a first reactivity strength value between atom a5 and atom a1, a first reactivity strength value between atom a5 and atom a2, a first reactivity strength value between atom a5 and atom a3, a first reactivity strength value between atom a5 and atom a4, a first reactivity strength value between atom a5 and atom a5, a first reactivity strength value between atom a5 and atom b1, a first reactivity strength value between atom a5 and atom b2, and a first reactivity strength value between atom a5 and atom b3.
[0124] The 6th row of the weight matrix can be a weight vector corresponding to atom b1, and the 6th row can include a second reaction intensity value between atom b1 and atom a1, a second reaction intensity value between atom b1 and atom a2, a second reaction intensity value between atom b1 and atom a3, a second reaction intensity value between atom b1 and atom a4, a second reaction intensity value between atom b1 and atom a5, a second reaction intensity value between atom b1 and atom b1, a second reaction intensity value between atom b1 and atom b2, and a second reaction intensity value between atom b1 and atom b3.
[0125] The 7th row of the weight matrix can be a weight vector corresponding to atom b2, and the 7th row can include a second reaction intensity value between atom b2 and atom a1, a second reaction intensity value between atom b2 and atom a2, a second reaction intensity value between atom b2 and atom a3, a second reaction intensity value between atom b2 and atom a4, a second reaction intensity value between atom b2 and atom a5, a second reaction intensity value between atom b2 and atom b1, a second reaction intensity value between atom b2 and atom b2, and a second reaction intensity value between atom b2 and atom b3.
[0126] The 8th row of the weight matrix can be a weight vector corresponding to atom b3, and the 8th row can include a second reaction intensity value between atom b3 and atom a1, a second reaction intensity value between atom b3 and atom a2, a second reaction intensity value between atom b3 and atom a3, a second reaction intensity value between atom b3 and atom a4, a second reaction intensity value between atom b3 and atom a5, a second reaction intensity value between atom b3 and atom b1, a second reaction intensity value between atom b3 and atom b2, and a second reaction intensity value between atom b3 and atom b3.
[0127] Optionally, only the encoder of the machine translation model is used in the embodiments of the present application, and the decoder of the machine translation model is not used. It can be known that there is a self-attention mechanism layer in the machine translation model, the self-attention mechanism layer includes a self-attention mechanism, and the machine translation model can generate the above weight matrix based on the self-attention mechanism layer included in the encoder, that is, the machine translation model can generate the first reaction intensity value between the first atom and each atom in the atom set and the second reaction intensity value between the second atom and each atom in the atom set based on the self-attention mechanism layer included in the encoder.
[0128] Optionally, the self-attention mechanism layer in the machine translation model can be a soft-attention mechanism. By using the soft-attention mechanism, the atomic distance information (reflected by the matrix D in the above formula (1)) can be introduced into the machine translation model. Generally, the closer the atomic distance, the stronger the reaction strength between the atoms. By introducing the soft-attention mechanism, the machine translation model can have the concept of a graph, so that the prediction of the substance affinity between the first molecular substance and the second molecular substance can be realized through the three-dimensional conformation (which can be reflected by the three-dimensional spatial position information of the atoms). This makes the predicted substance affinity have strong accuracy.
[0129] In step S104, the substance affinity between the first molecular substance and the second molecular substance is predicted according to the first reaction strength value between the first atom and each atom and the second reaction strength value between the second atom and each atom.
[0130] Specifically, the server can predict the affinity between the first molecular substance and the second molecular substance according to the first reaction strength value between the first atom and each atom in the atomic set and the second reaction strength value between the second atom and each atom in the atomic set. This affinity can be referred to as the substance affinity between the first molecular substance and the second molecular substance, which is also referred to as the chemical affinity. When the substance affinity is greater, it indicates that the reaction between the first molecular substance and the second molecular substance will be stronger. When the substance affinity is smaller, it indicates that the reaction between the first molecular substance and the second molecular substance will be weaker.
[0131] In the machine translation model, there is also a pooling layer. The server can perform a pooling operation on the first reaction strength value between the first atom and each atom in the atomic set and the second reaction strength value between the second atom and each atom in the atomic set through the pooling layer in the machine translation model. That is, the server can perform a pooling operation (which can be mean pooling) on the weight matrix containing the first reaction strength value and the second reaction strength value through the pooling layer in the machine translation model, so as to generate a vector. This vector is the feature vector recognized by the machine translation model for predicting the substance affinity between the first molecular substance and the second molecular substance, which can be referred to as a prediction feature vector.
[0132] More, the machine translation model further comprises an output layer, which can be a linear layer (Linear layer). In fact, the predicted feature vector is generated by the linear layer (Linear layer) according to the result of the pooling layer. The linear layer (Linear layer) can output the predicted substance affinity between the first molecule and the second molecule according to the generated predicted feature vector. The value range of the substance affinity can be 0 to 12, and thus the substance affinity between the first molecule and the second molecule output by the machine translation model can be any value between 0 and 12. The output of the linear layer (Linear layer) can be referred to as logits, which is the predicted substance affinity between the first molecule and the second molecule.
[0133] Optionally, the Linear layer in the machine translation model can also be a Position-wise Feed Forward layer (FFN, a fully connected layer). In other words, the FFN can also be used as the output layer of the machine translation model. The FFN can also generate the predicted feature vector, and can predict the substance affinity between the first molecule and the second molecule according to the predicted feature vector.
[0134] Generally, the number of layers and parameters of the FFN are more than those of the Linear layer. Therefore, the FFN can be used as the output layer for a prediction process involving complex and numerous prediction parameters (such as model parameters), which can improve the prediction accuracy, for example, the accuracy of the predicted substance affinity between molecules. Conversely, the Linear layer can be used as the output layer for a prediction process involving simple or small number of prediction parameters, which can improve the prediction efficiency, for example, the efficiency of predicting the substance affinity between molecules.
[0135] In addition, since the FFN is a nonlinear output layer and the Linear layer is a linear output layer, when the prediction process involves relatively regular prediction parameters, the linear Linear layer can be used as the output layer, and when the prediction process involves irregular prediction parameters, the nonlinear FFN can be used as the output layer. In this way, the prediction accuracy of the substance affinity between molecules can also be improved.
[0136] Please refer to Figure 5 , Figure 5 is a scenario diagram provided by the present application for predicting the affinity. As Figure 5As shown, the first molecular substance 100d can include 6 first atoms, which can be atom a1, atom a2, atom a3, atom a4, atom a5 and atom a6 respectively. Also, the server can further acquire atomic attribute information of the 6 first atoms, which includes atomic type information (containing atomic types of the first atoms) and spatial position information (containing three-dimensional coordinate information of the first atoms) of the 6 first atoms respectively.
[0137] As shown in block 102d, the atomic type information of the above-mentioned 6 first atoms specifically includes atomic type information a1 of atom a1, atomic type information a2 of atom a2, atomic type information a3 of atom a3, atomic type information a4 of atom a4, atomic type information a5 of atom a5 and atomic type information a6 of atom a6.
[0138] As shown in block 103d, the spatial position information of the above-mentioned 6 first atoms specifically includes spatial position information a1 of atom a1, spatial position information a2 of atom a2, spatial position information a3 of atom a3, spatial position information a4 of atom a4, spatial position information a5 of atom a5 and spatial position information a6 of atom a6.
[0139] The second molecular substance 101d can include 5 second atoms, which can be atom b1, atom b2, atom b3, atom b4 and atom b5 respectively. The server can also acquire atomic attribute information of the 5 second atoms, which includes atomic type information (containing atomic types of the second atoms) and spatial position information (containing three-dimensional coordinate information of the second atoms) of the 5 second atoms.
[0140] As shown in block 104d, the atomic type information of the above-mentioned 5 second atoms specifically can include atomic type information 1 of atom b1, atomic type information 2 of atom b2, atomic type information 3 of atom b3, atomic type information 4 of atom b4 and atomic type information 5 of atom b5.
[0141] As shown in block 105d, the spatial position information of the above-mentioned 5 second atoms specifically can include spatial position information 1 of atom b1, spatial position information 2 of atom b2, spatial position information 3 of atom b3, spatial position information 4 of atom b4 and spatial position information 5 of atom b5.
[0142] The server can input the obtained atomic attribute information of the first atom and the atomic attribute information of the second atom into the machine translation model 106b, and through the machine translation model 106b, the affinity 107d between the first molecular substance 100d and the second molecular substance 101d can be predicted according to the input atomic attribute information of the first atom and the atomic attribute information of the second atom.
[0143] Please refer to Figure 6 , Figure 6 is a scene schematic diagram provided by the present application for predicting affinity. As shown in Figure 6 , the machine translation model transformer is a deep learning model, and the machine translation model 100e can include an input layer, an encoder 104e, a pooling layer, and a linear layer (Linear), which can be the output layer of the machine translation model 100e. The encoder 104e can further include a self-attention layer (Multi-Head Attention, which can be soft-Attention), a linear residual layer (add&norm), and a feed forward neural network layer (feed forward).
[0144] First, the server can input the atomic attribute information of the first atom and the atomic attribute information of the second atom in the box 101e into the input layer of the machine translation model 100e. Further, as shown in box 102e, the server can generate the weight value (i.e. the first reaction intensity value) between the first atom and each atom in the atom set (including each first atom and each second atom), and the weight value (i.e. the second reaction intensity value) between the second atom and each atom in the atom set, through the self-attention layer (Multi-Head Attention), the linear residual layer (add&norm), and the feed forward neural network layer in the encoder 104e of the machine translation model 100e.
[0145] Then, the machine translation model 100e can perform a pooling operation on the first reaction intensity value corresponding to the first atom and the second reaction intensity value corresponding to the second atom in the box 102e through the pooling layer, and the result of the pooling operation is input into the linear layer (Linear), and the predicted feature vector 103e is obtained. The predicted feature vector is the feature generated by the machine translation model 100e for predicting the affinity between the first molecular substance and the second molecular substance.
[0146] The linear layer in the machine translation model can also predict the affinity between the first molecular substance and the second molecular substance according to the generated predicted feature vector 103e. Since the machine translation model 100e has been trained by the actual affinity between the sample molecular substances, it can be understood that the linear layer can identify which feature corresponds to what affinity. Therefore, the linear layer can perform feature recognition on the predicted feature vector 103e, that is, output the corresponding affinity. The output affinity is the predicted affinity 105e between the first molecular substance and the second molecular substance, and the value of the affinity 105e can be 6.
[0147] Further, when the first molecular substance is obtained by the docking pocket in the protein substance for the ligand substance, and the second molecular substance is the ligand substance, the affinity between the first molecular substance and the second molecular substance output by the machine translation model is the affinity between the protein substance and the ligand substance, which characterizes the strength of the reaction between the protein substance and the ligand substance.
[0148] It can be understood that the machine translation model can exist in the server, so the server can call the machine translation model to predict the affinity between the first molecular substance and the second molecular substance. Therefore, the operations performed by the machine translation model as the subject described above are all operations performed by the server.
[0149] Further, the greater the affinity between the macromolecular substances, the better the effect of combining the macromolecular substances, and the easier the combination. Therefore, after obtaining the affinity between the protein substance and the ligand substance, when the server detects that the affinity is greater than the affinity threshold, the server can combine the protein substance and the ligand substance to obtain the combined substance structure of the protein substance and the ligand substance. The value of the affinity threshold can be determined according to the actual application scenario, and is not limited.
[0150] The combined substance structure can be used in the field of drug research, for example, the combined substance structure can be applied to assist in the production and development of a certain drug (the structure of the drug can be referred to as a target drug structure). Therefore, the server can generate auxiliary analysis information for the target drug structure according to the obtained combined substance structure of the protein substance and the ligand substance, and the auxiliary analysis information is used to assist the researchers in further analyzing the target drug structure through the combined substance structure.
[0151] The first reaction intensity value between the first atom obtained in the step S103 and each atom in the atom set and the second reaction intensity value between the second atom and each atom in the atom set can be used to analyze the binding between the first atom and the second atom. As can be seen, the machine translation model can generate the weight value (such as the first reaction intensity value and the second reaction intensity value) between the first atom and the second atom, and the weight value between the first atom and the second atom can be used to analyze the affinity between the first molecular substance and the second molecular substance, which also makes the machine translation model have model interpretability for the predicted result (such as the substance affinity).
[0152] The first reaction intensity value between the first atom obtained in the step S103 and each atom in the atom set and the second reaction intensity value between the second atom and each atom in the atom set can be used to analyze the binding between the first atom and the second atom. As can be seen, the machine translation model can generate the weight value (such as the first reaction intensity value and the second reaction intensity value) between the first atom and the second atom, and the weight value between the first atom and the second atom can be used to analyze the affinity between the first molecular substance and the second molecular substance, which also makes the machine translation model have model interpretability for the predicted result (such as the substance affinity).
[0153] In other words, through the first reaction intensity value between the first atom and each atom in the atom set and the second reaction intensity value between the second atom and each atom in the atom set, the atom pair with weaker reaction and the atom pair with stronger reaction in the first molecular substance and the second molecular substance can be analyzed, which is the model interpretability of the machine translation model, which can help researchers to conduct more detailed and in-depth research and analysis on the affinity between the first molecular substance and the second molecular substance. For example, through the first reaction intensity value and the second reaction intensity value between the first atom and the second atom, the reasons and problems of the binding between the first molecular substance and the second molecular substance can also be better analyzed.
[0154] For example, the first atom can include atom 1, and the second atom can include atom 2. When the first reaction intensity value of atom 1 to atom 2 is greater than the intensity threshold value, and the second reaction intensity value of atom 2 to atom 1 is also greater than the intensity threshold value, it can be considered that the reaction between atom 1 and atom 2 is stronger, and therefore, atom 1 and atom 2 can constitute an atom reaction pair.
[0155] Therefore, when the server detects that the substance affinity between the protein substance and the ligand substance is greater than the affinity threshold value, the server can also obtain the atom reaction pair between the protein substance and the ligand substance, which is the atom reaction pair between the first molecular substance and the second molecular substance.
[0156] The server can bind the protein substance and the ligand substance by an atomic reaction pair between the protein substance and the ligand substance, which can be used for binding the protein substance and the ligand substance, for example, the atoms contained in the protein substance and the ligand substance can be bound according to the atomic reaction pair to achieve the binding between the protein substance and the ligand substance, thereby obtaining the above-mentioned binding substance structure.
[0157] Please refer to Figure 7 , Figure 7 is a scene diagram provided by the present application for weight visualization. As Figure 7 indicated, the first molecular substance can be a docking pocket of the protein substance for the ligand substance, and the second molecular substance can be the ligand substance, so the atoms contained in the docking pocket can be referred to as protein atoms, that is, the first atoms, and the atoms contained in the ligand can be referred to as ligand atoms, that is, the second atoms.
[0158] The weight values (i.e., the first reaction intensity values) between the first atoms generated by the machine translation model and each atom in the atomic set, and the weight values (i.e., the second reaction intensity values) between the second atoms and each atom in the atomic set can be visualized in the image 100f.
[0159] As Figure 7 indicated, the left area of the dividing line in the image 100f can display the weight values of the protein atoms, that is, the first reaction intensity values to which the first atoms belong. The right area of the dividing line in the image 100f can display the weight values of the ligand atoms, that is, the second reaction intensity values to which the second atoms belong.
[0160] From the weight values of the protein atoms and the weight values of the ligand atoms displayed in the image 100f, it can be analyzed that there is a weight alignment relationship between the weight values of the protein atoms and the weight values of the ligand atoms, which can directly help researchers analyze the binding position, binding reason and binding problem between the protein substance to which the protein atoms belong and the ligand substance to which the ligand atoms belong.
[0161] As shown in the image 100f, the weight alignment means that the weight values (the first reaction intensity values) of the protein atoms to the ligand atoms and the weight values (the second reaction intensity values) of the ligand atoms to the protein atoms can be approximately the same. For example, the weight value of the protein atom 1 to the ligand atom 1 and the weight value of the ligand atom 1 to the protein atom 1 can be approximately the same.
[0162] For example, for the weight value (the first reaction intensity value) to which the protein atom in the box 101f on the left side of the boundary line belongs, there is a corresponding weight value (the second reaction intensity value) to which the ligand atom in the box 102f on the right side of the boundary line belongs at the same height of the weight value size.
[0163] For another example, for the weight value (the first reaction intensity value) to which the protein atom in the box 103f on the left side of the boundary line belongs, there is a corresponding weight value (the second reaction intensity value) to which the ligand atom in the box 104f on the right side of the boundary line belongs at the same height of the weight value size.
[0164] Therefore, the first reaction intensity value to which the first atom belongs and the second reaction intensity value to which the second atom belongs have a weight alignment relationship generated by the machine translation model, and the weight alignment relationship can be used to analyze the “key” (as shown in the box 105f) between the first molecular substance and the second molecular substance, that is, the above-mentioned atomic reaction pair.
[0165] Experiments show that the machine translation model is used to predict the affinity between the first molecular substance and the second molecular substance, which can make the predicted affinity have model interpretability, which can be understood as being able to analyze and deduce the reason according to the model prediction result.
[0166] As can be seen from the above, the machine translation model can be used to accurately predict the affinity between the first molecular substance and the second molecular substance, and the affinity predicted by the machine translation model also has interpretability, which can be realized by the first reaction intensity value and the second reaction intensity value generated by the machine translation model.
[0167] Please refer to Figure 8 , Figure 8 is a scene diagram of a substance combination provided by the present application, as shown in Figure 8 The first molecular substance 100g can include six first atoms, as shown in the box 102g, which can include atom 1, atom 2, atom 3, atom 4, atom 5 and atom 6. The second molecular substance 101g can include five second atoms, as shown in the box 103g, which can include atom 7, atom 8, atom 9, atom 10 and atom 11.
[0168] Atom 1 has a first reaction strength value with atoms 7, 8, 9, 10, and 11, respectively; atom 2 has a first reaction strength value with atoms 7, 8, 9, 10, and 11, respectively; atom 3 has a first reaction strength value with atoms 7, 8, 9, 10, and 11, respectively; atom 4 has a first reaction strength value with atoms 7, 8, 9, 10, and 11, respectively; atom 5 has a first reaction strength value with atoms 7, 8, 9, 10, and 11, respectively; and atom 6 has a first reaction strength value with atoms 7, 8, 9, 10, and 11, respectively.
[0169] Similarly, atom 7 has a second reaction strength value with atom 1, atom 2, atom 3, atom 4, atom 5 and atom 6 respectively; atom 8 has a second reaction strength value with atom 1, atom 2, atom 3, atom 4, atom 5 and atom 6 respectively; atom 9 has a second reaction strength value with atom 1, atom 2, atom 3, atom 4, atom 5 and atom 6 respectively; atom 10 has a second reaction strength value with atom 1, atom 2, atom 3, atom 4, atom 5 and atom 6 respectively; and atom 11 has a second reaction strength value with atom 1, atom 2, atom 3, atom 4, atom 5 and atom 6 respectively.
[0170] Therefore, by using the first reaction intensity value of each first atom in box 102g and the second reaction intensity value of each second atom in box 103g, the atomic reaction pair between the first atom and the second atom can be obtained. An atomic reaction pair includes a first atom and a second atom. The weight value of the first atom relative to the second atom (first reaction intensity value) and the weight value of the second atom relative to the first atom in an atomic reaction pair can both be greater than the intensity threshold (which can be set by the user).
[0171] like Figure 8 As shown, it is assumed that the server obtained three atomic reaction pairs, namely atomic reaction pair ①, atomic reaction pair ②, and atomic reaction pair ③. Atomic reaction pair ① includes atoms 1 and 7, atomic reaction pair ② includes atoms 2 and 8, and atomic reaction pair ③ includes atoms 3 and 9.
[0172] Therefore, atoms 1 and 7 in atomic reaction pair ① can be connected together, atoms 2 and 8 in atomic reaction pair ② can be connected together, and atoms 3 and 9 in atomic reaction pair ③ can be connected together to obtain a combined material structure of 104g.
[0173] And, for the same dataset PDBind refined, the present application also conducts experiments on other neural network models, and it is found that the atomic-level machine translation model (Atomic-Transformer) provided by the present application has more accurate effect on the affinity between molecular substances.
[0174] Wherein, the Pearson correlation coefficient of difference (R2) can be used to measure the prediction accuracy of each neural network model for the affinity between molecular substances, the larger the R2, the higher the prediction accuracy, on the contrary, the smaller the R2, the lower the prediction accuracy, as shown in the following experimental results:
[0175] Model RF-Score K-Deep GRID-RF GRID-NN GCNN ACNN Transformer R2 0.58 0.60 0.54 0.53 0.50 0.51 0.62
[0176] Wherein, RF-Score is a model using random forest algorithm, its R2 is 0.58; K-Deep is a deep ranking model, its R2 is 0.60; GRID-RF is a grid model, its R2 is 0.54; GRID-NN is a grid model, its R2 is 0.53; GCNN is a deep neural network, its R2 is 0.50; ACNN is a deep neural network, its R2 is 0.51; Transformer is the machine translation model used by the present application, its R2 is 0.62. As can be seen, the Pearson correlation coefficient of difference (R2) of the Transformer model provided by the present application for predicting the affinity between molecular substances is greater than the Pearson correlation coefficient of difference (R2) of other commonly used models for predicting the affinity between molecular substances, which shows that the Transformer model can be used to more accurately predict the affinity between molecular substances than other models.
[0177] The application obtains atomic attribute information of first atoms contained in a first molecular substance and atomic attribute information of second atoms contained in a second molecular substance; inputs the atomic attribute information of the first atoms and the atomic attribute information of the second atoms into a machine translation model; in the machine translation model, generates first reaction intensity values between the first atoms and each atom in an atomic set respectively and second reaction intensity values between the second atoms and each atom respectively; the multiple atoms in the atomic set include the first atoms and the second atoms; and predicts a substance affinity between the first molecular substance and the second molecular substance according to the first reaction intensity values between the first atoms and each atom respectively and the second reaction intensity values between the second atoms and each atom respectively. As can be seen, the method provided in the application can predict the substance affinity between the first molecular substance and the second molecular substance through the machine translation model, thereby reducing the labor and material costs in detecting the substance affinity between the macromolecular substances. In addition, since the machine translation model has an attention mechanism, the first atoms of the first molecular substance and the second atoms of the second molecular substance can be focused on more strongly through the machine translation model, thereby realizing accurate prediction of the substance affinity between the first molecular substance and the second molecular substance.
[0178] See Figure 9 , Figure 9 is a flowchart of a biological data processing method provided by the application, as shown in Figure 9 , the method can include:
[0179] In step S201, atomic attribute information of first sample atoms contained in a first sample molecular substance and atomic attribute information of second sample atoms contained in a second sample molecular substance are obtained.
[0180] Specifically, the execution subject in the embodiment of the application can also be a computer device or a computer device cluster composed of multiple computer devices. The computer device can be a server or a terminal device. Therefore, the execution subject in the embodiment of the application can be a server, a terminal device, or a combination of a server and a terminal device. Here, the execution subject is taken as a server for describing the embodiment of the application.
[0181] The first sample molecular substance can be any macromolecular substance, and the second sample molecular substance can also be any macromolecular substance. The first sample molecular substance can include multiple atoms, and the atoms contained in the first sample molecular substance can be referred to as first sample atoms. The second sample molecular substance can also include multiple atoms, and the atoms contained in the second sample molecular substance can be referred to as second sample atoms.
[0182] The server can obtain atomic attribute information of the first sample atom and the second sample atom. Corresponding to the first atom, the atomic attribute information of the first sample atom may include the atom type information and spatial location information of the first sample atom. Corresponding to the second atom, the atomic attribute information of the second sample atom may include the atom type information and spatial location information of the second sample atom.
[0183] Step S202: Input the atomic attribute information of the first sample atom and the atomic attribute information of the second sample atom into the initial machine translation model;
[0184] Specifically, the server can input the atomic attribute information of the first sample atom and the atomic attribute information of the second sample atom into the initial machine translation model, which can be an initial transformer model. For example, the server can input the atomic attribute information of the first sample atom and the atomic attribute information of the second sample atom into the encoder of the initial machine translation model.
[0185] Step S203: In the initial machine translation model, generate the first sample reaction strength value between the first sample atom and each sample atom in the sample atom set, and the second sample reaction strength value between the second sample atom and each sample atom; the multiple sample atoms in the sample atom set include the first sample atom and the second sample atom.
[0186] Specifically, the set consisting of the first sample atom and the second sample atom can be called the sample atom set. Therefore, multiple atoms in the sample atom set can include the first sample atom and the second sample atom.
[0187] The server can generate weight values between the first sample atom and each sample atom in the sample atom set in the initial machine translation model. These weight values can be called the first sample response strength values. The server can also generate weight values between the second sample atom and each sample atom in the sample atom set in the initial machine translation model. These weight values can be called the second sample response strength values.
[0188] The method for generating the first sample reaction intensity value and the second sample reaction intensity value is the same as the method for generating the above. Figure 3 The first and second reaction intensity values in the corresponding embodiments are expressed in the same way, as detailed above. Figure 3 The corresponding step S103 in the embodiment.
[0189] Step S204, predicting the sample substance affinity between the first sample molecular substance and the second sample molecular substance according to the first sample reaction intensity value between the first sample atom and each sample atom, and the second sample reaction intensity value between the second sample atom and each sample atom;
[0190] Specifically, the server can predict the affinity between the first sample molecular substance and the second sample molecular substance according to the first sample reaction intensity value between the first sample atom and each sample atom in the sample atom set, and the second sample reaction intensity value between the second sample atom and each sample atom in the sample atom set. This affinity can be referred to as a sample substance affinity.
[0191] The process of predicting the sample substance affinity between the first sample molecular substance and the second sample molecular substance according to the first sample reaction intensity value and the second sample reaction intensity value is the same as the process of predicting the substance affinity between the first molecular substance and the second molecular substance in the corresponding embodiment, and can be referred to in the above Figure 3 The process of predicting the sample substance affinity between the first sample molecular substance and the second sample molecular substance according to the first sample reaction intensity value and the second sample reaction intensity value is the same as the process of predicting the substance affinity between the first molecular substance and the second molecular substance in the corresponding embodiment, and can be referred to in the above Figure 3 The process of predicting the sample substance affinity between the first sample molecular substance and the second sample molecular substance according to the first sample reaction intensity value and the second sample reaction intensity value is the same as the process of predicting the substance affinity between the first molecular substance and the second molecular substance in the corresponding embodiment, and can be referred to in the above
[0192] Step S205, correcting the model parameters of the initial machine translation model according to the actual substance affinity between the first sample molecular substance and the second sample molecular substance, and the sample substance affinity, to obtain a machine translation model;
[0193] Specifically, the server can also obtain the actual affinity between the first sample molecular substance and the second sample molecular substance, which can be referred to as the actual substance affinity between the first sample molecular substance and the second sample molecular substance.
[0194] The server can correct the model parameters of the initial machine translation model by the actual substance affinity and the sample substance affinity predicted by the initial machine translation model. The initial machine translation model whose model parameters are corrected can be referred to as a machine translation model, which can be used to implement the prediction of the substance affinity between the first molecular substance and the second molecular substance in the corresponding embodiment. Figure 3 The prediction of the substance affinity between the first molecular substance and the second molecular substance in the corresponding embodiment.
[0195] Optionally, when the server corrects the model parameters of the initial machine translation model to convergence by the actual substance affinity between the first sample molecular substance and the second sample molecular substance, and the sample substance affinity between the first sample molecular substance and the second sample molecular substance predicted by the initial machine translation model, it is considered that the correction of the model parameters of the initial translation model is completed, and the initial machine translation model at this time can be used as a machine translation model.
[0196] Optionally, the server can correct (i.e., adjust) the model parameters of the initial machine translation model by the actual substance affinity and the sample substance affinity, so that the sample substance affinity continuously approaches the actual substance affinity. The server can predict the affinity between the first sample molecular substance and the second sample molecular substance again by the initial machine translation model with the corrected model parameters, and the predicted affinity at this time can be referred to as the corrected substance affinity of the sample substance affinity.
[0197] When the server detects that the gap (which can be referred to as an affinity gap) between the corrected substance affinity and the actual substance affinity is less than a gap threshold (the numerical value can be set by itself), it can also be considered that the correction of the model parameters of the initial machine translation model is completed, and the initial machine translation model at this time can be used as the trained machine translation model.
[0198] Please refer to Figure 10 , Figure 10 is a scene diagram of model training provided by the present application. As shown in Figure 10 , the initial machine translation model 100h can include a plurality of structures 101h, and each structure 101h can include a self-attention layer (self-Attention) and a feed forward neural network layer (feed forward). The specific number of structures 101h included in the initial machine translation model 100h is determined according to the actual application scenario, and is not limited.
[0199] The server can input the atomic attribute information of the first sample atom contained in the first sample molecular substance and the atomic attribute information of the second sample atom contained in the second sample molecular substance into the initial machine translation model 100h, and then the initial machine translation model 100h can predict the affinity between the first sample molecular substance and the second sample molecular substance according to the input atomic attribute information of the first sample atom and the atomic attribute information of the second sample atom. This affinity can be referred to as a sample substance affinity 103h. The server can also obtain the actual affinity between the first sample molecular substance and the second sample molecular substance, which can be referred to as an actual substance affinity 104h.
[0200] Then the initial machine translation model 100h can perform back propagation on the prediction result by the sample substance affinity 103h and the actual substance affinity 104h to correct the model parameters of the initial machine translation model 100h. After the model parameters of the initial machine translation model 100h are corrected by a plurality of sample data (such as a plurality of first sample molecular substances and a plurality of second sample molecular substances), the initial machine translation model 100h with the corrected model parameters can be used as the final trained machine translation model 105h.
[0201] The method proposed in this application can predict the affinity between macromolecules by training an initial machine translation model, thereby reducing the human and material resources required for detecting the affinity between macromolecules. Furthermore, due to the attention mechanism inherent in the machine translation model, it can focus stronger attention on the atoms contained within the macromolecules, thus achieving accurate prediction of the affinity between macromolecules.
[0202] Please see Figure 11 , Figure 11 This is a schematic diagram of the structure of a biological data processing device provided in this application. Figure 11 It can be used to perform the above Figure 3 The steps in the corresponding application embodiments. For example... Figure 11 As shown, the biological data processing device 1 may include: an attribute acquisition module 11, an attribute input module 12, an intensity generation module 13, and an affinity prediction module 14;
[0203] The attribute acquisition module 11 is used to acquire the atomic attribute information of the first atom contained in the first molecule and the atomic attribute information of the second atom contained in the second molecule.
[0204] The attribute input module 12 is used to input the atomic attribute information of the first atom and the atomic attribute information of the second atom into the machine translation model;
[0205] The intensity generation module 13 is used in the machine translation model to generate a first reaction intensity value between the first atom and each atom in the atom set, and a second reaction intensity value between the second atom and each atom in the atom set; the multiple atoms in the atom set include the first atom and the second atom;
[0206] The affinity prediction module 14 is used to predict the material affinity between the first molecule and the second molecule based on the first reaction strength value between the first atom and each atom respectively and the second reaction strength value between the second atom and each atom respectively.
[0207] For details on the implementation of the specific functions of the attribute acquisition module 11, attribute input module 12, intensity generation module 13, and affinity prediction module 14, please refer to [link to relevant documentation]. Figure 3 Steps S101-S104 in the corresponding embodiments will not be described again here.
[0208] The attribute acquisition module 11 includes: a category acquisition unit 111, a location acquisition unit 112, and an attribute generation unit 113.
[0209] The category acquisition unit 111 is configured to acquire atomic category information of the first atom and atomic category information of the second atom.
[0210] The position acquisition unit 112 is configured to scan the first molecular substance based on the electromagnetic wave ray to obtain spatial position information of the first atom, and scan the second molecular substance based on the electromagnetic wave ray to obtain spatial position information of the second atom.
[0211] The attribute generation unit 113 is configured to generate atomic attribute information of the first atom according to the atomic category information and the spatial position information of the first atom, and generate atomic attribute information of the second atom according to the atomic category information and the spatial position information of the second atom.
[0212] The specific function implementation manners of the category acquisition unit 111, the position acquisition unit 112, and the attribute generation unit 113 can be referred to the Figure 3 The step S101 in the corresponding embodiment will not be repeated here.
[0213] The intensity generation module 13 includes an atomic distance determination unit 131 and an intensity value generation unit 132.
[0214] The atomic distance determination unit 131 is configured to determine, in the machine translation model, atomic distances between the first atom and each atom and atomic distances between the second atom and each atom according to the spatial position information of the first atom and the spatial position information of the second atom.
[0215] The intensity value generation unit 132 is configured to generate first reaction intensity values of the first atom with each atom and second reaction intensity values of the second atom with each atom according to the atomic category information of the first atom, the atomic distances between the first atom and each atom, the atomic category information of the second atom, and the atomic distances between the second atom and each atom.
[0216] The specific function implementation manners of the atomic distance determination unit 131 and the intensity value generation unit 132 can be referred to the Figure 3 The step S103 in the corresponding embodiment will not be repeated here.
[0217] The device 1 further includes a substance acquisition module 15 and a pocket acquisition module 16.
[0218] The substance acquisition module 15 is configured to acquire the protein substance and the ligand substance.
[0219] The pocket acquisition module 16 is configured to acquire a docking pocket in the protein substance for the ligand substance. The docking pocket includes at least two atoms in the protein substance.
[0220] The substance determining module 17 is configured to determine the ligand substance as the second molecular substance according to the determination of the first molecular substance.
[0221] The specific function implementation of the substance obtaining module 15 and the pocket obtaining module 16 can be referred to Figure 3 The step S101 in the corresponding embodiment will not be repeated here.
[0222] The affinity prediction module 14 includes a vector generating unit 141, an affinity prediction unit 142, and an affinity determining unit 143.
[0223] The vector generating unit 141 is configured to perform a pooling operation on the first reaction intensity value between each atom and the first atom and the second reaction intensity value between each atom and the second atom, to generate a prediction feature vector.
[0224] The affinity prediction unit 142 is configured to predict the substance affinity between the first molecular substance and the second molecular substance according to the prediction feature vector.
[0225] The affinity determining unit 143 is configured to determine the substance affinity as the affinity between the protein substance and the ligand substance.
[0226] The specific function implementation of the vector generating unit 141, the affinity prediction unit 142, and the affinity determining unit 143 can be referred to Figure 3 The step S104 in the corresponding embodiment will not be repeated here.
[0227] The device 1 further includes a combination module 18 and an auxiliary module 19.
[0228] The combination module 18 is configured to combine the protein substance and the ligand substance to obtain a combined substance structure when the substance affinity between the protein substance and the ligand substance is greater than the affinity threshold.
[0229] The auxiliary module 19 is configured to generate auxiliary analysis information for a target drug structure based on the combined substance structure.
[0230] The specific function implementation of the combination module 18 and the auxiliary module 19 can be referred to Figure 3 The step S104 in the corresponding embodiment will not be repeated here.
[0231] The number of the first atoms is at least two, and the number of the second atoms is at least two.
[0232] The combination module 18 includes a reaction pair determining unit 181 and a substance combining unit 182.
[0233] The reaction pair determination unit 181 is configured to determine, according to the first reaction intensity value between each first atom and each atom and the second reaction intensity value between each second atom and each atom, an atomic reaction pair between each first atom and each second atom when the substance affinity between the protein substance and the ligand substance is greater than the affinity threshold value; one atomic reaction pair includes one first atom and one second atom.
[0234] The substance combination unit 182 is configured to combine, according to the atomic reaction pair, the protein substance and the ligand substance to obtain a combined substance structure; the atomic reaction pair is used to determine a combination manner of the protein substance and the ligand substance.
[0235] For details of the specific function implementation of the reaction pair determination unit 181 and the substance combination unit 182, please refer to Figure 3 For details of the specific function implementation of the reaction pair determination unit 181 and the substance combination unit 182, please refer to
[0236] The attribute input module 11 is configured to:
[0237] The atomic attribute information of the first atom and the atomic attribute information of the second atom are input into an encoder of the machine translation model;
[0238] The intensity generation module 13 is configured to:
[0239] The first reaction intensity value between the first atom and each atom and the second reaction intensity value between the second atom and each atom are generated based on an attention mechanism included in the encoder.
[0240] The application obtains atomic attribute information of a first atom included in a first molecular substance and atomic attribute information of a second atom included in a second molecular substance; inputs the atomic attribute information of the first atom and the atomic attribute information of the second atom into a machine translation model; in the machine translation model, generates a first reaction intensity value between the first atom and each atom in an atomic set and a second reaction intensity value between the second atom and each atom; the multiple atoms in the atomic set include the first atom and the second atom; and predicts a substance affinity between the first molecular substance and the second molecular substance according to the first reaction intensity value between the first atom and each atom and the second reaction intensity value between the second atom and each atom. As can be seen, the device provided in the application can predict the substance affinity between the first molecular substance and the second molecular substance through the machine translation model, thereby reducing the labor and material costs in detecting the substance affinity between the macromolecular substances. In addition, since the machine translation model has an attention mechanism, the first atom of the first molecular substance and the second atom of the second molecular substance can be focused on more strongly through the machine translation model, thereby realizing accurate prediction of the substance affinity between the first molecular substance and the second molecular substance.
[0241] See Figure 12 , Figure 12 is a structural schematic diagram of a biological data processing device provided by the application. The biological data processing device can be used to execute the steps in the above Figure 9 application embodiments. As Figure 12 indicated, the biological data processing device 1 can include a sample attribute acquisition module 21, a sample attribute input module 22, a sample intensity generation module 23, a sample affinity prediction module 24, and a parameter correction module 25;
[0242] The sample attribute acquisition module 21 is configured to acquire atomic attribute information of a first sample atom included in a first sample molecular substance and atomic attribute information of a second sample atom included in a second sample molecular substance;
[0243] The sample attribute input module 22 is configured to input the atomic attribute information of the first sample atom and the atomic attribute information of the second sample atom into an initial machine translation model;
[0244] The sample intensity generation module 23 is configured to generate, in the initial machine translation model, a first sample reaction intensity value between the first sample atom and each sample atom in a sample atomic set and a second sample reaction intensity value between the second sample atom and each sample atom; the multiple sample atoms in the sample atomic set include the first sample atom and the second sample atom;
[0245] The sample affinity prediction module 24 is used to predict the sample affinity between the first sample molecule and the second sample molecule based on the first sample reaction strength value between the first sample atom and each sample atom, and the second sample reaction strength value between the second sample atom and each sample atom.
[0246] The parameter correction module 25 is used to correct the model parameters of the initial machine translation model based on the actual material affinity between the first sample molecular substance and the second sample molecular substance, as well as the sample material affinity, to obtain the machine translation model.
[0247] For details on the implementation of the sample attribute acquisition module 21, sample attribute input module 22, sample strength generation module 23, sample affinity prediction module 24, and parameter correction module 25, please refer to [link to relevant documentation]. Figure 9 Steps S201-S205 in the corresponding embodiments will not be described again here.
[0248] The parameter correction module 25 includes: a parameter correction unit 251, a correction affinity acquisition unit 252, and a model determination unit 253.
[0249] The parameter correction unit 251 is used to correct the model parameters of the initial machine translation model based on the actual material affinity and the sample material affinity.
[0250] The modified affinity acquisition unit 252 is used to acquire the modified material affinity between the first sample molecular substance and the second sample molecular substance based on the initial machine translation model after model parameter correction;
[0251] The model determination unit 253 is used to determine the initial machine translation model after model parameter correction as the machine translation model when the affinity difference between the corrected affinity and the actual material affinity is less than the gap threshold.
[0252] For details on the specific functional implementation of the parameter correction unit 251, the corrected affinity acquisition unit 252, and the model determination unit 253, please refer to [link to relevant documentation]. Figure 9 Step S205 in the corresponding embodiment will not be described again here.
[0253] The device proposed in this application can predict the affinity between macromolecules by training an initial machine translation model and then using that model, thus reducing the human and material resources required for detecting the affinity between macromolecules. Furthermore, because the machine translation model incorporates an attention mechanism, it can focus more intently on the atoms contained within the macromolecules, thereby achieving accurate prediction of the affinity between macromolecules.
[0254] Please seeFigure 13 , Figure 13 This is a schematic diagram of the structure of a computer device provided in this application. For example... Figure 13 As shown, the computer device 1000 may include a processor 1001, a network interface 1004, and a memory 1005. Furthermore, the computer device 1000 may also include a user interface 1003 and at least one communication bus 1002. The communication bus 1002 is used to enable communication between these components. The user interface 1003 may include a display screen and a keyboard; optionally, the user interface 1003 may also include a standard wired interface or a wireless interface. The network interface 1004 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface). The memory 1005 may be high-speed RAM or non-volatile memory, such as at least one disk storage device. Optionally, the memory 1005 may also be at least one storage device located remotely from the aforementioned processor 1001. Figure 13 As shown, the memory 1005, which serves as a computer storage medium, may include an operating system, a network communication module, a user interface module, and a device control application program.
[0255] exist Figure 13 In the computer device 1000 shown, the network interface 1004 provides network communication functionality; the user interface 1003 is mainly used to provide an input interface for the user; and the processor 1001 can be used to call the device control application stored in the memory 1005 to achieve:
[0256] Obtain the atomic attribute information of the first atom contained in the first molecule, and obtain the atomic attribute information of the second atom contained in the second molecule;
[0257] Input the atomic attribute information of the first atom and the atomic attribute information of the second atom into the machine translation model;
[0258] In the machine translation model, the first reaction strength value between the first atom and each atom in the atom set, and the second reaction strength value between the second atom and each atom in the atom set are generated; the multiple atoms in the atom set include the first atom and the second atom.
[0259] Based on the first reaction strength values between the first atom and each atom respectively, and the second reaction strength values between the second atom and each atom respectively, the material affinity between the first molecule and the second molecule is predicted.
[0260] In one embodiment, when the processor 1001 invokes the device control application stored in the memory 1005, it specifically performs the following steps:
[0261] In one embodiment, when the processor 1001 invokes the device control application stored in the memory 1005, it specifically performs the following steps:
[0262] Obtain the atom type information of the first atom and the atom type information of the second atom;
[0263] The spatial position information of the first atom is obtained by scanning the first molecule of matter with electromagnetic waves, and the spatial position information of the second atom is obtained by scanning the second molecule of matter with electromagnetic waves.
[0264] Based on the atom type and spatial location information of the first atom, the atomic attribute information of the first atom is generated; based on the atom type and spatial location information of the second atom, the atomic attribute information of the second atom is generated.
[0265] In one embodiment, when the processor 1001 invokes the device control application stored in the memory 1005, it specifically performs the following steps:
[0266] In the machine translation model, the atomic distances between the first atom and each atom, and between the second atom and each atom, are determined based on the spatial position information of the first atom and the second atom, respectively.
[0267] Based on the atom type information of the first atom, the atomic distance between the first atom and each atom, the atom type information of the second atom, and the atomic distance between the second atom and each atom, the first reaction strength value between the first atom and each atom and the second reaction strength value between the second atom and each atom are generated.
[0268] In one embodiment, when the processor 1001 invokes the device control application stored in the memory 1005, it specifically performs the following steps:
[0269] To obtain protein and ligand substances;
[0270] Obtain docking pockets for ligands in a protein; the docking pockets comprise at least two atoms in the protein.
[0271] The first molecule is determined based on the docking pocket, and the ligand is determined as the second molecule.
[0272] In one embodiment, when the processor 1001 invokes the device control application stored in the memory 1005, it specifically performs the following steps:
[0273] Pooling operations are performed on the first reaction intensity values between the first atom and each atom, and the second reaction intensity values between the second atom and each atom to generate a predicted feature vector;
[0274] Based on the predicted feature vector, predict the material affinity between the first molecule and the second molecule;
[0275] Material affinity is defined as the affinity between protein substances and ligand substances.
[0276] In one embodiment, when the processor 1001 invokes the device control application stored in the memory 1005, it specifically performs the following steps:
[0277] When the affinity between a protein and a ligand is greater than the affinity threshold, the protein and ligand bind together to obtain a bound substance structure.
[0278] Based on the combined material structure, auxiliary analytical information targeting the structure of the drug is generated.
[0279] The number of the first atom is at least two; the number of the second atom is at least two;
[0280] In one embodiment, when the processor 1001 invokes the device control application stored in the memory 1005, it specifically performs the following steps:
[0281] When the affinity between a protein and a ligand is greater than the affinity threshold, atomic reaction pairs between each first atom and each second atom are determined based on the first reaction strength values between each first atom and each second atom, respectively; an atomic reaction pair includes one first atom and one second atom.
[0282] Based on atomic reaction pairs, protein and ligand substances are combined to obtain the structure of the combined substance; atomic reaction pairs are used to determine the way protein and ligand substances are combined.
[0283] In one embodiment, when the processor 1001 invokes the device control application stored in the memory 1005, it specifically performs the following steps:
[0284] The atomic attribute information of the first atom and the atomic attribute information of the second atom are input into the encoder of the machine translation model;
[0285] In one embodiment, when the processor 1001 invokes the device control application stored in the memory 1005, it specifically performs the following steps:
[0286] Based on the attention mechanism contained in the encoder, a first reaction strength value between the first atom and each atom, and a second reaction strength value between the second atom and each atom are generated.
[0287] Please see Figure 14 , Figure 14 This is a schematic diagram of the structure of a computer device provided in this application. For example... Figure 14 As shown, computer device 2000 may include: processor 2001, network interface 2004, and memory 2005. Furthermore, computer device 2000 may also include: user interface 2003, and at least one communication bus 2002. The communication bus 2002 is used to implement communication between these components. The user interface 2003 may include a display screen and a keyboard; optionally, the user interface 2003 may also include a standard wired interface or a wireless interface. The network interface 2004 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface). The memory 2005 may be high-speed RAM or non-volatile memory, such as at least one disk storage device. Optionally, the memory 2005 may also be at least one storage device located remotely from the aforementioned processor 2001. Figure 14 As shown, the memory 2005, which is a computer storage medium, may include an operating system, a network communication module, a user interface module, and a device control application program.
[0288] exist Figure 14 In the computer device 2000 shown, the network interface 2004 provides network communication functionality; the user interface 2003 is mainly used to provide an input interface for the user; and the processor 2001 can be used to call the device control application program stored in the memory 2005 to achieve:
[0289] Obtain atomic attribute information of the first sample atoms contained in the first sample molecule, and obtain atomic attribute information of the second sample atoms contained in the second sample molecule;
[0290] Input the atomic attribute information of the first sample atom and the atomic attribute information of the second sample atom into the initial machine translation model;
[0291] In the initial machine translation model, a first sample reaction strength value is generated between a first sample atom and each sample atom in the sample atom set, and a second sample reaction strength value is generated between a second sample atom and each sample atom in the sample atom set; the multiple sample atoms in the sample atom set include the first sample atom and the second sample atom.
[0292] Based on the first sample reaction strength values between the first sample atom and each sample atom, and the second sample reaction strength values between the second sample atom and each sample atom, the sample material affinity between the first sample molecule and the second sample molecule is predicted;
[0293] Based on the actual material affinity between the first sample molecules and the second sample molecules, as well as the sample material affinity, the model parameters of the initial machine translation model are corrected to obtain the machine translation model.
[0294] In one implementation, when the processor 2001 calls the device control application stored in the memory 2005, it specifically performs the following steps:
[0295] The model parameters of the initial machine translation model are corrected based on the actual material affinity and the sample material affinity.
[0296] Based on the initial machine translation model after model parameter correction, the corrected material affinity between the first sample molecular substance and the second sample molecular substance is obtained;
[0297] When the affinity difference between the corrected affinity and the actual material affinity is less than the difference threshold, the initial machine translation model with corrected model parameters is determined as the machine translation model.
[0298] Furthermore, it should be noted that this application also provides a computer-readable storage medium storing a computer program executed by the aforementioned biological data processing device 1 and biological data processing device 2. The computer program includes program instructions, which, when executed by the processor, enable the execution of the aforementioned... Figure 3 and Figure 9 The description of the biological data processing method in any of the corresponding embodiments is already provided, and therefore will not be repeated here. Furthermore, the beneficial effects of using the same method will also not be repeated. For technical details not disclosed in the computer storage medium embodiments related to this application, please refer to the description of the method embodiments of this application.
[0299] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.
[0300] The above-disclosed embodiments are merely preferred embodiments of this application and should not be construed as limiting the scope of this application. Therefore, any equivalent variations made in accordance with the claims of this application shall still fall within the scope of this application.
Claims
1. A biological data processing method, characterized in that, include: Obtain the atomic attribute information of the first atom contained in the first molecule, and obtain the atomic attribute information of the second atom contained in the second molecule; The atomic attribute information of the first atom and the atomic attribute information of the second atom are input into the machine translation model; In the machine translation model, a first reaction strength value is generated between the first atom and each atom in the atom set, and a second reaction strength value is generated between the second atom and each atom. The set of atoms includes the first atom and the second atom; wherein, the first reaction intensity value between the first atom and each atom is the weight value between the first atom and each atom generated by the machine translation model based on the included attention mechanism, and the second reaction intensity value between the second atom and each atom is the weight value between the second atom and each atom generated by the machine translation model based on the included attention mechanism; The material affinity between the first molecule and the second molecule is predicted based on the first reaction strength value between the first atom and each of the atoms and the second reaction strength value between the second atom and each of the atoms.
2. The method according to claim 1, characterized in that, The step of obtaining atomic attribute information of the first atom contained in the first molecule and obtaining atomic attribute information of the second atom contained in the second molecule includes: Obtain the atom type information of the first atom and the atom type information of the second atom; The first molecule is scanned using electromagnetic waves to obtain the spatial position information of the first atom, and the second molecule is scanned using the same electromagnetic waves to obtain the spatial position information of the second atom. Based on the atom type information and spatial location information of the first atom, the atomic attribute information of the first atom is generated; based on the atom type information and spatial location information of the second atom, the atomic attribute information of the second atom is generated.
3. The method according to claim 2, characterized in that, In the machine translation model, generating a first reaction strength value between the first atom and each atom in the atom set, and a second reaction strength value between the second atom and each atom, includes: In the machine translation model, the atomic distances between the first atom and each of the atoms, and the atomic distances between the second atom and each of the atoms, are determined based on the spatial position information of the first atom and the second atom. Based on the atom type information of the first atom, the atomic distance between the first atom and each of the atoms, the atom type information of the second atom, and the atomic distance between the second atom and each of the atoms, a first reaction strength value between the first atom and each of the atoms and a second reaction strength value between the second atom and each of the atoms are generated.
4. The method according to claim 1, characterized in that, The method further includes: To obtain protein and ligand substances; Obtain a docking pocket in the protein material for the ligand substance; the docking pocket comprises at least two atoms in the protein material. The first molecule is determined based on the docking pocket, and the ligand is determined as the second molecule.
5. The method according to claim 4, characterized in that, The step of predicting the material affinity between the first molecule and the second molecule based on the first reaction strength value between the first atom and each of the atoms respectively and the second reaction strength value between the second atom and each of the atoms respectively includes: Pooling operations are performed on the first reaction intensity values between the first atom and each of the atoms respectively, and the second reaction intensity values between the second atom and each of the atoms respectively, to generate a predicted feature vector; Based on the predicted feature vector, predict the material affinity between the first molecule and the second molecule; The affinity of the substance is defined as the affinity between the protein substance and the ligand substance.
6. The method according to claim 5, characterized in that, The method further includes; When the affinity between the protein and the ligand is greater than the affinity threshold, the protein and the ligand are combined to obtain a combined substance structure. Based on the structure of the bound material, auxiliary analytical information targeting the structure of the drug is generated.
7. The method according to claim 6, characterized in that, The number of the first atom is at least two; the number of the second atom is at least two; When the affinity between the protein and the ligand is greater than an affinity threshold, the protein and the ligand are bound together to obtain a bound substance structure, including: When the affinity between the protein and the ligand is greater than the affinity threshold, an atomic reaction pair between each first atom and each second atom is determined based on the first reaction strength value between each first atom and each atom and the second reaction strength value between each second atom and each atom; an atomic reaction pair includes one first atom and one second atom. The protein and the ligand are combined according to the atomic reaction pairs to obtain the structure of the combined substance; the atomic reaction pairs are used to determine the mode of combination of the protein and the ligand.
8. The method according to claim 1, characterized in that, The step of inputting the atomic attribute information of the first atom and the atomic attribute information of the second atom into the machine translation model includes: The atomic attribute information of the first atom and the atomic attribute information of the second atom are input into the encoder of the machine translation model; In the machine translation model, generating a first reaction strength value between the first atom and each atom in the atom set, and a second reaction strength value between the second atom and each atom, includes: Based on the attention mechanism included in the encoder, a first reaction intensity value between the first atom and each of the atoms, and a second reaction intensity value between the second atom and each of the atoms are generated.
9. A biological data processing method, characterized in that, include: Obtain atomic attribute information of the first sample atoms contained in the first sample molecule, and obtain atomic attribute information of the second sample atoms contained in the second sample molecule; The atomic attribute information of the first sample atom and the atomic attribute information of the second sample atom are input into the initial machine translation model; In the initial machine translation model, a first sample response strength value is generated between the first sample atom and each sample atom in the sample atom set, and a second sample response strength value is generated between the second sample atom and each sample atom; the sample atom set includes the first sample atom and the second sample atom; wherein, the first sample response strength value between the first sample atom and each sample atom is a weight value generated by the initial machine translation model based on the included attention mechanism between the first sample atom and each sample atom, and the second sample response strength value between the second sample atom and each sample atom is a weight value generated by the machine translation model based on the included attention mechanism between the second sample atom and each sample atom; Based on the first sample reaction strength value between the first sample atom and each of the sample atoms, and the second sample reaction strength value between the second sample atom and each of the sample atoms, the sample affinity between the first sample molecule and the second sample molecule is predicted; Based on the actual material affinity between the first sample molecule and the second sample molecule, and the material affinity of the sample, the model parameters of the initial machine translation model are corrected to obtain the machine translation model.
10. The method according to claim 9, characterized in that, The step of correcting the model parameters of the initial machine translation model based on the actual material affinity between the first sample molecule and the second sample molecule, and the material affinity of the samples, to obtain the machine translation model includes: Based on the actual material affinity and the sample material affinity, the model parameters of the initial machine translation model are corrected; Based on the initial machine translation model after model parameter correction, the corrected material affinity between the first sample molecular substance and the second sample molecular substance is obtained; When the affinity difference between the modified affinity and the actual material affinity is less than a threshold, the initial machine translation model after the model parameters are corrected is determined as the machine translation model.
11. A biological data processing device, characterized in that, include: The attribute acquisition module is used to acquire the atomic attribute information of the first atom contained in the first molecule and the atomic attribute information of the second atom contained in the second molecule. The attribute input module is used to input the atomic attribute information of the first atom and the atomic attribute information of the second atom into the machine translation model; An intensity generation module is used to generate, in the machine translation model, a first reaction intensity value between the first atom and each atom in the atom set, and a second reaction intensity value between the second atom and each atom. The set of atoms includes the first atom and the second atom; wherein, the first reaction intensity value between the first atom and each atom is the weight value between the first atom and each atom generated by the machine translation model based on the included attention mechanism, and the second reaction intensity value between the second atom and each atom is the weight value between the second atom and each atom generated by the machine translation model based on the included attention mechanism; An affinity prediction module is used to predict the material affinity between the first molecule and the second molecule based on the first reaction strength value between the first atom and each of the atoms respectively and the second reaction strength value between the second atom and each of the atoms respectively.
12. The apparatus according to claim 11, characterized in that, The attribute acquisition module includes: A type acquisition unit is used to acquire the atom type information of the first atom and the atom type information of the second atom; The position acquisition unit is used to scan the first molecule based on electromagnetic wave rays to obtain the spatial position information of the first atom, and to scan the second molecule based on the electromagnetic wave rays to obtain the spatial position information of the second atom. The attribute generation unit is used to generate atomic attribute information of the first atom based on the atom type information and spatial location information of the first atom, and to generate atomic attribute information of the second atom based on the atom type information and spatial location information of the second atom.
13. A biological data processing device, characterized in that, include: The sample attribute acquisition module is used to acquire the atomic attribute information of the first sample atoms contained in the first sample molecule and to acquire the atomic attribute information of the second sample atoms contained in the second sample molecule. The sample attribute input module is used to input the atomic attribute information of the first sample atom and the atomic attribute information of the second sample atom into the initial machine translation model; A sample intensity generation module is used in the initial machine translation model to generate a first sample response intensity value between the first sample atom and each sample atom in the sample atom set, and a second sample response intensity value between the second sample atom and each sample atom; the sample atom set includes the first sample atom and the second sample atom; wherein, the first sample response intensity value between the first sample atom and each sample atom is a weight value between the first sample atom and each sample atom generated by the initial machine translation model based on the included attention mechanism, and the second sample response intensity value between the second sample atom and each sample atom is a weight value between the second sample atom and each sample atom generated by the machine translation model based on the included attention mechanism; The sample affinity prediction module is used to predict the sample material affinity between the first sample molecule and the second sample molecule based on the first sample reaction strength value between the first sample atom and each of the sample atoms, and the second sample reaction strength value between the second sample atom and each of the sample atoms. The parameter correction module is used to correct the model parameters of the initial machine translation model based on the actual material affinity between the first sample molecule and the second sample molecule, as well as the material affinity of the sample, to obtain the machine translation model.
14. A computer device, characterized in that, It includes a memory and a processor, the memory storing a computer program that, when executed by the processor, causes the processor to perform the steps of the method according to any one of claims 1-10.
15. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program adapted to be loaded by a processor and executed as described in any one of claims 1-10.
Citation Information
Patent Citations
Molecule generation method and device, computer readable storage medium and terminal equipment
CN111508568A
Method and apparatus for analysis of molecular configurations and combinations
US20050119837A1