Object molecule sequence data processing method and model training method
Through the combination of feature extraction model and molecular interaction prediction model, the interaction between biological molecules is predicted, which solves the problem of prediction in the prior art and achieves efficient and accurate molecular interaction prediction.
Patent Information
- Application Number
- CN202311793908.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-22
- Publication Date
- 2025-06-24
AI Technical Summary
The prior art is difficult to effectively predict the interaction between biomolecules, resulting in difficulties in drug screening and drug development.
By determining at least two object molecular sequence data, inputting them into feature extraction models, obtaining predicted molecular sequence characteristics, and inputting these characteristics into the molecular interaction prediction model to predict the interaction results between molecules.
The molecular interaction prediction based on artificial intelligence is realized, which improves the accuracy and efficiency of prediction, and can be applied to a variety of downstream tasks.
Smart Images

Figure CN120199329A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this specification relate to the field of computer technology, and particularly to a method for processing object molecular sequence data. Background Art
[0002] In the biological field, each biomolecule that constitutes an organism does not exist in isolation, but builds the entire living body and its life activities through the interactions that occur between them. Microscopically, interactions between biomolecules are constantly taking place within each living body. These biomolecules can be generated within the living body, or come from the external natural world or be synthetic. Generally, by studying the interactions between biomolecules, the mechanisms of biological reactions and the essence of life phenomena can be revealed, which is particularly important in aspects such as understanding life, mechanism research, drug screening, and drug development. Therefore, there is an urgent need for an effective technical solution to predict the interactions between biomolecules. Summary of the Invention
[0003] In view of this, the embodiments of this specification provide a method for object molecular sequence data. One or more embodiments of this specification also relate to an object molecular sequence data device, a model training method, a model training device, a computing device, a computer-readable storage medium, and a computer program to solve the technical defects existing in the prior art.
[0004] According to the first aspect of the embodiments of this specification, a method for processing object molecular sequence data is provided, including:
[0005] Determine at least two object molecular sequence data;
[0006] Input the at least two object molecular sequence data into a feature extraction model to obtain the predicted molecular sequence features of each object molecular sequence data among the at least two object molecular sequence data, where the predicted molecular sequence features contain element vectors of molecular elements, and the object molecular sequence data is composed according to the molecular elements;
[0007] Input the predicted molecular sequence features of each object molecular sequence data into a molecular interaction prediction model to obtain the interaction results between the at least two object molecular sequence data, where the molecular interaction prediction model outputs the interaction results between the at least two object molecular sequence data according to the correlation between the element vectors contained in the predicted molecular sequence features of each object molecular sequence data.
[0008] According to the second aspect of the embodiments of this specification, an object molecular sequence data processing device is provided, including:
[0009] A determination module, configured to determine at least two object molecular sequence data;
[0010] A first input module, configured to input the at least two object molecular sequence data into a feature extraction model, and obtain predicted molecular sequence features of each of the at least two object molecular sequence data, wherein the predicted molecular sequence features include element vectors of molecular elements, and the object molecular sequence data is composed according to the molecular elements;
[0011] A second input module, configured to input the predicted molecular sequence features of each object molecular sequence data into a molecular interaction prediction model, and obtain interaction results between the at least two object molecular sequence data, wherein the molecular interaction prediction model outputs the interaction results between the at least two object molecular sequence data according to the correlation between the element vectors included in the predicted molecular sequence features of each object molecular sequence data.
[0012] According to a third aspect of the embodiments of the present specification, there is provided a model training method, including:
[0013] Determine an object molecular sequence data sample;
[0014] Train a feature extraction model and a molecular interaction prediction model according to the object molecular sequence data sample;
[0015] Wherein, the feature extraction model is used to predict the predicted molecular sequence features of the object molecular sequence data, the predicted molecular sequence features include element vectors of molecular elements, the object molecular sequence data is composed according to the molecular elements, and the molecular interaction prediction model is used to output the interaction results between the at least two object molecular sequence data according to the correlation between the element vectors included in the predicted molecular sequence features of each object molecular sequence data.
[0016] According to a fourth aspect of the embodiments of the present specification, there is provided a model training device, including:
[0017] A determination module, configured to determine an object molecular sequence data sample;
[0018] A training module, configured to train a feature extraction model and a molecular interaction prediction model according to the object molecular sequence data sample;
[0019] Among them, the feature extraction model is used to predict the predicted molecular sequence features of the object molecular sequence data. The predicted molecular sequence features include the element vectors of the molecular elements. The object molecular sequence data is composed according to the molecular elements. The molecular interaction prediction model is used to output the interaction results between the at least two object molecular sequence data according to the correlation between the element vectors included in the predicted molecular sequence features of each object molecular sequence data.
[0020] According to a fifth aspect of the embodiments of the present specification, there is provided a computing device, including:
[0021] a memory and a processor;
[0022] The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the above object molecular sequence data processing method or model training method are implemented.
[0023] According to a sixth aspect of the embodiments of the present specification, there is provided a computer-readable storage medium storing computer-executable instructions, and when the instructions are executed by a processor, the steps of the above object molecular sequence data processing method or model training method are implemented.
[0024] According to a seventh aspect of the embodiments of the present specification, there is provided a computer program, wherein when the computer program is executed on a computer, the computer is made to execute the steps of the above object molecular sequence data processing method or model training method.
[0025] An embodiment of the present specification provides an object molecular sequence data processing method, which determines at least two object molecular sequence data; inputs the at least two object molecular sequence data into a feature extraction model to obtain the predicted molecular sequence features of each object molecular sequence data among the at least two object molecular sequence data. Among them, the predicted molecular sequence features include the element vectors of the molecular elements, and the object molecular sequence data is composed according to the molecular elements; inputs the predicted molecular sequence features of each object molecular sequence data into a molecular interaction prediction model to obtain the interaction results between the at least two object molecular sequence data. Among them, the molecular interaction prediction model outputs the interaction results between the at least two object molecular sequence data according to the correlation between the element vectors included in the predicted molecular sequence features of each object molecular sequence data.
[0026] The above method extracts the predicted molecular sequence features of each object molecular sequence data by using a feature extraction model. Moreover, since the predicted molecular sequence features contain the element vectors of molecular elements, feature prediction at the element level is achieved. Then, based on the correlation between the element vectors, a molecular interaction prediction model predicts the interaction results between at least two object molecular sequence data, realizing artificial intelligence-based molecular interaction prediction. Further, based on the predicted molecular sequence features at the element level, the predicted interaction results are made more accurate. Molecular interaction prediction is realized through the combination of an upstream pre-trained large model (i.e., the feature extraction model) and a downstream specific task network (i.e., the molecular interaction prediction model), while ensuring the efficiency and accuracy of the prediction results. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 FIG. is a schematic diagram of an application scenario of a method for processing object molecular sequence data provided by an embodiment of the present specification;
[0028] Figure 2 FIG. is a flowchart of a method for processing object molecular sequence data provided by an embodiment of the present specification;
[0029] Figure 3 FIG. is a schematic diagram of a feature extraction model in the method for processing object molecular sequence data provided by an embodiment of the present specification;
[0030] Figure 4 FIG. is a flowchart of a processing procedure for training a molecular interaction prediction model in the method for processing object molecular sequence data provided by an embodiment of the present specification;
[0031] Figure 5 FIG. is a schematic diagram of the structure of an apparatus for processing object molecular sequence data provided by an embodiment of the present specification;
[0032] Figure 6 FIG. is a flowchart of a model training method provided by an embodiment of the present specification;
[0033] Figure 7 FIG. is a schematic diagram of the structure of a model training apparatus provided by an embodiment of the present specification;
[0034] Figure 8 FIG. is a schematic diagram of a method for processing object molecular sequence data and a model training method provided by an embodiment of the present specification;
[0035] Figure 9 FIG. is a structural block diagram of a computing device provided by an embodiment of the present specification. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0036] Numerous specific details are set forth in the following description to facilitate a full understanding of this specification. However, this specification can be implemented in many other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the connotation of this specification. Therefore, this specification is not limited by the specific implementations disclosed below.
[0037] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a", "the", and "said" used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0038] It should be understood that although the terms first, second, etc. may be used in one or more embodiments of this specification to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".
[0039] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data that have been authorized by the user or fully authorized by all parties. And the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or refuse.
[0040] In one or more embodiments of this specification, a large model refers to a deep learning model with a large number of model parameters, usually containing hundreds of millions, tens of billions, hundreds of billions, trillions, or even more than one quadrillion model parameters. A large model can also be referred to as a Foundation Model. Through pre-training of a large model with a large amount of unlabeled corpus, a pre-trained model with more than hundreds of millions of parameters is produced. This kind of model can adapt to a wide range of downstream tasks and has good generalization ability, such as large language models (LLMs), multi-modal pre-training models, etc.
[0041] When large models are applied in practice, only a small number of samples are needed to fine-tune the pre-trained models, which can then be applied to different tasks. Large models can be widely used in fields such as Natural Language Processing (NLP) and computer vision. Specifically, they can be applied to tasks in the field of computer vision such as Visual Question Answering (VQA), Image Caption (IC), and image generation, as well as tasks in the field of natural language processing such as text-based sentiment classification, text summary generation, and machine translation. The main application scenarios of large models include digital assistants, intelligent robots, search, online education, office software, e-commerce, intelligent design, etc.
[0042] First, the noun terms involved in one or more embodiments of this specification are explained.
[0043] Transformer: A neural network model based on the attention mechanism, consisting of two parts: an encoder and a decoder.
[0044] Intra-Attention: Internal attention mechanism.
[0045] Inter-Attention: An inter-attention mechanism used to process the interactions between different molecules.
[0046] RoPE: A position encoding method.
[0047] FFN: Fully Connected Feedforward Neural Network, fully connected layer.
[0048] LayerNorm: Layer normalization, a normalization technique used in deep neural networks. It can normalize the output of each neuron in the network, so that the output of each layer in the network has a similar distribution.
[0049] Softmax: Also known as the normalized exponential function. It is a generalization of the binary classification function sigmoid to multi-classification. Its purpose is to present the results of multi-classification in the form of probabilities. It maps the outputs of multiple neurons to the interval (0,1) for multi-classification.
[0050] sigmoid: A classification function used to determine whether an object belongs to a certain category based on the probability value of belonging to that category. It is usually used for binary classification.
[0051] Small molecule compounds: In the field of molecular biology and pharmacology, small molecule compounds can be understood as non-peptide low molecular weight organic compounds used to regulate biological processes. Small molecule compounds as research materials have many advantages, such as selectivity, water solubility and cell permeability. Based on these characteristics, most drugs are small molecules. Small molecule compounds are common ingredients in medicines, foods, cosmetics, pesticides, anesthetics and research reagents, and can also exist in the human body and the environment. The basic building blocks of small molecule compounds are atoms, such as hydrogen (H), carbon (C), oxygen (O), nitrogen (N), sulfur (S), phosphorus (P), fluorine (F), chlorine (Cl), bromine (Br), iodine (I), etc. Small molecule compounds are small molecules with active substances. Small molecule components can be made into tablets or capsules that are easily absorbed by the body. For gastrointestinal dissolving tablets, the dissolved active substances will be absorbed by the body through the intestinal wall and enter the blood. Due to their small size, small molecules can reach almost any target in the body. In addition, with their compact structure and chemical composition, small molecules can generally easily penetrate the cell membrane to reach the designated location to produce corresponding biochemical reactions and achieve the designated purpose, such as treating diseases.
[0052] Biological macromolecules include carbohydrates, proteins, lipids (such as fats) and nucleic acids (DNA and RNA).
[0053] Protein macromolecule: Protein is a biological macromolecule, the embodiment of all life activities. It is composed of multiple amino acids through dehydration condensation (amino acid residues) to form a long chain sequence and fold into a macromolecule with a spatial conformation. It is also one of the most important components of the human body. The basic unit of protein is amino acid. Protein is the material basis of life. It is an organic macromolecule composed of many different amino acids. It is the basic substance of cell composition and the most important carrier of life activities and processes. It performs various important functions of biological activities in living bodies.
[0054] Nucleic acid macromolecule: a biological macromolecule usually located in cells, mainly responsible for carrying and transmitting genetic information of organisms. There are two major types of nucleic acids, deoxyribonucleic acid (DNA) and ribonucleic acid (RNA).
[0055] Amino acids are the basic elements of protein. There are about 20 common amino acids and 2 rare ones. In terms of representation, each amino acid is generally represented by one of the 26 letters, namely "ARNDCEQGHILKMFPSTWYVUO". For example, "A" represents alanine, "R" represents arginine, and "O" represents pyridine.
[0056] Nucleotide: The basic building block of nucleic acids. A nucleotide consists of a nitrogenous base at its core, along with a pentose sugar and one or more phosphate groups. There are five types of nitrogenous bases, namely adenine (A), guanine (G), cytosine (C), thymine (T), and uracil (U), and one of these five letters is commonly used to represent them.
[0057] Small molecule SMILES: The full English name is Simplified Molecular Input Line Entry System, which is a common simplified representation of small molecules. SMILES is a string (sequence) representation method used to describe the chemical structure of molecules. For example, "(CH3)2CHCH2OH".
[0058] Nucleotide sequence: DNA mainly consists of two complementary sequences composed of four nucleotides A / T / C / G. A pairs with T, and C pairs with G. Therefore, a nucleotide sequence composed of the four letters ATCG is commonly used to represent the DNA sequence. For example, "ATCGGTAACGTC". The composition of RNA mainly includes a sequence composed of four nucleotides A / U / C / G. Therefore, a nucleotide sequence of the four letters AUCG is commonly used to represent the RNA sequence. For example, "AUCGGUAACGUC".
[0059] Amino acid sequence: A protein is composed of one or more protein sequences, and each sequence is an amino acid sequence, which is a sequence composed of 20 to 22 amino acid letters. For example, "AEGDKLMGFA".
[0060] Protein-protein interaction: Protein interaction refers to the process in which two or more proteins bind, usually aiming to perform their biochemical functions. In cells, a large number of protein components form molecular machines, and most important molecular processes and system functions within the cell are executed through protein interactions. For example, signal molecules transmit external signals into the cell through protein interactions, and the interaction between antibodies and antigens in the immune system. Most cell functions are performed by proteins. To perform their functions, proteins need to interact with other molecules (including proteins).
[0061] Nucleic acid-protein interaction: The interaction between nucleic acids and proteins includes two types: DNA nucleic acid-protein interaction and RNA nucleic acid-protein interaction. The interaction between proteins and DNA plays a key role in biological processes. For example, transcription factors (one of the most important types of DNA-binding proteins) can specifically recognize open chromatin regions or accessible chromatin DNA sequences, thereby regulating gene transcription and expression. Precise analysis of protein-DNA interaction can reveal the mutual recognition mechanism and dynamic changes between the two, which is crucial for a deep understanding of the gene regulation mechanism under physiological and pathological conditions. RNA and proteins are biomolecules that can interact with each other. Through physical interactions, they regulate each other's life cycles and functions. For instance, the coding sequence of mRNA guides the synthesis of proteins and some regulatory sequences, while the untranslated region of mRNA affects the fate of the encoded protein by regulating protein translation, localization, and interaction with other proteins. On the other hand, during the process from RNA synthesis to degradation, proteins can in turn bind to and regulate the expression and function of RNA.
[0062] Small molecule compound-protein interaction: Generally refers to the interaction between drugs (small molecule compounds) and target proteins. Identifying drug-target interactions (DTIs) can greatly narrow the search scope of candidate drugs and thus play a key role in drug discovery. Drugs usually interact with one or more proteins to achieve their functions.
[0063] Pre-trained model: Use a large amount of molecular data (including small molecules and macromolecules) to train a pre-trained model, so as to represent small molecules (small molecule compounds) and macromolecules (proteins, DNA, RNA) as vectors of a specified dimension. Some downstream tasks are modeled and calculated based on this vector.
[0064] The prediction of molecular interactions is a very important and challenging issue in the field of life sciences, and it has important practical value in aspects such as understanding life, mechanism research, drug screening, and drug development. Each molecule in every organism does not exist in isolation, but rather constructs the entire living organism and its life activities through the interactions that occur between them. Although each organism can be considered a relatively static object when observed macroscopically, at the microscopic level, molecular interactions are occurring constantly within each living organism. These molecules can be generated within the living organism, or they can come from the external natural world or be synthetic. The study of biomolecular interactions is of great significance for clarifying the mechanisms of biological reactions and revealing the essence of life phenomena, and it is also a very important link in pharmaceuticals, which includes the interactions between proteins and proteins, proteins and nucleic acids, proteins and ligands (small molecule compounds), nucleic acids and ligands (small molecule compounds), etc. Therefore, how to predict whether molecules can interact with each other has very important research and practical significance.
[0065] Currently, it is usually possible to judge whether molecules can interact based on aspects such as molecular dynamics, molecular structure, thermodynamics, and biochemical properties. These methods generally include empirical rules, spatial configurations, methods based on molecular mechanics simulations, spectroscopic methods, thermodynamic methods, crystallographic methods, and density functional theory. Specifically, empirical rules can be understood as the empirical rules summarized through experiments and observations. For example, characteristics such as the shape, charge distribution, and polarity of molecules affect intermolecular interactions. Spatial configuration can be understood as obtaining the spatial configuration of molecules through experiments or calculations and inferring the nature of intermolecular interactions from it. Methods based on molecular mechanics simulations can be understood as using classical force fields to simulate intermolecular interactions, including molecular dynamics simulations and molecular mechanics simulations. Spectroscopic methods can be understood as inferring the nature of intermolecular interactions by analyzing the spectral data of molecules, such as infrared spectra and Raman spectra. Thermodynamic methods can be understood as understanding the strength and nature of intermolecular interactions by measuring thermodynamic properties, such as heat capacity and entropy change. Crystallographic methods can be understood as understanding the nature of intermolecular interactions by studying the arrangement and interactions of molecules in crystal structures. Density functional theory can be understood as predicting the nature of intermolecular interactions by calculating the electronic structure and energy of molecules.
[0066] However, these methods are usually based on biological, physical, and chemical principles, relying on the professional knowledge, experience of operators, and good experimental environments and equipment, and they have a long time cycle and low efficiency, unable to perform large-scale predictions. Different types of molecular interactions also require different experimental environments and equipment due to different principles, experiences, and knowledge, etc. Therefore, it is impossible to achieve generalization and large-scale application, resulting in limited applications. Therefore, there is an urgent need for an effective technical solution to solve the above problems.
[0067] In this specification, a method for object molecular sequence data is provided. One or more embodiments of this specification are also related to an object molecular sequence data device, a model training method, a model training device, a computing device, and a computer-readable storage medium, which will be described in detail one by one in the following embodiments.
[0068] See Figure 1 , Figure 1 shows a schematic diagram of an application scenario of an object molecular sequence data processing method provided according to an embodiment of this specification.
[0069] Figure 1 includes a client 102, a feature extraction model 104, and a molecular interaction prediction model 106. Among them, the feature extraction model 104 and the molecular interaction prediction model 106 are deployed on the client 102.
[0070] Specifically in implementation, in the biological field, R & D personnel can input or select two object molecular sequence data for which interaction prediction is required in the client 102. In response to the input or selection instruction of the R & D personnel, the client 102 calls the feature extraction model 104 to perform feature extraction on the two object molecular sequence data, obtains the characterization matrix of each object molecular sequence data output by the feature extraction model 104, and calls the molecular interaction prediction model 106. Input the characterization matrix of each object molecular sequence data into the molecular interaction prediction model 106, obtain the interaction result between the two object molecular sequence data, and display the interaction result to the R & D personnel through the display interface of the client 102. Thus, the prediction of molecular interaction is realized, which is convenient for subsequent drug R & D, etc. according to the interaction result.
[0071] In addition, the object molecular sequence data method provided in the embodiments of this specification can also be applied to the server. In this case, the feature extraction model and the molecular interaction prediction model can be deployed on the server, and the server makes model calls.
[0072] See Figure 2 , Figure 2 shows a flowchart of an object molecular sequence data processing method provided according to an embodiment of this specification, which specifically includes the following steps.
[0073] Step 202: Determine at least two object molecular sequence data.
[0074] Specifically, the object molecular sequence data processing method provided in the embodiments of this specification can be applied to the biological field and can specifically be used to predict whether biological molecules interact with each other.
[0075] In specific implementation, the object molecular sequence data processing method provided in the embodiments of this specification can be applied to the client or the server. When the object molecular sequence data processing method is applied to the client, at least two object molecular sequence data that need to be predicted for molecular interaction (i.e., molecular interaction) can be determined in response to the input instruction or click instruction of the researcher. When the object molecular sequence data processing method is applied to the server, at least two object molecular sequence data that need to be predicted for molecular interaction sent by the client can be received.
[0076] Among them, the object molecular sequence data can be understood as biomolecular sequences, such as small molecule compounds, DNA nucleic acid macromolecules, RNA nucleic acid macromolecules, protein macromolecules, etc. The object molecular sequence data processing method provided in the embodiments of this specification can predict the interactions between these four types of object molecular sequence data. For example, it can predict the interactions between protein macromolecules and protein macromolecules, nucleic acid macromolecules (i.e., including DNA nucleic acid macromolecules and RNA nucleic acid macromolecules) and protein macromolecules, nucleic acid macromolecules and small molecule compounds, and protein macromolecules and small molecule compounds. The basic calculation unit of a small molecule compound is an atom, so the molecular sequence of a small molecule compound is composed of atoms. The basic calculation unit of DNA nucleic acid macromolecules and RNA nucleic acid macromolecules is a nucleotide, so the molecular sequences of DNA nucleic acid macromolecules and RNA nucleic acid macromolecules can be composed of nucleotide sequences. The basic calculation unit of a protein macromolecule is an amino acid, so the molecular sequence of a protein macromolecule can be composed of an amino acid sequence. It can be understood that those skilled in the art can select any biomolecular sequence that needs to be predicted according to actual needs.
[0077] Based on this, at least two biomolecular sequences that need to be predicted for molecular interaction can be determined. For example, a small molecule compound sequence 1 and a small molecule compound sequence 2 can be determined, or a DNA nucleic acid macromolecule sequence and a protein macromolecule sequence can be determined, or a DNA nucleic acid macromolecule sequence 1, a DNA nucleic acid macromolecule sequence 2, and a protein macromolecule sequence, etc.
[0078] For example, object molecular sequence data A and object molecular sequence data B that need to be predicted for molecular interaction can be determined.
[0079] In practical applications, when predicting molecular interactions for small molecule compounds, DNA nucleic acid macromolecules, RNA nucleic acid macromolecules, and protein macromolecules, since the letters of the basic calculation units of these four types of molecules are repeated. For example, A in nucleotides and A in amino acids represent different objects. Therefore, it is necessary to process the representation methods of these four types of molecules to ensure that there are no repetitions in the sequence data of these four types of target molecule. For example, in protein macromolecules, amino acids that make up the protein macromolecule can be represented by English letters. For example, A represents alanine, R represents arginine, and O represents pyrroline, etc. In DNA nucleic acid macromolecules and RNA nucleic acid macromolecules, nucleotides can be represented by Arabic numerals. For example, 1 represents adenine, 2 represents thymine, 3 represents uracil, 4 represents cytosine, and 5 represents guanine. For small molecule compounds, atoms that make up the small molecule compound can be represented by Greek letters. For example, α represents hydrogen, β represents carbon, and γ represents oxygen, etc.
[0080] It can be understood that the above representation methods are only for examples, and other representation methods can also be selected to represent these four types of molecules to achieve the distinction of each molecule.
[0081] Step 204: Input the at least two target molecule sequence data into a feature extraction model to obtain the predicted molecular sequence features of each target molecule sequence data among the at least two target molecule sequence data, where the predicted molecular sequence features include element vectors of molecular elements, and the target molecule sequence data is composed of the molecular elements.
[0082] Specifically, after determining the at least two target molecule sequence data for which molecular interaction prediction is required, the at least two target molecule sequence data can be input into a feature extraction model to obtain the predicted molecular sequence features of each target molecule sequence data output by the feature extraction model among the at least two target molecule sequence data.
[0083] Among them, the predicted molecular sequence features can be understood as the characterization matrix of the target molecule sequence data, and this characterization matrix is composed of the element vectors of the molecular elements contained in the target molecule sequence data. In practical applications, the size of this characterization matrix is seq_len×dim, where seq_len is the length of the target molecule sequence data, and dim is the preset encoding length of each molecular element. Molecular elements can be understood as the basic building blocks that make up the target molecule sequence data. For example, for a protein macromolecule sequence, the amino acids that make up the protein macromolecule sequence are the molecular elements of the protein macromolecule sequence. For a nucleic acid macromolecule sequence, the nucleotides that make up the nucleic acid macromolecule sequence are the molecular elements of the nucleic acid macromolecule sequence.
[0084] In specific implementation, the feature extraction model can be implemented based on the encoding layer architecture of the Transformer model, or can also be implemented based on other architectures modified from the encoding layer of the Transformer model. Based on this, the feature extraction model includes an encoding layer, and at least two object molecular sequence data can be input into the feature extraction model. In the feature extraction model, the encoding layer is used to predict the predicted molecular sequence features of each object molecular sequence data.
[0085] For example, object molecular sequence data A and object molecular sequence data can be input into the feature extraction model to obtain the representation matrix A11 of object molecular sequence data A and the representation matrix B11 of object molecular sequence data B output by the feature extraction model.
[0086] In practical applications, the feature extraction model can be a pre-trained large model. Since the above four types of molecules can all be represented as molecular sequences, the feature extraction model can take the sequences of molecules as input. However, pre-training the feature extraction model requires a large amount of training data, and it is difficult to find a large amount of labeled training data of various molecules. Therefore, in order to ensure the generality of the feature extraction model, self-supervised training can be carried out using the information of the object molecular sequence data samples themselves. The specific implementation method is as follows:
[0087] The training steps of the feature extraction model include:
[0088] Determine the object molecular sequence data samples;
[0089] According to the object molecular sequence data samples, perform self-supervised training on the feature extraction model until a feature extraction model that meets the training stop condition is obtained.
[0090] Among them, the object molecular sequence data samples can include small molecule compound samples, DAN nucleic acid macromolecule samples, RNA nucleic acid macromolecule samples, protein macromolecule samples, etc. During the self-supervised training process, a large number of object molecular sequence data samples can be obtained to perform self-supervised training on the feature extraction model. The training stop condition can be, for example, that the model loss value reaches a preset loss value threshold and / or the number of training times reaches a preset number threshold. The embodiments of this specification do not limit this.
[0091] In an embodiment of this specification, for multiple types of object molecular sequence data samples, a corresponding feature extraction model can be trained for each type. For example, according to the small molecule compound samples, the feature extraction model corresponding to the small molecule compound type can be trained. When subsequent molecular interaction prediction between small molecule compounds is required, the feature extraction model corresponding to the small molecule compound type can be used for feature extraction.
[0092] In another embodiment of this specification, for multiple types of object molecule sequence data samples, a unified feature extraction model corresponding to these multiple types can be trained. That is to say, the feature extraction model can be self-supervised trained based on small molecule compound samples, DAN nucleic acid macromolecule samples, RNA nucleic acid macromolecule samples, and protein macromolecule samples.
[0093] In another embodiment of this specification, multiple feature extraction models can be trained according to the correlation between multiple types of object molecule sequence data samples. For example, corresponding feature extraction models can be trained based on DAN nucleic acid macromolecule samples and RNA nucleic acid macromolecule samples. When subsequent molecular interaction predictions between DAN nucleic acid macromolecules, between RNA nucleic acid macromolecules, or between DNA nucleic acid macromolecules and RNA nucleic acid macromolecules are required, these feature extraction models can be used for feature extraction. Or, due to the correlation between DNA, RNA, and proteins, corresponding feature extraction models can also be trained based on DAN nucleic acid macromolecule samples, RNA nucleic acid macromolecule samples, and protein macromolecule samples.
[0094] It can be understood that the number of feature extraction models trained in the embodiments of this specification is not limited. The number of feature extraction models to be trained can be determined according to the types and characteristics of the object molecule sequence data samples for which molecular predictions are needed. Based on this, the interaction prediction tasks between different types of molecules can be unified to ensure generality.
[0095] In practical applications, when performing self-supervised training on the feature extraction model, the object molecule sequence data samples for the molecular interaction prediction task can be determined as the training data. Or, a large number of object molecule sequence data samples irrelevant to the specific molecular interaction prediction task can also be used for training, so that the feature extraction model can self-learn from a large number of object molecule sequence data samples, enabling the generalization and universalization of the feature extraction model, and enabling the feature extraction model to output a general and universal sequence characterization region, so that it can be applicable to multiple downstream tasks.
[0096] In summary, through self-supervised training of the feature extraction model, a large amount of labeled training data is not required, reducing the labeling cost, and enabling the model to provide a good starting point for the next step of molecular interaction prediction. Moreover, by training with object molecule sequence data samples, it only depends on the sequence of the molecule and does not require the input of the molecular structure, thus avoiding the problems of less accurate structure data obtained by experimental methods for existing macromolecules and the difficulty of structure acquisition.
[0097] Specifically, when implementing, the pre-training task for self-supervised training of the feature extraction model can be a random masking task. The specific implementation method is as follows:
[0098] Performing self-supervised training on the feature extraction model according to the object molecular sequence data sample until a feature extraction model that meets the training stop condition is obtained, including:
[0099] Inputting the object molecular sequence data sample into the feature extraction model, and in the feature extraction model, performing masking processing on the object molecular sequence data sample to obtain a masked object molecular sequence data;
[0100] Performing prediction on the masked object molecular sequence data to obtain predicted object molecular sequence data;
[0101] Performing self-supervised training on the feature extraction model according to the object molecular sequence data sample and the predicted object molecular sequence data until a feature extraction model that meets the training stop condition is obtained.
[0102] Among them, performing masking processing on the object molecular sequence data sample can be understood as masking the characters of the molecular elements in the object molecular sequence data sample. For example, the molecular element can be replaced with a masking character, so that the feature extraction model predicts the character of the masked molecular element according to the masked object molecular sequence data.
[0103] In an embodiment of this specification, when performing masking processing on the object molecular sequence data sample, masking processing can be performed on some molecular elements in the object molecular sequence data sample, and these partial molecular elements can be randomly determined from the object molecular sequence data sample.
[0104] Based on this, the object molecular sequence data sample can be input into the feature extraction model. In the feature extraction model, use masking characters to replace the characters of some molecular elements in the object molecular sequence data sample, so as to obtain masked object molecular sequence data, and use the encoding layer to predict the masked object molecular sequence data to obtain predicted object molecular sequence data. Calculate the model loss value according to the object molecular sequence data sample and the predicted object molecular sequence data, and train the feature extraction model according to the model loss value until a feature extraction model that meets the training stop condition is obtained.
[0105] In specific implementation, when using masking characters to replace the characters of some molecular elements in the object molecular sequence data sample, the characters of a specified proportion of molecular elements in the object molecular sequence data sample can be replaced.
[0106] For example, the object molecular sequence data sample can be, for instance, "ARNDCEQGHILKMFPSTWYVUO". The characters "N", "H", "F", and "W" of the molecular elements in the object molecular sequence data sample can be replaced with the mask character [MASK] to obtain the masked object molecular sequence data "AR[MASK]DCEQG[MASK]ILKM[MASK]PST[MASK]YVUO", and the characters at these four [MASK] positions can be predicted based on this masked object molecular sequence data. The predicted object molecular sequence data can be, for example, "ARCDCEQGHILKMFPSTWYVUO".
[0107] In summary, by randomly masking the object molecular sequence data sample to train the model's ability to predict the characters at the masked positions, it is convenient to subsequently improve the prediction accuracy of the feature extraction model.
[0108] In specific implementation, the pre-training task for self-supervised training of the feature extraction model can also be other pre-training tasks. The specific implementation method is as follows:
[0109] The self-supervised training of the feature extraction model based on the object molecular sequence data sample until a feature extraction model that meets the training stop condition is obtained includes:
[0110] Input the object molecular sequence data sample into the feature extraction model, and in the feature extraction model, process the object molecular sequence data sample to obtain a first object molecular sequence data sample and a second object molecular sequence data sample;
[0111] Create a target object molecular sequence data sample based on the first object molecular sequence data sample and the second object molecular sequence data sample;
[0112] Predict the target object molecular sequence data sample to obtain a prediction result, where the prediction result is used to predict whether the first object molecular sequence data sample and the second object molecular sequence data sample belong to the same object molecular sequence data sample;
[0113] Based on the target object molecular sequence data sample and the prediction result, perform self-supervised training on the feature extraction model until a feature extraction model that meets the training stop condition is obtained.
[0114] Among them, processing the object molecular sequence data sample may include splitting the object molecular sequence data sample. At this time, the first object molecular sequence data sample and the second object molecular sequence data sample obtained are subsequences of the object molecular sequence data sample. Then, the target object molecular sequence data sample created based on the first object molecular sequence data sample and the second object molecular sequence data sample is a positive sample, and the target object molecular sequence data sample can be obtained by inserting characters at the front end, the rear end, and the intermediate insertion positions of the first object molecular sequence data sample and the second object molecular sequence data sample.
[0115] Correspondingly, processing the object molecular sequence data sample further includes splitting and replacing the object molecular sequence data sample. Specifically, splitting the object molecular sequence data sample to obtain two subsequences of the object molecular sequence data sample, and randomly replacing these two subsequences. For example, they can be randomly replaced with subsequences of other object molecular sequence data, so as to obtain the first object molecular sequence data sample and the second object molecular sequence data sample. Then, at this time, the target object molecular sequence data sample created based on the first object molecular sequence data sample and the second object molecular sequence data sample is a negative sample, and the target object molecular sequence data sample can be obtained by inserting characters at the front end of the first object molecular sequence data sample, the rear end of the second object molecular sequence data sample, and the intermediate insertion positions of the first object molecular sequence data sample and the second object molecular sequence data sample.
[0116] Based on this, after creating the target object molecular sequence data sample, the target object molecular sequence data sample can be predicted to obtain a prediction result. The prediction result can be used to predict whether the first object molecular sequence data sample and the second object molecular sequence data sample belong to the same object molecular sequence data sample, and calculate the model loss value according to whether the target object molecular sequence data sample is a positive sample or a negative sample and the prediction result. The feature extraction model is trained according to the model loss value until a feature extraction model that meets the training stop condition is obtained.
[0117] In practical applications, [CLS] can be used to represent the front end (i.e., the start), [EOS] can be used to represent the rear end (i.e., the end), and [SEP] can be used to represent the middle.
[0118] For example, split the object molecular sequence data sample 1 "ARNDCEQGHILKMFPSTWYVUO" to obtain subsequence 1 "ARNDCEQGH" and subsequence 2 "ILKMFPSTWYVUO". Split the object molecular sequence data sample 2 (as another object molecular sequence data sample) "HAGEYGAEALERMFLSFPTTKTYFPHFDLSHGSAQV" to obtain subsequence 1 "HAGEYGAEALERMFLSFPTTK" and subsequence 2 "TYFPHFDLSHGSAQV". To create a positive sample, the two subsequences of the object molecular sequence data sample 1 can be determined as the first object molecular sequence data sample and the second object molecular sequence data sample. The target object molecular sequence data sample created based on the first object molecular sequence data sample and the second object molecular sequence data sample is the positive sample "[CLS]ARNDCEQGH[SEP]ILKMFPSTWYVUO[EOS], True". To create a negative sample, the subsequence 1 in the object molecular sequence data sample 1 can be replaced with the subsequence 1 in the object molecular sequence data sample 2. That is to say, the subsequence 1 of the object molecular sequence data sample 1 and the subsequence 1 of the object molecular sequence data sample 2 are determined as the first object molecular sequence data sample and the second object molecular sequence data sample. The target object molecular sequence data sample created based on the first object molecular sequence data sample and the second object molecular sequence data sample is the negative sample "[CLS]ARNDCEQGH[SEP]HAGEYGAEALERMFLSFPTTK[EOS], False".
[0119] In summary, by splitting and replacing the object molecular sequence data samples, the ability of the training model to predict whether the subsequences belong to the same object molecular sequence data sample is trained, which is convenient for improving the prediction accuracy of the feature extraction model subsequently.
[0120] In one embodiment of this specification, the above two pre-training tasks can be combined. Specifically, training object molecular sequence data samples can be created according to the target object molecular sequence data samples and the masked object molecular sequence data. Based on the training object molecular sequence data samples, self-supervised training of the feature extraction model is performed to achieve simultaneous training of the prediction ability of the model.
[0121] Specifically, refer to Figure 3 , Figure 3 shows a schematic diagram of the feature extraction model in the object molecular sequence data processing method provided according to an embodiment of this specification. As Figure 3As shown, in the supervised training stage of the feature extraction model, the object molecular sequence data sample is input into the feature extraction model. In the feature extraction model, the object molecular sequence data sample is processed to obtain a training object molecular sequence data sample. According to the encoding layer, the training object molecular sequence data sample is predicted to determine whether the first object molecular sequence data sample and the second object molecular sequence data sample contained in the training object molecular sequence data sample belong to the same object molecular sequence data sample, obtaining a prediction result. Also, the character of the molecular element at the predicted mask position is predicted to obtain predicted object molecular sequence data. Based on the predicted object molecular sequence data and the object molecular sequence data sample, the feature extraction model is trained. In the inference stage of the feature extraction model, after the object molecular sequence data is input into the feature extraction model, the encoding layer can be used for prediction to obtain the representation matrix of the output object molecular sequence data.
[0122] For example, the training object molecular sequence data sample can be, for instance, "[CLS]AR[MASK]DC[MASK]QGH[SEP]ILK[MASK]FP[MASK]TWYVUO[EOS]True", "[CLS]H[MASK]GEYG[MASK]EALERM[MASK]LSFPTTK[SEP]ILKM[MASK]PSTW[MASK]VU[MASK][EOS]False", "[CLS]H[MASK]GEYG[MASK]EALERM[MASK]LSFPTTK[SEP]ILKM[MASK]PSTW[MASK]VU[MASK][EOS]False", "[CLS]HAGEYGAEAL[MASK]RMFLSFPTTK[SEP]ILKMFPS[MASK]WYVUO[EOS]False", etc.
[0123] Step 206: Input the predicted molecular sequence features of each object molecular sequence data into the molecular interaction prediction model to obtain the interaction result between the at least two object molecular sequence data, where the molecular interaction prediction model outputs the interaction result between the at least two object molecular sequence data according to the correlation between the element vectors included in the predicted molecular sequence features of each object molecular sequence data.
[0124] Specifically, after obtaining the predicted molecular sequence features of each object molecular sequence data output by the feature extraction model, the output of the feature extraction model can be used as the input of the molecular interaction prediction model. The molecular interaction prediction model can output the interaction result between the at least two object molecular sequence data according to the correlation of the element vectors included in the predicted molecular sequence features.
[0125] In specific implementation, inputting the at least two object molecular sequence data into the feature extraction model to obtain the predicted molecular sequence features of each object molecular sequence data among the at least two object molecular sequence data includes:
[0126] Selecting any two object molecular sequence data from the at least two object molecular sequence data as the first object molecular sequence data and the second object molecular sequence data;
[0127] Inputting the first object molecular sequence data and the second object molecular sequence data into the feature extraction model to obtain the first predicted molecular sequence feature of the first object molecular sequence data and the second predicted molecular sequence feature of the second object molecular sequence data;
[0128] Correspondingly, inputting the predicted molecular sequence features of each object molecular sequence data into the molecular interaction prediction model to obtain the interaction result between the at least two object molecular sequence data includes:
[0129] Inputting the first predicted molecular sequence feature and the second predicted molecular sequence feature into the molecular interaction prediction model to obtain the interaction result between the first object molecular sequence data and the second object molecular sequence data.
[0130] Specifically, when performing molecular interaction prediction on at least two object molecular sequence data, any two object molecular sequence data can be selected from the at least two object molecular sequence data as the first object molecular sequence data and the second object molecular sequence data, and the first object molecular sequence data and the second object molecular sequence data are input into the feature extraction model to obtain the first predicted molecular sequence feature of the first object molecular sequence data and the second predicted molecular sequence feature of the second object molecular sequence data, and the first predicted molecular sequence feature and the second predicted molecular sequence feature are input into the molecular interaction prediction model, so as to obtain the interaction result between the first object molecular sequence data and the second object molecular sequence data output by the molecular interaction prediction model. Repeat the above steps until the interaction results of any two object molecular sequence data among the at least two object molecular sequence data are obtained.
[0131] For example, when performing molecular interaction prediction on three object molecular sequence data 1, 2, and 3, the interaction result between object molecular sequence data 1 and 2, the interaction result between object molecular sequence data 1 and 3, and the interaction result between object molecular sequence data 2 and 3 can be determined in sequence according to the above steps, so as to obtain the interaction result between the three object molecular sequence data.
[0132] In another embodiment of this specification, when predicting molecular interactions for three object molecular sequence data 1, 2, and 3, the representation matrices of the object molecular sequence data 1, the object molecular sequence data 2, and the object molecular sequence data 3 can be directly input into the molecular interaction prediction model to obtain the interaction results between the three object molecular sequence data.
[0133] In practical applications, the molecular interaction prediction model may include a correlation processing layer, an element position encoding layer, and an output layer. Due to the interactions between molecules, macroscopically, it is whether there is an interaction between molecule levels. Microscopically, it may be an interaction at the element level of two molecules. For example, the interaction between protein A and protein B may only be the interaction between several amino acids in a certain fragment of protein A and several amino acids in a certain fragment of protein B, and even the interacting amino acids are several discontinuous ones. That is to say, when molecules interact, not only do the molecular elements within the same molecule interact, but also the elements between different molecules interact. Therefore, the correlation processing layer can achieve fine-grained (element-level) feature extraction.
[0134] Correspondingly, the step of inputting the first predicted molecular sequence feature and the second predicted molecular sequence feature into the molecular interaction prediction model to obtain the interaction result between the first object molecular sequence data and the second object molecular sequence data includes:
[0135] Inputting the first predicted molecular sequence feature and the second predicted molecular sequence feature into the correlation processing layer in the molecular interaction prediction model to obtain the first correlation feature between the first molecular elements included in the first object molecular sequence data, the second correlation feature between the second molecular elements included in the second object molecular sequence data, and the third correlation feature between the first molecular elements and the second molecular elements;
[0136] Inputting the first correlation feature, the second correlation feature, and the third correlation feature into the element position encoding layer in the molecular interaction prediction model to obtain the encoded first correlation feature, second correlation feature, and third correlation feature;
[0137] Inputting the encoded first correlation feature, second correlation feature, and third correlation feature into the output layer in the molecular interaction prediction model to obtain the interaction result between the first object molecular sequence data and the second object molecular sequence data.
[0138] Among them, the correlation processing layer may include a first Intra-Attention module and a second Inter-Attention module. The first Intra-Attention module can be used to capture the correlation matrix of molecular elements within the object molecular sequence data, and the second Inter-Attention module can be used to capture the correlation matrix between molecular elements across object molecular sequence data. The first correlation feature can be understood as the correlation matrix between the first molecular elements included in the first object molecular sequence data, the second correlation feature can be understood as the correlation matrix between the second molecular elements included in the second object molecular sequence data, and the third correlation feature may include the correlation matrix of the first molecular elements relative to the second molecular elements and the correlation matrix of the second molecular elements relative to the first molecular elements.
[0139] Based on this, the first predicted molecular sequence feature and the second predicted molecular sequence feature can be input into the correlation processing layer. The first Intra-Attention module is used to extract the correlation between the element vectors of each first molecular element included in the first object molecular sequence data according to the first predicted molecular sequence feature, and obtain the first correlation matrix between the molecular elements within the first object molecular sequence data. Similarly, the second correlation matrix between the molecular elements within the second object molecular sequence data is obtained using the first Intra-Attention module. And the second Inter-Attention module is used to extract the correlation of the first molecular elements relative to the second molecular elements and the correlation of the second molecular elements relative to the first molecular elements, so as to obtain the third correlation matrix of the first molecular elements relative to the second molecular elements and the third correlation matrix of the second molecular elements relative to the first molecular elements. The first correlation matrix, the second correlation matrix, and the third correlation matrix are input into the element position encoding layer. The element position encoding layer adds a position encoding matrix to each correlation matrix to obtain the encoded first correlation matrix, second correlation matrix, and third correlation matrix, and the encoded first correlation matrix, second correlation matrix, and third correlation matrix are input into the output layer, so as to obtain the interaction result between the first object molecular sequence data and the second object molecular sequence data.
[0140] In practical applications, since each molecular element in the object molecular sequence data has a position attribute, a position encoding matrix can be added to each correlation matrix. Among them, the position encoding matrix added to the correlation matrix can be a relative position encoding matrix, such as a RoPE rotation encoding matrix.
[0141] Continuing with the above example, after obtaining the representation matrix A11 of the target molecular sequence data A and the representation matrix B11 of the target molecular sequence data B, the representation matrix A11 and the representation matrix B11 can be input into the correlation processing layer. According to the representation matrix A11, the correlation matrix AA between the first molecular elements contained in the target molecular sequence data A is determined. According to the representation matrix B11, the correlation matrix BB between the second molecular elements contained in the target molecular sequence data B is determined, and the correlation matrix AB of the first molecular elements relative to the second molecular elements, as well as the correlation matrix BA of the second molecular elements relative to the first molecular elements, are determined. The correlation matrices AA, BB, AB, and BA are input into the position encoding layer, a position encoding matrix is added to each correlation matrix to obtain the encoded correlation matrices AA, BB, AB, and BA, and the encoded correlation matrices AA, BB, AB, and BA are input into the output layer to obtain the interaction result between the target molecular sequence data A and the target molecular sequence data B.
[0142] In summary, by capturing the correlations of the molecular elements within the target molecular sequence data and the correlations between the molecular elements across the target molecular sequence data, the molecular interaction prediction model can capture the characteristics of the molecular interaction itself, that is, a fine-grained learning network within the same molecule and between different molecules, enabling the molecular interaction prediction model to have the prediction ability at the fine-grained molecular element level. Adding relative position encoding to the correlation matrix can not only bring the position information of each molecular element in the target molecular sequence data but also enable the molecular interaction prediction model to process longer target molecular sequence data at the prediction interface, further enhancing the prediction ability of the molecular interaction prediction model.
[0143] In specific implementation, the molecular interaction prediction model further includes a fully connected layer and a normalization layer. Correspondingly, before inputting the first predicted molecular sequence feature and the second predicted molecular sequence feature into the correlation processing layer in the molecular interaction prediction model, it further includes:
[0144] Inputting the first predicted molecular sequence feature and the second predicted molecular sequence feature into the fully connected layer in the molecular interaction prediction model to obtain the first predicted molecular sequence feature and the second predicted molecular sequence feature after spatial transformation processing;
[0145] Inputting the first predicted molecular sequence feature and the second predicted molecular sequence feature after the spatial transformation processing into the normalization layer in the molecular interaction prediction model to obtain the first predicted molecular sequence feature and the second predicted molecular sequence feature after normalization processing;
[0146] Correspondingly, inputting the first predicted molecular sequence feature and the second predicted molecular sequence feature into the correlation processing layer in the molecular interaction prediction model includes:
[0147] Input the first predicted molecular sequence feature and the second predicted molecular sequence feature after the normalization process into the correlation processing layer in the molecular interaction prediction model.
[0148] Specifically, the fully connected layer can be used to perform a spatial transformation on the predicted molecular sequence feature. Specifically, it can perform a non-linear spatial transformation on the predicted molecular sequence feature, thereby improving the model representation ability to enhance the model processing ability. The normalization layer can be used to ensure the stability of the molecular interaction prediction model.
[0149] In one embodiment of this specification, the way the fully connected layer performs a spatial transformation on the predicted molecular sequence feature can be, for example, a non-linear spatial transformation.
[0150] In practical applications, the fully connected layer can be an FFN, and the normalization layer can be a LayerNorm.
[0151] Continuing with the above example, before inputting the representation matrices A11 and B11 into the correlation processing layer, the representation matrices A11 and B11 can be first input into the fully connected layer, and the fully connected layer is used to perform a spatial transformation process on the representation matrices A11 and B11 to obtain the representation matrices A11 and B11 after the spatial transformation process. Then, the representation matrices A11 and B11 after the spatial transformation process are input into the normalization layer, and the normalization layer is used to perform a normalization process on the representation matrices A11 and B11 after the spatial transformation process to obtain the representation matrices A11 and B11 after the normalization process. At this time, the representation matrices A11 and B11 after the normalization process are input into the correlation processing layer.
[0152] In summary, by setting a fully connected layer and a normalization layer before the correlation processing layer, the training stability and inference stability of the molecular interaction prediction model can be ensured, and the model converges better.
[0153] In practical applications, before inputting the encoded first correlation feature, second correlation feature, and third correlation feature into the output layer in the molecular interaction prediction model, it further includes:
[0154] Input the first predicted molecular sequence feature and the second predicted molecular sequence feature after the spatial transformation process, the encoded first correlation feature, second correlation feature, and third correlation feature into the normalization layer in the molecular interaction prediction model to obtain the first correlation feature, second correlation feature, and third correlation feature after the normalization process;
[0155] Input the first correlation feature, second correlation feature, and third correlation feature after the normalization process into the fully connected layer in the molecular interaction prediction model to obtain the first correlation feature, second correlation feature, and third correlation feature after the spatial transformation process;
[0156] Accordingly, inputting the encoded first correlation feature, second correlation feature, and third correlation feature into the output layer of the molecular interaction prediction model to obtain the interaction result between the first object molecular sequence data and the second object molecular sequence data includes:
[0157] Inputting the encoded first correlation feature, second correlation feature, and third correlation feature, and the first correlation feature, second correlation feature, and third correlation feature after spatial transformation processing into the output layer of the molecular interaction prediction model to obtain the interaction result between the first object molecular sequence data and the second object molecular sequence data.
[0158] Specifically, continuing with the above example, before inputting the encoded correlation matrices AA, BB, AB, and BA into the output layer, the encoded correlation matrices AA, BB, AB, and BA, and the representation matrices A11 and B11 after spatial transformation processing can also be input into the normalization layer to obtain the normalized correlation matrices AA, BB, AB, and BA, and input the normalized correlation matrices AA, BB, AB, and BA into the fully connected layer to obtain the correlation matrices AA, BB, AB, and BA after spatial transformation processing. At this time, input the encoded correlation matrices AA, BB, AB, and BA, and the correlation matrices AA, BB, AB, and BA after spatial transformation processing into the output layer to obtain the interaction result.
[0159] In summary, by setting a fully connected layer and a normalization layer before the output layer, the fully connected layer can perform non-linear transformation to improve the representation ability of the molecular interaction prediction model, and the normalization layer can make the training of the deep network of the molecular interaction prediction model more stable.
[0160] In specific implementation, inputting the encoded first correlation feature, second correlation feature, and third correlation feature, and the first correlation feature, second correlation feature, and third correlation feature after spatial transformation processing into the output layer of the molecular interaction prediction model to obtain the interaction result between the first object molecular sequence data and the second object molecular sequence data includes:
[0161] Performing fusion processing on the encoded first correlation feature, second correlation feature, and third correlation feature, and the first correlation feature, second correlation feature, and third correlation feature after spatial transformation processing to obtain a first fused correlation feature, a second fused correlation feature, and a third fused correlation feature;
[0162] Input the first fusion correlation feature, the second fusion correlation feature, and the third fusion correlation feature into the output layer of the molecular interaction prediction model to obtain the interaction result between the first object molecular sequence data and the second object molecular sequence data.
[0163] Specifically, performing a fusion process on the encoded first correlation feature, second correlation feature, and third correlation feature, and the first correlation feature, second correlation feature, and third correlation feature after spatial transformation processing can be understood as adding the encoded first correlation feature and the first correlation feature after spatial transformation processing to obtain the first fusion correlation feature, adding the encoded second correlation feature and the second correlation feature after spatial transformation processing to obtain the second fusion correlation feature, and adding the encoded third correlation feature and the third correlation feature after spatial transformation processing to obtain the third fusion correlation feature.
[0164] Continuing with the above example, when inputting the encoded correlation matrices AA, BB, AB, and BA, and the correlation matrices AA, BB, AB, and BA after spatial transformation processing into the output layer, the encoded correlation matrix AA and the correlation matrix AA after spatial transformation processing can be added to obtain the fused correlation matrix AA2, the encoded correlation matrix BB and the correlation matrix BB after spatial transformation processing can be added to obtain the fused correlation matrix BB2, the encoded correlation matrix AB and the correlation matrix AB after spatial transformation processing can be added to obtain the fused correlation matrix AB2, the encoded correlation matrix BA and the correlation matrix BA after spatial transformation processing can be added to obtain the fused correlation matrix BA2, and the fused correlation matrices AA2, BB2, AB2, and BA2 are input into the output layer.
[0165] In summary, through the fusion of the encoded correlation features and the correlation features after spatial transformation processing, the molecular interaction prediction model (i.e., the deep learning model) converges better, ensuring the existence of the derivative gradient during the backpropagation of the molecular interaction prediction model.
[0166] During specific implementation, the output layer includes a dimensionality reduction processing unit;
[0167] Correspondingly, the step of inputting the first fusion correlation feature, the second fusion correlation feature, and the third fusion correlation feature into the output layer of the molecular interaction prediction model to obtain the interaction result between the first object molecular sequence data and the second object molecular sequence data includes:
[0168] Input the first fusion correlation feature, the second fusion correlation feature, and the third fusion correlation feature into the output layer of the molecular interaction prediction model;
[0169] In the output layer, the dimensionality reduction processing unit performs dimensionality reduction processing on the first fusion correlation feature, the second fusion correlation feature, and the third fusion correlation feature to obtain a first prediction vector, a second prediction vector, and a third prediction vector;
[0170] According to the first prediction vector, the second prediction vector, and the third prediction vector, an interaction result between the first object molecular sequence data and the second object molecular sequence data is output.
[0171] Among them, performing dimensionality reduction processing on the fusion correlation feature is to reduce the representation matrix to a vector. The way of dimensionality reduction processing can be, for example, a pooling method, or the maximum value, average value, weighted average value, etc. corresponding to the positions in all rows can be selected for each column in the representation matrix.
[0172] In practical applications, the dimensionality reduction processing unit can be a pooling layer, and the pooling method is used to implement dimensionality reduction processing. The specific pooling methods can include: using the element vector corresponding to the position identifier [CLS] indicating the start added to the front end of the object molecular sequence data (i.e., the first row vector of the representation matrix), or the maximum value or average value of each column in all rows can be selected for each column in the representation matrix, or a weighted method can be used to set weights for each row of the representation matrix, and the values of each column are weighted according to the weights of each row to obtain a prediction vector with a length of dim.
[0173] Moreover, when outputting the interaction result according to the first prediction vector, the second prediction vector, and the third prediction vector, the first prediction vector, the second prediction vector, and the third prediction vector can be concatenated to obtain a target prediction vector. The target prediction vector is input into a linear layer and a multi-classification layer (softmax), and the obtained interaction result can be the degree level of the interaction between the first object molecular sequence data and the second object molecular sequence data, so as to achieve multi-classification of multiple levels; or the target prediction vector is input into a linear layer and a binary classification layer (sigmoid), and the obtained interaction result can be whether the first object molecular sequence data and the second object molecular sequence data interact; or the target prediction vector is input into a linear layer (regression), and the obtained interaction result can be the intensity value of the interaction between the first object molecular sequence data and the second object molecular sequence data.
[0174] In summary, by setting a dimensionality reduction processing unit in the output layer, the subsequent prediction of the interaction result can be further realized.
[0175] In addition, the output layer further includes a normalization processing unit;
[0176] Accordingly, before the dimensionality reduction processing unit performs dimensionality reduction processing on the first fused correlation feature, the second fused correlation feature, and the third fused correlation feature to obtain a first prediction vector, a second prediction vector, and a third prediction vector, it further includes:
[0177] According to the normalization processing unit, perform normalization processing on the first fused correlation feature, the second fused correlation feature, and the third fused correlation feature to obtain the first fused correlation feature, the second fused correlation feature, and the third fused correlation feature after normalization processing;
[0178] Accordingly, the dimensionality reduction processing unit performing dimensionality reduction processing on the first fused correlation feature, the second fused correlation feature, and the third fused correlation feature to obtain a first prediction vector, a second prediction vector, and a third prediction vector includes:
[0179] According to the dimensionality reduction processing unit, perform dimensionality reduction processing on the first fused correlation feature, the second fused correlation feature, and the third fused correlation feature after normalization processing to obtain a first prediction vector, a second prediction vector, and a third prediction vector.
[0180] In practical applications, according to the above content, the molecular interaction prediction model sequentially includes a fully connected layer, a normalization layer, a correlation processing layer, a position encoding layer, a normalization layer, and a fully connected layer. The sub-module composed of these layers can be stacked N times, that is, the molecular interaction prediction model can include multiple sub-modules, and each sub-module includes a fully connected layer, a normalization layer, a correlation processing layer, a position encoding layer, a normalization layer, and a fully connected layer, belonging to a deep network. Therefore, setting a normalization processing unit in the output layer can make the training of this molecular interaction prediction model more stable. When inputting the representation matrices of the characterization of 2 object molecular sequence data, 4 correlation characterization matrices are obtained through the above sub-module. When inputting the representation matrices of k (k>2) object molecular sequence data, there are k intra-attention characterization matrices + A(k, 2) = k(k - 1) inter-attention characterization matrices.
[0181] Specifically, when implementing, inputting the first predicted molecular sequence feature and the second predicted molecular sequence feature after the spatial transformation processing, and the first correlation feature, the second correlation feature, and the third correlation feature after encoding into the normalization layer in the molecular interaction prediction model includes:
[0182] According to the preset fusion rule, perform a fusion process on the first predicted molecular sequence feature and the second predicted molecular sequence feature after the spatial transformation process, the encoded first correlation feature, second correlation feature, and third correlation feature, to obtain the fused first correlation feature, second correlation feature, and third correlation feature;
[0183] Input the fused first correlation feature, second correlation feature, and third correlation feature into the normalization layer in the molecular interaction prediction model.
[0184] Among them, the preset fusion rule can be understood as a rule for fusion according to molecular elements.
[0185] Specifically, when performing a fusion process on the first predicted molecular sequence feature and the second predicted molecular sequence feature after the spatial transformation process, the encoded first correlation feature, second correlation feature, and third correlation feature, according to the first molecular element and the second molecular element, fuse the first predicted molecular sequence feature and the first correlation feature corresponding to the first molecular element to obtain the fused first correlation feature; fuse the second predicted molecular sequence feature and the second correlation feature corresponding to the second molecular element to obtain the fused third correlation feature, fuse the first predicted molecular sequence feature and the third correlation feature of the second molecular element relative to the first molecular element, and fuse the second predicted molecular sequence feature and the third correlation feature of the first molecular element relative to the second molecular element, so as to obtain the fused third correlation feature.
[0186] In summary, by performing fusion according to molecular elements, the convergence effect of the molecular interaction prediction model is ensured.
[0187] In practical applications, the training steps of the molecular interaction prediction model include:
[0188] Determine the molecular sequence feature samples of at least two object molecular sequence data samples, and the interaction result labels between the at least two object molecular sequence data samples;
[0189] Input the molecular sequence feature samples of the at least two object molecular sequence data samples and the interaction result labels into the molecular interaction prediction model to obtain the predicted interaction results;
[0190] According to the predicted interaction results and the interaction result labels, train the molecular interaction prediction model until a molecular interaction prediction model that meets the training stop condition is obtained.
[0191] Among them, the molecular sequence feature samples of at least two object molecular sequence data samples can be understood as the molecular sequence feature samples of each object molecular sequence data sample among the at least two object molecular sequence data samples.
[0192] Specifically, when determining the molecular sequence feature samples of each object molecular sequence data sample, the molecular sequence feature samples of each object molecular sequence data sample can be obtained according to the feature extraction model. In one embodiment of this specification, for the molecular sequence sample features of each object molecular sequence data sample output by the feature extraction model, they remain unchanged during the training process of the molecular interaction prediction model, that is, the molecular sequence sample features are not used as parameters of the molecular interaction prediction model. In another embodiment of this specification, during the training process of the molecular interaction prediction model, the molecular interaction prediction model can be initialized according to the molecular sequence sample features, and the molecular sequence sample features can be adjusted, that is, the molecular sequence sample features are used as parameters of the molecular interaction prediction model. That is to say, the molecular sequence feature samples in the first training method described above do not change with the training of the molecular interaction prediction model, while the molecular sequence feature samples in the second training method described above are used as parameters of the molecular interaction prediction model and will be learned and updated with the training of the molecular interaction prediction model. This second training method can incorporate relevant information on molecular interactions according to the training and update of the training data set of molecular interactions, making the molecular interaction prediction model more robust.
[0193] Based on this, the molecular interaction prediction model can be supervised trained according to the molecular sequence feature samples of at least two object molecular sequence data samples and the interaction result labels between the at least two object molecular sequence data samples.
[0194] It can be understood that the training process of the molecular interaction prediction model is similar to the foregoing application process, and specific details can be found in the following Figure 4 .
[0195] In summary, based on a large model, training a model for a specific downstream task (i.e., the molecular interaction prediction model) based on a small amount of labeled data to predict unknown data is an end-to-end method that can quickly predict in large quantities and has high-throughput computing characteristics.
[0196] In summary, the above method extracts the predicted molecular sequence features of each object molecular sequence data by using the feature extraction model. And since the predicted molecular sequence features contain the element vectors of molecular elements, element-level feature prediction is achieved. Then, the molecular interaction prediction model predicts the interaction results between at least two object molecular sequence data according to the correlation between the element vectors, realizing molecular interaction prediction based on artificial intelligence. And based on the element-level predicted molecular sequence features, the predicted interaction results are further made more accurate. Molecular interaction prediction is achieved through the combination of the upstream pre-trained large model (i.e., the feature extraction model) and the downstream specific task network (i.e., the molecular interaction prediction model), while ensuring the efficiency and accuracy of the prediction results.
[0197] The following, in combination with the attached Figure 4 , taking the training process of the molecular interaction prediction model in the object molecular sequence data processing method provided in this specification as an example, the object molecular sequence data processing method will be further described. Among them, Figure 4 FIG. shows a flowchart of the processing procedure for training a molecular interaction prediction model in an object molecular sequence data processing method provided by an embodiment of this specification, which specifically includes the following steps.
[0198] Step 402: Input the first object molecular sequence data sample and the second object molecular sequence data sample into the feature extraction model to obtain the first molecular sequence feature sample and the second molecular sequence feature sample.
[0199] Specifically, the first object molecular sequence data sample X and the second object molecular sequence data sample Y can be input into the feature extraction model to obtain the first molecular sequence feature sample X1 and the second molecular sequence feature sample Y1.
[0200] Among them, the first molecular sequence feature sample can be understood as the representation matrix of the first object molecular sequence data sample, and the second molecular sequence feature sample can be understood as the representation matrix of the second object molecular sequence data sample.
[0201] Step 404: Input the first molecular sequence feature sample, the first molecular sequence feature sample, and the interaction result label into the fully connected layer in the molecular interaction prediction model, perform spatial transformation processing on the first molecular sequence feature sample and the first molecular sequence feature sample, and obtain the first molecular sequence feature sample and the first molecular sequence feature sample after spatial transformation processing.
[0202] Specifically, input the first molecular sequence feature sample X1 and the second molecular sequence feature sample Y1 into the fully connected layer to obtain the first molecular sequence feature sample X2 and the first molecular sequence feature sample Y2 after spatial transformation processing.
[0203] Step 406: Input the first molecular sequence feature sample and the first molecular sequence feature sample after spatial transformation processing into the normalization layer to obtain the first molecular sequence feature sample and the first molecular sequence feature sample after normalization processing.
[0204] Specifically, input the first molecular sequence feature sample X2 and the first molecular sequence feature sample Y2 after spatial transformation processing into the normalization layer to obtain the first molecular sequence feature sample X3 and the first molecular sequence feature sample Y3 after normalization processing.
[0205] Step 408: Input the normalized first molecular sequence feature sample and the first molecular sequence feature sample into the correlation processing layer to obtain the first correlation feature sample, the second correlation feature sample, and the third correlation feature sample.
[0206] Specifically, input the normalized first molecular sequence feature sample X3 and the first molecular sequence feature sample Y3 into the correlation processing layer to obtain the first correlation feature sample XX1, the second correlation feature sample YY1, and the third correlation feature samples XY1 and YX1.
[0207] Step 410: Input the first correlation feature sample, the second correlation feature sample, and the third correlation feature sample into the element position encoding layer to obtain the encoded first correlation feature sample, the second correlation feature sample, and the third correlation feature sample.
[0208] Specifically, input the first correlation feature sample XX1, the second correlation feature sample YY1, and the third correlation feature samples XY1 and YX1 into the element position encoding layer to obtain the encoded first correlation feature sample XX2, the second correlation feature sample YY2, and the third correlation feature samples XY2 and YX2.
[0209] Step 412: Input the spatially transformed first molecular sequence feature sample and the first molecular sequence feature sample, the encoded first correlation feature sample, the second correlation feature sample, and the third correlation feature sample into the normalization layer to obtain the normalized first correlation feature sample, the second correlation feature sample, and the third correlation feature sample.
[0210] Specifically, input the spatially transformed first molecular sequence feature sample X2 and the first molecular sequence feature sample Y2, the encoded first correlation feature sample XX2, the second correlation feature sample YY2, and the third correlation feature samples XY2 and YX2 into the normalization layer. It is possible to add X2 and XX2, add Y2 and YY2, add X2 and YX2, add Y2 and XY2, and input the obtained first correlation feature sample XXX2, the second correlation feature sample YYY2, and the third correlation feature samples XYY2 and YXX2 into the normalization layer to obtain the normalized first correlation feature sample XXX2, the second correlation feature sample YYY2, and the third correlation feature samples XYY2 and YXX2.
[0211] Step 414: Input the normalized first correlation feature sample, the second correlation feature sample, and the third correlation feature sample into the fully connected layer to obtain the spatially transformed first correlation feature sample, the second correlation feature sample, and the third correlation feature sample.
[0212] Specifically, the normalized first correlation feature sample XXX2, second correlation feature sample YYY2, and third correlation feature samples XYY2 and YXX2 are input into the fully connected layer to obtain the first correlation feature sample XXX3, second correlation feature sample YYY3, and third correlation feature samples XYY3 and YXX3 after spatial transformation processing.
[0213] Step 416: The first correlation feature sample, second correlation feature sample, and third correlation feature sample after spatial transformation processing, and the encoded first correlation feature sample, second correlation feature sample, and third correlation feature sample are input into the normalization processing unit in the output layer to obtain the normalized first correlation feature sample, second correlation feature sample, and third correlation feature sample.
[0214] Specifically, the first correlation feature sample XXX3, second correlation feature sample YYY3, and third correlation feature samples XYY3 and YXX3 after spatial transformation processing, and the encoded first correlation feature sample XX2, second correlation feature sample YY2, and third correlation feature samples XY2 and YX2 are input into the normalization processing unit to obtain the normalized first correlation feature sample XXX4, second correlation feature sample YYY4, and third correlation feature samples XYY4 and YXX4.
[0215] Step 418: The normalized first correlation feature sample, second correlation feature sample, and third correlation feature sample are input into the dimensionality reduction processing unit in the output layer to obtain the first prediction vector, second prediction vector, and third prediction vector.
[0216] Step 420: The first prediction vector, second prediction vector, and third prediction vector are concatenated to obtain the target prediction vector, and based on the target prediction vector, the predicted interaction result between the first object molecular sequence data sample and the second object molecular sequence data sample is output.
[0217] Step 422: According to the predicted interaction result and the interaction result label, the molecular interaction prediction model is trained until a molecular interaction prediction model that meets the training stop condition is obtained.
[0218] In summary, the above method extracts the predicted molecular sequence features of each object molecular sequence data by using a feature extraction model. Moreover, since the predicted molecular sequence features contain the element vectors of molecular elements, feature prediction at the element level is achieved. Then, based on the correlation between the element vectors, a molecular interaction prediction model predicts the interaction results between at least two object molecular sequence data, realizing molecular interaction prediction based on artificial intelligence. Furthermore, based on the predicted molecular sequence features at the element level, the predicted interaction results are made more accurate. Molecular interaction prediction is achieved through the combination of an upstream pre-trained large model (i.e., the feature extraction model) and a downstream specific task network (i.e., the molecular interaction prediction model), while ensuring the efficiency and accuracy of the prediction results.
[0219] Corresponding to the above method embodiment, this specification also provides an embodiment of an object molecular sequence data processing device. Figure 5 The structural schematic diagram of an object molecular sequence data processing device provided by an embodiment of this specification is shown. As Figure 5 shown, the device includes:
[0220] A determination module 502, configured to determine at least two object molecular sequence data;
[0221] A first input module 504, configured to input the at least two object molecular sequence data into a feature extraction model to obtain the predicted molecular sequence features of each object molecular sequence data among the at least two object molecular sequence data, where the predicted molecular sequence features contain the element vectors of molecular elements, and the object molecular sequence data is composed of the molecular elements;
[0222] A second input module 506, configured to input the predicted molecular sequence features of each object molecular sequence data into a molecular interaction prediction model to obtain the interaction results between the at least two object molecular sequence data, where the molecular interaction prediction model outputs the interaction results between the at least two object molecular sequence data according to the correlation between the element vectors included in the predicted molecular sequence features of each object molecular sequence data.
[0223] In summary, the above device extracts the predicted molecular sequence features of each object molecular sequence data by using a feature extraction model. Moreover, since the predicted molecular sequence features contain the element vectors of molecular elements, feature prediction at the element level is achieved. Then, based on the correlation between the element vectors, a molecular interaction prediction model predicts the interaction results between at least two object molecular sequence data, realizing artificial-intelligence-based molecular interaction prediction. Furthermore, based on the predicted molecular sequence features at the element level, the predicted interaction results are made more accurate. Molecular interaction prediction is realized through the combination of an upstream pre-trained large model (i.e., the feature extraction model) and a downstream specific-task network (i.e., the molecular interaction prediction model), while ensuring the efficiency and accuracy of the prediction results.
[0224] The above is a schematic solution of an object molecular sequence data processing device according to this embodiment. It should be noted that the technical solution of this object molecular sequence data processing device and the technical solution of the above object molecular sequence data processing method belong to the same concept. For the details not described in the technical solution of the object molecular sequence data processing device, reference can be made to the description of the technical solution of the above object molecular sequence data processing method.
[0225] See Figure 6 , Figure 6 which shows a flowchart of a model training method according to an embodiment of this specification, specifically including the following steps.
[0226] Step 602: Determine an object molecular sequence data sample;
[0227] Step 604: Train a feature extraction model and a molecular interaction prediction model according to the object molecular sequence data sample;
[0228] Among them, the feature extraction model is used to predict the predicted molecular sequence features of object molecular sequence data. The predicted molecular sequence features contain the element vectors of molecular elements. The object molecular sequence data is composed of the molecular elements. The molecular interaction prediction model is used to output the interaction results between at least two object molecular sequence data according to the correlation between the element vectors included in the predicted molecular sequence features of each object molecular sequence data.
[0229] In summary, the above method extracts the predicted molecular sequence features of each object molecular sequence data by using a feature extraction model. Moreover, since the predicted molecular sequence features contain the element vectors of molecular elements, feature prediction at the element level is achieved. Then, based on the correlation between the element vectors, a molecular interaction prediction model predicts the interaction results between at least two object molecular sequence data, realizing artificial-intelligence-based molecular interaction prediction. Furthermore, based on the predicted molecular sequence features at the element level, the predicted interaction results are made more accurate. Molecular interaction prediction is realized through the combination of an upstream pre-trained large model (i.e., the feature extraction model) and a downstream specific task network (i.e., the molecular interaction prediction model), while ensuring the efficiency and accuracy of the prediction results.
[0230] Corresponding to the above method embodiment, this specification also provides a model training device embodiment. Figure 7 The structural schematic diagram of a model training device provided by an embodiment of this specification is shown. As Figure 7 shown, the device includes:
[0231] A determination module 702, configured to determine an object molecular sequence data sample;
[0232] A training module 704, configured to train a feature extraction model and a molecular interaction prediction model according to the object molecular sequence data sample;
[0233] Wherein, the feature extraction model is used to predict the predicted molecular sequence features of object molecular sequence data. The predicted molecular sequence features contain the element vectors of molecular elements. The object molecular sequence data is composed of the molecular elements. The molecular interaction prediction model is used to output the interaction results between at least two object molecular sequence data according to the correlation between the element vectors included in the predicted molecular sequence features of each object molecular sequence data.
[0234] In summary, the above device extracts the predicted molecular sequence features of each object molecular sequence data by using a feature extraction model. Moreover, since the predicted molecular sequence features contain the element vectors of molecular elements, feature prediction at the element level is achieved. Then, based on the correlation between the element vectors, a molecular interaction prediction model predicts the interaction results between at least two object molecular sequence data, realizing artificial-intelligence-based molecular interaction prediction. Furthermore, based on the predicted molecular sequence features at the element level, the predicted interaction results are made more accurate. Molecular interaction prediction is realized through the combination of an upstream pre-trained large model (i.e., the feature extraction model) and a downstream specific task network (i.e., the molecular interaction prediction model), while ensuring the efficiency and accuracy of the prediction results.
[0235] The above is a schematic solution of a model training device according to this embodiment. It should be noted that the technical solution of this model training device and the technical solution of the above model training method belong to the same concept. For the details not described in detail in the technical solution of the model training device, reference can be made to the description of the technical solution of the above model training method.
[0236] See Figure 8 , Figure 8 shows a schematic diagram of a method for processing object molecule sequence data and a model training method according to an embodiment of this specification.
[0237] As Figure 8 shown, according to what is shown in 802, a large amount of unlabeled molecular data (i.e., a molecular pre-training dataset) can be used to train a pre-trained large model (i.e., a feature extraction model) to obtain one or more trained pre-trained large models. Then, according to what is shown in 804, a constructed molecular interaction prediction model can be trained using a labeled molecular interaction dataset (usually a relatively small dataset). Specifically, during the training process, each molecular sequence will call the trained pre-trained large model to obtain the representation matrix of the sequence. The molecular interaction prediction model uses this representation matrix as an input or one of the inputs to obtain a trained molecular interaction prediction model. Finally, according to what is shown in 806, in the application stage, for two or more object molecule sequence data of unknown interactions (such as object molecule sequence data 1 and object molecule sequence data 2), the trained pre-trained large model is called respectively to obtain their representation matrices. The molecular sequence pair (or multiple molecular sequences) and their large model representation matrices are used as inputs to the trained molecular interaction prediction model, and the interaction result (binary classification / multi-classification / regression) of this molecular pair (or multiple molecules) is output.
[0238] Figure 9 shows a structural block diagram of a computing device 900 according to an embodiment of this specification. The components of this computing device 900 include but are not limited to a memory 910 and a processor 920. The processor 920 is connected to the memory 910 through a bus 930, and a database 950 is used to store data.
[0239] The computing device 900 also includes an access device 940, which enables the computing device 900 to communicate via one or more networks 960. Examples of such networks include the Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or a combination of communication networks such as the Internet. The access device 940 may include one or more of any type of wired or wireless network interface (e.g., network interface card (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, Worldwide Interoperability for Microwave Access (Wi-MAX) interface, Ethernet interface, Universal Serial Bus (USB) interface, cellular network interface, Bluetooth interface, Near Field Communication (NFC) interface, and so on.
[0240] In one embodiment of the present application, the above components of the computing device 900 and Figure 9 other components not shown may also be connected to each other, for example, via a bus. It should be understood that Figure 9 the block diagram of the computing device shown is for illustrative purposes only and is not a limitation on the scope of the present application. Those skilled in the art can add or replace other components as needed.
[0241] The computing device 900 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs). The computing device 900 can also be a mobile or stationary server.
[0242] Among them, the processor 920 is used to execute the following computer-executable instructions, and when the computer-executable instructions are executed by the processor, the steps of the above object molecular sequence data processing method and model training method are implemented.
[0243] The above is a schematic solution of a computing device according to this embodiment. It should be noted that the technical solution of this computing device and the technical solutions of the above object molecular sequence data processing method and model training method belong to the same concept. For the detailed content not described in the technical solution of the computing device, reference can be made to the descriptions of the technical solutions of the above object molecular sequence data processing method and model training method.
[0244] An embodiment of this specification further provides a computer-readable storage medium, which stores computer-executable instructions. When the computer-executable instructions are executed by a processor, the steps of the above object molecular sequence data processing method and model training method are implemented.
[0245] The above is a schematic solution of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium and the technical solutions of the above object molecular sequence data processing method and model training method belong to the same concept. For the detailed content not described in the technical solution of the storage medium, reference can be made to the descriptions of the technical solutions of the above object molecular sequence data processing method and model training method.
[0246] An embodiment of this specification further provides a computer program. When the computer program is executed on a computer, the computer is made to execute the steps of the above object molecular sequence data processing method and model training method.
[0247] The above is a schematic solution of a computer program according to this embodiment. It should be noted that the technical solution of this computer program and the technical solutions of the above object molecular sequence data processing method and model training method belong to the same concept. For the detailed content not described in the technical solution of the computer program, reference can be made to the descriptions of the technical solutions of the above object molecular sequence data processing method and model training method.
[0248] The above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be executed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain implementations, multitasking and parallel processing are also possible or may be advantageous.
[0249] The computer instructions include computer program code, which may be in the form of source code, object code, executable files, or some intermediate forms, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of patent practice. For example, in some regions, according to patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0250] It should be noted that for the foregoing method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the embodiments of this specification are not limited by the described action sequence, because according to the embodiments of this specification, some steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential for the embodiments of this specification.
[0251] In the above embodiments, the descriptions of each embodiment have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0252] The preferred embodiments of this specification disclosed above are only used to help explain this specification. The optional embodiments do not describe all the details in detail, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and changes can be made according to the content of the embodiments of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the embodiments of this specification, so that those skilled in the art can understand and utilize this specification well. This specification is only limited by the claims and their full scope and equivalents.
Claims
1. A method for processing object molecule sequence data, comprising: Determine at least two pieces of object molecule sequence data; Input the at least two pieces of object molecule sequence data into a feature extraction model to obtain the predicted molecular sequence features of each piece of object molecule sequence data among the at least two pieces of object molecule sequence data, wherein the predicted molecular sequence features include element vectors of molecular elements, and the object molecule sequence data is composed of the molecular elements; Input the predicted molecular sequence features of each piece of object molecule sequence data into a molecular interaction prediction model to obtain the interaction results between the at least two pieces of object molecule sequence data, wherein the molecular interaction prediction model outputs the interaction results between the at least two pieces of object molecule sequence data according to the correlation between the element vectors included in the predicted molecular sequence features of each piece of object molecule sequence data.
2. The method for processing object molecule sequence data according to claim 1, wherein the step of inputting the at least two pieces of object molecule sequence data into a feature extraction model to obtain the predicted molecular sequence features of each piece of object molecule sequence data among the at least two pieces of object molecule sequence data comprises: Select any two pieces of object molecule sequence data from the at least two pieces of object molecule sequence data as the first object molecule sequence data and the second object molecule sequence data; Input the first object molecule sequence data and the second object molecule sequence data into the feature extraction model to obtain the first predicted molecular sequence feature of the first object molecule sequence data and the second predicted molecular sequence feature of the second object molecule sequence data; Correspondingly, the step of inputting the predicted molecular sequence features of each piece of object molecule sequence data into a molecular interaction prediction model to obtain the interaction results between the at least two pieces of object molecule sequence data comprises: Input the first predicted molecular sequence feature and the second predicted molecular sequence feature into the molecular interaction prediction model to obtain the interaction results between the first object molecule sequence data and the second object molecule sequence data.
3. The method for processing object molecule sequence data according to claim 2, wherein the step of inputting the first predicted molecular sequence feature and the second predicted molecular sequence feature into the molecular interaction prediction model to obtain the interaction results between the first object molecule sequence data and the second object molecule sequence data comprises: Input the first predicted molecular sequence feature and the second predicted molecular sequence feature into the correlation processing layer in the molecular interaction prediction model to obtain the first correlation feature between the first molecular elements included in the first object molecule sequence data, the second correlation feature between the second molecular elements included in the second object molecule sequence data, and the third correlation feature between the first molecular elements and the second molecular elements; Input the first correlation feature, the second correlation feature, and the third correlation feature into the element position encoding layer in the molecular interaction prediction model to obtain the encoded first correlation feature, second correlation feature, and third correlation feature; Input the encoded first correlation feature, second correlation feature, and third correlation feature into the output layer of the molecular interaction prediction model to obtain the interaction result between the first object molecular sequence data and the second object molecular sequence data.
4. The method for processing object molecular sequence data according to claim 3, before inputting the first predicted molecular sequence feature and the second predicted molecular sequence feature into the correlation processing layer of the molecular interaction prediction model, further comprising: Input the first predicted molecular sequence feature and the second predicted molecular sequence feature into the fully connected layer of the molecular interaction prediction model to obtain the first predicted molecular sequence feature and the second predicted molecular sequence feature after spatial transformation processing; Input the first predicted molecular sequence feature and the second predicted molecular sequence feature after spatial transformation processing into the normalization layer of the molecular interaction prediction model to obtain the first predicted molecular sequence feature and the second predicted molecular sequence feature after normalization processing; Correspondingly, the inputting the first predicted molecular sequence feature and the second predicted molecular sequence feature into the correlation processing layer of the molecular interaction prediction model includes: Input the first predicted molecular sequence feature and the second predicted molecular sequence feature after normalization processing into the correlation processing layer of the molecular interaction prediction model.
5. The method for processing object molecular sequence data according to claim 4, before inputting the encoded first correlation feature, second correlation feature, and third correlation feature into the output layer of the molecular interaction prediction model, further comprising: Input the first predicted molecular sequence feature and the second predicted molecular sequence feature after spatial transformation processing, the encoded first correlation feature, second correlation feature, and third correlation feature into the normalization layer of the molecular interaction prediction model to obtain the first correlation feature, second correlation feature, and third correlation feature after normalization processing; Input the first correlation feature, second correlation feature, and third correlation feature after normalization processing into the fully connected layer of the molecular interaction prediction model to obtain the first correlation feature, second correlation feature, and third correlation feature after spatial transformation processing; Correspondingly, the inputting the encoded first correlation feature, second correlation feature, and third correlation feature into the output layer of the molecular interaction prediction model to obtain the interaction result between the first object molecular sequence data and the second object molecular sequence data includes: Input the encoded first correlation feature, second correlation feature, and third correlation feature, the first correlation feature, second correlation feature, and third correlation feature after spatial transformation processing into the output layer of the molecular interaction prediction model to obtain the interaction result between the first object molecular sequence data and the second object molecular sequence data.
6. The method for processing object molecule sequence data according to claim 5, wherein inputting the encoded first correlation feature, second correlation feature, and third correlation feature, and the first correlation feature, second correlation feature, and third correlation feature after the spatial transformation processing into the output layer of the molecule interaction prediction model to obtain the interaction result between the first object molecule sequence data and the second object molecule sequence data includes: Performing a fusion process on the encoded first correlation feature, second correlation feature, and third correlation feature, and the first correlation feature, second correlation feature, and third correlation feature after the spatial transformation processing to obtain a first fused correlation feature, a second fused correlation feature, and a third fused correlation feature; Inputting the first fused correlation feature, the second fused correlation feature, and the third fused correlation feature into the output layer of the molecule interaction prediction model to obtain the interaction result between the first object molecule sequence data and the second object molecule sequence data.
7. The method for processing object molecule sequence data according to claim 6, wherein the output layer includes a dimensionality reduction processing unit; Correspondingly, the step of inputting the first fused correlation feature, the second fused correlation feature, and the third fused correlation feature into the output layer of the molecule interaction prediction model to obtain the interaction result between the first object molecule sequence data and the second object molecule sequence data includes: Inputting the first fused correlation feature, the second fused correlation feature, and the third fused correlation feature into the output layer of the molecule interaction prediction model; In the output layer, performing dimensionality reduction processing on the first fused correlation feature, the second fused correlation feature, and the third fused correlation feature according to the dimensionality reduction processing unit to obtain a first prediction vector, a second prediction vector, and a third prediction vector; Outputting the interaction result between the first object molecule sequence data and the second object molecule sequence data according to the first prediction vector, the second prediction vector, and the third prediction vector.
8. The method for processing object molecule sequence data according to claim 7, wherein the output layer further includes a normalization processing unit; Correspondingly, before performing dimensionality reduction processing on the first fused correlation feature, the second fused correlation feature, and the third fused correlation feature according to the dimensionality reduction processing unit to obtain a first prediction vector, a second prediction vector, and a third prediction vector, it further includes: Performing normalization processing on the first fused correlation feature, the second fused correlation feature, and the third fused correlation feature according to the normalization processing unit to obtain a normalized first fused correlation feature, a normalized second fused correlation feature, and a normalized third fused correlation feature; Correspondingly, the step of performing dimensionality reduction processing on the first fused correlation feature, the second fused correlation feature, and the third fused correlation feature according to the dimensionality reduction processing unit to obtain a first prediction vector, a second prediction vector, and a third prediction vector includes: According to the dimensionality reduction processing unit, perform dimensionality reduction processing on the first fused correlation feature, the second fused correlation feature, and the third fused correlation feature after normalization processing to obtain a first prediction vector, a second prediction vector, and a third prediction vector.
9. The method for processing object molecular sequence data according to claim 5, wherein inputting the first predicted molecular sequence feature and the second predicted molecular sequence feature after the spatial transformation processing, the first correlation feature, the second correlation feature, and the third correlation feature after encoding into the normalization layer in the molecular interaction prediction model includes: According to a preset fusion rule, perform fusion processing on the first predicted molecular sequence feature and the second predicted molecular sequence feature after the spatial transformation processing, the first correlation feature, the second correlation feature, and the third correlation feature after encoding to obtain a first fused correlation feature, a second fused correlation feature, and a third fused correlation feature; Input the first fused correlation feature, the second fused correlation feature, and the third fused correlation feature into the normalization layer in the molecular interaction prediction model.
10. The method for processing object molecular sequence data according to claim 1, wherein the training steps of the molecular interaction prediction model include: Determine molecular sequence feature samples of at least two object molecular sequence data samples and interaction result labels between the at least two object molecular sequence data samples; Input the molecular sequence feature samples of the at least two object molecular sequence data samples and the interaction result labels into the molecular interaction prediction model to obtain a predicted interaction result; According to the predicted interaction result and the interaction result label, train the molecular interaction prediction model until a molecular interaction prediction model that meets the training stop condition is obtained.
11. The method for processing object molecular sequence data according to claim 1, wherein the training steps of the feature extraction model include: Determine an object molecular sequence data sample; According to the object molecular sequence data sample, perform self-supervised training on the feature extraction model until a feature extraction model that meets the training stop condition is obtained.
12. The method for processing object molecular sequence data according to claim 11, wherein according to the object molecular sequence data sample, performing self-supervised training on the feature extraction model until a feature extraction model that meets the training stop condition is obtained, includes: Input the object molecular sequence data sample into the feature extraction model, and in the feature extraction model, perform masking processing on the object molecular sequence data sample to obtain masked object molecular sequence data; Perform prediction on the masked object molecular sequence data to obtain predicted object molecular sequence data; According to the object molecular sequence data sample and the predicted object molecular sequence data, perform self-supervised training on the feature extraction model until a feature extraction model that meets the training stop condition is obtained.
13. The method for processing object molecular sequence data according to claim 11, wherein according to the object molecular sequence data sample, performing self-supervised training on the feature extraction model until a feature extraction model that meets the training stop condition is obtained, includes: Input the object molecular sequence data sample into the feature extraction model. In the feature extraction model, process the object molecular sequence data sample to obtain a first object molecular sequence data sample and a second object molecular sequence data sample; Create a target object molecular sequence data sample according to the first object molecular sequence data sample and the second object molecular sequence data sample; Perform a prediction on the target object molecular sequence data sample to obtain a prediction result, where the prediction result is used to predict whether the first object molecular sequence data sample and the second object molecular sequence data sample belong to the same object molecular sequence data sample; Perform self-supervised training on the feature extraction model according to the target object molecular sequence data sample and the prediction result until a feature extraction model that meets the training stop condition is obtained.
14. A model training method, comprising: Determine an object molecular sequence data sample; Train a feature extraction model and a molecular interaction prediction model according to the object molecular sequence data sample; Wherein, the feature extraction model is used to predict the predicted molecular sequence features of the object molecular sequence data, the predicted molecular sequence features include element vectors of molecular elements, the object molecular sequence data is composed of the molecular elements, and the molecular interaction prediction model is used to output the interaction result between at least two object molecular sequence data according to the correlation between the element vectors included in the predicted molecular sequence features of each object molecular sequence data.
15. The model training method according to claim 14, wherein the training step of the feature extraction model comprises: Determine an object molecular sequence data sample; Perform self-supervised training on the feature extraction model according to the object molecular sequence data sample until a feature extraction model that meets the training stop condition is obtained.
16. The model training method according to claim 15, wherein performing self-supervised training on the feature extraction model according to the object molecular sequence data sample until a feature extraction model that meets the training stop condition is obtained comprises: Input the object molecular sequence data sample into the feature extraction model. In the feature extraction model, perform a masking process on the object molecular sequence data sample to obtain a masked object molecular sequence data; Perform a prediction on the masked object molecular sequence data to obtain a predicted object molecular sequence data; Perform self-supervised training on the feature extraction model according to the object molecular sequence data sample and the predicted object molecular sequence data until a feature extraction model that meets the training stop condition is obtained.
17. The model training method according to claim 15, wherein performing self-supervised training on the feature extraction model according to the object molecular sequence data sample until a feature extraction model that meets the training stop condition is obtained comprises: Input the object molecular sequence data sample into the feature extraction model. In the feature extraction model, process the object molecular sequence data sample to obtain a first object molecular sequence data sample and a second object molecular sequence data sample; Create a target object molecular sequence data sample based on the first object molecular sequence data sample and the second object molecular sequence data sample; Perform prediction on the target object molecular sequence data sample to obtain a prediction result, where the prediction result is used to predict whether the first object molecular sequence data sample and the second object molecular sequence data sample belong to the same object molecular sequence data sample; Perform self-supervised training on the feature extraction model based on the target object molecular sequence data sample and the prediction result until a feature extraction model that meets the training stop condition is obtained.
18. The model training method according to claim 14, wherein the training steps of the molecular interaction prediction model include: Determine the molecular sequence feature samples of at least two object molecular sequence data samples and the interaction result labels between the at least two object molecular sequence data samples; Input the molecular sequence feature samples of the at least two object molecular sequence data samples and the interaction result labels into the molecular interaction prediction model to obtain a predicted interaction result; Train the molecular interaction prediction model based on the predicted interaction result and the interaction result label until a molecular interaction prediction model that meets the training stop condition is obtained.
19. A computing device, comprising: A memory and a processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the method according to any one of claims 1 to 18 are implemented.
20. A computer-readable storage medium storing computer-executable instructions, which when executed by a processor implement the steps of the method according to any one of claims 1 to 18.