Protein matching method, device, equipment, storage medium and program product
By constructing a target evolutionary tree and querying the matching information of neighboring proteins, the problem of low antibody screening efficiency in existing technologies is solved, and more efficient and accurate protein matching is achieved.
Patent Information
- Application Number
- CN202410258455.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-06
- Publication Date
- 2025-09-09
AI Technical Summary
The existing technology has low efficiency and high complexity in antibody screening, resulting in insufficient protein matching efficiency and accuracy.
By constructing a target evolutionary tree, querying neighbor proteins based on the distance of the candidate protein, and using the matching information of the neighbor proteins to predict the matching results of the candidate protein, the matching process is simplified.
The matching efficiency and accuracy of candidate proteins with the second type of proteins are improved, the matching process is simplified, and the evolutionary relationship and functional similarity between proteins can be understood more accurately.
Smart Images

Figure CN120613002A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a protein matching method, apparatus, device, storage medium, and program product. Background Art
[0002] With the rapid development of artificial intelligence (AI), this technology is being applied in a variety of fields. Antibodies in the biomedical field, due to their ability to bind to protein antigens with high affinity and specificity, have become one of the most commonly used protein therapeutics. The use of AI to screen specific antibodies is an important tool in drug development, auxiliary diagnostics, and other fields.
[0003] In related technologies, taking the screening of specific antibodies as an example, antibody testers can obtain candidate antibodies through the antibody library, and then evaluate the affinity of the candidate antibodies for antigens through a computational method based on the antibody sequence prediction of antibodies that specifically bind to antigens.
[0004] However, due to the large number of possible antibodies and the overly complex related screening methods, the efficiency of antibody matching is low. Summary of the Invention
[0005] The present invention provides a protein matching method, apparatus, device, storage medium, and program product, which can improve the efficiency and accuracy of matching candidate proteins. The technical solution is as follows:
[0006] In one aspect, a protein matching method is provided, comprising:
[0007] Obtaining a target evolutionary tree, the target evolutionary tree being an evolutionary tree constructed based on specified amino acid sequences in a plurality of first-type proteins; leaf nodes of the evolutionary tree being the first-type proteins; and some of the first-type proteins in the evolutionary tree having matching information, the matching information being used to indicate a degree of matching between the first-type proteins and second-type proteins;
[0008] For a candidate protein, querying neighbor proteins of a target number of neighbors from the target evolutionary tree based on the distance to the candidate protein; the candidate protein is other first-type proteins other than the first-type proteins having the matching information; the neighbor protein is the first-type protein having the matching information;
[0009] obtaining predicted matching information of the candidate protein based on the matching information of the neighbor proteins of a target number of neighbors;
[0010] Based on the predicted matching information of the candidate protein, a matching result of the candidate protein is obtained.
[0011] In another aspect, a protein matching device is provided, comprising:
[0012] a target evolutionary tree acquisition module, configured to acquire a target evolutionary tree, wherein the target evolutionary tree is an evolutionary tree constructed based on specified amino acid sequences in a plurality of first-type proteins; leaf nodes of the evolutionary tree are the first-type proteins; and some of the first-type proteins in the evolutionary tree have matching information, wherein the matching information is used to indicate the degree of matching between the first-type proteins and the second-type proteins;
[0013] a query module configured to query neighbor proteins having a target number of neighbors from the target evolutionary tree for a candidate protein based on a distance from the candidate protein; the candidate protein being other proteins of the first type other than the first type protein having the matching information; and the neighbor protein being the first type protein having the matching information;
[0014] a predicted matching information acquisition module, configured to acquire predicted matching information of the candidate protein based on the matching information of the neighbor proteins of the target number of neighbors;
[0015] A matching result acquisition module is used to acquire the matching result of the candidate protein based on the predicted matching information of the candidate protein.
[0016] In one possible implementation, the predicted matching information acquisition module is used to obtain the predicted matching information of the candidate protein based on the matching information of the neighbor protein with a target number of neighbors, and the distance between the neighbor protein with a target number of neighbors and the candidate protein in the target evolutionary tree.
[0017] In one possible implementation, the predicted matching information acquisition module is used to scale the matching information of the neighbor protein based on the distance between the neighbor protein and the candidate protein in the target evolutionary tree, and obtain the predicted matching sub-information of the candidate protein corresponding to the neighbor protein; and obtain the predicted matching information of the candidate protein based on the predicted matching sub-information of the neighbor protein corresponding to the target number of neighbors of the candidate protein.
[0018] In one possible implementation, the coefficient for scaling the matching information of the neighbor protein is the inverse of the distance between the neighbor protein and the candidate protein in the target evolutionary tree; or, the coefficient for scaling the matching information of the neighbor protein is determined by the distance interval between the distance between the neighbor protein and the candidate protein in the target evolutionary tree.
[0019] In a possible implementation, the apparatus further includes:
[0020] An evolutionary tree acquisition module, configured to acquire a plurality of the evolutionary trees before the target evolutionary tree acquisition module acquires the target evolutionary tree; each of the plurality of evolutionary trees corresponds to one of the specified amino acid sequences, or corresponds to a combination of multiple of the specified amino acid sequences;
[0021] The combination acquisition module is used to obtain a combination of the target evolutionary tree and the target number of neighbors by taking each of the first type proteins having the matching information in the plurality of evolutionary trees as a sample.
[0022] In one possible implementation, the combination acquisition module is used to divide each of the first-type proteins having the matching information into a training sample set and a test sample set; for a candidate protein sample, based on the distance from the candidate protein sample, query a first number of neighbor protein samples from a first evolutionary tree; the candidate protein sample is a protein in the test sample set, and the neighbor protein sample is a protein in the training sample set; the first evolutionary tree is one of multiple evolutionary trees; based on the matching information of the neighbor protein samples of the first number of neighbors, obtain the predicted matching information of the candidate protein sample; based on the predicted matching information of the candidate protein sample and the matching information of the candidate protein sample, obtain a correlation coefficient corresponding to a combination of the first evolutionary tree and the first number of neighbors; based on the correlation coefficients corresponding to multiple evolutionary trees and various combinations of multiple numbers of neighbors, obtain a combination of the target evolutionary tree and the target number of neighbors.
[0023] In a possible implementation, the combination acquisition module is used to obtain the predicted matching information of each protein in the test sample set, the average value of the predicted matching information of each protein in the test sample set, the matching information of each protein in the test sample set, and the average value of the matching information of each protein in the test sample set; based on the predicted matching information of each protein in the test sample set, the average value of the predicted matching information of each protein in the test sample set, the matching information of each protein in the test sample set, and the average value of the matching information of each protein in the test sample set, calculate the Pearson correlation coefficient to obtain the correlation coefficient corresponding to the combination of the first evolutionary tree and the first number of neighbors.
[0024] In one possible implementation, the combination acquisition module is used to obtain the correlation coefficient of the combination of the first evolutionary tree and the first number of neighbors relative to the candidate protein sample based on the predicted matching information of the candidate protein sample, the matching information of the candidate protein sample, the average value of the predicted matching information of each protein in the test sample set, and the average value of the matching information of each protein in the test sample set; and take the average or median of the correlation coefficients of the combination of the first evolutionary tree and the first number of neighbors relative to each protein in the test sample set to obtain the correlation coefficient corresponding to the combination of the first evolutionary tree and the first number of neighbors.
[0025] In one possible implementation, the combination acquisition module is used to arrange various combinations of multiple evolutionary trees and multiple numbers of neighbors in order from high to low according to the corresponding correlation coefficients; determine one or more combinations arranged in front of the various combinations of multiple evolutionary trees and multiple numbers of neighbors as the combination of the target evolutionary tree and the target number of neighbors; or, determine the combination whose corresponding correlation coefficient is higher than the correlation coefficient threshold among the various combinations of multiple evolutionary trees and multiple numbers of neighbors as the combination of the target evolutionary tree and the target number of neighbors.
[0026] In a possible implementation, each combination of the plurality of evolutionary trees and the plurality of neighbor numbers corresponds to a plurality of correlation coefficients; each of the plurality of correlation coefficients corresponds to a different sample partitioning method, wherein the sample partitioning method is a method of partitioning each of the first type proteins having the matching information into the training sample set and the test sample set;
[0027] The combination acquisition module is used to take the median of the multiple correlation coefficients of each combination of the multiple evolutionary trees and the multiple numbers of neighbors to obtain the median correlation coefficient of the combination; arrange the various combinations of the multiple evolutionary trees and the multiple numbers of neighbors in ascending order according to the difference between the corresponding median correlation coefficient and the target correlation coefficient; and determine one or more combinations arranged in front of the various combinations of the multiple evolutionary trees and the multiple numbers of neighbors as the combination of the target evolutionary tree and the target number of neighbors.
[0028] In a possible implementation, the specified amino acid sequence includes one or more of the following amino acid sequences:
[0029] Heavy chain amino acid sequence, light chain amino acid sequence, amino acid sequence corresponding to the heavy chain V gene, amino acid sequence corresponding to the light chain V gene, amino acid sequence corresponding to the heavy chain J gene, and amino acid sequence corresponding to the light chain J gene.
[0030] On the other hand, a computer device is provided, comprising a processor and a memory, wherein the memory stores at least one computer program, and the processor loads and executes the at least one computer program to implement the above protein matching method.
[0031] On the other hand, a computer-readable storage medium is provided, wherein at least one computer program is stored in the computer-readable storage medium, and the computer program is loaded and executed by a processor to implement the above protein matching method.
[0032] In another aspect, a computer program product is provided, comprising a computer program stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium and executes the computer program, causing the computer device to perform the protein matching method provided in the various optional implementations described above.
[0033] The technical solution provided by this application may have the following beneficial effects:
[0034] Different predicted matching information is obtained by using the distance and matching information of neighbor proteins with different target neighbor numbers of the candidate protein in the target evolutionary tree, and then the matching results of the candidate protein and the second type of protein are obtained based on the different predicted matching information; in this scheme, the distance information between the candidate protein and the neighbor protein in the target evolutionary tree provides important information about the evolutionary relationship of proteins. The distance information can be used to more accurately understand the evolutionary relationship between proteins, thereby more accurately understanding the structural and functional similarities between different proteins. In addition, the function of a protein is usually related to the relationship with other similar proteins near its evolutionary tree position. Obtaining neighbor proteins helps to more accurately predict the function of the candidate protein. Based on the target neighbors and the target evolutionary tree, the matching results of the candidate protein can be determined to complete the prediction of the candidate protein without the need for complex parameters, simplifying the prediction and matching process of the candidate protein and the second type of protein, thereby improving the prediction efficiency and accuracy of the candidate protein and the second type of protein. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0036] Figure 1 is a schematic diagram of a system involved in an exemplary embodiment of the present application;
[0037] Figure 2 is a flow chart of a protein matching method shown in an exemplary embodiment of the present application;
[0038] Figure 3 is a flow chart of a protein matching method shown in an exemplary embodiment of the present application;
[0039] Figure 4 is a schematic diagram of a target evolutionary tree provided by an exemplary embodiment of the present application;
[0040] Figure 5 is a flow chart of a protein matching method shown in an exemplary embodiment of the present application;
[0041] Figure 6 is a flow chart of a protein matching method shown in an exemplary embodiment of the present application;
[0042] Figure 7 is a framework diagram of an antibody detection scheme involved in an exemplary embodiment of the present application;
[0043] Figure 8 is an evolutionary tree structure diagram involved in an exemplary embodiment of the present application;
[0044] Figure 9 It is a box plot of 12 evolutionary trees with 1-7 neighbors involved in an exemplary embodiment of the present application;
[0045] Figure 10 is a block diagram of a protein matching device shown in an exemplary embodiment of the present application;
[0046] Figure 11 A structural block diagram of a computer device shown in an exemplary embodiment of the present application is shown. DETAILED DESCRIPTION
[0047] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with the present application. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present application, as detailed in the appended claims.
[0048] It should be understood that the term "plurality" used herein refers to two or more. "And / or" describes a relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can mean: A exists alone, A and B exist simultaneously, or B exists alone. The character " / " generally indicates an "or" relationship between the associated objects.
[0049] The present application provides an object classification method for detecting the concentration level of a second substance triggered by another substance. This method can be used to reduce the difficulty and requirements of substance concentration level detection and improve the efficiency and accuracy of substance concentration level detection. To facilitate understanding, several terms used in this application are explained below.
[0050] 1) Artificial Intelligence (AI)
[0051] Artificial intelligence (AI) refers to the theories, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI also involves studying the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making.
[0052] Artificial intelligence (AI) technology is a comprehensive discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, big data processing, pre-trained models, operating / interaction systems, and mechatronics. Pre-trained models, also known as large models or basic models, can be fine-tuned and widely applied to downstream tasks across various AI disciplines. AI software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning / deep learning.
[0053] 2) Machine Learning (ML)
[0054] Machine learning is a multidisciplinary field that encompasses probability theory, statistics, approximation theory, convex analysis, and algorithmic complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is at the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and self-learning. Pretrained models are the latest development in deep learning, integrating these techniques.
[0055] 3) Antibodies
[0056] Antibodies are large Y-shaped proteins secreted by plasma cells (effector B cells) and used by the immune system to identify and neutralize foreign substances such as bacteria and viruses. They are only found in the blood and other body fluids of vertebrates and on the cell membrane surface of their B cells.
[0057] 4) ANARCI (Antibody Numbering and Antigen Receptor Classification)
[0058] ANARCI is a program for numbering antibody and T-cell receptor sequences.
[0059] 5) HMMER (Hidden Markov Model based on Evolutionary Relationships)
[0060] It is mainly used to search for sequences containing specific structural domains in protein databases. It uses hidden Markov models to represent the sequence characteristics of protein families and can capture the probability distribution and evolutionary relationships of protein sequences. The output of HMMER usually contains an expected value, which indicates the expected probability of the search result occurring randomly. The smaller the expected value, the better the match.
[0061] 6) IMGT (ImMunoGeneTics information system)
[0062] A database and analysis tool focused on information related to the immune system and antibodies. IMGT provides a standardized scheme for describing and naming antibody genes, antibody domains, and other molecules related to the immune system. A database and information system for storing and analyzing immunoglobulins (antibodies) and T cell receptors (TCRs). IMGT provides a standardized naming and numbering system for these molecules to facilitate sharing and comparing data across different studies. In IMGT, the domains of antibodies and TCRs are defined as specific sequence regions and are identified using specific symbols and numbering systems.
[0063] 7) KNN (k-Nearest Neighbor) algorithm
[0064] If most of the k nearest neighboring samples of a sample in the feature space belong to a certain category, then the sample also belongs to this category and has the characteristics of the samples in this category.
[0065] The solution provided in the embodiments of the present application involves an artificial intelligence hidden Markov model based on evolutionary relationships and a K-nearest neighbor algorithm, which is specifically illustrated by the following embodiments.
[0066] Figure 1 A schematic diagram of a system used in a protein matching method provided by an exemplary embodiment of the present application is shown. Figure 1 As shown, the system includes: a server 110 and a terminal 120.
[0067] Among them, the above-mentioned server 110 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), as well as big data and artificial intelligence platforms.
[0068] The server 110 may include a server that deploys an object classification system and provides object classification services to users through the object classification system. Alternatively, the server 110 may also include a server that deploys an object classification system and trains or updates the object classification system.
[0069] The terminal 120 may be a terminal device having network connection and data processing functions, for example, a smart phone, a tablet computer, an e-book reader, smart glasses, a smart watch, a smart TV, a laptop computer, a desktop computer, etc.
[0070] The terminal 120 may include a user terminal for receiving an object classification service, or the terminal 120 may include a development terminal used by developers of the object classification system.
[0071] Optionally, an object classification system may also be deployed in the terminal 120 .
[0072] Optionally, the above system includes one or more servers 110 and multiple terminals 120. The embodiment of the present application does not limit the number of servers 110 and terminals 120.
[0073] The terminal and the server are connected via a communication network. Optionally, the communication network is a wired network or a wireless network.
[0074] Optionally, the above-mentioned wireless network or wired network uses standard communication technology and / or protocol. The network is typically the Internet, but it can also be any network, including but not limited to a local area network (LAN), a metropolitan area network (MAN), a wide area network (WAN), a mobile, wired or wireless network, a private network or any combination of a virtual private network. In some embodiments, the data exchanged over the network is represented using technologies and / or formats such as Hyper Text Markup Language (HTML) and Extensible Markup Language (XML). Conventional encryption technologies such as Secure Socket Layer (SSL), Transport Layer Security (TLS), Virtual Private Network (VPN), and Internet Protocol Security (IPsec) can also be used to encrypt all or some links. In other embodiments, customized and / or dedicated data communication technologies can also be used to replace or supplement the above-mentioned data communication technologies. This application is not limited here.
[0075] Figure 2 A flowchart of a protein matching method according to an exemplary embodiment of the present invention is shown. The method is executed by a computer device. The computer device can be implemented as a terminal or a server. The terminal or server can be Figure 1 The terminal or server shown, such as Figure 2 As shown, the object classification method includes the following steps:
[0076] Step 210: Obtain a target evolutionary tree, which is an evolutionary tree constructed based on specified amino acid sequences in multiple first-type proteins; the leaf nodes of the evolutionary tree are first-type proteins; some of the first-type proteins in the evolutionary tree have matching information, and the matching information is used to indicate the degree of matching between the first-type proteins and the second-type proteins.
[0077] Among them, the above-mentioned evolutionary tree is used to indicate the evolutionary relationship between multiple proteins. The evolutionary tree consists of nodes, multiple leaf nodes and branches; each node represents a protein, the leaf node represents the protein molecular sequence, and the branch is a line segment connecting different nodes, representing the path or relationship of the protein in the evolutionary process.
[0078] The first type of protein may be an antibody, the second type of protein may be an antigen, and the antibody may be an antibody that can bind to a specific position of the antigen.
[0079] The matching information may be information that is pre-labeled by the user for some of the first type of proteins; for example, the matching information may be in the form of a score, indicating the degree of matching between the first type of protein and the second type of protein, or the matching information may be in the form of a classification label, indicating the type of matching between the first type of protein and the second type of protein, such as high matching, medium matching, and low matching.
[0080] Step 220: For the candidate protein, query neighbor proteins with a target number of neighbors from the target evolutionary tree based on the distance to the candidate protein; the candidate protein is other first type proteins other than the first type proteins with matching information; the neighbor protein is the first type protein with matching information.
[0081] The candidate protein is a first type of protein that has no matching information.
[0082] Among them, the above distance is the distance between the neighbor protein and the candidate protein, which is used to represent the similarity between the neighbor protein and the candidate protein in the protein space. The above distance can be any one or more of sequence similarity distance, structural similarity distance, functional similarity distance and binding affinity distance.
[0083] Step 230: Based on the matching information of neighbor proteins with the target number of neighbors, obtain the predicted matching information of the candidate protein.
[0084] In one possible implementation, there may be multiple combinations of target evolutionary trees and target numbers of neighbors. Corresponding to the multiple combinations, the computer device may obtain multiple predicted matching information and obtain the average value of the multiple predicted matching information as the final predicted matching information of the candidate protein.
[0085] Step 240: Based on the predicted matching information of the candidate protein, obtain the matching result of the candidate protein.
[0086] For example, the above matching result can be used to indicate whether the candidate protein matches the second type of protein.
[0087] Among them, the above matching results can be displayed as one of "yes" or "no", "yes" indicates that the candidate protein matches the second type of protein, and "no" indicates that the candidate protein does not match the second type of protein; the above matching results can also be displayed in the form of a percentage, and the larger the percentage value, the higher the degree of matching between the candidate protein and the second type of protein; the above matching results can also be displayed in the form of a score, and the larger the score, the higher the degree of matching between the candidate protein and the second type of protein.
[0088] In an embodiment of the present application, a computer device obtains different predicted matching information through the distance and matching information of neighbor proteins with different target numbers of neighbors in the target evolutionary tree of the candidate protein, and then obtains the matching results of the candidate protein and the second type of protein based on the different predicted matching information; in this scheme, the distance information between the candidate protein and the neighbor protein in the target evolutionary tree provides important information about the evolutionary relationship of proteins. The distance information can be used to more accurately understand the evolutionary relationship between proteins, thereby more accurately understanding the structural and functional similarities between different proteins. In addition, the function of a protein is usually related to the relationship with other similar proteins near its evolutionary tree position. Obtaining neighbor proteins helps to more accurately predict the function of the candidate protein. Based on the target neighbors and the target evolutionary tree, the matching results of the candidate protein can be determined to complete the prediction of the candidate protein without the need for complex parameters, simplifying the prediction and matching process of the candidate protein and the second type of protein, thereby improving the prediction efficiency and accuracy of the candidate protein and the second type of protein.
[0089] Based on the above Figure 2 The scheme shown, Figure 3 A flowchart of a protein matching method provided by an exemplary embodiment of the present application is shown. The method can be executed by a computer device. The above step 230 can be implemented as step 230a:
[0090] Step 230a: Based on the matching information of the neighbor proteins with the target number of neighbors and the distance between the neighbor proteins with the target number of neighbors and the candidate protein in the target evolutionary tree, the predicted matching information of the candidate protein is obtained.
[0091] In some embodiments, the distance can be obtained by adding the lengths of all branches corresponding to the shortest path between two nodes. One of the two nodes represents a neighbor protein with a target number of neighbors, and the other of the two nodes represents a candidate protein. The branch is a line segment connecting the two nodes, representing a segment of a protein's evolutionary path. The length of the branch represents the evolutionary distance between the two nodes. The distance between the two nodes can be obtained by adding the lengths of all branches along the path connecting the two nodes.
[0092] For example, please refer to Figure 4 , which shows a schematic diagram of a target evolutionary tree provided by an exemplary embodiment of the present application.
[0093] In the target evolutionary tree 400, there are nodes A 41, B 42, C 43, D 44 and E 45. Different nodes represent different proteins, and straight edges represent the evolutionary relationship between them. Assuming that node B 42 is a neighbor protein and node D 44 is a candidate protein, the distance of the shortest path from the neighbor protein to the candidate protein is the distance from node B 42 to node D 44. The branch length from node B 42 to node D 44 is obtained by adding the length of each branch on the path from node B 42 to node D 44. Assuming that the length from node B 42 to node C 43 is 1 and the length from node C 43 to node D 44 is 2, the path length between node B 42 and node D 44 is 1+2=3.
[0094] In the embodiments of the present application, the distance information between the candidate protein and the neighboring proteins in the target evolutionary tree provides important information about the evolutionary relationship of the proteins. The distance information can be used to more accurately understand the evolutionary relationship between proteins, thereby more accurately understanding the structural and functional similarities between different proteins. Through the above matching information and the above distance, the predicted matching information of the candidate protein is obtained, which simplifies the matching process between the candidate protein and the second type of protein, thereby improving the matching efficiency and accuracy of the candidate protein and the second type of protein.
[0095] In one possible implementation, a computer device scales the matching information of neighbor proteins based on the distance between the neighbor proteins and the candidate protein in a target evolutionary tree to obtain predicted matching sub-information of the candidate protein corresponding to the neighbor proteins; and obtains predicted matching information of the candidate protein based on the predicted matching sub-information of the neighbor proteins corresponding to the target number of neighbors of the candidate protein.
[0096] The scaling is an operation performed by the computer device to adjust or modify the matching information based on the distance. The computer device may perform a mathematical operation on the matching information based on the distance to obtain another value as the predicted matching sub-information, where the operation includes, but is not limited to, any one or more of addition, subtraction, multiplication, and division.
[0097] In some embodiments, the scaling method may be to perform mathematical operations on a specified scaling factor, or to magnify smaller related information.
[0098] In the embodiment of the present application, scaling the above matching information can improve the reliability of the predicted matching sub-information at different distances, so that the obtained predicted matching sub-information better meets the expected analysis requirements.
[0099] Among them, the above-mentioned predicted matching sub-information is used to indicate the degree of matching between the candidate protein and the second type of protein. The above-mentioned predicted matching sub-information can be in the form of a score, indicating the degree of matching between the candidate protein and the second type of protein, or it can be in the form of a classification label, indicating the type of matching between the candidate protein and the second type of protein, such as high matching, medium matching, and low matching.
[0100] In an embodiment of the present application, the matching information of the neighboring proteins is scaled based on the distance in the above-mentioned target evolutionary tree, which can achieve more precise control over the similarity of the candidate proteins, thereby better reflecting the biological relationship of the proteins, and obtaining more accurate predicted matching sub-information. On the basis of improving the accuracy of the predicted matching sub-information, the predicted matching information of the candidate protein is obtained, which helps to better understand the degree of matching between the neighboring protein and the candidate protein, and improves the accuracy of the prediction.
[0101] In some embodiments, the coefficient for scaling the matching information of the neighbor protein is the inverse of the distance between the neighbor protein and the candidate protein in the target evolutionary tree; or, the coefficient for scaling the matching information of the neighbor protein is determined by the distance interval between the neighbor protein and the candidate protein in the target evolutionary tree.
[0102] The distance interval may be a specified threshold interval. When the distance between the neighbor protein and the candidate protein in the target evolutionary tree is within the specified threshold interval, the coefficient corresponding to the threshold interval may be specified as the scaling coefficient.
[0103] For example, assume that the distance between candidate protein X1 and neighbor protein Y1 in the target evolutionary tree is 2. Then the scaling factor based on the reciprocal distance is 1 / 2 = 0.5, which means that the matching information of neighbor protein Y1 will be scaled by 0.5.
[0104] Please refer to formula 1, and the computer equipment obtains the distance between the neighbor protein and the candidate protein in the target evolutionary tree i , scale the matching information Score of the neighbor protein to obtain the predicted matching sub-information Score of the candidate protein corresponding to the neighbor protein i , where n is the number of neighbor proteins.
[0105]
[0106] For example, assuming that a distance threshold interval [1, 3] is specified, the corresponding scaling factor is 0.5. If the distance between the candidate protein X2 and the neighbor protein Y2 in the target evolutionary tree is within this interval, the scaling factor is 0.5; otherwise, the scaling factor is 1.
[0107] In the embodiments of the present application, when the reciprocal is used as the scaling factor, the smaller the distance, the larger the scaling factor, thereby enhancing the matching information of neighboring proteins, which can reflect that proteins that are closer in the target evolutionary tree may have higher similarity, emphasizing the relationship between adjacent proteins in the evolutionary tree, and using the distance interval in which the distance is located to determine the scaling factor can more flexibly consider the changes in protein similarity at different evolutionary distances. The above two scaling methods help to better utilize the evolutionary information of proteins and improve the accuracy of predicted matching information for candidate proteins.
[0108] Based on the above Figure 2 The scheme shown, Figure 5 FIG1 shows a flow chart of a protein matching method provided by an exemplary embodiment of the present application. The method can be executed by a computer device. Figure 2 In the illustrated embodiment, step 204 and step 208 may be included before step 210:
[0109] Step 204: Obtain multiple evolutionary trees; each of the multiple evolutionary trees corresponds to a specified amino acid sequence, or corresponds to a combination of multiple specified amino acid sequences.
[0110] In the embodiment of the present application, multiple evolutionary trees may be constructed in advance based on a certain specified amino acid sequence of the first type of protein sample, or a combination of multiple specified amino acid sequences.
[0111] In one possible implementation, the above-mentioned evolutionary tree can be constructed using the Maximum Parsimony method, which selects the tree that minimizes the number of events (e.g., changes in bases or amino acids in the sequence) that occur throughout the evolutionary process as the optimal tree. In other words, the Maximum Parsimony method assumes that the most likely evolutionary path is the simplest path, that is, the path that requires the fewest changes. This method is also called the Minimum Evolution Method, and the evolutionary tree obtained using the ProtPars tool.
[0112] Step 208: Taking each first type protein with matching information in the plurality of evolutionary trees as a sample, a combination of a target evolutionary tree and a target number of neighbors is obtained.
[0113] In an embodiment of the present application, a computer device uses each first type of protein with matching information as a sample, determines the accuracy of multiple evolutionary trees in predicting the matching information of the first type of protein under various numbers of neighbors, and then determines a combination of a target evolutionary tree and a target number of neighbors according to the accuracy of multiple evolutionary trees in predicting the matching information of the first type of protein under various numbers of neighbors; for example, one or more combinations of evolutionary trees and the number of neighbors that have a higher accuracy in predicting the matching information of the first type of protein are determined as the combination of the above-mentioned target evolutionary tree and the target number of neighbors.
[0114] In an embodiment of the present application, a computer device uses each first type of protein with matching information as a sample to train multiple evolutionary trees, and can screen out a suitable combination of target evolutionary trees and target number of neighbors. Since different evolutionary trees correspond to different evolutionary information, using multiple evolutionary trees for training can introduce more evolutionary information, thereby enriching the training data, ensuring the training results, and improving the prediction accuracy when predicting candidate proteins.
[0115] Based on the above Figure 5 The scheme shown, Figure 6 FIG1 shows a flow chart of a protein matching method provided by an exemplary embodiment of the present application. The method can be executed by a computer device, that is, Figure 5 Step 208 in the embodiment can also be implemented as step 208a, step 208b, step 208c, step 208d and step 208e:
[0116] Step 208a: Divide each first type protein with matching information into a training sample set and a test sample set.
[0117] In an embodiment of the present application, each first type protein with matching information can be divided into a training sample set and a test sample set according to a preset ratio; optionally, the same ratio can be randomly divided multiple times to obtain multiple pairs of training sample sets and test sample sets, and each division method corresponds to a pair of training sample sets and test sample sets.
[0118] Exemplarily, the train_test_split of the python package sklearn is used to divide each first type protein with matching information. The division is performed based on the same setting of the ratio of positive and negative samples in the training and test sample sets. The training sample set has the following four ratios (0.5, 0.6, 0.7, 0.8) (trainratio=4), and the rest is the test sample set. Each ratio is randomly divided 10 times (rand=10). The data in the test sample set are predicted according to the average value of the nearest neighbors in the tree in the training sample set. The number of neighbors is set to 7 (1, 2, 3, 4, 5, 6, 7), and neignum=7.
[0119] Step 208b: For the candidate protein sample, query a first number of neighbor protein samples from the first evolutionary tree based on the distance to the candidate protein sample; the candidate protein sample is a protein in the test sample set, and the neighbor protein sample is a protein in the training sample set; the first evolutionary tree is one of the multiple evolutionary trees.
[0120] In an embodiment of the present application, the matching information of the first type of protein samples in the training sample set can be used as a known condition. For each first type of protein sample in the test sample set (i.e., the above-mentioned candidate protein sample), the neighbor protein samples (i.e., the proteins in the training sample set) with various numbers of neighbors closest to the candidate protein sample in each evolutionary tree are searched.
[0121] In some embodiments, the computer device may first search the evolutionary tree, obtain the distance between each protein sample in the training sample set and the candidate protein sample in the evolutionary tree, and then arrange the distances in ascending order, and obtain protein samples with various numbers of neighbors as neighbor protein samples.
[0122] For example, assume there is an evolutionary tree Z, a candidate protein sample X, and the training sample set includes protein sample A, protein sample B, protein sample C, protein sample D, and protein sample E. After the computer device searches the evolutionary tree Z, the distances between the candidate protein sample X and protein samples A to E are ranked as 2.5, 3.0, 4.2, 5.1, and 6.0, respectively. When the number of neighbors of candidate protein sample X is 1, its nearest neighbor protein sample is protein sample A, and the distance is 2.5; when the number of neighbors of candidate protein sample X is 2, its two nearest neighbor protein samples are protein sample A and protein sample B, and the distances are 2.5 and 3.0 respectively; when the number of neighbors of candidate protein sample X is 3, its three nearest neighbor protein samples are protein sample A, protein sample B and protein sample C, and the distances are 2.5, 3.0 and 4.2 respectively; when the number of neighbors of candidate protein sample X is 4, its four nearest neighbor protein samples are protein sample A, protein sample B, protein sample C and protein sample D, and the distances are 2.5, 3.0, 4.2 and 5.1 respectively; when the number of neighbors of candidate protein sample X is 5, its five nearest neighbor protein samples are protein sample A, protein sample B, protein sample C, protein sample D and protein sample E, and the distances are 2.5, 3.0, 4.2, 5.1 and 6.0 respectively.
[0123] Step 208c: Based on the matching information of the neighbor protein samples of the first number of neighbors, obtain the predicted matching information of the candidate protein sample.
[0124] In an embodiment of the present application, for each first type protein sample in the test sample set and neighbor protein samples with various numbers of neighbors in each evolutionary tree, predicted matching information is calculated for each first type protein sample in the test sample set corresponding to each combination of evolutionary tree and number of neighbors.
[0125] The computer device may input a specified number of neighbors, the matching information of the neighboring protein samples corresponding to the specified number, and the distance of the neighboring protein samples relative to a single protein sample of the first type into a preset model or formula to obtain a single piece of predicted matching information corresponding to the first type of protein sample. The computer device may input different specified numbers of neighbors, the matching information of the neighboring protein samples corresponding to the specified number, and the distance into the preset model or formula multiple times to obtain multiple pieces of predicted matching information corresponding to the first type of protein sample.
[0126] Step 208d: Based on the predicted matching information of the candidate protein sample and the matching information of the candidate protein sample, obtain a correlation coefficient corresponding to the combination of the first evolutionary tree and the first number of neighbors.
[0127] The correlation coefficient is the correlation degree between the predicted matching information and the actual matching information of the candidate protein sample under the combination of the first evolutionary tree and the first number of neighbors.
[0128] The computer device can input the predicted matching information of the candidate protein sample corresponding to the specified evolutionary tree and the specified number of neighbors, and the matching information of the candidate protein sample into a preset model or formula respectively, and obtain a similarity through calculation by the preset model or formula. The above similarity is used to indicate the degree of correlation between the predicted matching information corresponding to the current candidate protein under the specified evolutionary tree and the specified number of neighbors, and the actual matching information corresponding to the current candidate protein under the specified evolutionary tree and the specified number of neighbors. The above actual matching information refers to the actual matching result or label of the current candidate protein sample under the specified evolutionary tree and the specified number of neighbors.
[0129] In an embodiment of the present application, since the above steps calculate the predicted matching information for each first type protein sample in the test sample set corresponding to each combination of evolutionary tree and number of neighbors, and each first type protein sample in the test sample set has actual matching information, the correlation coefficient corresponding to each combination of evolutionary tree and number of neighbors can be obtained by the difference between the predicted matching information for each first type protein sample in the test sample set corresponding to each combination of evolutionary tree and number of neighbors and the actual matching information of each first type protein sample in the test sample set; for example, the correlation coefficient corresponding to each combination of evolutionary tree and number of neighbors can be represented by the degree of correlation between the predicted matching information for each first type protein sample in the test sample set under the combination of the evolutionary tree and number of neighbors and the actual matching information of each first type protein sample in the test sample set.
[0130] Step 208e: Based on the correlation coefficients corresponding to various combinations of multiple evolutionary trees and multiple numbers of neighbors, obtain a combination of a target evolutionary tree and a target number of neighbors.
[0131] The higher the correlation coefficient, the higher the correlation between the predicted match information and the actual match information for the candidate protein sample under the combination of the first evolutionary tree and the first number of neighbors. In other words, the higher the correlation coefficient, the higher the accuracy of subsequent prediction of match information between the candidate protein and the second type of protein using the combination of the first evolutionary tree and the first number of neighbors. The computer device obtains the correlation coefficients corresponding to various combinations of multiple evolutionary trees and multiple numbers of neighbors, compares the multiple correlation coefficients, and selects the combination of the evolutionary tree and the number of neighbors with the largest or greater correlation coefficient as the target evolutionary tree and target number of neighbors.
[0132] In an embodiment of the present application, by training the first type of protein samples with matching information, the combination of the first evolutionary tree and the first number of neighbors with the highest correlation between the predicted matching information and the actual matching information is selected for subsequent prediction of candidate proteins. The combination with a high degree of correlation can better capture the evolutionary relationship and functional similarity between proteins, thereby ensuring the reliability of the prediction results and effectively improving the prediction accuracy of candidate proteins.
[0133] In one possible implementation, the predicted matching information of each protein in the test sample set, the average value of the predicted matching information of each protein in the test sample set, the matching information of each protein in the test sample set, and the average value of the matching information of each protein in the test sample set are obtained; based on the predicted matching information of each protein in the test sample set, the average value of the predicted matching information of each protein in the test sample set, the matching information of each protein in the test sample set, and the average value of the matching information of each protein in the test sample set, the Pearson correlation coefficient is calculated to obtain the correlation coefficient corresponding to the combination of the first evolutionary tree and the first number of neighbors.
[0134] For example, please refer to Formula 2, which shows a calculation method for a Pearson correlation coefficient, where x j is the predicted value of the jth protein sample with predicted matching information in the test sample set, is the average of the predicted values of the m protein samples with predicted matching information in the test sample set, y j is the true value of m proteins with matching information in the test sample set, is the average of the true values of the m proteins with matching information in the test sample set, r xy is the correlation coefficient corresponding to the combination of the first evolutionary tree and the first number of neighbors.
[0135]
[0136] In the embodiments of the present application, by calculating the Pearson correlation coefficient, the degree of correlation between the combinations corresponding to different evolutionary trees and numbers of neighbors and the prediction results can be understood, which helps to screen out combinations with higher correlation for protein matching information prediction, and use them in subsequent practical applications of protein matching prediction to improve the accuracy of the prediction results.
[0137] In one possible implementation, based on the predicted matching information of the candidate protein sample, the matching information of the candidate protein sample, the average of the predicted matching information of each protein in the test sample set, and the average of the matching information of each protein in the test sample set, a correlation coefficient of the combination of the first evolutionary tree and the first number of neighbors relative to the candidate protein sample is obtained; the average or median of the correlation coefficients of the combination of the first evolutionary tree and the first number of neighbors relative to each protein in the test sample set is taken to obtain the correlation coefficient corresponding to the combination of the first evolutionary tree and the first number of neighbors.
[0138] The above-mentioned correlator coefficient is used to indicate the degree of correlation between the combination of the first evolutionary tree and the first number of neighbors and the single protein sample. The computer device obtains the correlator coefficients of the combination of the first evolutionary tree and the first number of neighbors relative to multiple candidate protein samples, and averages the multiple correlator coefficients or selects the median of the multiple correlator coefficients as the correlation coefficient corresponding to the combination of the first evolutionary tree and the first number of neighbors.
[0139] It should be noted that, when the data distribution of the multiple correlation sub-coefficients is relatively symmetrical and there are no obvious outliers or extreme values, the computer device obtains the average value of the multiple correlation sub-coefficients as the above-mentioned correlation coefficient; when there are extreme values or outliers in the multiple correlation sub-coefficients, the computer device obtains the median of the multiple correlation sub-coefficients as the above-mentioned correlation coefficient.
[0140] Exemplarily, the correlation sub-coefficient of the combination of the first evolutionary tree and the first number of neighbors relative to candidate protein sample A is 0.3, the correlation sub-coefficient relative to candidate protein sample B is 0.2, and the correlation sub-coefficient relative to candidate protein sample C is 0.33. The computer device obtains three correlation sub-coefficients, takes the average of the three correlation sub-coefficients, and obtains 0.44, which is obtained as the correlation coefficient corresponding to the combination of the first evolutionary tree and the first number of neighbors.
[0141] Exemplarily, the correlation coefficients of the combination of the first evolutionary tree and the first number of neighbors relative to candidate protein sample B to candidate protein sample E are (0.2, 0.8, 0.3, 0.45) respectively, and the computer device obtains the median as (0.3+0.45) / 2=0.375, and obtains 0.375 as the correlation coefficient corresponding to the combination of the first evolutionary tree and the first number of neighbors.
[0142] In an embodiment of the present application, the correlation sub-coefficient of the combination of the first evolutionary tree and the first number of neighbors relative to the candidate protein sample is first obtained, and then the average or median of multiple correlation sub-coefficients is taken to understand the degree of correlation between the combinations corresponding to different evolutionary trees and the number of neighbors and the prediction results. This can reduce the impact of single samples or outliers on the coefficient results, improve the stability of the prediction, and ensure the reliability of the prediction results.
[0143] In one possible implementation, various combinations of multiple evolutionary trees and multiple numbers of neighbors are arranged in descending order according to the corresponding correlation coefficients; one or more combinations arranged at the top of the various combinations of multiple evolutionary trees and multiple numbers of neighbors are determined as the combination of the target evolutionary tree and the target number of neighbors; or, among the various combinations of multiple evolutionary trees and multiple numbers of neighbors, the combination whose corresponding correlation coefficient is higher than the correlation coefficient threshold is determined as the combination of the target evolutionary tree and the target number of neighbors.
[0144] Among them, the above-mentioned single evolutionary tree and a single combination of the specified number of neighbors have a corresponding correlation coefficient. The computer device obtains multiple correlation coefficients corresponding to different combinations of single evolutionary trees and the specified number of neighbors, sorts the multiple correlation coefficients from large to small, and selects the combination of one or more evolutionary trees and the number of neighbors corresponding to the one or more correlation coefficients with the larger correlation coefficient as the combination of the target evolutionary tree and the target number of neighbors; or presets a specified threshold (that is, the above-mentioned correlation coefficient threshold) to compare with multiple correlation coefficients, and selects the combination of one or more evolutionary trees and the number of neighbors corresponding to the correlation coefficient with a coefficient value higher than the specified threshold as the combination of the target evolutionary tree and the target number of neighbors.
[0145] Exemplarily, the computer device obtains correlation coefficient Z1 0.2, correlation coefficient Z2 0.3, and correlation coefficient Z3 0.5. The above three correlation coefficients correspond to the combination of evolutionary tree T1 and 3 neighbors, the combination of evolutionary tree T2 and the number of 8 neighbors, and the combination of evolutionary tree T3 and 6 neighbors, respectively. At this time, the largest correlation coefficient of 0.5 is selected, which corresponds to the combination of evolutionary tree T3 and 6 neighbors. The target evolutionary tree is T3, and the target number of neighbors is 6. In another implementation method, the computer device presets a correlation coefficient threshold of 0.3. At this time, the correlation coefficient Z3 is higher than the correlation coefficient threshold, and the combination of the evolutionary tree and the number of neighbors corresponding to the correlation coefficient Z3 is determined as the combination of the target evolutionary tree and the target number of neighbors.
[0146] In an embodiment of the present application, by arranging the correlation coefficients in order, or combining the correlation coefficient threshold, determining one or more combinations with high correlation coefficients as the combination of the target evolutionary tree and the target number of neighbors, the evolutionary tree and neighbor number combination that is strongly correlated with the prediction result can be identified, thereby improving the screening efficiency of the target evolutionary tree and the target number of neighbors combination.
[0147] In one possible implementation, among various combinations of multiple evolutionary trees and multiple numbers of neighbors, each combination corresponds to multiple correlation coefficients; the multiple correlation coefficients correspond to different sample division methods, and the sample division method is a method of dividing each first type protein with matching information into a training sample set and a test sample set; for each combination of various combinations of multiple evolutionary trees and multiple numbers of neighbors, the median of the multiple correlation coefficients of the combination is taken to obtain the median correlation coefficient of the combination; the various combinations of multiple evolutionary trees and multiple numbers of neighbors are arranged in ascending order according to the difference between the corresponding median correlation coefficient and the target correlation coefficient; among the various combinations of multiple evolutionary trees and multiple numbers of neighbors, one or more combinations arranged at the top are determined as the combination of the target evolutionary tree and the target number of neighbors.
[0148] In some embodiments, the computer device can determine the combination of the target evolutionary tree and the target number of neighbors as the combination of the target evolutionary tree and the target number of neighbors, in which the difference between the corresponding median correlation coefficient and the target correlation coefficient is less than a difference threshold.
[0149] Exemplarily, a difference threshold preset by the computer device is 0.05, and 4 correlation coefficients (0.9, 0.83, 0.76, 0.73) are obtained. The median correlation coefficient is 0.7955. The differences between the 4 correlation coefficients and the median correlation coefficient are (0.1, 0.0345, 0.0355, 0.0655) respectively. The difference of 0.0345 and the difference of 0.0355 are both less than the difference threshold. The combination of the evolutionary tree and the number of neighbors corresponding to the correlation coefficient of 0.83 and the correlation coefficient of 0.76 is determined as the combination of the target evolutionary tree and the target number of neighbors.
[0150] In some embodiments, when there are multiple combinations in which the difference between the corresponding median correlation coefficient and the target correlation coefficient is less than the difference threshold, the differences less than the difference threshold can be arranged from small to large, and the combinations corresponding to the correlation coefficients corresponding to one or more of the differences arranged in front are selected and determined as the combination of the target evolutionary tree and the target number of neighbors.
[0151] In an embodiment of the present application, the median of multiple correlation coefficients is used, and the difference between the median correlation coefficient and the target correlation coefficient is combined for sorting and selection, which can improve the accuracy, stability and reliability of determining the combination of the target evolutionary tree and the target number of neighbors.
[0152] In one possible implementation, the designated amino acid sequence includes one or more of the following amino acid sequences: a heavy chain amino acid sequence, a light chain amino acid sequence, an amino acid sequence corresponding to the heavy chain V gene, an amino acid sequence corresponding to the light chain V gene, an amino acid sequence corresponding to the heavy chain J gene, and an amino acid sequence corresponding to the light chain J gene.
[0153] The heavy chain has a longer amino acid sequence; the light chain has a shorter amino acid sequence than the heavy chain.
[0154] In the embodiments of the present application, the diversity of amino acid sequences is expanded, the diversity of sample data is guaranteed, and through training with multiple types of samples, the accuracy, stability and reliability of determining the combination of the target evolutionary tree and the number of target neighbors can be improved, thereby improving the matching effect of candidate proteins.
[0155] The solutions shown in the above embodiments of this application can be applied in the development of antibody drugs to screen antibodies that can bind to specific locations on antigens. Figure 7 , which shows Figure 7 This is a framework diagram of an antibody detection scheme involved in an exemplary embodiment of the present application. Figure 7 As shown in the figure, the detection scheme includes three stages: sample collection, model training, and model application.
[0156] S1, sample collection, obtains combined samples of multiple evolutionary trees and multiple numbers of neighbors.
[0157] First, all antibody sequences can be aligned using ANARCI. ANARCI is a program for numbering antibody and T-cell receptor sequences. HMMER is used to search for the immunoglobulin superfamily (IgSF) domain in the input sequence, and then the found domains are numbered according to the user-specified numbering scheme (such as IMGT, AHo, Martin, etc.). Taking the IMGT scheme as an example, when ANARCI uses the "imgt" rule for antibody sequence alignment, HMMER is first used to search for the IgSF domain in the input sequence. The found domains are then numbered according to the IMGT rules, including determining the start and end positions of each domain, as well as the positions of the complementarity determining regions (CDRs) and framework regions (FRs) of each domain.
[0158] The purpose of numbering is to clearly divide and identify different regions of the antibody, and different regions play a key role in the structure and function of the antibody. When using the "imgt" rule, ANARCI will number the antibody sequence according to the rules of IMGT (International Immunogenetic Database). In the IMGT numbering system, each amino acid is assigned a unique position number in its corresponding V, D or J gene segment. For the V gene segment, the numbering starts from 1 and goes to the end of the gene segment. The numbering system pays special attention to the position of CDR and FR. For example, the CDR1, CDR2 and CDR3 of the V gene segments of the light chain and heavy chain are numbered as: CDR1: 24-34, CDR2: 50-65, respectively.
[0159] Then, based on the different regions of the antibody obtained by ANARCI, the alignment sequences of different structural regions of the antibody can be obtained. The amino acid sequences corresponding to the V / J genes in the heavy and light chains can be clipped using the VDJ amino acid sequences in IMGT. The specific steps of clipping are as follows:
[0160] a) Determine the V / J gene regions of the heavy and light chains
[0161] The starting and ending positions of the V / J gene regions of the antibody heavy and light chains are determined by looking at the numbering given by ANARCI. For example, the V gene region usually starts at position 1, while the end position of the J gene region can be inferred from the end position of CDR3.
[0162] b) Trimming of the V / J gene region
[0163] This can be achieved by selecting the portion of the sequence that corresponds to the starting and ending positions. For example, a V gene region starts at position 1 and ends at position 105 (which is where CDR3 begins), and this portion of the sequence can be selected as the V gene region.
[0164] c) Comparison of V / J gene regions
[0165] The cropped V / J gene region can be compared with the VDJ amino acid sequence in the IMGT database to identify its possible source gene. This step can be achieved using an alignment tool.
[0166] After trimming, 12 different amino acid sequences and their different combinations were obtained, and these sequence combinations were used to construct evolutionary trees:
[0167] 1) Heavy chain / light chain sequence, 2 evolutionary trees
[0168] 2) Heavy chain + light chain sequence, 1 evolutionary tree
[0169] 3) Amino acid sequences corresponding to heavy chain / light chain V genes, 2 evolutionary trees
[0170] 4) Amino acid sequence corresponding to heavy chain V gene + light chain V gene, 1 evolutionary tree
[0171] 5) Amino acid sequences corresponding to heavy chain / light chain J genes, 2 evolutionary trees
[0172] 6) Amino acid sequences corresponding to heavy chain / light chain VJ genes, 2 evolutionary trees
[0173] 7) Amino acid sequences corresponding to heavy chain V gene + light chain VJ gene, 2 evolutionary trees
[0174] The tree structure is as follows Figure 8 As shown, Figure 8 Shown is an evolutionary tree structure diagram involved in an exemplary embodiment of the present application.
[0175] S2, model training, obtains the best combination of evolutionary tree and number of neighbors.
[0176] After constructing the phylogenetic tree, the candidate antibody samples can be trained using the nearest neighbor method using antibody samples with known labels. The prediction model (Protree) 701 includes Formula 1 and Formula 2. This step can be applied to steps 208c to 208e of the above embodiment. In other words, the above training process corresponds to the implementation process of steps 208c to 208e of the above embodiment.
[0177] Use formula 1 to select the number of neighbors. Assume n = 10 and find the 10 labeled antibody samples closest to the candidate antibody sample. The fraction of labeled antibody samples Socre i It can be a classification label (active score = 1, no activity = 0, weak activity = 0.5) or a continuous label (affinity value), and then calculate the phylogenetic distance between the antibody sample and the candidate antibody sample. i , the score of the candidate antibody sample can be obtained according to the formula, and finally the correlation between the score and the true label can be calculated by formula 2.
[0178]
[0179] The correlation coefficient 702 corresponding to the combination of evolutionary tree samples and the number of neighbors is obtained by formula 2, where x j is the predicted value of the jth labeled antibody sample in the test sample set, is the average of the predicted values of the m labeled antibody samples in the test sample set, y j is the true value of the m labeled antibody samples in the test sample set, is the average of the true values of the m labeled antibody samples in the test sample set, rxy is the correlation coefficient 702 corresponding to the combination of evolutionary tree samples and the number of neighbors.
[0180]
[0181] In this way, 12 evolutionary trees can be obtained after training. The box plot of the correlation coefficient obtained by selecting different numbers of neighbors is as follows: Figure 9 As shown, Figure 9 A box plot of 12 evolutionary trees with 1-7 neighbors, according to an exemplary embodiment of the present application, is shown. The vertical axis represents the correlation coefficient, and the horizontal axis represents the number of neighbors. By comparing the median of each box, the optimal combination 703 is selected, i.e., the most effective combination of evolutionary trees and number of neighbors.
[0182] S3, model application, obtains the correlation coefficient through the prediction model and,optimal combination.
[0183] In the embodiment of the present application, the optimal combination 703 in the training process can be selected to score the candidate antibodies, and the different results can be integrated, such as Figure 8 The hl_vj tree (heavy and light chain VJ gene corresponding amino acid sequence) and the heavy_light tree perform best with 4 and 5 neighbors, respectively. The scores of these two numbers under these two types of neighbors are integrated. The integration method is as shown in Formula 3, averaging the predictions of multiple models. Assuming there are n models, and the predictions of each model for the same data point are p1, p2, ..., pn, then the average prediction of these models is:
[0184]
[0185] In the examples of the present application, more specifically binding antibodies are discovered in a faster and more economical manner, and the affinity of antibodies to antigens can also be predicted, which greatly improves the efficiency of functional antibody screening.
[0186] Figure 10 A block diagram of a protein matching device according to an exemplary embodiment of the present application is shown. The device can be used to perform the following steps: Figures 2 to 5 All or part of the steps in the method shown; Figure 10 As shown, the device includes:
[0187] A target evolutionary tree acquisition module 1001 is configured to acquire a target evolutionary tree, wherein the target evolutionary tree is an evolutionary tree constructed based on specified amino acid sequences in a plurality of first-type proteins; the leaf nodes of the evolutionary tree are first-type proteins; and some of the first-type proteins in the evolutionary tree have matching information, the matching information being used to indicate the degree of matching between the first-type proteins and the second-type proteins.
[0188] A query module 1002 is configured to query neighbor proteins having a target number of neighbors from a target evolutionary tree for a candidate protein based on a distance from the candidate protein; the candidate protein is a first type protein other than the first type protein having matching information; and the neighbor protein is a first type protein having matching information;
[0189] A predicted matching information acquisition module 1003 is used to acquire predicted matching information of a candidate protein based on matching information of neighbor proteins of a target number of neighbors;
[0190] The matching result acquisition module 1004 is used to acquire the matching result of the candidate protein based on the predicted matching information of the candidate protein.
[0191] In a possible implementation, the predicted matching information acquisition module 1003 is used to acquire the predicted matching information of the candidate protein based on the matching information of the neighbor proteins with the target number of neighbors and the distance between the neighbor proteins with the target number of neighbors and the candidate protein in the target evolutionary tree.
[0192] In one possible implementation, the predicted matching information acquisition module 1003 is used to scale the matching information of the neighbor protein based on the distance between the neighbor protein and the candidate protein in the target evolutionary tree, and obtain the predicted matching sub-information of the candidate protein corresponding to the neighbor protein; based on the predicted matching sub-information of the neighbor protein corresponding to the target number of neighbors of the candidate protein, obtain the predicted matching information of the candidate protein.
[0193] In one possible implementation, the coefficient for scaling the matching information of the neighbor protein is the inverse of the distance between the neighbor protein and the candidate protein in the target evolutionary tree; or, the coefficient for scaling the matching information of the neighbor protein is determined by the distance interval between the neighbor protein and the candidate protein in the target evolutionary tree.
[0194] In a possible implementation, the apparatus further includes:
[0195] The phylogenetic tree acquisition module is used to obtain multiple phylogenetic trees before the target phylogenetic tree acquisition module 1001 obtains the target phylogenetic tree; each phylogenetic tree in the multiple phylogenetic trees corresponds to a specified amino acid sequence, or corresponds to a combination of multiple specified amino acid sequences; the combination acquisition module is used to obtain a combination of the target phylogenetic tree and the number of target neighbors using each first type protein with matching information in the multiple phylogenetic trees as a sample.
[0196] In one possible implementation, a combination acquisition module is used to divide each first type protein with matching information into a training sample set and a test sample set; for a candidate protein sample, based on the distance to the candidate protein sample, query a first number of neighbor protein samples from a first evolutionary tree; the candidate protein sample is a protein in the test sample set, and the neighbor protein sample is a protein in the training sample set; the first evolutionary tree is one of multiple evolutionary trees; based on the matching information of the first number of neighbor protein samples, the predicted matching information of the candidate protein sample is obtained; based on the predicted matching information of the candidate protein sample and the matching information of the candidate protein sample, the correlation coefficient corresponding to the combination of the first evolutionary tree and the first number of neighbors is obtained; based on the correlation coefficients corresponding to various combinations of multiple evolutionary trees and multiple numbers of neighbors, the combination of a target evolutionary tree and a target number of neighbors is obtained.
[0197] In one possible implementation, a combination acquisition module is used to obtain predicted matching information of each protein in the test sample set, the average value of the predicted matching information of each protein in the test sample set, the matching information of each protein in the test sample set, and the average value of the matching information of each protein in the test sample set; based on the predicted matching information of each protein in the test sample set, the average value of the predicted matching information of each protein in the test sample set, the matching information of each protein in the test sample set, and the average value of the matching information of each protein in the test sample set, calculate the Pearson correlation coefficient to obtain the correlation coefficient corresponding to the combination of the first evolutionary tree and the first number of neighbors.
[0198] In one possible implementation, the combination acquisition module is used to obtain the correlation coefficient of the combination of the first evolutionary tree and the first number of neighbors relative to the candidate protein sample based on the predicted matching information of the candidate protein sample, the matching information of the candidate protein sample, the average of the predicted matching information of each protein in the test sample set, and the average of the matching information of each protein in the test sample set; and take the average or median of the correlation coefficients of the combination of the first evolutionary tree and the first number of neighbors relative to each protein in the test sample set to obtain the correlation coefficient corresponding to the combination of the first evolutionary tree and the first number of neighbors.
[0199] In one possible implementation, the combination acquisition module is used to arrange various combinations of multiple evolutionary trees and multiple numbers of neighbors in order from high to low according to the corresponding correlation coefficients; determine one or more combinations arranged in front of the various combinations of multiple evolutionary trees and multiple numbers of neighbors as the combination of the target evolutionary tree and the target number of neighbors; or, determine the combination of the multiple evolutionary trees and multiple numbers of neighbors whose corresponding correlation coefficient is higher than the correlation coefficient threshold as the combination of the target evolutionary tree and the target number of neighbors.
[0200] In one possible implementation, each combination of multiple evolutionary trees and multiple neighbor numbers corresponds to multiple correlation coefficients; each of the multiple correlation coefficients corresponds to a different sample partitioning method, where the sample partitioning method is a method of partitioning each first type protein having matching information into a training sample set and a test sample set;
[0201] The combination acquisition module is used to take the median of the multiple correlation coefficients of each combination of multiple evolutionary trees and multiple neighbor numbers to obtain the median correlation coefficient of the combination; arrange the various combinations of multiple evolutionary trees and multiple neighbor numbers in ascending order according to the difference between the corresponding median correlation coefficient and the target correlation coefficient; and determine one or more combinations that are arranged in the front among the various combinations of multiple evolutionary trees and multiple neighbor numbers as the combination of the target evolutionary tree and the target neighbor number.
[0202] In one possible implementation, the designated amino acid sequence includes one or more of the following amino acid sequences:
[0203] Heavy chain amino acid sequence, light chain amino acid sequence, amino acid sequence corresponding to the heavy chain V gene, amino acid sequence corresponding to the light chain V gene, amino acid sequence corresponding to the heavy chain J gene, and amino acid sequence corresponding to the light chain J gene.
[0204] Figure 11 The block diagram of a computer device 1100 shown in an exemplary embodiment of the present application is shown. The computer device can be implemented as a server or terminal in the above-mentioned solution of the present application. In this embodiment, the computer device is described as a server. The computer device 1100 includes a central processing unit (CPU) 1101, a system memory 1104 including a random access memory (RAM) 1102 and a read-only memory (ROM) 1103, and a system bus 1105 connecting the system memory 1104 and the central processing unit 1101. The computer device 1100 also includes a mass storage device 1106 for storing an operating system 1109, application programs 1110, and other program modules 1111.
[0205] The mass storage device 1106 is connected to the central processing unit 1101 via a mass storage controller (not shown) connected to the system bus 1105. The mass storage device 1106 and its associated computer-readable media provide non-volatile storage for the computer device 1100. In other words, the mass storage device 1106 may include a computer-readable medium (not shown) such as a hard disk or a compact disc read-only memory (CD-ROM) drive.
[0206] Without loss of generality, the computer-readable medium may include computer storage media and communication media. Computer storage media include volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules or other data. Computer storage media include RAM, ROM, Erasable Programmable Read Only Memory (EPROM), Electronically Erasable Programmable Read-Only Memory (EEPROM), flash memory or other solid-state storage technology, CD-ROM, Digital Versatile Disc (DVD) or other optical storage, tape cassettes, magnetic tape, disk storage or other magnetic storage devices. Of course, those skilled in the art will appreciate that the computer storage medium is not limited to the above-mentioned ones. The above-mentioned system memory 1104 and mass storage device 1106 can be collectively referred to as memory.
[0207] According to various embodiments of the present disclosure, the computer device 1100 can also be connected to a remote computer on a network such as the Internet for operation. That is, the computer device 1100 can be connected to the network via the network interface unit 1107 connected to the system bus 1105, or the network interface unit 1107 can be used to connect to other types of networks or remote computer systems (not shown).
[0208] The memory further includes at least one computer program, which is stored in the memory. The central processing unit 1101 implements all or part of the steps in the methods shown in the above embodiments by executing the at least one computer program.
[0209] In an exemplary embodiment, a computer-readable storage medium is further provided, storing at least one computer program. The at least one computer program is loaded and executed by a processor to implement all or part of the steps of the methods described in the above embodiments. For example, the computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc (CD-ROM), a magnetic tape, a floppy disk, or an optical data storage device.
[0210] In an exemplary embodiment, a computer program product is also provided, comprising a computer program stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium and executes the computer program, causing the computer device to perform all or part of the steps of the methods described in the various embodiments above.
[0211] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, and the true scope and spirit of the present application are indicated by the following claims.
[0212] It will be understood that the present application is not limited to the exact construction that has been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof.
Claims
1. A protein matching method, characterized in that: The method comprises: Obtaining a target evolutionary tree, the target evolutionary tree being an evolutionary tree constructed based on specified amino acid sequences in a plurality of first-type proteins; leaf nodes of the evolutionary tree being the first-type proteins; and some of the first-type proteins in the evolutionary tree having matching information, the matching information being used to indicate a degree of matching between the first-type proteins and second-type proteins; For a candidate protein, querying neighbor proteins of a target number of neighbors from the target evolutionary tree based on the distance to the candidate protein; the candidate protein is other first-type proteins other than the first-type proteins having the matching information; the neighbor protein is the first-type protein having the matching information; obtaining predicted matching information of the candidate protein based on the matching information of the neighbor proteins of a target number of neighbors; Based on the predicted matching information of the candidate protein, a matching result of the candidate protein is obtained.
2. The method according to claim 1, characterized in that The step of obtaining predicted matching information of the candidate protein based on the matching information of the neighbor proteins of the target neighbor quantity includes: Based on the matching information of the neighbor proteins with a target number of neighbors and the distance between the neighbor proteins with a target number of neighbors and the candidate protein in the target evolutionary tree, the predicted matching information of the candidate protein is obtained.
3. The method according to claim 2, characterized in that The step of obtaining predicted matching information of the candidate protein based on the matching information of the neighbor protein of the target number of neighbors and the distance between the neighbor protein of the target number of neighbors and the candidate protein in the target evolutionary tree comprises: Scaling the matching information of the neighbor protein based on the distance between the neighbor protein and the candidate protein in the target evolutionary tree to obtain predicted matching sub-information of the candidate protein corresponding to the neighbor protein; Based on the predicted matching sub-information of the neighbor proteins corresponding to the target number of neighbors of the candidate protein, the predicted matching information of the candidate protein is obtained.
4. The method according to claim 3, characterized in that The coefficient for scaling the matching information of the neighbor protein is the inverse of the distance between the neighbor protein and the candidate protein in the target evolutionary tree; or, The coefficient for scaling the matching information of the neighbor protein is determined by the distance interval between the neighbor protein and the candidate protein in the target evolutionary tree.
5. The method according to any one of claims 1 to 4, characterized in that: Before obtaining the target evolutionary tree, the method further includes: Obtaining a plurality of the phylogenetic trees; each of the plurality of phylogenetic trees corresponds to one of the specified amino acid sequences, or corresponds to a combination of multiple of the specified amino acid sequences; Taking each of the first type proteins having the matching information in a plurality of the evolutionary trees as a sample, a combination of the target evolutionary tree and the target number of neighbors is obtained.
6. The method according to claim 5, characterized in that The step of taking each of the first type proteins having the matching information in the plurality of evolutionary trees as a sample and obtaining a combination of the target evolutionary tree and the target number of neighbors includes: dividing each of the first-type proteins having the matching information into a training sample set and a test sample set; For a candidate protein sample, querying a first number of neighbor protein samples from a first evolutionary tree based on a distance from the candidate protein sample; the candidate protein sample is a protein in the test sample set, and the neighbor protein sample is a protein in the training sample set; the first evolutionary tree is one of the multiple evolutionary trees; acquiring the predicted matching information of the candidate protein sample based on the matching information of the neighbor protein samples of the first number of neighbors; Based on the predicted matching information of the candidate protein sample and the matching information of the candidate protein sample, obtaining a correlation coefficient corresponding to a combination of the first evolutionary tree and the first number of neighbors; Based on the correlation coefficients corresponding to various combinations of the plurality of evolutionary trees and the plurality of neighbor numbers, a combination of the target evolutionary tree and the target neighbor number is obtained.
7. The method according to claim 6, characterized in that The obtaining, based on the predicted matching information of the candidate protein sample and the matching information of the candidate protein sample, a correlation coefficient corresponding to a combination of the first evolutionary tree and the first number of neighbors includes: Obtaining the predicted matching information of each protein in the test sample set, the average of the predicted matching information of each protein in the test sample set, the matching information of each protein in the test sample set, and the average of the matching information of each protein in the test sample set; Based on the predicted matching information of each protein in the test sample set, the average of the predicted matching information of each protein in the test sample set, the matching information of each protein in the test sample set, and the average of the matching information of each protein in the test sample set, the Pearson correlation coefficient is calculated to obtain the correlation coefficient corresponding to the combination of the first evolutionary tree and the first number of neighbors.
8. The method according to claim 6, characterized in that The obtaining, based on the predicted matching information of the candidate protein sample and the matching information of the candidate protein sample, a correlation coefficient corresponding to a combination of the first evolutionary tree and the first number of neighbors includes: Obtaining a correlation sub-coefficient of a combination of the first evolutionary tree and the first number of neighbors relative to the candidate protein sample based on the predicted matching information of the candidate protein sample, the matching information of the candidate protein sample, an average of the predicted matching information of each protein in the test sample set, and an average of the matching information of each protein in the test sample set; The correlation coefficients corresponding to the combination of the first evolutionary tree and the first number of neighbors are obtained by taking an average or median of the correlation sub-coefficients of the combination of the first evolutionary tree and the first number of neighbors relative to each protein in the test sample set.
9. The method according to claim 6, characterized in that The obtaining of the combination of the target evolutionary tree and the target number of neighbors based on the correlation coefficients corresponding to various combinations of the plurality of evolutionary trees and the plurality of neighbor numbers includes: Arrange various combinations of the plurality of evolutionary trees and the number of neighbors in descending order according to the corresponding correlation coefficients; determine one or more combinations ranked first among the various combinations of the plurality of evolutionary trees and the number of neighbors as the combination of the target evolutionary tree and the target number of neighbors; or Among various combinations of the plurality of evolutionary trees and the plurality of neighbor numbers, a combination whose corresponding correlation coefficient is higher than a correlation coefficient threshold is determined as a combination of the target evolutionary tree and the target neighbor number.
10. The method according to claim 6, characterized in that Each combination of the plurality of evolutionary trees and the plurality of neighbor numbers corresponds to a plurality of correlation coefficients; each of the plurality of correlation coefficients corresponds to a different sample partitioning method, wherein the sample partitioning method is a method of partitioning each of the first type proteins having the matching information into the training sample set and the test sample set; The obtaining of the combination of the target evolutionary tree and the target number of neighbors based on the correlation coefficients corresponding to various combinations of the plurality of evolutionary trees and the plurality of neighbor numbers includes: For each combination of the plurality of evolutionary trees and the plurality of neighbor numbers, taking the median of the plurality of correlation coefficients of the combination to obtain a median correlation coefficient of the combination; Arrange various combinations of the plurality of evolutionary trees and the plurality of neighbor numbers in ascending order according to the difference between the corresponding median correlation coefficient and the target correlation coefficient; Among various combinations of the plurality of evolutionary trees and the plurality of neighbor numbers, one or more combinations ranked first are determined as combinations of the target evolutionary tree and the target neighbor number.
11. The method according to any one of claims 1 to 4, characterized in that: The specified amino acid sequence includes one or more of the following amino acid sequences: Heavy chain amino acid sequence, light chain amino acid sequence, amino acid sequence corresponding to the heavy chain V gene, amino acid sequence corresponding to the light chain V gene, amino acid sequence corresponding to the heavy chain J gene, and amino acid sequence corresponding to the light chain J gene.
12. A protein matching device, characterized in that: The device comprises: a target evolutionary tree acquisition module, configured to acquire a target evolutionary tree, wherein the target evolutionary tree is an evolutionary tree constructed based on specified amino acid sequences in a plurality of first-type proteins; leaf nodes of the evolutionary tree are the first-type proteins; and some of the first-type proteins in the evolutionary tree have matching information, wherein the matching information is used to indicate the degree of matching between the first-type proteins and the second-type proteins; a query module configured to query neighbor proteins having a target number of neighbors from the target evolutionary tree for a candidate protein based on a distance from the candidate protein; the candidate protein being other proteins of the first type other than the first type protein having the matching information; and the neighbor protein being the first type protein having the matching information; a predicted matching information acquisition module, configured to acquire predicted matching information of the candidate protein based on the matching information of the neighbor proteins of the target number of neighbors; A matching result acquisition module is used to acquire the matching result of the candidate protein based on the predicted matching information of the candidate protein.
13. A computer device, characterized in that: The computer device includes a processor and a memory, wherein the memory stores at least one computer program, and the at least one computer program is loaded and executed by the processor to implement the protein matching method according to any one of claims 1 to 11.
14. A computer-readable storage medium, characterized in that The computer-readable storage medium stores at least one computer program, and the computer program is loaded and executed by a processor to implement the protein matching method according to any one of claims 1 to 11.
15. A computer program product, characterized in that The computer program product includes a computer program stored in a computer-readable storage medium; the computer program is read and executed by a processor of a computer device to implement the protein matching method according to any one of claims 1 to 11.