Reasonable information determination method and device, computer equipment and storage medium
By determining the residue structure of the predicted binding site in the complex, the problem of inaccurate judgment of the rationality of antigen and antibody complex structures in the prior art is solved, more accurate acquisition of rationality information is achieved, and the process of antibody drug and vaccine design is promoted.
Patent Information
- Application Number
- CN202311576950.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-22
- Publication Date
- 2025-05-23
AI Technical Summary
It is difficult for the prior art to accurately determine whether the complex structure predicted by antigens and antibodies is reasonable, especially since antigens and antibodies are only a small category of proteins, and the analysis of protein-based properties is too broad to accurately reflect the characteristics of antigens and antibodies.
By obtaining the complex, multiple predicted binding sites exist on the complex, each predicted binding site includes residues on the antigen and residues on the antibody, and the rationality information of the complex is determined based on the structure of these residues.
It can obtain reasonable information more accurately, improve the accuracy of judging the rationality of the complex structure, thereby facilitating the research and development of antibody drugs and vaccine design.
Smart Images

Figure CN120032707A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a method, apparatus, computer equipment and storage medium for determining rationality information. Background Art
[0002] With the development of biomedicine, it is crucial to figure out which region of the antigen the antibody binds to for the development of antibody drugs. Generally, the structure of the complex of antigen and antibody is predicted by deep learning models. However, it is also very important whether the structure of the predicted complex is reasonable. "Reasonable" means that the antigen and antibody can combine into the complex in a natural environment without interference factors.
[0003] At present, the commonly used approach is to treat the predicted complex as a "macromolecule protein-macromolecule protein" complex, and then analyze the predicted complex based on the properties of the protein to determine whether the predicted complex is reasonable. For example, the pTM (predicted Template Modeling) score of the complex is calculated.
[0004] However, antigens and antibodies are only a small category of proteins and are special cases. The above technical solution is too broad to accurately determine whether the structure of the complex is reasonable. For example, if a predicted complex meets the characteristics of proteins but does not meet the characteristics of antigens and antibodies, the above technical solution cannot accurately analyze whether the structure of the complex is reasonable. Summary of the invention
[0005] The embodiments of the present application provide a method, apparatus, computer device and storage medium for determining rationality information, which takes into account the residues and residue structures contained in each predicted binding site, so that rationality information can be obtained more accurately. The technical solution is as follows:
[0006] In one aspect, a method for determining rationality information is provided, the method comprising:
[0007] obtaining a complex, the complex comprising an antigen and an antibody, the complex being predicted based on the antigen and the antibody;
[0008] Based on the position of the antigen and the position of the antibody in the complex, determining a plurality of predicted binding sites present on the complex, each predicted binding site comprising a residue on the antigen and a residue on the antibody;
[0009] Based on the structures of multiple residues in the multiple predicted binding sites, reasonableness information of the complex is determined, and the reasonableness information is used to indicate whether the complex predicted based on the antigen and the antibody is reasonable.
[0010] In another aspect, a device for determining rationality information is provided, the device comprising:
[0011] An acquisition module, which acquires a complex, wherein the complex includes an antigen and an antibody, and the complex is predicted based on the antigen and the antibody;
[0012] A first determination module is used to determine a plurality of predicted binding sites on the complex based on the position of the antigen and the position of the antibody in the complex, each predicted binding site comprising a residue on the antigen and a residue on the antibody;
[0013] The second determination module is used to determine the rationality information of the complex based on the structures of multiple residues in the multiple predicted binding sites, wherein the rationality information is used to indicate whether the complex predicted based on the antigen and the antibody is reasonable.
[0014] In some embodiments, the first determining module includes:
[0015] A first acquisition unit, configured to acquire positions of a plurality of first residues based on the antigen in the complex, wherein the plurality of first residues are residues located on the surface of the antigen;
[0016] a second acquisition unit, configured to acquire positions of a plurality of second residues based on the antibody in the complex, wherein the plurality of second residues include a site capable of binding to the antigen;
[0017] The first determining unit is used to determine a plurality of predicted binding sites between the antigen and the antibody based on the positions of the plurality of first residues and the positions of the plurality of second residues.
[0018] In some embodiments, the first acquisition unit is used to acquire multiple first residues based on the antigen in the complex; for any first residue, the position of the center of mass of the multiple atoms is calculated based on the positions of the multiple atoms in the first residue; and the position of the center of mass is used as the position of the first residue.
[0019] In some embodiments, the first determination unit is used to determine, for any first residue among the multiple first residues, the distance between the first residue and each second residue among the multiple second residues based on the position of the first residue and the positions of the multiple second residues; sort the multiple second residues in order of distance from near to far; and determine the predicted binding site corresponding to the first residue based on the first residue and a preset number of second residues that are ranked first.
[0020] In some embodiments, the second determining module includes:
[0021] A third acquisition unit is used to acquire a comparison database, wherein the comparison database contains multiple categories of reference binding sites, wherein the reference binding sites are binding sites in a complex synthesized by an antigen and an antibody in a real situation;
[0022] A second determining unit is used to determine, for any predicted binding site among the multiple predicted binding sites, a sequence similarity between the predicted binding site and the reference binding sites of each category in the comparison database based on multiple residues in the predicted binding site and multiple residues in the reference binding sites of each category;
[0023] A third determining unit is used to determine the structural similarity between the predicted binding site and the reference binding sites of each category based on the structure of multiple residues in the predicted binding site and the structure of multiple residues in the reference binding sites of each category in the comparison database;
[0024] The fourth determination unit is used to determine the rationality information of the complex based on the sequence similarity and structural similarity corresponding to the multiple predicted binding sites.
[0025] In some embodiments, the third determination unit is used to use an initial rotation matrix to process the positions of multiple residues in the reference binding site of any category in the comparison database to obtain an intermediate site; determine the gap between the intermediate site and the predicted binding site; adjust the initial rotation matrix with the goal of minimizing the gap; and determine the structural similarity between the predicted binding site and the reference binding site based on the minimum gap.
[0026] In some embodiments, the apparatus further comprises:
[0027] A processing unit, for obtaining a residue correspondence relationship by pairwise matching multiple residues in the reference binding sites with multiple residues between the binding sites for any category of reference binding sites in the comparison database;
[0028] The third determination unit is used to determine, for the third residue in the middle part, based on the residue correspondence, from the predicted binding site, a fourth residue corresponding to the third residue, wherein the third residue is any residue in the middle part; determine the residue gap between the third residue and the fourth residue based on the position of the third residue and the position of the fourth residue; and determine the gap between the middle part and the binding site based on multiple residue gaps.
[0029] In some embodiments, the third acquisition unit is used to acquire multiple sample complexes, where the multiple sample complexes are complexes synthesized by antigens and antibodies in real situations; for any sample complex, based on the sample complex, multiple sample binding sites between the antigen and the antibody are determined; based on the multiple sample binding site structures corresponding to the multiple sample complexes, the multiple sample binding sites corresponding to the multiple sample complexes are clustered to obtain the multiple categories of sample binding sites; for any category, the central binding site of the multiple sample binding sites in the category is used as the reference binding site corresponding to the category, and the sum of the gaps between the central binding site and other sample binding sites in the category is the smallest; based on the reference binding sites corresponding to the multiple categories, the comparison database is constructed.
[0030] In some embodiments, the apparatus further comprises:
[0031] A third determination module is used to determine the residue sequence of any sample complex among the multiple sample complexes based on the residue sequence of the antibody heavy chain, the residue sequence of the antibody light chain and the residue sequence of the antigen in the sample complex;
[0032] An encoding module, used for encoding the residue sequence of the sample complex to obtain a sequence identifier of the sample complex;
[0033] The processing module is configured to delete, from the plurality of sample complexes, a sample complex whose sequence identifier has a similarity with the sequence identifier of the sample complex that satisfies a first condition.
[0034] In some embodiments, the fourth determination unit is used to determine, for any predicted binding site, the category of the binding site as a positive class if the sequence similarity between the predicted binding site and any reference binding site satisfies the second condition; if the structural similarity between the predicted binding site and the reference binding site satisfies the third condition, and the positive class is used to indicate that the predicted binding site is consistent with the binding situation of the antigen and antibody under actual conditions; based on the categories of the multiple binding sites, determine the rationality information of the complex.
[0035] In some embodiments, the fourth determination unit is also used to, for any binding site, when the sequence similarity between the binding site and any reference binding site satisfies the second condition, determine that the category of the binding site is a negative class, and the negative class is used to indicate that the binding site does not conform to the binding situation of the antigen and antibody under actual conditions; for any binding site, when the sequence similarity between the binding site and any reference binding site satisfies the second condition, if the structural similarity between the binding site and the reference binding site does not satisfy the third condition, determine that the category of the binding site is a negative class.
[0036] In some embodiments, the fourth determination unit is used to determine that the rationality information of the complex is reasonable when the number of predicted binding sites of the positive class meets the fourth condition, and the rationality is used to indicate that the complex conforms to the binding situation of the antigen and the antibody under actual circumstances; when the number of binding sites of the positive class does not meet the fourth condition, the rationality information of the complex is determined to be unreasonable.
[0037] On the other hand, a computer device is provided, which includes a processor and a memory, wherein the memory is used to store at least one computer program, and the at least one computer program is loaded and executed by the processor to implement the method for determining rationality information in an embodiment of the present application.
[0038] On the other hand, a computer-readable storage medium is provided, in which at least one computer program is stored. The at least one computer program is loaded and executed by a processor to implement a method for determining rationality information as in an embodiment of the present application.
[0039] On the other hand, a computer program product is provided, including a computer program, which is stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the computer device executes the method for determining the rationality information provided in the above-mentioned various aspects or various optional implementations of various aspects.
[0040] The embodiment of the present application provides a method for determining rationality information. For a complex predicted based on an antigen and an antibody, multiple predicted binding positions existing on the complex can be determined according to the positions of the antigen and the antibody in the complex. Since each predicted binding position contains residues on the antigen and residues on the antibody, the predicted binding position can accurately reflect the predicted binding between the antigen and the antibody. Then, the rationality information of the complex is determined according to the structures of multiple residues in the multiple predicted binding sites. Since the residues contained in each predicted binding site and the structure of the residues are taken into consideration, the rationality information can be obtained more accurately. That is, the rationality information of the complex is determined from two aspects, namely, the residue sequence and the structure of the residues in the predicted binding site of the antigen and the antibody, which can improve the accuracy of the rationality information. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0042] Figure 1 It is a schematic diagram of an implementation environment of a method for determining rationality information provided in an embodiment of the present application;
[0043] Figure 2 is a flow chart of a method for determining rationality information provided according to an embodiment of the present application;
[0044] Figure 3 is a flowchart of another method for determining rationality information provided in an embodiment of the present application;
[0045] Figure 4 It is a flow chart of building a comparison database provided in an embodiment of the present application;
[0046] Figure 5 It is a framework diagram of a method for determining rationality information provided according to an embodiment of the present application;
[0047] Figure 6 is a block diagram of a device for determining rationality information provided according to an embodiment of the present application;
[0048] Figure 7 is a block diagram of another device for determining rationality information provided according to an embodiment of the present application;
[0049] Figure 8 is a structural block diagram of a terminal provided according to an embodiment of the present application;
[0050] Fig. 9 It is a structural diagram of a server provided according to an embodiment of the present application. DETAILED DESCRIPTION
[0051] In order to make the objectives, technical solutions and advantages of the present application clearer, the implementation methods of the present application will be further described in detail below with reference to the accompanying drawings.
[0052] In this application, the terms "first", "second", etc. are used to distinguish identical or similar items with basically the same effects and functions. It should be understood that there is no logical or temporal dependency between "first", "second", and "nth", nor is there any limitation on quantity and execution order.
[0053] In the present application, the term "at least one" means one or more, and the term "plurality" means two or more.
[0054] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.) and signals involved in this application are all authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards of relevant countries and regions. For example, the antigens, antibodies and complexes involved in this application are all obtained with full authorization.
[0055] For ease of understanding, the terms involved in this application are explained below.
[0056] Artificial Intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines so that machines have the functions of perception, reasoning and decision-making.
[0057] Artificial intelligence technology is a comprehensive discipline that covers a wide range of fields, including both hardware-level and software-level technologies. Basic artificial intelligence technologies generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, mechatronics, etc. Among them, pre-trained models are also called large models and basic models. After fine-tuning, they can be widely used in downstream tasks in various major directions of artificial intelligence. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0058] With the research and progress of artificial intelligence technology, artificial intelligence technology has been studied and applied in many fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, unmanned driving, automatic driving, drones, digital twins, virtual humans, robots, artificial intelligence generated content (AIGC), conversational interaction, smart medical care, smart customer service, game AI, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role. The method for determining the rationality information provided in the embodiment of the present application can be applied to the field of smart medical care in artificial intelligence.
[0059] The method for determining rationality information provided in the embodiment of the present application can be executed by a computer device. In some embodiments, the computer device is a terminal or a server. The following first takes the computer device as an example to introduce the implementation environment of the method for determining rationality information provided in the embodiment of the present application. Figure 1 Schematic diagram of an implementation environment of a method for determining rationality information provided in an embodiment of the present application. Figure 1 The implementation environment includes a terminal 101 and a server 102. The terminal 101 and the server 102 can be directly or indirectly connected via wired or wireless communication, which is not limited in this application.
[0060] In some embodiments, terminal 101 is a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, an intelligent voice interaction device, a smart home appliance, a vehicle-mounted terminal, etc., but is not limited thereto. Terminal 101 runs an application that supports protein detection. The application may be a medical application or a detection application, and the embodiments of the present application are not limited thereto. Schematically, terminal 101 is a terminal used by a user. Terminal 101 may send an antigen identifier and an antibody identifier to server 102. Then, server 102 may obtain corresponding antigens and antibodies based on the antigen identifier and the antibody identifier, and predict a complex based on the antigen and the antibody. Then, server 102 may determine the rationality information of the complex based on the position of the antigen and the antibody in the complex. Then, server 102 may send the rationality information to terminal 101 so that terminal 101 may display the rationality information to the user.
[0061] Those skilled in the art will appreciate that the number of the above terminals may be more or less. For example, the above terminal may be only one, or the above terminals may be dozens or hundreds, or more. The embodiment of the present application does not limit the number of terminals and device types.
[0062] In some embodiments, server 102 is an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), big data and artificial intelligence platforms. Server 102 is used to provide background services for applications that support protein detection. In some embodiments, server 102 undertakes the main computing work and terminal 101 undertakes the secondary computing work; or, server 102 undertakes the secondary computing work and terminal 101 undertakes the main computing work; or, server 102 and terminal 101 use a distributed computing architecture for collaborative computing.
[0063] Figure 2 is a flowchart of a method for determining rationality information provided in an embodiment of the present application, see Figure 2 In the present application embodiment, the method for determining the rationality information is described by taking the computer device as an example. The method for determining the rationality information comprises the following steps:
[0064] 201. A computer device obtains a complex, the complex includes an antigen and an antibody, and the complex is predicted based on the antigen and the antibody.
[0065] In the embodiments of the present application, antigens and antibodies may be artificially synthesized proteins or proteins already existing in nature, and the embodiments of the present application do not limit the structures of antigens and antibodies. The complex is predicted based on antigens and antibodies. The prediction method may be a prediction method based on deep learning or a prediction method based on molecular dynamics simulation, and the embodiments of the present application do not limit the prediction method of the complex. The computer device may obtain the complex from other computer devices, or may obtain the complex based on the prediction of the antigen and antibody themselves, and the embodiments of the present application do not limit this. Among them, the complex obtained by the computer device refers to obtaining the structural information of the complex. The structural information of the complex includes the identification of multiple residues in the complex and the positions of the multiple residues.
[0066] 202. The computer device determines a plurality of predicted binding sites on the complex based on the position of the antigen and the position of the antibody in the complex, each predicted binding site comprising a residue on the antigen and a residue on the antibody.
[0067] In an embodiment of the present application, a computer device obtains the position of the antigen and the position of the antibody in the complex. Then, the computer device determines the position of the residues on the antigen based on the position of the antigen in the complex; and determines the position of the residues on the antibody based on the position of the antibody in the complex. Then, the computer device determines multiple predicted binding sites on the complex based on the positions of the residues on the antigen and the positions of the residues on the antibody. The predicted binding site is used to indicate the site where the antigen and the antibody bind on the complex. That is, the predicted binding site is used to indicate the site where the predicted antigen and the antibody will bind. The embodiment of the present application does not limit the number of residues in each predicted binding site.
[0068] For example, each predicted binding site includes 1 residue on the antigen and 3 residues on the antibody, and each predicted binding site indicates that 1 residue on the antigen binds to 3 residues on the antibody. Since both antigens and antibodies are proteins, the binding between antigens and antibodies is a three-dimensional structure, so the computer equipment can obtain more residues to obtain the predicted binding site, so as to more accurately reflect the binding between antigens and antibodies from the perspective of three-dimensional structure.
[0069] 203. The computer device determines the rationality information of the complex based on the structure of multiple residues in multiple predicted binding sites, and the rationality information is used to indicate whether the complex predicted based on the antigen and the antibody is reasonable.
[0070] In an embodiment of the present application, for each predicted binding site, the computer device can determine the multiple residues present in the predicted binding site. On this basis, the computer device can also obtain the structure of each residue in each predicted binding site. Then, the computer device determines the rationality information of the complex based on the multiple residues in the multiple predicted binding sites and the structures of the multiple residues. The rationality information can indicate that the complex is reasonable or unreasonable. Among them, "reasonable" means that the combination of antigen and antibody into the complex conforms to the laws of nature, that is, in a natural environment without interference factors, the antigen and antibody can spontaneously combine to form the complex.
[0071] The embodiment of the present application provides a method for determining rationality information. For a complex predicted based on an antigen and an antibody, multiple predicted binding positions existing on the complex can be determined according to the positions of the antigen and the antibody in the complex. Since each predicted binding position contains residues on the antigen and residues on the antibody, the predicted binding position can accurately reflect the predicted binding between the antigen and the antibody. Then, the rationality information of the complex is determined according to the structures of multiple residues in the multiple predicted binding sites. Since the residues contained in each predicted binding site and the structure of the residues are taken into consideration, the rationality information can be obtained more accurately. That is, the rationality information of the complex is determined from two aspects, namely, the residue sequence and the structure of the residues in the predicted binding site of the antigen and the antibody, which can improve the accuracy of the rationality information.
[0072] Figure 3 is a flowchart of another method for determining rationality information provided in an embodiment of the present application, see Figure 3 In the present application embodiment, the method for determining the rationality information is described by taking the computer device as an example. The method for determining the rationality information comprises the following steps:
[0073] 301. A computer device obtains a complex, the complex includes an antigen and an antibody, and the complex is predicted based on the antigen and the antibody.
[0074] In the embodiment of the present application, the complex can be predicted by other computer devices based on antigens and antibodies, or by the current computer device based on antigens and antibodies, and the embodiment of the present application does not limit this. Accordingly, the computer device can obtain the complex from other computer devices, or predict the complex by itself, and the embodiment of the present application does not limit this.
[0075] The antibody includes a heavy chain and a light chain. The computer device can determine the structure of the heavy chain and the structure of the light chain in the antibody according to the identifier of the heavy chain and the identifier of the light chain in the antibody. The identifier can be an ID (Identity Document), which is not limited in the present embodiment of the application. Then, the computer device can determine the residue sequence in the heavy chain according to the identifier of the heavy chain; and determine the residue sequence in the light chain according to the identifier of the light chain. Then, the computer device can encode the residue sequence in the heavy chain and the residue sequence in the light chain respectively, thereby obtaining a first string and a second string. The first string is used to represent the residue sequence in the heavy chain; the second string is used to represent the residue sequence in the light chain. "First string + second string" is used to represent the residue sequence in the antibody. "+" is used to represent splicing. The computer device can also determine the structure of the antigen according to the identifier of the antigen. Then, the computer device can determine the residue sequence in the antigen according to the structure of the antigen. Then, the computer device encodes the residue sequence in the antigen and the residue sequence in the antigen to obtain a third string. "First string + second string + third string" is used to represent the complex predicted based on the antigen and the antibody. That is, the computer device can predict the structure of the complex based on the residue sequence represented by "the first character string + the second character string + the third character string".
[0076] The computer device may be encoded by adopting IMGT (immunogenetics) encoding, which is not limited in the present embodiment. Each residue may be encoded into at least one character. The characters corresponding to the multiple residues are arranged according to the connection relationship (position relationship) between the multiple residues to form a character string.
[0077] 302. The computer device determines a plurality of predicted binding sites on the complex based on the position of the antigen and the position of the antibody in the complex, each predicted binding site comprising a residue on the antigen and a residue on the antibody.
[0078] In the present application embodiment, the binding between the antigen and the antibody refers to the binding between the residues on the antigen and the residues on the antibody. The computer device can determine the position of the residues on the antigen and the position of the residues on the antibody based on the position of the antigen and the position of the antibody in the complex. Then, the computer device determines multiple predicted binding sites present on the complex based on the position of the residues on the antigen and the position of the residues on the antibody. In the present application embodiment, the binding site on the complex can be referred to as a "local gripper" between the antigen and the antibody.
[0079] In some embodiments, the residues on the antigen that can bind to the antibody are located on the surface of the antigen. The heavy chain and light chain in the antibody both include side chains. The position on the antibody that can bind to the antigen is located on the side chain of the residue. Accordingly, the process of the computer device determining the multiple predicted binding sites present on the complex includes: the computer device obtains the positions of multiple first residues based on the antigen in the complex. The computer device obtains the positions of multiple second residues based on the antibody in the complex. Then, the computer device determines multiple predicted binding sites between the antigen and the antibody based on the positions of the multiple first residues and the positions of the multiple second residues. Among them, the multiple first residues are residues located on the surface of the antigen. The multiple second residues include a site that can bind to the antigen. The site is located on the side chain of multiple second residues. The embodiment of the present application does not limit the number of second residues. The scheme provided in the embodiment of the present application determines the predicted binding site between the antigen and the antibody according to the positions of the residues that can bind on each structure of the antibody and the antigen, thereby improving the accuracy of determining the predicted binding site.
[0080] In some embodiments, residue refers to the remaining part after the dehydration condensation of the amino group and carboxyl group of an amino acid. The residue is composed of multiple atoms. Multiple atoms may include hydrogen atoms, oxygen atoms, and carbon atoms, etc., which are not limited by the embodiments of the present application. In the process of obtaining the position of the residue, the computer device can determine the position of the residue according to the position of the atoms in the residue. Accordingly, the process of obtaining the position of multiple first residues based on the antigen in the complex by the computer device includes: the computer device obtains multiple first residues based on the antigen in the complex. Then, for any first residue, the computer device calculates the position of the center of mass of multiple atoms based on the position of multiple atoms in the first residue. Then, the computer device uses the position of the center of mass as the position of the first residue. The solution provided in the embodiments of the present application uses the position of the center of mass of multiple atoms in the residue as the position of the residue, so that the position of the entire residue can be accurately reflected, thereby improving the accuracy of obtaining the position of the residue.
[0081] In some embodiments, in the process of obtaining the first residue, the computer device can obtain multiple first residues from multiple residues of the antigen based on the solvent accessible surface area of the residues in the antigen. Accordingly, the process of the computer device obtaining multiple first residues includes: the computer device calculates the solvent accessible surface area of each residue in the antigen. Then, the computer device selects the residue whose solvent accessible surface area is not less than the area threshold as the first residue. The area threshold can be (square angstrom), which is not limited in the embodiments of the present application. The first residue may also be referred to as a surface residue. The solution provided in the embodiments of the present application, by calculating the solvent accessible surface area of the residue, enables accurate determination of the spatial position of the residue in the antigen, whether it is located on the surface of the antigen or inside the antigen; the solvent accessible surface area of the residue located on the surface of the antigen is generally larger than the solvent accessible surface area of the residue located inside the antigen, and by determining the residue whose solvent accessible surface area is not less than the area threshold as a surface residue, the accuracy of obtaining the surface residue is improved.
[0082] Then, the computer device can obtain the position of the second residue in the same manner as that of obtaining the position of the first residue, which will not be described in detail here.
[0083] In some embodiments, the distance between the antigen and the antibody is close enough to have an interaction force. That is, the distance between the antigen and the antibody is close enough to cause a binding reaction between the two to form a complex. The computer device can determine multiple predicted binding sites between the antigen and the antibody on the complex based on the positional relationship between the surface residues on the antigen and the side chains of the residues in the antibody. Accordingly, the process of determining multiple predicted binding sites between the antigen and the antibody based on the positions of multiple first residues and the positions of multiple second residues by the computer device includes: for any first residue in the multiple first residues, the computer device determines the distance between the first residue and each second residue in the multiple second residues based on the position of the first residue and the position of the multiple second residues. Then, the computer device sorts the multiple second residues in order from near to far. Then, the computer device determines the predicted binding site corresponding to the first residue based on the first residue and the preset number of second residues ranked first. The preset number K can be equal to 3, 4 or 5, etc., and the embodiment of the present application does not limit the preset number. The solution provided in the embodiment of the present application is that since the distance between the antigen and the antibody when they bind is relatively close in real situations, based on the position of the first residue on the antigen that can bind and the second residue in the antibody that can bind, the first residue and the second residue that are relatively close are used as predicted binding sites. That is, the position where the antigen and the antibody in the predicted complex are relatively close is used as the predicted binding site, thereby improving the accuracy of determining the predicted binding site.
[0084] Before sorting the plurality of second residues according to the distance, the computer device may also screen the plurality of second residues according to the distance. Accordingly, the computer device obtains the second residues whose corresponding distance exceeds the distance threshold from the plurality of second residues. Then, the computer device sorts the second residues whose distance exceeds the distance threshold in the order of distance from near to far. Then, the computer device determines the predicted binding site corresponding to the first residue based on the first residue and the preset number of second residues ranked first. The distance threshold may be (Angstroms), which is not limited in the present embodiment. The scheme provided in the present embodiment screens multiple second residues by using a distance threshold, so as to avoid the situation where the first residue in the predicted binding site is far away from the multiple second residues. Since the first residue is far away from the multiple second residues, it indicates that there is a high probability that there is no binding at the site, so the above method further ensures the accuracy of determining the predicted binding site.
[0085] 303. The computer device obtains a comparison database, which contains multiple categories of reference binding sites, and the reference binding sites are binding sites in the complex synthesized by the antigen and the antibody under real conditions.
[0086] In an embodiment of the present application, a comparison database includes reference binding sites of multiple categories. Each reference binding site is used to represent the binding situation of the corresponding category under the actual situation (natural environment). The binding situation refers to the residues bound by the antigen and the antibody in the complex synthesized by the antigen and the antibody in the natural environment. Each reference binding site includes the residues on the antigen and the residues on the antibody in the synthesized complex. That is, the reference binding site is used to represent the real binding situation between the antigen and the antibody in the natural environment. The sequence and structure of the residues between different reference binding sites are not exactly the same. The computer device can obtain the comparison database from other computer devices, or it can construct the comparison database by itself. The embodiment of the present application does not limit the acquisition method of the comparison database.
[0087] In some embodiments, the process of building a comparison database by a computer device includes the following steps 3031 to 3035. Figure 4 , Figure 4 It is a flowchart of building a comparison database provided in an embodiment of the present application.
[0088] 3031. The computer device obtains multiple sample complexes.
[0089] The plurality of sample complexes are complexes synthesized by antigens and antibodies in real situations. The computer device can obtain the plurality of sample complexes based on at least one preset database. The preset database can be a SAbDab database or a RCSB PDB database, etc., which is not limited in the embodiments of the present application.
[0090] In the process of obtaining any sample complex, the computer device can obtain basic information such as the identification of the heavy chain of the antibody, the identification of the light chain, and the identification of the ligand of the antibody from the preset database. The ligand of the antibody is the antigen that can bind to the antibody. Since the sample complex used in this application is a structure that has been published and experimentally analyzed based on a docking tool (such as haddock), it can be regarded as the real antigen-antibody binding situation in the natural environment. Therefore, the computer device can directly obtain the corresponding sample complex based on the identification of the heavy chain, the identification of the light chain, and the identification of the ligand, providing a basis for subsequent judgment on whether the structure of the predicted complex is reasonable.
[0091] In some embodiments, there may be duplicate antibodies and antigens in the preset database, thereby obtaining duplicate sample complexes. The computer device can also remove duplicates from the above-mentioned multiple sample complexes. Accordingly, for any sample complex among the multiple sample complexes, the computer device determines the residue sequence of the sample complex based on the residue sequence of the antibody heavy chain, the residue sequence of the antibody light chain, and the residue sequence of the antigen in the sample complex. Then, the computer device encodes the residue sequence of the sample complex to obtain the sequence identifier of the sample complex. Then, the computer device deletes the sample complex whose sequence identifier and the sequence identifier of the sample complex meet the first condition in similarity from the multiple sample complexes. Among them, the encoding can be an IMGT encoding method, which is not limited in the embodiments of the present application. The solution provided in the embodiment of the present application determines the sequence identifier of the sample complex through the residue sequence of the antibody heavy chain, the residue sequence of the antibody light chain and the residue sequence of the antigen, so that the structure of the sample complex can be reflected by the sequence identifier, and then deletes the sample complex whose sequence identifier and the sequence identifier of the sample complex have similarity that meets the first condition, thereby filtering out the sample complexes with a large degree of structural similarity from multiple sample complexes, reducing the amount of repeated data, and facilitating the subsequent faster construction of the comparison database.
[0092] Among them, the computer device determines the residue sequence in the heavy chain according to the identifier of the heavy chain; determines the residue sequence in the light chain according to the identifier of the light chain; and determines the residue sequence in the antigen according to the identifier of the antigen. The residue sequence can not only reflect the residues contained in the structure, but also reflect the arrangement order between the residues. Then, the computer device encodes the above-mentioned residue sequences to obtain the character strings corresponding to the heavy chain, light chain and antigen respectively. Then, the computer device splices the character strings corresponding to the heavy chain, light chain and antigen respectively together to obtain the character string corresponding to the sample complex. The character string is the sequence identifier of the sample complex. Then, the computer device uses the sequence identifier of the sample complex as a reference to detect whether the sequence identifiers of other sample complexes are repeated with the sequence identifier.
[0093] Accordingly, the computer device calculates the string edit distance between the sequence identifier of the first sample complex and the sequence identifier of the second sample complex. Then, the computer device compares the string edit distance with the length of the sequence identifier of the first sample complex to obtain a ratio. When the ratio is less than a preset value, the computer device determines that the first sample complex is repeated with the second sample complex and deletes the second sample complex. When the ratio is not less than the preset value, the computer device determines that the first sample complex is not repeated with the second sample complex and retains the second sample complex. The preset value may be equal to 0.2, and the embodiment of the present application is not limited to this.
[0094] 3032. For any sample complex, the computer device determines multiple sample binding sites between the antigen and the antibody based on the sample complex.
[0095] The principle of the computer device executing step 3032 is similar to the principle of determining the predicted binding site in step 302, and will not be repeated here.
[0096] 3033. The computer device clusters the multiple sample binding sites corresponding to the multiple sample complexes based on the multiple sample binding site structures corresponding to the multiple sample complexes to obtain multiple categories of sample binding sites.
[0097] Wherein, for any two sample binding sites among the multiple sample binding sites, the computer device determines the sequence similarity between the two sample binding sites based on the multiple residues in the two sample binding sites. Then, the computer device determines the structural similarity between the two sample binding sites based on the structures of the multiple residues in the two sample binding sites. The calculation principle of sequence similarity and structural similarity can be found in steps 304 and 305, which will not be repeated here. Then, the computer device clusters the multiple sample binding sites based on the sequence similarity and structural similarity between the sample binding sites, thereby obtaining multiple categories of sample binding sites.
[0098] For any sample binding site, the computer device can determine the similarity list of the sample binding site. The similarity list includes the sequence similarity and structural similarity between each other sample binding site and the sample binding site. Before clustering, the computer device can filter the data in the similarity list of each sample binding site based on the sequence similarity and structural similarity. Optionally, for any similarity list, the computer device can first filter out the sample binding sites whose corresponding sequence similarity reaches the sequence threshold; then, the computer device filters out the sample binding sites whose structural similarity reaches the structural threshold from the sample binding sites whose sequence similarity reaches the sequence threshold, thereby obtaining a new similarity list. The embodiment of the present application does not limit the size of the sequence threshold and the structural threshold. Alternatively, the computer device can also perform weighted summation of the sequence similarity and the structural similarity to obtain the total similarity value of the sample binding site; then, the computer device deletes the sample binding sites whose total similarity value does not reach the similarity threshold from the similarity list. The embodiment of the present application does not limit the screening method. For example, for the similarity list of each sample binding site, the computer device only retains 1% of the data in the similarity list.
[0099] Then, the computer device can use the DBSCAN algorithm to cluster the sample binding sites after screening, so that similar sample binding sites can be classified into one category as much as possible. The present application does not limit the clustering algorithm used, which can be a density clustering method or a hierarchical clustering method.
[0100] 3034. For any category, the computer device uses the central binding site of multiple sample binding sites in the category as the reference binding site corresponding to the category.
[0101] Among them, the sum of the gaps between the central binding site and other sample binding sites in the category is the smallest. In other words, the sum of the similarities between the central binding site and other sample binding sites in the category is the largest. For any sample binding site in the category, the computer device can calculate the sum of the sequence similarities and the sum of the structural similarities between the sample binding site and other sample binding sites in the category. The central binding site can be the sample binding site with the largest sum of sequence similarities, or it can be the sample binding site with the largest sum of the sum of sequence similarities and the sum of structural similarities, or it can be the sample binding site with the largest weighted sum of the sum of sequence similarities and the sum of structural similarities, and the embodiment of the application does not limit this. Then, the computer device uses the central binding site as the reference binding site of the category. The reference binding site is the characterization of the category.
[0102] 3035. The computer device constructs a comparison database based on the reference binding sites corresponding to the multiple categories.
[0103] The computer device can record information such as the sequence and structure of the residues in each reference binding site, and construct a comparison database. The structure can include information such as the position of the residues and the connection relationship between atoms in the residues. The computer device can also record basic information such as the identifier of the heavy chain of the antibody where the residue is located, the identifier of the light chain, and the identifier of the antigen, which is not limited in the present embodiment.
[0104] 304. For any predicted binding site among the multiple predicted binding sites, the computer device determines the sequence similarity between the predicted binding site and the reference binding sites of each category based on multiple residues in the predicted binding site and multiple residues in the reference binding sites of each category in the comparison database.
[0105] In an embodiment of the present application, for any predicted binding site among the multiple predicted binding sites, the computer device can use the BLOSUM scoring matrix to process the sequence of multiple residues in the predicted binding site and the sequence of multiple residues in each reference binding site in the comparison database, so as to determine the sequence similarity between the predicted binding site and the reference binding sites of each category.
[0106] The BLOSUM scoring matrix used in the present application may be BLOSUM62 or BLOSUM80, and the embodiments of the present application are not limited thereto.
[0107] 305. The computer device determines the structural similarity between the predicted binding site and the reference binding sites of each category based on the structure of multiple residues in the predicted binding site and the structure of multiple residues in the reference binding sites of each category in the comparison database.
[0108] In an embodiment of the present application, for any predicted binding site among multiple predicted binding sites, the computer device can determine the structural similarity between the predicted binding site and the reference binding sites of each category based on the position of the residues in the predicted binding site and the positions of multiple residues in the reference binding sites of each category. Accordingly, the process of the computer device determining the structural similarity between the predicted binding site and the reference binding sites of each category includes: for any category of reference binding sites in the comparison database, the computer device uses an initial rotation matrix to process the positions of multiple residues in the reference binding site to obtain an intermediate site. Then, the computer device determines the gap between the intermediate site and the predicted binding site. Then, the computer device adjusts the initial rotation matrix with the goal of minimizing the gap. Then, the computer device determines the structural similarity between the predicted binding site and the reference binding site based on the minimum gap.
[0109] Among them, the computer device can use the Kabsch algorithm to calculate the optimal rotation matrix between the predicted binding site and any reference binding site. The optimal rotation matrix is the rotation matrix when the gap between the middle site and the predicted binding site is the smallest. The optimal rotation matrix can be obtained by adjusting the initial rotation matrix, and the embodiment of the present application does not limit the initial rotation matrix. In the process of calculating the optimal rotation matrix, the computer device can obtain the smallest gap between the predicted binding site and the middle site, and determine the structural similarity between the predicted binding site and the reference binding site based on the gap.
[0110] In certain embodiments, the computer device can calculate the gap between the middle part and the predicted binding site based on the position of multiple residues in the middle part and the position of multiple residues in the predicted binding site. The gap can be RMSD (Root Mean Square Deviation), which is not limited in the present embodiment. Accordingly, for the reference binding site of any category in the comparison database, the computer device will correspond the multiple residues in the reference binding site to the multiple residues between the binding sites in pairs to obtain the residue correspondence. Then, for the third residue in the middle part, the computer device determines the fourth residue corresponding to the third residue from the predicted binding site based on the residue correspondence. The third residue is any residue in the middle part. Then, the computer device determines the residue gap between the third residue and the fourth residue based on the position of the third residue and the position of the fourth residue. Then, the computer device determines the gap between the middle part and the binding site based on multiple residue gaps.
[0111] That is, the computer device first matches the residues in the predicted binding site with the residues in the reference binding site in pairs. Then, the computer device uses a rotation matrix to process the positions of multiple residues in the reference binding site to obtain the middle site. The residues in the middle site also correspond one to one with the residues in the reference binding site. Then, the computer device calculates the residue gap between the corresponding residues in the predicted binding site and the middle site. Then, the computer device determines the gap between the middle site and the binding site based on the multiple residue gaps between the predicted binding site and the middle site.
[0112] In some embodiments, the computer device can determine multiple residue correspondences. The multiple residue correspondences cover the correspondences of all residues in the predicted binding site and the middle site (reference binding site). For any residue correspondence, the computer device can calculate the smallest gap between the middle site and the predicted binding site under the residue correspondence. Then, the computer device can select the smallest gap from the smallest gaps corresponding to the multiple residue correspondences, and determine the structural similarity between the predicted binding site and the reference binding site based on the smallest gap.
[0113] The correspondences contained in different types of residue correspondences are not exactly the same. For example, in the first residue correspondence, residue A in the predicted binding site corresponds to residue a in the middle site, residue B in the predicted binding site corresponds to residue b in the middle site; residue C in the predicted binding site corresponds to residue c in the middle site. In the second residue correspondence, residue A in the predicted binding site corresponds to residue b in the middle site, residue B in the predicted binding site corresponds to residue a in the middle site; residue C in the predicted binding site corresponds to residue c in the middle site. In the third residue correspondence, residue A in the predicted binding site corresponds to residue c in the middle site, residue B in the predicted binding site corresponds to residue a in the middle site; residue C in the predicted binding site corresponds to residue b in the middle site, and the embodiments of the present application are not limited to this.
[0114] 306. The computer device determines the rationality information of the complex based on the sequence similarity and structural similarity corresponding to the multiple predicted binding sites, and the rationality information is used to indicate whether the complex predicted based on the antigen and the antibody is reasonable.
[0115] In an embodiment of the present application, for any predicted binding site, when the sequence similarity between the predicted binding site and any reference binding site meets the second condition, if the structural similarity between the predicted binding site and the reference binding site meets the third condition, the computer device determines that the category of the predicted binding site is a positive class. Then, the computer device determines the rationality information of the complex based on the categories of multiple predicted binding sites. Among them, the positive class is used to indicate that the predicted binding site conforms to the binding situation of the antigen and the antibody under real conditions. The second condition may be that the sequence similarity reaches a first threshold value, or that the sequence similarity is within a first range, which is not limited in the embodiment of the present application. The third condition may be that the structural similarity reaches a second threshold value, or that the structural similarity is within a second range, which is not limited in the embodiment of the present application.
[0116] In some embodiments, for any predicted binding site, if the sequence similarity between the predicted binding site and any reference binding site does not satisfy the second condition, the computer device determines that the category of the predicted binding site is a negative class. The negative class is used to indicate that the binding site does not conform to the binding of the antigen and the antibody in the actual situation. Alternatively, for any predicted binding site, if the sequence similarity between the predicted binding site and any reference binding site satisfies the second condition, if the structural similarity between the predicted binding site and the reference binding site does not satisfy the third condition, the computer device determines that the category of the predicted binding site is a negative class.
[0117] In some embodiments, the process of determining the rationality information of the complex based on the categories of multiple predicted binding sites by the computer device includes: when the number of predicted binding sites of the positive class meets the fourth condition, the rationality information of the complex is determined to be reasonable. Reasonable is used to indicate that the complex conforms to the binding situation of the antigen and the antibody under the actual situation. When the number of predicted binding sites of the positive class does not meet the fourth condition, the computer device determines that the rationality information of the complex is unreasonable. The fourth condition can be that the number of predicted binding sites of the positive class reaches the third threshold, or the number of predicted binding sites of the positive class is within the third range, which is not limited in the embodiment of the present application. The scheme provided in the embodiment of the present application classifies the predicted binding sites of the predicted complex that conform to the characteristics of the reference binding sites in the comparison database into positive classes, and classifies the predicted binding sites that do not conform to the characteristics of the reference binding sites in the comparison database into negative classes, and then determines the rationality information of the complex according to the proportion of predicted binding sites that conform to the actual situation among all predicted binding sites on the predicted complex, so that the rationality information can accurately reflect the binding situation on the complex, and improves the accuracy of the rationality information.
[0118] In order to more clearly describe the method for determining the rationality information provided in the embodiment of the present application, the method for determining the rationality information is further described below in conjunction with the accompanying drawings. Figure 5 , Figure 5 It is a framework diagram of a method for determining rationality information provided according to an embodiment of the present application. Before determining whether the predicted complex is reasonable, the computer device can pre-construct a comparison database. In the process of constructing the comparison database, the computer device constructs a sample complex based on basic information such as the identifier of the heavy chain, the identifier of the light chain, and the identifier of the antigen in the antibody. Among them, the computer device removes duplicates from multiple sample complexes. Then, for any sample complex, the computer device determines multiple sample binding sites between the antigen and the antibody based on the sample complex. Then, based on the multiple sample binding site structures corresponding to the multiple sample complexes, the computer device clusters the multiple sample binding sites corresponding to the multiple sample complexes to obtain multiple categories of sample binding sites. For any category, the computer device uses the central binding site of multiple sample binding sites in the category as the reference binding site corresponding to the category. The computer device constructs a comparison database based on the reference binding sites corresponding to multiple categories. The computer equipment then compares the predicted binding site in the predicted complex with the reference binding sites of each category in the comparison database, and searches the comparison database for reference binding sites similar to the predicted binding site by calculating parameters such as sequence similarity and structural similarity, thereby determining the predicted complex.
[0119] For example, when the sequence similarity between a predicted binding site and any reference binding site meets H1, if the structural similarity between the predicted binding site and the reference binding site meets H2, the computer device determines that the class of the predicted binding site is a positive class. Otherwise, the computer device determines that the class of the predicted binding site is a negative class. When the proportion of the number of predicted binding sites with a positive class reaches H3, the computer device determines that the rationality information of the complex is reasonable.
[0120] The method for determining rationality information provided in the embodiments of this application can be applied to the following fields:
[0121] Antibody drug research and development field: In the process of antibody drug research and development, one of the bases for judging whether an antibody drug has a competitive relationship with a ligand on an antigen is to observe whether they overlap in terms of spatial structure. For example, when it is known that a ligand binds to a specific region of an antigen, several candidate antibodies may also compete with the ligand at this position. After obtaining their complexes through methods such as simulated docking or prediction, and then by using the method for determining rationality information provided in the embodiments of this application to judge the rationality of the complexes, and thus the sequence similarity and structural similarity, the order of antibodies most likely to compete with the ligand can be generated. According to this order, experimental verification of competitiveness can be carried out, which can greatly reduce the research and development time and trial-and-error cost of antibody drugs.
[0122] Vaccine design field: In vaccine design, understanding the rationality of the complex of an antigen and an antibody helps to identify effective antigenic epitopes, thereby designing more effective vaccines. For different mutation sites, the method for determining rationality information provided in the embodiments of this application can also be used to judge the stability of the binding of various mutant antigens to antibodies, thus providing strong support for vaccine design and improving the protective effect of vaccines.
[0123] Bioinformatics tool development field: The method for determining rationality information provided in the embodiments of this application can be used as a bioinformatics tool to provide services for researchers to judge the rationality of the complex of an antigen and an antibody. This can help researchers screen out complexes with potential value from a large amount of experimental data and improve research efficiency.
[0124] The embodiment of the present application provides a method for determining rationality information. For a complex predicted based on an antigen and an antibody, multiple predicted binding positions existing on the complex can be determined according to the positions of the antigen and the antibody in the complex. Since each predicted binding position contains residues on the antigen and residues on the antibody, the predicted binding position can accurately reflect the predicted binding between the antigen and the antibody. Then, the rationality information of the complex is determined according to the structures of multiple residues in the multiple predicted binding sites. Since the residues contained in each predicted binding site and the structure of the residues are taken into consideration, the rationality information can be obtained more accurately. That is, the rationality information of the complex is determined from two aspects, namely, the residue sequence and the structure of the residues in the predicted binding site of the antigen and the antibody, which can improve the accuracy of the rationality information and thus help to suppress the situation of misjudgment. Moreover, by comparing the database, the reference binding site can be directly obtained, so as to judge whether the predicted complex is reasonable based on the reference binding site, thereby reducing the time cost, helping to determine the rationality information of the complex more quickly, and improving the efficiency of obtaining the rationality information.
[0125] Figure 6 is a block diagram of a device for determining rationality information provided according to an embodiment of the present application. The device for determining rationality information is used to execute the steps of the method for determining rationality information described above, see Figure 6 The device for determining rationality information includes: an acquisition module 601, a first determination module 602 and a second determination module 603.
[0126] An acquisition module 601 is used to acquire a complex, where the complex includes an antigen and an antibody, and the complex is predicted based on the antigen and the antibody;
[0127] A first determination module 602, for determining a plurality of predicted binding sites on the complex based on the position of the antigen and the position of the antibody in the complex, each predicted binding site comprising a residue on the antigen and a residue on the antibody;
[0128] The second determination module 603 is used to determine the rationality information of the complex based on the structures of multiple residues in multiple predicted binding sites, and the rationality information is used to indicate whether the complex predicted based on the antigen and the antibody is reasonable.
[0129] In some embodiments, Figure 7 is a block diagram of another device for determining rationality information provided in accordance with an embodiment of the present application. Figure 7 , the first determining module 602 includes:
[0130] A first acquisition unit 6021 is used to acquire positions of a plurality of first residues based on the antigen in the complex, where the plurality of first residues are residues located on the surface of the antigen;
[0131] A second acquisition unit 6022, configured to acquire positions of a plurality of second residues based on the antibody in the complex, wherein the plurality of second residues include a site capable of binding to the antigen;
[0132] The first determining unit 6023 is used to determine a plurality of predicted binding sites between the antigen and the antibody based on the positions of the plurality of first residues and the positions of the plurality of second residues.
[0133] In some embodiments, see Figure 7 The first acquisition unit 6021 is used to acquire multiple first residues based on the antigens in the complex; for any first residue, based on the positions of multiple atoms in the first residue, calculate the positions of the center of mass of the multiple atoms; and use the position of the center of mass as the position of the first residue.
[0134] In some embodiments, see Figure 7 The first determination unit 6023 is used to determine, for any first residue among the multiple first residues, the distance between the first residue and each second residue among the multiple second residues based on the position of the first residue and the positions of the multiple second residues; sort the multiple second residues in order of distance from near to far; and determine the predicted binding site corresponding to the first residue based on the first residue and a preset number of second residues that are ranked first.
[0135] In some embodiments, see Figure 7 , the second determining module 603 includes:
[0136] The third acquisition unit 6031 is used to acquire a comparison database, the comparison database contains multiple categories of reference binding sites, and the reference binding sites are binding sites in the complex synthesized by the antigen and the antibody in real conditions;
[0137] A second determining unit 6032 is used to determine, for any predicted binding site among the multiple predicted binding sites, the sequence similarity between the predicted binding site and the reference binding sites of each category in the comparison database based on multiple residues in the predicted binding site and multiple residues in the reference binding sites of each category;
[0138] The third determining unit 6033 is used to determine the structural similarity between the predicted binding site and the reference binding sites of each category based on the structure of multiple residues in the predicted binding site and the structure of multiple residues in the reference binding sites of each category in the comparison database;
[0139] The fourth determining unit 6034 is used to determine the rationality information of the complex based on the sequence similarity and structural similarity corresponding to the multiple predicted binding sites.
[0140] In some embodiments, see Figure 7The third determination unit 6033 is used to use an initial rotation matrix to process the positions of multiple residues in the reference binding site of any category in the comparison database to obtain an intermediate site; determine the gap between the intermediate site and the predicted binding site; adjust the initial rotation matrix with the goal of minimizing the gap; and determine the structural similarity between the predicted binding site and the reference binding site based on the minimum gap.
[0141] In some embodiments, see Figure 7 , the device further comprises:
[0142] A processing unit 6035 is used for, for any category of reference binding sites in the comparison database, pairwise matching of multiple residues in the reference binding sites with multiple residues between the binding sites to obtain a residue correspondence relationship;
[0143] The third determination unit 6033 is used to determine, for the third residue in the middle part, based on the residue correspondence, from the predicted binding site, a fourth residue corresponding to the third residue, where the third residue is any residue in the middle part; determine the residue gap between the third residue and the fourth residue based on the position of the third residue and the position of the fourth residue; and determine the gap between the middle part and the binding site based on multiple residue gaps.
[0144] In some embodiments, see Figure 7 The third acquisition unit 6031 is used to acquire multiple sample complexes, where the multiple sample complexes are complexes synthesized by antigens and antibodies in real situations; for any sample complex, based on the sample complex, multiple sample binding sites between the antigen and the antibody are determined; based on the structures of the multiple sample binding sites corresponding to the multiple sample complexes, the multiple sample binding sites corresponding to the multiple sample complexes are clustered to obtain multiple categories of sample binding sites; for any category, the central binding site of the multiple sample binding sites in the category is used as the reference binding site corresponding to the category, and the sum of the gaps between the central binding site and other sample binding sites in the category is the smallest; based on the reference binding sites corresponding to the multiple categories, a comparison database is constructed.
[0145] In some embodiments, see Figure 7 , the device further comprises:
[0146] A third determination module 604 is used to determine the residue sequence of any sample complex among the multiple sample complexes based on the residue sequence of the antibody heavy chain, the residue sequence of the antibody light chain and the residue sequence of the antigen in the sample complex;
[0147] An encoding module 605 is used to encode the residue sequence of the sample complex to obtain a sequence identifier of the sample complex;
[0148] The processing module 606 is configured to delete, from the plurality of sample complexes, the sample complexes whose sequence identifiers have a similarity with the sequence identifier of the sample complex that meets a first condition.
[0149] In some embodiments, see Figure 7 The fourth determination unit 6034 is used to determine, for any predicted binding site, that the category of the binding site is a positive category if the sequence similarity between the predicted binding site and any reference binding site satisfies the second condition and the structural similarity between the predicted binding site and the reference binding site satisfies the third condition. The positive category is used to indicate that the predicted binding site is consistent with the binding situation of the antigen and the antibody under actual conditions; based on the categories of multiple binding sites, determine the rationality information of the complex.
[0150] In some embodiments, see Figure 7 The fourth determination unit 6034 is also used for, for any binding site, when the sequence similarity between the binding site and any reference binding site satisfies the second condition, determining that the category of the binding site is a negative class, where the negative class is used to indicate that the binding site does not conform to the binding situation of the antigen and the antibody under actual conditions; for any binding site, when the sequence similarity between the binding site and any reference binding site satisfies the second condition, if the structural similarity between the binding site and the reference binding site does not satisfy the third condition, determining that the category of the binding site is a negative class.
[0151] In some embodiments, see Figure 7 The fourth determination unit 6034 is used to determine that the rationality information of the complex is reasonable when the number of predicted binding sites of the positive class meets the fourth condition, and reasonable is used to indicate that the complex conforms to the binding situation of the antigen and the antibody under actual conditions; when the number of binding sites of the positive class does not meet the fourth condition, the rationality information of the complex is determined to be unreasonable.
[0152] The embodiment of the present application provides a device for determining rationality information. For a complex predicted based on an antigen and an antibody, the device can determine multiple predicted binding positions on the complex according to the positions of the antigen and the antibody in the complex. Since each predicted binding position contains residues on the antigen and residues on the antibody, the predicted binding position can accurately reflect the predicted binding between the antigen and the antibody. Then, the rationality information of the complex is determined according to the structures of multiple residues in the multiple predicted binding sites. Since the residues contained in each predicted binding site and the structure of the residues are taken into consideration, the rationality information can be obtained more accurately. That is, the rationality information of the complex is determined from two aspects, namely, the residue sequence and the structure of the residues in the predicted binding site of the antigen and the antibody, which can improve the accuracy of the rationality information.
[0153] It should be noted that: the device for determining rationality information provided in the above embodiment only uses the division of the above functional modules as an example when running an application program. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the device for determining rationality information provided in the above embodiment and the method for determining rationality information belong to the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.
[0154] In the embodiments of the present application, the computer device can be configured as a terminal or a server. When the computer device is configured as a terminal, the terminal can be used as the execution subject to implement the technical solution provided in the embodiments of the present application. When the computer device is configured as a server, the server can be used as the execution subject to implement the technical solution provided in the embodiments of the present application. The technical solution provided in the present application can also be implemented through interaction between the terminal and the server. The embodiments of the present application are not limited to this.
[0155] Figure 8 800 is a block diagram of a terminal 800 according to an embodiment of the present application. The terminal 800 may be a portable mobile terminal, such as a smart phone, a tablet computer, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 player (Moving Picture Experts Group Audio Layer IV), a laptop computer or a desktop computer. The terminal 800 may also be referred to as a user device, a portable terminal, a laptop terminal, a desktop terminal or other names.
[0156] Typically, the terminal 800 includes a processor 801 and a memory 802 .
[0157] The processor 801 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 801 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 801 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 801 may be integrated with a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 801 may also include an AI (Artificial Intelligence) processor, which is used to process computing operations related to machine learning.
[0158] The memory 802 may include one or more computer-readable storage media, which may be non-transitory. The memory 802 may also include a high-speed random access memory and a non-volatile memory, such as one or more disk storage devices, flash memory storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 802 is used to store at least one computer program, which is used to be executed by the processor 801 to implement the method for determining the rationality information provided in the method embodiment of the present application.
[0159] In some embodiments, the terminal 800 may further optionally include: a peripheral device interface 803 and at least one peripheral device. The processor 801, the memory 802 and the peripheral device interface 803 may be connected via a bus or a signal line. Each peripheral device may be connected to the peripheral device interface 803 via a bus, a signal line or a circuit board. Specifically, the peripheral device includes: at least one of a radio frequency circuit 804, a display screen 805, a camera assembly 806, an audio circuit 807 and a power supply 808.
[0160] The peripheral device interface 803 can be used to connect at least one I / O (Input / Output) related peripheral device to the processor 801 and the memory 802. In some embodiments, the processor 801, the memory 802, and the peripheral device interface 803 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 801, the memory 802, and the peripheral device interface 803 can be implemented on a separate chip or circuit board, and this embodiment does not limit this.
[0161] The radio frequency circuit 804 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 804 communicates with a communication network and other communication devices through electromagnetic signals. The radio frequency circuit 804 converts an electrical signal into an electromagnetic signal for transmission, or converts a received electromagnetic signal into an electrical signal. In some embodiments, the radio frequency circuit 804 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, and so on. The radio frequency circuit 804 can communicate with other terminals through at least one wireless communication protocol. The wireless communication protocol includes but is not limited to: the World Wide Web, a metropolitan area network, an intranet, each generation of mobile communication networks (2G, 3G, 4G, and 5G), a wireless local area network, and / or a WiFi (Wireless Fidelity) network. In some embodiments, the radio frequency circuit 804 may further include a circuit related to NFC (Near Field Communication), and this application does not limit this.
[0162] The display screen 805 is used to display the UI (User Interface). The UI may include graphics, text, icons, videos, and any combination thereof. When the display screen 805 is a touch display screen, the display screen 805 also has the ability to collect touch signals on the surface or above the surface of the display screen 805. The touch signal can be input to the processor 801 as a control signal for processing. At this time, the display screen 805 can also be used to provide virtual buttons and / or virtual keyboards, also known as soft buttons and / or soft keyboards. In some embodiments, the display screen 805 can be one, set on the front panel of the terminal 800; in other embodiments, the display screen 805 can be at least two, respectively set on different surfaces of the terminal 800 or in a folding design; in other embodiments, the display screen 805 can be a flexible display screen, set on the curved surface or folding surface of the terminal 800. Even, the display screen 805 can also be set to a non-rectangular irregular shape, that is, a special-shaped screen. The display screen 805 can be made of materials such as LCD (Liquid Crystal Display), OLED (Organic Light-Emitting Diode, organic light-emitting diode).
[0163] The camera assembly 806 is used to capture images or videos. In some embodiments, the camera assembly 806 includes a front camera and a rear camera. Typically, the front camera is disposed on the front panel of the terminal, and the rear camera is disposed on the back of the terminal. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth of field camera, a wide-angle camera, and a telephoto camera, so as to realize the fusion of the main camera and the depth of field camera to realize the background blur function, the fusion of the main camera and the wide-angle camera to realize panoramic shooting and VR (Virtual Reality) shooting function or other fusion shooting functions. In some embodiments, the camera assembly 806 may also include a flash. The flash may be a monochrome temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to a combination of a warm light flash and a cold light flash, which can be used for light compensation at different color temperatures.
[0164] The audio circuit 807 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, and convert the sound waves into electrical signals and input them into the processor 801 for processing, or input them into the radio frequency circuit 804 to achieve voice communication. For the purpose of stereo acquisition or noise reduction, there may be multiple microphones, which are respectively arranged at different parts of the terminal 800. The microphone may also be an array microphone or an omnidirectional acquisition microphone. The speaker is used to convert the electrical signal from the processor 801 or the radio frequency circuit 804 into sound waves. The speaker may be a traditional film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert the electrical signal into sound waves audible to humans, but also convert the electrical signal into sound waves inaudible to humans for purposes such as ranging. In some embodiments, the audio circuit 807 may also include a headphone jack.
[0165] The power supply 808 is used to power various components in the terminal 800. The power supply 808 can be an alternating current, a direct current, a disposable battery, or a rechargeable battery. When the power supply 808 includes a rechargeable battery, the rechargeable battery can be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery charged through a wired line, and a wireless rechargeable battery is a battery charged through a wireless coil. The rechargeable battery can also be used to support fast charging technology.
[0166] In some embodiments, the terminal 800 further includes one or more sensors 809 , including but not limited to: an acceleration sensor 810 , a gyroscope sensor 811 , a pressure sensor 812 , an optical sensor 813 , and a proximity sensor 814 .
[0167] The acceleration sensor 810 can detect the magnitude of acceleration on the three coordinate axes of the coordinate system established by the terminal 800. For example, the acceleration sensor 810 can be used to detect the components of gravity acceleration on the three coordinate axes. The processor 801 can control the display screen 805 to display the user interface in a horizontal view or a vertical view according to the gravity acceleration signal collected by the acceleration sensor 810. The acceleration sensor 810 can also be used to collect game or user motion data.
[0168] The gyro sensor 811 can detect the body direction and rotation angle of the terminal 800, and the gyro sensor 811 can cooperate with the acceleration sensor 810 to collect the user's 3D actions on the terminal 800. The processor 801 can implement the following functions based on the data collected by the gyro sensor 811: motion sensing (such as changing the UI according to the user's tilt operation), image stabilization during shooting, game control, and inertial navigation.
[0169] The pressure sensor 812 can be set on the side frame of the terminal 800 and / or the lower layer of the display screen 805. When the pressure sensor 812 is set on the side frame of the terminal 800, it can detect the user's holding signal of the terminal 800, and the processor 801 performs left and right hand recognition or shortcut operation according to the holding signal collected by the pressure sensor 812. When the pressure sensor 812 is set on the lower layer of the display screen 805, the processor 801 controls the operability controls on the UI interface according to the user's pressure operation on the display screen 805. The operability controls include at least one of a button control, a scroll bar control, an icon control, and a menu control.
[0170] The optical sensor 813 is used to collect the ambient light intensity. In one embodiment, the processor 801 can control the display brightness of the display screen 805 according to the ambient light intensity collected by the optical sensor 813. Specifically, when the ambient light intensity is high, the display brightness of the display screen 805 is increased; when the ambient light intensity is low, the display brightness of the display screen 805 is reduced. In another embodiment, the processor 801 can also dynamically adjust the shooting parameters of the camera component 806 according to the ambient light intensity collected by the optical sensor 813.
[0171] The proximity sensor 814, also called a distance sensor, is usually disposed on the front panel of the terminal 800. The proximity sensor 814 is used to collect the distance between the user and the front of the terminal 800. In one embodiment, when the proximity sensor 814 detects that the distance between the user and the front of the terminal 800 is gradually decreasing, the processor 801 controls the display screen 805 to switch from the screen-on state to the screen-off state; when the proximity sensor 814 detects that the distance between the user and the front of the terminal 800 is gradually increasing, the processor 801 controls the display screen 805 to switch from the screen-off state to the screen-on state.
[0172] Those skilled in the art will understand that Figure 8 The structure shown in the figure does not constitute a limitation on the terminal 800, and the terminal 800 may include more or less components than those shown in the figure, or combine some components, or adopt a different component arrangement.
[0173] Fig. 9It is a structural diagram of a server provided according to an embodiment of the present application. The server 900 may have relatively large differences due to different configurations or performances, and may include one or more processors (Central Processing Units, CPU) 901 and one or more memories 902, wherein the memory 902 stores at least one computer program, and the at least one computer program is loaded and executed by the processor 901 to implement the determination method of rationality information provided by the above-mentioned various method embodiments. Of course, the server 900 may also have components such as a wired or wireless network interface, a keyboard, and an input and output interface for input and output. The server 900 may also include other components for realizing device functions, which will not be repeated here.
[0174] The embodiment of the present application also provides a computer-readable storage medium, in which at least one computer program is stored, and the at least one computer program is loaded and executed by a processor of a computer device to implement the operation performed by the computer device in the method for determining the rationality information of the above embodiment. For example, the computer-readable storage medium can be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), a magnetic tape, a floppy disk, an optical data storage device, etc.
[0175] The embodiment of the present application also provides a computer program product, including a computer program, which is stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the computer device executes the method for determining rationality information provided in the above various optional implementations.
[0176] A person skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware or by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, and the above-mentioned storage medium may be a read-only memory, a disk or an optical disk, etc.
[0177] The above description is only an optional embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.
Claims
1. A method for determining reasonableness information, It is characterized in that The method comprises: obtaining a complex, the complex comprising an antigen and an antibody, the complex being predicted based on the antigen and the antibody; Based on the position of the antigen and the position of the antibody in the complex, determining a plurality of predicted binding sites present on the complex, each predicted binding site comprising a residue on the antigen and a residue on the antibody; Based on the structures of multiple residues in the multiple predicted binding sites, reasonableness information of the complex is determined, and the reasonableness information is used to indicate whether the complex predicted based on the antigen and the antibody is reasonable.
2. The method according to claim 1, It is characterized in that The step of determining a plurality of predicted binding sites on the complex based on the position of the antigen and the position of the antibody in the complex comprises: Based on the antigen in the complex, obtaining positions of a plurality of first residues, wherein the plurality of first residues are residues located on the surface of the antigen; Based on the antibody in the complex, obtaining positions of a plurality of second residues, wherein the plurality of second residues include a site capable of binding to the antigen; Based on the positions of the plurality of first residues and the positions of the plurality of second residues, a plurality of predicted binding sites between the antigen and the antibody are determined.
3. The method according to claim 2, It is characterized in that The obtaining positions of a plurality of first residues based on the antigen in the complex comprises: obtaining a plurality of first residues based on the antigen in the complex; For any first residue, based on the positions of the plurality of atoms in the first residue, calculating the position of the center of mass of the plurality of atoms; The position of the centroid is taken as the position of the first residue.
4. The method according to claim 2, It is characterized in that The step of determining a plurality of predicted binding sites between the antigen and the antibody based on the positions of the plurality of first residues and the positions of the plurality of second residues comprises: For any first residue in the plurality of first residues, determining a distance between the first residue and each second residue in the plurality of second residues based on a position of the first residue and positions of the plurality of second residues; Sorting the plurality of second residues in order of distance from nearest to farthest; Based on the first residue and a preset number of second residues ranked at the top, a predicted binding site corresponding to the first residue is determined.
5. The method according to claim 1, It is characterized in that The determining of the rationality information of the complex based on the structures of the multiple residues in the multiple predicted binding sites comprises: Obtaining a comparison database, wherein the comparison database contains multiple categories of reference binding sites, wherein the reference binding sites are binding sites in a complex synthesized by an antigen and an antibody under real conditions; For any predicted binding site among the plurality of predicted binding sites, determining the sequence similarity between the predicted binding site and the reference binding sites of each category in the comparison database based on a plurality of residues in the predicted binding site and a plurality of residues in the reference binding sites of each category; Determining the structural similarity between the predicted binding site and the reference binding sites of each category based on the structures of the plurality of residues in the predicted binding site and the structures of the plurality of residues in the reference binding sites of each category in the comparison database; Based on the sequence similarities and structural similarities corresponding to the multiple predicted binding sites, the rationality information of the complex is determined.
6. The method according to claim 5, It is characterized in that Determining the structural similarity between the predicted binding site and the reference binding sites of each category based on the structure of the multiple residues in the predicted binding site and the structures of the multiple residues in the reference binding sites of each category in the comparison database comprises: For any category of reference binding sites in the comparison database, using an initial rotation matrix, processing the positions of multiple residues in the reference binding site to obtain an intermediate site; determining a distance between the intermediate site and the predicted binding site; Adjusting the initial rotation matrix with the goal of minimizing the gap; Based on the smallest of the gaps, the structural similarity between the predicted binding site and the reference binding site is determined.
7. The method according to claim 6, It is characterized in that The method further comprises: For any category of reference binding sites in the comparison database, multiple residues in the reference binding sites are matched with multiple residues between the binding sites in pairs to obtain a residue correspondence relationship; Determining the gap between the middle site and the predicted binding site comprises: For the third residue in the middle part, based on the residue correspondence, determining a fourth residue corresponding to the third residue from the predicted binding site, wherein the third residue is any residue in the middle part; Determining a residue distance between the third residue and the fourth residue based on the position of the third residue and the position of the fourth residue; Based on a plurality of said residue gaps, the gap between said middle portion and said binding portion is determined.
8. The method according to claim 5, It is characterized in that The process of constructing the comparison database includes: Acquire a plurality of sample complexes, wherein the plurality of sample complexes are complexes synthesized by antigens and antibodies under real conditions; For any sample complex, determining a plurality of sample binding sites between the antigen and the antibody based on the sample complex; Based on the structures of the multiple sample binding sites corresponding to the multiple sample complexes, clustering the multiple sample binding sites corresponding to the multiple sample complexes to obtain the multiple categories of sample binding sites; For any category, the central binding site of multiple sample binding sites in the category is used as the reference binding site corresponding to the category, and the sum of the gaps between the central binding site and other sample binding sites in the category is the smallest; The comparison database is constructed based on the reference binding sites corresponding to the multiple categories.
9. The method according to claim 8, It is characterized in that The method further comprises: For any sample complex among the plurality of sample complexes, determining the residue sequence of the sample complex based on the residue sequence of the antibody heavy chain, the residue sequence of the antibody light chain and the residue sequence of the antigen in the sample complex; Encoding the residue sequence of the sample complex to obtain a sequence identifier of the sample complex; From the plurality of sample complexes, sample complexes whose sequence identifiers have a similarity with the sequence identifier of the sample complex that satisfies a first condition are deleted.
10. The method according to claim 5, It is characterized in that The determining of the rationality information of the complex based on the sequence similarity and structural similarity corresponding to the plurality of predicted binding sites comprises: For any predicted binding site, when the sequence similarity between the predicted binding site and any reference binding site satisfies the second condition, if the structural similarity between the predicted binding site and the reference binding site satisfies the third condition, the category of the predicted binding site is determined to be a positive category, and the positive category is used to indicate that the predicted binding site conforms to the binding situation of the antigen and the antibody under real conditions; Based on the categories of the plurality of predicted binding sites, plausibility information of the complex is determined.
11. The method according to claim 10, It is characterized in that The method further comprises: For any predicted binding site, if the sequence similarity between the predicted binding site and any reference binding site does not satisfy the second condition, the category of the predicted binding site is determined to be a negative category, and the negative category is used to indicate that the binding site does not conform to the binding situation of the antigen and the antibody in the actual situation; or, For any predicted binding site, when the sequence similarity between the predicted binding site and any reference binding site satisfies the second condition, if the structural similarity between the predicted binding site and the reference binding site does not satisfy the third condition, the category of the predicted binding site is determined to be a negative class.
12. The method according to claim 10, It is characterized in that The determining of the rationality information of the complex based on the categories of the plurality of predicted binding sites comprises: When the number of predicted binding sites of the positive class satisfies the fourth condition, determining the rationality information of the complex as reasonable, wherein the rationality is used to indicate that the complex conforms to the binding situation of the antigen and the antibody under real conditions; When the number of predicted binding sites of the positive class does not satisfy the fourth condition, the rationality information of the complex is determined to be unreasonable.
13. A device for determining rationality information, It is characterized in that The device comprises: An acquisition module, which acquires a complex, wherein the complex includes an antigen and an antibody, and the complex is predicted based on the antigen and the antibody; A first determination module is used to determine a plurality of predicted binding sites on the complex based on the position of the antigen and the position of the antibody in the complex, each predicted binding site comprising a residue on the antigen and a residue on the antibody; The second determination module is used to determine the rationality information of the complex based on the structures of multiple residues in the multiple predicted binding sites, wherein the rationality information is used to indicate whether the complex predicted based on the antigen and the antibody is reasonable.
14. A computer device, It is characterized in that The computer device includes a processor and a memory, the memory is used to store at least one computer program, and the at least one computer program is loaded by the processor and executes the method for determining rationality information as described in any one of claims 1 to 12.
15. A computer-readable storage medium, It is characterized in that The computer-readable storage medium is used to store at least one computer program, and the at least one computer program is used to execute the method for determining rationality information as described in any one of claims 1 to 12.
16. A computer program product comprising a computer program, It is characterized in that When the computer program is executed by a processor, the method for determining rationality information according to any one of claims 1 to 12 is implemented.