Method, equipment, medium and program product for determining binding protein information with pH sensitivity
By generating pH-sensitive binding protein information through flow matching prediction models, the issues of targeting specificity and normal cell toxicity of binding proteins in the tumor microenvironment are resolved, thereby improving design efficiency and safety.
Patent Information
- Application Number
- CN202511736628.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-24
- Publication Date
- 2026-02-24
AI Technical Summary
Existing technologies struggle to effectively distinguish between the acidic conditions of the tumor microenvironment and normal cells when designing binding proteins with high affinity and specificity, leading to adverse effects on normal tissues and cytotoxicity, resulting in low design efficiency.
By using a trained binding protein generation model, based on target protein information, binding region amino acid information, and noise information, a flow matching prediction model is used to generate acid-base sensitive binding protein information. The noise information is gradually updated to meet preset conditions, thereby generating an ideal binding protein sequence and structure.
It improves the targeting specificity of binding proteins in the tumor microenvironment, reduces toxicity to normal cells, enhances the efficiency and safety of immunotherapy design, and lowers the professional requirements for designers.
Smart Images

Figure CN121565264A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of bioinformatics technology, and more particularly to a technique for determining information about binding proteins that are sensitive to acidity and alkalinity. Background Technology
[0002] Designing binding proteins with high affinity and specificity, enabling them to precisely recognize and target receptors within complex proteomes, is crucial for regulating immune responses or proliferation signaling, and holds broad application prospects in tumor immunotherapy. However, some receptors are also expressed in normal cells; if binding proteins bind extensively to receptors on non-target cells, it may lead to adverse effects on normal tissues and cytotoxicity.
[0003] The tumor microenvironment, due to its active aerobic glycolysis, often accumulates lactic acid and CO2, creating a locally weakly acidic environment. Specifically, the pH of normal tissue is approximately 7.2–7.4, while the pH of the tumor microenvironment is typically between 6.5–6.8, and can even be as low as 6.2. Constructing binding proteins with pH-sensitive properties could potentially further improve targeting specificity; while enhancing anti-tumor immune responses, it could significantly reduce toxicity to normal cells, thereby improving the efficiency and safety of immunotherapy design. Existing methods mainly involve protein backbone sampling, protein sequence design, and validation screening of the designed sequences and structures. The parameter settings and screening criteria involved in each step rely on the designer's experience and judgment. This requires a high level of expertise from the designer and results in relatively low design efficiency. Summary of the Invention
[0004] One object of this application is to provide a method, apparatus, medium, and procedure for determining information about binding proteins that are sensitive to acidity and alkalinity.
[0005] According to one aspect of this application, a method for determining information about acid-base sensitive binding proteins is provided, the method comprising:
[0006] Based on target protein information, the target amino acid corresponding to the amino acid in the binding region of the target protein, and the noise information of the binding protein corresponding to the target protein in the initial time step, the binding protein information corresponding to the target protein is generated through the trained first binding protein generation model. The target protein information includes target protein sequence information and target protein structure information. The target amino acid and the amino acid in the binding region of the target protein match the preset acid-base sensitivity information. The binding protein information includes binding protein sequence information and binding protein structure information. The binding protein sequence information contains the target amino acid.
[0007] The first binding protein generation model includes a first-order matching prediction model, and the generation of binding protein information corresponding to the target protein through the trained first binding protein generation model includes:
[0008] Based on the target protein information and the noise information of the binding protein corresponding to the target protein in the initial time step, the first velocity field corresponding to the initial time step is generated by the first flow matching prediction model.
[0009] Based on the first velocity field, determine the noise information corresponding to the binding protein in the next time step;
[0010] Based on the noise information of the binding protein in the next time step, and in conjunction with the target amino acid, update the noise information of the binding protein in the next time step; determine whether a preset first stopping condition is met;
[0011] If the first stopping condition is not met, based on the noise information of the binding protein and the target protein information in the updated next time step, the first velocity field corresponding to the next time step is generated again through the first flow matching prediction model, and steps i and j are repeated.
[0012] If the first stopping condition is met, the binding protein information is determined.
[0013] According to one aspect of this application, a computer device for determining information about binding proteins with acid-base sensitivity is provided, comprising a memory, a processor, and a computer program stored in the memory, characterized in that the processor executes the computer program to implement the steps of any of the methods described above.
[0014] According to one aspect of this application, a computer-readable storage medium is provided having a computer program stored thereon, characterized in that the computer program, when executed by a processor, implements the steps of any of the methods described above.
[0015] According to one aspect of this application, a computer program product is provided, comprising a computer program, characterized in that, when executed by a processor, the computer program implements the steps of any of the methods described above.
[0016] According to one aspect of this application, an apparatus is provided for determining information about binding proteins that are sensitive to acidity and alkalinity, the apparatus comprising:
[0017] A module is used to generate binding protein information corresponding to the target protein based on target protein information, the target amino acid corresponding to the binding region of the target protein, and the noise information of the binding protein corresponding to the target protein in the initial time step, through a trained first binding protein generation model. The target protein information includes target protein sequence information and target protein structure information. The target amino acid and the amino acid of the binding region of the target protein are matched with preset acid-base sensitivity information. The binding protein information includes binding protein sequence information and binding protein structure information. The binding protein sequence information contains the target amino acid.
[0018] The first binding protein generation model includes a first-order matching prediction model, and the generation of binding protein information corresponding to the target protein through the trained first binding protein generation model includes:
[0019] Based on the target protein information and the noise information of the binding protein corresponding to the target protein in the initial time step, the first velocity field corresponding to the initial time step is generated by the first flow matching prediction model.
[0020] Based on the first velocity field, determine the noise information corresponding to the binding protein in the next time step;
[0021] Based on the noise information of the binding protein in the next time step, and in conjunction with the target amino acid, update the noise information of the binding protein in the next time step; determine whether a preset first stopping condition is met;
[0022] If the first stopping condition is not met, based on the noise information of the binding protein and the target protein information in the updated next time step, the first velocity field corresponding to the next time step is generated again through the first flow matching prediction model, and steps i and j are repeated.
[0023] If the first stopping condition is met, the binding protein information is determined.
[0024] Compared with existing technologies, this application generates binding protein information corresponding to the target protein based on target protein information, the target amino acid corresponding to the binding region of the target protein, and the noise information of the binding protein corresponding to the target protein in the initial time step, through a trained first binding protein generation model. The target protein information includes target protein sequence information and target protein structure information; the target amino acid and the amino acid of the binding region of the target protein match preset acid-base sensitivity information; the binding protein information includes binding protein sequence information and binding protein structure information; the binding protein sequence information contains the target amino acid; and the first binding protein generation model includes a first-stream matching prediction model. The training of the first binding protein generation model generates the binding protein information corresponding to the target protein. The process includes: i) generating a first velocity field corresponding to the initial time step based on target protein information and noise information of the binding protein corresponding to the target protein in the initial time step using a first flow matching prediction model; j) determining the noise information corresponding to the binding protein in the next time step based on the first velocity field; j) updating the noise information of the binding protein in the next time step based on the noise information of the binding protein in the next time step, combined with the target amino acid; determining whether a preset first stopping condition is met; if the first stopping condition is not met, generating the first velocity field corresponding to the next time step again using the first flow matching prediction model based on the updated noise information of the binding protein in the next time step and the target protein information, and repeating steps i and j; if the first stopping condition is met, determining the binding protein information. The flow matching method predicts the velocity field from noise data to binding protein information, generating ideal binding protein information, and introducing corresponding target amino acids in each prediction step to improve the pH sensitivity of the binding protein, thereby obtaining a pH-sensitive binding protein. This reduces the experience requirements for designers designing pH-sensitive binding proteins and improves design efficiency. Attached Figure Description
[0025] Other features, objects, and advantages of this application will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings:
[0026] Figure 1 A flowchart illustrating a method for determining information about acid-base sensitive binding proteins according to an embodiment of this application is shown.
[0027] Figure 2 This diagram illustrates a device structure for determining information about a binding protein with pH sensitivity, according to one embodiment of this application.
[0028] Figure 3 Exemplary systems that can be used to implement the various embodiments described in this application are shown.
[0029] The same or similar reference numerals in the accompanying drawings represent the same or similar parts. Detailed Implementation
[0030] The present application will now be described in further detail with reference to the accompanying drawings.
[0031] In a typical configuration of this application, the terminal, the device of the service network, and the trusted party all include one or more processors (e.g., a central processing unit (CPU)), input / output interfaces, network interfaces, and memory.
[0032] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash memory. Memory is an example of computer-readable media.
[0033] Computer-readable media, including both permanent and non-permanent, removable and non-removable media, can store information using any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PCM), programmable random access memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.
[0034] The devices referred to in this application include, but are not limited to, user equipment, network equipment, or devices composed of user equipment and network equipment integrated through a network. The user equipment includes, but is not limited to, any mobile electronic product capable of human-computer interaction (e.g., via a touchpad), such as smartphones and tablets. These mobile electronic products can use any operating system, such as Android or iOS. The network equipment includes an electronic device capable of automatically performing numerical calculations and information processing according to pre-set or stored instructions. Its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), and embedded devices. The network equipment includes, but is not limited to, computers, network hosts, single network servers, multiple network server clusters, or clouds composed of multiple servers. Here, a cloud consists of a large number of computers or network servers based on cloud computing, where cloud computing is a type of distributed computing, consisting of a virtual supercomputer composed of a group of loosely coupled computer clusters. The network includes, but is not limited to, the Internet, wide area network, metropolitan area network, local area network, VPN network, wireless ad hoc network, etc. Preferably, the device can also be a program running on the user equipment, network device, or a device formed by integrating user equipment and network device, network device, touch terminal, or network device and touch terminal through a network.
[0035] Of course, those skilled in the art should understand that the above-described devices are merely examples, and other existing or future devices that are applicable to this application should also be included within the scope of protection of this application, and are hereby incorporated by reference.
[0036] In the description of this application, "multiple" means two or more, unless otherwise expressly and specifically defined.
[0037] Figure 1The flowchart illustrates a method for determining acid-base sensitive binding protein information according to an embodiment of this application. The method includes step S11. In step S11, based on target protein information, the target amino acid corresponding to the binding region amino acid of the target protein, and noise information of the binding protein corresponding to the target protein in the initial time step, a trained first binding protein generation model is used to generate binding protein information corresponding to the target protein. The target protein information includes target protein sequence information and target protein structure information. The target amino acid and the binding region amino acid of the target protein match preset acid-base sensitivity information. The binding protein information includes binding protein sequence information and binding protein structure information. The binding protein sequence information contains the target amino acid. The first binding protein generation model includes a first flow matching prediction model. Generating the binding protein information corresponding to the target protein using the trained first binding protein generation model includes: Based on the target protein information and the noise information of the binding protein corresponding to the target protein in the initial time step, a first velocity field corresponding to the initial time step is generated through a first-flow matching prediction model; i. Based on the first velocity field, the noise information corresponding to the binding protein in the next time step is determined; j. Based on the noise information of the binding protein in the next time step, combined with the target amino acid, the noise information of the binding protein in the next time step is updated; it is determined whether a preset first stopping condition is met; if the first stopping condition is not met, based on the updated noise information of the binding protein in the next time step and the target protein information, the first velocity field corresponding to the next time step is generated again through the first-flow matching prediction model, and steps i and j are repeated; if the first stopping condition is met, the binding protein information is determined.
[0038] In some embodiments, the target amino acid is determined based on the binding region amino acids of the target protein and preset pH sensitivity information. The preset pH sensitivity information is the user-expected pH sensitivity direction between the target protein and the designed binding protein. The preset pH sensitivity information includes pH information indicating a relatively high affinity between the target protein and the binding protein, and pH information indicating a relatively low affinity between the target protein and the binding protein. This pH information can be a specific pH value or a corresponding pH range. For example, pH 7.4 can be set as high affinity and pH 6.5 as low affinity. Alternatively, pH 7.35–7.45 can be set as high affinity and pH 6.0–7.0 as low affinity.
[0039] In some embodiments, the method further includes: step S12 (not shown), determining an amino acid pair set based on the preset acid-base sensitivity information, wherein the amino acid pair set includes one or more amino acid pairs that match the acid-base sensitivity information; step S13 (not shown), determining a corresponding target amino acid based on the binding region amino acid of the target protein, wherein the binding region amino acid of the target protein and the target amino acid belong to an amino acid pair in the amino acid pair set.
[0040] For example, based on the pH sensitivity information, one or more amino acid pairs that can promote proton transfer under pH conditions with relatively high affinity or inhibit proton transfer under pH conditions with relatively low affinity can be identified to form a corresponding amino acid pair set.
[0041] In some embodiments, step S12 includes: determining one or more first amino acids whose acidity coefficients match the preset acid-base sensitivity information based on the acidity coefficients corresponding to each amino acid; determining one or more second amino acids that match each first amino acid based on the acid-base sensitivity information, and then determining an amino acid pair set, wherein the amino acid pair set includes amino acid pairs composed of each first amino acid and a matching second amino acid.
[0042] Specifically, the preset pH sensitivity information includes two pH values: one for the target protein with a relatively high affinity to the binding protein, and the other for the target protein with a relatively low affinity to the binding protein. Based on the acidity coefficient (pKa) of each amino acid, one or more first amino acids whose pKa is closest to or between these two pH values can be determined. This first amino acid can serve as a protonation / deprotonation source for providing pH sensitivity. For example, consider pH 7.4 as high affinity and pH 6.0 as low affinity. The pKa values for each amino acid are: aspartic acid (ASP): 3.9, glutamic acid (GLU): 4.3, histidine (HIS): 6.0, cysteine (CYS): 8.3, tyrosine (TYR): 10.1, lysine (LYS): 10.5, and arginine (ARG): 12.5. Therefore, for the above pH sensitivity information, histidine, whose pKa is closest to the low affinity pH value, can be used as the first amino acid.
[0043] Based on the pH sensitivity information, specifically the pH information indicating a high affinity between the target protein and its binding protein, and the pH information indicating a low affinity between the target protein and its binding protein, corresponding neutral conditions and protonation / deprotonation conditions are determined. For example, the pH information closer to neutral can be used as the neutral condition, and the other pH information can be used as the protonation / deprotonation condition. A binding interface for the first amino acid under neutral conditions is constructed, and a corresponding second amino acid is selected to enhance or weaken the affinity under protonation / deprotonation conditions. If the pH information used as the protonation / deprotonation condition indicates a high affinity between the target protein and its binding protein, then a second amino acid that enhances the affinity under protonation / deprotonation conditions needs to be determined. That is, a second amino acid that can form an amino acid pair that promotes proton transfer needs to be determined. If the pH information used as the protonation / deprotonation condition indicates a low affinity between the target protein and its binding protein, then a second amino acid that weakens the affinity under protonation / deprotonation conditions needs to be determined. This involves identifying the second amino acid that can form an amino acid pair that inhibits proton transfer. For example, for histidine identified in the previous example, under protonation / deprotonation conditions (i.e., pH=6.0), histidine is positively charged. To weaken the affinity, a positively charged amino acid needs to be introduced opposite it, i.e., arginine or lysine. Thus, the amino acid pairs histidine and arginine, and histidine and lysine, can be identified. As another example, if pH 7.4 is set as low affinity and pH 6.0 as high affinity, for the previously identified first amino acid histidine, under protonation / deprotonation conditions (i.e., pH=6.0), to enhance the affinity, a negatively charged amino acid needs to be introduced opposite it, i.e., aspartic acid or glutamic acid. Thus, the amino acid pairs histidine and aspartic acid, and histidine and glutamic acid, can be identified.
[0044] Based on the amino acid in the binding region of the target protein, a search is conducted within a set of amino acid pairs to determine the corresponding amino acid pair, where the amino acid in the binding region of the target protein belongs to that pair. The other amino acid in that pair is then identified as the target amino acid corresponding to the amino acid in the binding region of the target protein. For example, suppose the set of amino acid pairs determined in the preceding steps includes histidine and arginine, or histidine and lysine. If the amino acid in the binding region of the target protein is determined to be histidine, then the target amino acid can be determined to be either arginine or lysine. If the amino acid in the binding region of the target protein is determined to be either arginine or lysine, then the target amino acid can be determined to be histidine.
[0045] In some embodiments, the first binding protein generation model includes a first flow matching prediction model. The first binding protein generation model is trained based on flow matching and corresponding sample data. Flow matching is a widely used generative model in recent years, which learns a deterministic velocity field from a noise distribution to a data distribution through a neural network, thereby generating ideal data prediction results. For the first binding protein generation model, it can gradually generate noise-free binding protein information from the pure noise information of the binding protein corresponding to the initial time step, combined with target protein information and the target amino acid information.
[0046] In some embodiments, the target protein information includes target protein structural information and target protein sequence information. The target protein structural information can be obtained by experimentally determining the protein crystal structure or by predicting it using second protein structure prediction models such as Alphafold3 or Boltz-2. The target protein structural information includes the coordinate information corresponding to each amino acid in the target protein. (For example, the three-dimensional coordinates of the C-alpha atoms corresponding to each amino acid can be used), and the rotation matrix between the local reference frame of each amino acid (the reference frame formed by the main chain atoms of that amino acid) and the global reference frame. Where i represents the i-th amino acid in the target protein. The overall reference frame is a reference frame consisting of the main chain atoms of one amino acid in the target protein. Typically, the first amino acid is chosen, but any amino acid in the target protein can be selected based on user settings. SO(3) refers to a combination of rotations in three-dimensional space, i.e., multiple consecutive rotations. This is because in the flow matching model, a rotation occurs at each time step, meaning a combination of multiple rotations. The target protein sequence information includes the amino acid types corresponding to each amino acid in the target protein. The information pertains to the relative positions of amino acid sequences. The amino acid type information includes one of 20 amino acids and a mask type M. The relative positions of the amino acid sequences are determined by sinusoidal position encoding, showing the relative positions of each amino acid within its respective chain.
[0047] In some embodiments, similar to the target protein information, the noise information of the binding protein includes binding protein structural noise information and binding protein sequence noise information. The binding protein structural noise information includes coordinate noise information corresponding to each amino acid in the binding protein. Rotation matrix between the local reference frame and the global reference frame for each amino acid The binding protein sequence noise information includes the noise information of the amino acid types corresponding to each amino acid in the binding protein. The relative position information of the amino acid sequence. Here, j represents the j-th amino acid in the binding protein. For the initial time step t=0, the noise information of the binding protein is pure noise data. The coordinate noise information corresponding to each amino acid in the binding protein can be obtained by sampling from white noise at the initial time step. For the rotation matrix between the local reference frame and the global reference frame of each amino acid in the binding protein, it can be obtained by sampling from the uniform distribution of the rotation matrix of SO(3) to obtain the rotation matrix between the local reference frame and the global reference frame of each amino acid in the binding protein in the initial time step. For the amino acid type noise information corresponding to each amino acid in the binding protein in the initial time step, the amino acid type noise information corresponding to each amino acid can be set to mask type M based on the expected sequence length.
[0048] The target protein information, the target amino acids corresponding to the binding region of the target protein, and the noise information of the binding protein in the initial time step are input into the first binding protein generation model. The model progressively removes noise to generate a noise-free result corresponding to the binding protein, i.e., the corresponding binding protein information. Specifically, the first flow matching prediction model in the first binding protein generation model can generate a first velocity field corresponding to the current time step based on the target protein information and the noise information of the binding protein corresponding to the target protein in the initial time step. This first velocity field includes the coordinate velocity fields of each amino acid corresponding to the current time step t. And the rotation matrix velocity field corresponding to each amino acid. Therefore, based on this first velocity field, the coordinate noise information of each amino acid in the binding protein at the next time step t+Δt can be calculated. And the rotation matrix between the local reference frame and the global reference frame for each amino acid. Δt is the set time step size. The amino acid type noise information corresponding to each amino acid in the binding protein is a discrete feature. The first binding protein generation model predicts the conditional probability distribution of amino acid types in the binding protein after noise reduction (i.e., at t=1, i.e., the noise-free time) based on the target protein information and the noise information of the binding protein at the current time step t. .in, The binding protein noise information corresponding to the current time step t, This refers to the number of amino acids contained in the binding protein; Corresponding target protein information, This represents the number of amino acids contained in the target protein. The conditional probability distribution of each amino acid type corresponds to a 21-dimensional probability vector, representing the probability of each amino acid being represented. This can be based on the conditional probability distribution of each amino acid in the denoised binding protein. With mask distribution PM (i.e., mask type probability is 1, other types are 0), determine the velocity field of the probability distribution evolution of each amino acid in the binding protein at the current time step t. = -P M The first velocity field also includes the velocity field evolved from this probability distribution, and the conditional probability distribution of each amino acid in the binding protein at the corresponding next time step. The coordinate noise information corresponding to each amino acid in the binding protein in the next time step, the rotation matrix between the local reference frame and the global reference frame of each amino acid, the conditional probability distribution of the amino acid type of each amino acid, and the corresponding relative position of the amino acid sequence are used as the noise information corresponding to the binding protein in the next time step.
[0049] To make the model's predicted binding proteins acid-base sensitive, in each prediction step, the corresponding target amino acid is introduced at the corresponding position of the binding protein corresponding to the amino acid in the target protein's binding region, thereby updating the noise information of the binding protein in the next time step. This means that the amino acid in the binding protein closest to the target protein's binding region is forcibly modified to be the target amino acid. For example, if the target amino acid is arginine, then the amino acid type information of the amino acid in the binding protein closest to the target protein's binding region is changed to arginine. That is, the probability of arginine is 1, and other types are 0. As in the example above, if histidine is found in the binding region of the target protein, the corresponding target amino acid can be determined to be either arginine or lysine. For cases where there are multiple selectable target amino acids, any one can be selected for introduction. Furthermore, the type of the selected target amino acid is not changed in subsequent prediction steps. For example, if arginine is selected for introduction, then arginine will be selected for introduction in subsequent time steps.
[0050] Then, it is determined whether the first stopping condition is met. In some embodiments, the first stopping condition includes at least one of the following: the number of running time steps reaches a preset number of time steps; the difference between the determined first velocity field and the first velocity field determined in the previous time step is less than a first threshold; the difference between the noise information of the updated binding protein and the noise information of the updated binding protein in the previous time step is less than a second threshold. Whether to stop is determined by the number of iterations or whether the predicted first velocity field or the prediction result converges. If the first stopping condition is met, the noise information of the currently updated binding protein is determined as the binding protein information output. If the first stopping condition is not met, the first velocity field corresponding to the next time step is generated by the first flow matching prediction model, and the above calculation process is repeated to obtain the updated noise information of the binding protein.
[0051] In some embodiments, the method further includes: step S14 (not shown), generating corresponding sample binding protein prediction information based on the noise information of the sample binding protein in the sampling time step and the corresponding sample target protein information through the first binding protein generation model; determining a first loss between the sample binding protein information corresponding to the sample binding protein and the sample binding protein prediction information based on a preset first loss function; and updating the first binding protein generation model based on the first loss through a backpropagation algorithm to obtain a trained first binding protein generation model.
[0052] The noise information of the sample-binding protein is similar to that of the aforementioned binding protein, and also includes structural noise information and sequence noise information of the sample-binding protein. The structural noise information of the sample-binding protein includes the coordinate noise information corresponding to each amino acid in the sample-binding protein, and the rotation matrix between the local reference frame and the global reference frame for each amino acid. The sequence noise information of the sample-binding protein includes the amino acid type noise information and the relative position information of the amino acid sequence for each amino acid in the sample-binding protein. The sample target protein information and sample-binding protein information are similar to the aforementioned target protein information and binding protein information, respectively, and therefore will not be repeated hereafter, but are included by reference.
[0053] It is generally assumed that the data at t=1 is pure and noise-free, while the data at t=0 is pure noise. For model training, one can... Random sampling is performed within a uniform distribution as the sampling time step. Then, based on the sampling time, combined with the corresponding sample-binding protein information and the pure noise information corresponding to the sample-binding protein, corresponding noise is introduced to determine the noise information of the sample-binding protein in that sampling time step. The determination of the pure noise information corresponding to the sample-binding protein is similar to the determination of the noise information of the binding protein in the initial time step in step S11 above. Therefore, it will not be repeated here, but is included by reference. The pure noise information corresponding to the sample-binding protein includes the coordinate pure noise information corresponding to each amino acid in the sample-binding protein. Pure noise information of the rotation matrix between the local reference frame and the global reference frame of each amino acid in the sample-binding protein. Accordingly, the sample-binding protein information includes the coordinate information corresponding to each amino acid in the sample-binding protein. Rotation matrix between the local reference frame and the global reference frame for each amino acid in the sample-binding protein. Information on the amino acid types corresponding to each amino acid in the protein that binds to the sample. For sampling time step t, the coordinate noise information corresponding to each amino acid in the sample-binding protein is as follows:
[0054]
[0055] The rotation matrix between the local reference frame and the global reference frame for each amino acid is as follows:
[0056]
[0057] The amino acid type noise information corresponding to each amino acid in the sample-binding protein is as follows:
[0058]
[0059] The δ function satisfies the property that if a = b, then δ(a,b) = 1; otherwise, δ(a,b) = 0. Furthermore, by combining the relative positions of the amino acid sequences, the noise information of the protein binding to the sample in that sampling time step is determined and input into the model for prediction.
[0060] Here, the process of generating corresponding sample binding protein prediction information through the first binding protein generation model is similar to the process of generating binding protein information in step S11 above. However, since introducing target amino acids during model prediction exposes the amino acid sequence information to be learned, and considering that most of the obtained sample target protein / sample binding protein complex interfaces do not have strong pH sensitivity, the step of introducing target amino acids is not included in the model training process to ensure the quality of model learning. That is, during training, the model does not execute the process of binding target amino acids and updating the noise information of the sample binding protein in the next time step as described in step j above, but directly determines whether the first stopping condition is met. Other processes are similar to step S11 above, and therefore will not be repeated here, but are included by reference.
[0061] For the predicted sample-binding protein information, a first loss relative to the sample-binding protein information can be calculated using a first loss function. The gradient of the model parameters is then calculated using a backpropagation algorithm, and subsequently, the first binding protein generation model is updated using a corresponding optimization algorithm (e.g., stochastic gradient descent, AdaGrad, or RMSProp). In some embodiments, the first loss function includes the mean squared error (MSE) loss function of the coordinates corresponding to each amino acid in the binding protein, the mean squared error (MSE) loss function of the rotation matrix corresponding to each amino acid in the binding protein, and the cross-entropy loss function of the amino acid types in the binding protein.
[0062] In some embodiments, the method further includes: step S15 (not shown), generating binding protein structure information and a corresponding quality score under the pH information based on the corresponding pH information and the binding protein structure information using a trained second binding protein generation model; wherein, the second binding protein generation model includes a second flow matching prediction model, and generating binding protein structure information and a corresponding quality score under the pH information using the trained second binding protein generation model includes: generating a second velocity field corresponding to the current time step using the second flow matching prediction model based on the pH information and the binding protein structure information in the current time step; determining the binding protein structure information corresponding to the next time step based on the second velocity field; determining whether a preset second stopping condition is met; if the second stopping condition is not met, continuing to generate the second velocity field corresponding to the next time step using the second flow matching prediction model based on the binding protein structure information corresponding to the next time step and the pH information; repeating step k; if the second stopping condition is met, determining the binding protein structure information and the corresponding quality score under the pH information.
[0063] The first binding protein generation model predicts the structure of the binding protein unaffected by pH. It can incorporate pH information that allows protonation / deprotonation of the target amino acid in the binding protein sequence. Based on the second binding protein generation model, the predicted binding protein structure information is optimized to generate corresponding binding protein structure information and a quality score. Similar to the first model, the second binding protein generation model is trained using flow matching with relevant sample data. It includes a second flow matching prediction model. Its prediction process is similar to the first model, predicting the corresponding second velocity field. This second velocity field includes the coordinate velocity field of each amino acid and the rotation matrix velocity field corresponding to each amino acid. The binding protein structure information for the next time step is determined based on the second velocity field, and it is determined whether the second stopping condition is met. The second stopping condition is similar to the first stopping condition, including but not limited to: the number of running time steps reaching a preset number of time steps; the difference between the determined second velocity field and the second velocity field determined in the previous time step being less than a third threshold; and the difference between the determined binding protein structure information and the binding protein structure information determined in the previous time step being less than a fourth threshold. If the second stopping condition is met, the currently determined binding protein structure information is determined as the binding protein structure information under the pH information, and the corresponding quality score is determined. If the second stopping condition is not met, the second velocity field corresponding to the next time step is generated through the second flow matching prediction model, and the above calculation process is repeated to determine the binding protein structure information.
[0064] In some embodiments, generating the binding protein information corresponding to the target protein based on target protein information, the target amino acid corresponding to the binding region of the target protein, and the noise information of the binding protein corresponding to the target protein in the initial time step, using a trained first binding protein generation model, includes: determining one or more sequence lengths corresponding to the binding protein based on the target protein information; for each sequence length, generating the binding protein information corresponding to that sequence length using a trained first binding protein generation model based on the target protein information, the target amino acid corresponding to the binding region of the target protein, and the noise information of the binding protein corresponding to the target protein in the initial time step; the step based on the corresponding pH information, and the... The method further includes: generating binding protein structure information and corresponding quality scores under the pH information using a trained second binding protein generation model, based on the corresponding pH information and the binding protein structure information corresponding to each sequence length, using a trained second binding protein generation model; the method also includes: step S16 (not shown), based on the quality scores, preferentially determining corresponding binding protein sequence information and corresponding binding protein structure information from the binding protein sequence information corresponding to one or more sequence lengths and the corresponding binding protein structure information under the pH information.
[0065] Specifically, for the target protein, users can try applying different binding protein sequence lengths for prediction. Since the sequence length of the binding protein must be sufficient to form a complete interface for interaction with the target protein, one or more sequence lengths corresponding to the binding protein can be determined based on the sequence length of the binding region corresponding to the target protein. For each sequence length, the binding protein sequence information, corresponding binding protein structural information, and quality score can be predicted according to the methods described in steps S11 and S15 above. Then, the binding protein sequence information and corresponding binding protein structural information corresponding to the sequence length with the optimal quality score can be selected.
[0066] In some embodiments, the method further includes: step S17 (not shown), based on the sample binding protein structure information in the sample binding protein information and the corresponding sample pH information, generating sample binding protein structure prediction information and corresponding prediction quality scores under the sample pH information using a second binding protein generation model; generating sample binding protein structure information under the sample pH information using a first protein structure prediction model based on the sample binding protein structure information and the sample pH information; determining the corresponding quality score based on the sample binding protein structure information under the sample pH information and the sample binding protein structure prediction information under the sample pH information; determining a second loss between the sample binding protein structure prediction information under the sample pH information and the sample binding protein structure information under the sample pH information, and between the prediction quality score and the quality score based on a preset second loss function; updating the second binding protein generation model using a backpropagation algorithm based on the second loss to obtain a trained second binding protein generation model.
[0067] Here, the process of generating the sample binding protein structure prediction information and the corresponding prediction quality score under the sample pH information using the second binding protein generation model is similar to the process in step S15 described above. Therefore, it will not be repeated here, but is incorporated herein by reference. In some embodiments, the first protein structure prediction model includes a Rosetta model. The sample binding protein structure information under the corresponding sample pH information is obtained using the Rosetta model. That is, protonated / deprotonated binding protein structures are introduced. The corresponding quality score can also be determined based on the similarity between the sample binding protein structure prediction information under the sample pH information and the sample binding protein structure information under the sample pH information. The higher the similarity, the higher the quality score. Then, the second loss of the sample binding protein structure prediction information and the corresponding prediction quality score under the sample pH information relative to the Rosetta model generation result can be determined based on the second loss function. The gradient of the model parameters is calculated by the backpropagation algorithm, and then the second binding protein generation model is updated by combining the corresponding optimization algorithm. The second loss function includes the mean squared error loss function.
[0068] In some embodiments, the method further includes: step S18 (not shown), obtaining complex information corresponding to the sample protein complex, wherein the sample protein complex contains multiple chains; taking each chain in the sample protein complex as a sample binding protein, and taking the remaining chain in the sample protein complex as the sample target protein corresponding to the sample binding protein, determining the corresponding sample binding protein information and the corresponding sample target protein information. For example, a user can obtain the complex information corresponding to the sample protein complex from a protein data bank (PDB). These sample protein complexes contain at least two chains. For each sample protein complex, each chain can be traversed and taken as the sample binding protein, and the remaining chain as the sample target protein to obtain the corresponding sample binding protein information and the corresponding sample target protein information. For example, if a sample protein complex has three chains: chain A, chain B, and chain C, then three pairs of samples can be obtained, namely, sample target protein chain A + chain B and sample binding protein chain C, sample target protein chain A + chain C and sample binding protein chain B, and sample target protein chain B + chain C and sample binding protein chain A. This allows the model to handle cases where the target protein has multiple chains, resulting in better model generalization performance.
[0069] In some embodiments, the complex information includes complex sequence information and complex structure information. Obtaining the complex information corresponding to the sample protein complex includes: determining the corresponding complex structure information based on the complex sequence information using a second protein structure prediction model. For cases where only the complex sequence information is available, the corresponding complex structure information can be predicted using second protein structure prediction models such as Alphafold3 or Boltz-2 to obtain the complex information, thereby determining the corresponding sample-binding protein information and the corresponding sample target protein information.
[0070] Figure 2This diagram illustrates a device structure for determining acid-base sensitive binding protein information according to an embodiment of this application. The device 1 includes a module 11. The module 11 generates binding protein information corresponding to the target protein based on target protein information, target amino acids corresponding to the binding region amino acids of the target protein, and noise information of the binding protein corresponding to the target protein in the initial time step, using a trained first binding protein generation model. The target protein information includes target protein sequence information and target protein structure information; the target amino acids and the binding region amino acids of the target protein match preset acid-base sensitivity information; the binding protein information includes binding protein sequence information and binding protein structure information; the binding protein sequence information contains the target amino acids; the first binding protein generation model includes a first flow matching prediction model; generating the binding protein information corresponding to the target protein using the trained first binding protein generation model includes: Based on the target protein information and the noise information of the binding protein corresponding to the target protein in the initial time step, a first velocity field corresponding to the initial time step is generated through a first-flow matching prediction model; i. Based on the first velocity field, the noise information corresponding to the binding protein in the next time step is determined; j. Based on the noise information of the binding protein in the next time step, combined with the target amino acid, the noise information of the binding protein in the next time step is updated; it is determined whether a preset first stopping condition is met; if the first stopping condition is not met, based on the updated noise information of the binding protein in the next time step and the target protein information, the first velocity field corresponding to the next time step is generated again through the first-flow matching prediction model, and steps i and j are repeated; if the first stopping condition is met, the binding protein information is determined. Here, the... Figure 2 The specific implementation of module 11 shown is the same as or similar to the specific embodiment of step S11 described above, so it will not be repeated here, but is included by reference.
[0071] In some embodiments, the device 1 further includes a second module 12 (not shown) and a third module 13 (not shown). The second module 12 determines an amino acid pair set based on the preset pH sensitivity information, wherein the amino acid pair set includes one or more amino acid pairs matching the pH sensitivity information; the third module 13 determines the corresponding target amino acid based on the binding region amino acid of the target protein, wherein the binding region amino acid of the target protein and the target amino acid belong to the amino acid pair in the amino acid pair set. Here, the specific implementations of the second module 12 and the third module 13 are the same as or similar to the specific embodiments of steps S12 and S13 described above, and therefore will not be repeated here, but are incorporated herein by reference.
[0072] In some embodiments, the device 1 further includes a four-module 14 (not shown). The four-module 14 generates corresponding sample-binding protein prediction information based on noise information of the sample-binding protein in the sampling time step and the corresponding sample target protein information, using the first binding protein generation model; determines a first loss between the sample-binding protein information corresponding to the sample-binding protein and the sample-binding protein prediction information based on a preset first loss function; and updates the first binding protein generation model based on the first loss using a backpropagation algorithm to obtain a trained first binding protein generation model. Here, the specific implementation of the four-module 14 is the same as or similar to the specific embodiment of the aforementioned step S14, and therefore will not be repeated, but is incorporated herein by reference.
[0073] In some embodiments, the device 1 further includes a five-module 15 (not shown). The five-module 15, based on the corresponding pH information and the binding protein structure information, generates binding protein structure information and a corresponding quality score under the pH information using a trained second binding protein generation model. The second binding protein generation model includes a second flow matching prediction model. Generating the binding protein structure information and the corresponding quality score under the pH information using the trained second binding protein generation model includes: generating a second velocity field corresponding to the current time step based on the pH information and the binding protein structure information in the current time step using the second flow matching prediction model; determining the binding protein structure information corresponding to the next time step based on the second velocity field; determining whether a preset second stopping condition is met; if the second stopping condition is not met, continuing to generate the second velocity field corresponding to the next time step using the second flow matching prediction model based on the binding protein structure information corresponding to the next time step and the pH information; repeating step k; if the second stopping condition is met, determining the binding protein structure information and the corresponding quality score under the pH information. Here, the specific implementation of module 15 is the same as or similar to the specific embodiment of step S15 described above, so it will not be repeated here, but is included by reference.
[0074] In some embodiments, generating the binding protein information corresponding to the target protein based on target protein information, the target amino acid corresponding to the binding region of the target protein, and the noise information of the binding protein corresponding to the target protein in the initial time step, using a trained first binding protein generation model, includes: determining one or more sequence lengths corresponding to the binding protein based on the target protein information; for each sequence length, generating the binding protein information corresponding to that sequence length using a trained first binding protein generation model based on the target protein information, the target amino acid corresponding to the binding region of the target protein, and the noise information of the binding protein corresponding to the target protein in the initial time step; generating the binding protein structure information and corresponding quality score under the corresponding pH information using a trained second binding protein generation model, based on the corresponding pH information and the binding protein structure information, includes: generating the binding protein structure information and corresponding quality score for each sequence length under the corresponding pH information using a trained second binding protein generation model. The device 1 also includes a six-module 16 (not shown). Based on the quality score, module 16 preferentially determines the corresponding binding protein sequence information and the corresponding binding protein structure information from the binding protein sequence information corresponding to the one or more sequence lengths and the corresponding binding protein structure information under the pH information. Here, the specific implementation of module 16 is the same as or similar to the specific embodiment of step S16 described above, and therefore will not be repeated here, but is incorporated herein by reference.
[0075] In some embodiments, the device 1 further includes a seven-module 17 (not shown). The seven-module 17, based on the sample-binding protein structure information and the corresponding sample pH information in the sample-binding protein information, generates sample-binding protein structure prediction information and a corresponding prediction quality score under the sample pH information using a second binding protein generation model; based on the sample-binding protein structure information and the sample pH information, it generates sample-binding protein structure information under the sample pH information using a first protein structure prediction model; based on the sample-binding protein structure information and the sample-binding protein structure prediction information under the sample pH information, it determines the corresponding quality score; based on a preset second loss function, it determines a second loss between the sample-binding protein structure prediction information under the sample pH information and the sample-binding protein structure information under the sample pH information, and between the predicted quality score and the quality score; based on the second loss, it updates the second binding protein generation model using a backpropagation algorithm to obtain a trained second binding protein generation model. Here, the specific implementation of the seven-module 17 is the same as or similar to the specific embodiment of the aforementioned step S17, and therefore will not be repeated, but is included herein by reference.
[0076] In some embodiments, the device 1 further includes an A / B module 18 (not shown). The A / B module 18 acquires complex information corresponding to a sample protein complex, wherein the sample protein complex comprises multiple chains; each chain in the sample protein complex is designated as a sample binding protein, and the remaining chains in the sample protein complex are designated as the sample target proteins corresponding to the sample binding proteins, thereby determining the corresponding sample binding protein information and the corresponding sample target protein information. Here, the specific implementation of the A / B module 18 is the same as or similar to the specific embodiment of the aforementioned step S18, and therefore will not be repeated here, but is incorporated herein by reference.
[0077] Figure 3 Exemplary systems that can be used to implement the various embodiments described in this application are shown; such as Figure 3 As shown in some embodiments, system 300 can function as any of the devices described in each of the embodiments. In some embodiments, system 300 may include one or more computer-readable media having instructions (e.g., system memory or NVM / storage device 320) and one or more processors (e.g., one or more processors 305) coupled to the one or more computer-readable media and configured to execute the instructions to implement the module and thus perform the actions described in this application.
[0078] In one embodiment, the system control module 310 may include any suitable interface controller to provide any suitable interface to at least one of the processors 305 and / or any suitable device or component communicating with the system control module 310.
[0079] The system control module 310 may include a memory controller module 330 to provide an interface to the system memory 315. The memory controller module 330 may be a hardware module, a software module, and / or a firmware module.
[0080] System memory 315 can be used, for example, to load and store data and / or instructions for system 300. In one embodiment, system memory 315 may include any suitable volatile memory, such as suitable DRAM. In some embodiments, system memory 315 may include double data rate type quad synchronous dynamic random access memory (DDR4 SDRAM).
[0081] In one embodiment, the system control module 310 may include one or more input / output (I / O) controllers to provide interfaces to the NVM / storage device 320 and (one or more) communication interfaces 325.
[0082] For example, NVM / storage device 320 may be used to store data and / or instructions. NVM / storage device 320 may include any suitable non-volatile memory (e.g., flash memory) and / or may include any suitable (one or more) non-volatile storage devices (e.g., one or more hard disk drive (HDD), one or more optical disc (CD) drives, and / or one or more digital universal optical disc (DVD) drives).
[0083] NVM / storage device 320 may include storage resources that are physically part of a device on which system 300 is mounted, or that can be accessed by the device without necessarily being part of it. For example, NVM / storage device 320 may be accessed via a network through one or more communication interfaces 325.
[0084] One or more communication interfaces 325 may provide the system 300 with an interface to communicate over one or more networks and / or with any other suitable device. The system 300 may wirelessly communicate with one or more components of a wireless network in accordance with any of one or more wireless network standards and / or protocols.
[0085] In one embodiment, at least one of the processors 305 may be logically packaged with one or more controllers of the system control module 310 (e.g., memory controller module 330). In one embodiment, at least one of the processors 305 may be logically packaged with one or more controllers of the system control module 310 to form a system-in-package (SiP). In one embodiment, at least one of the processors 305 may be integrated with the logic of one or more controllers of the system control module 310 on the same die. In one embodiment, at least one of the processors 305 may be integrated with the logic of one or more controllers of the system control module 310 on the same die to form a system-on-a-chip (SoC).
[0086] In various embodiments, system 300 may be, but is not limited to, a server, workstation, desktop computing device, or mobile computing device (e.g., laptop computing device, handheld computing device, tablet computer, netbook, etc.). In various embodiments, system 300 may have more or fewer components and / or different architectures. For example, in some embodiments, system 300 includes one or more cameras, a keyboard, a liquid crystal display (LCD) screen (including a touchscreen display), a non-volatile memory port, multiple antennas, a graphics chip, an application-specific integrated circuit (ASIC), and a speaker.
[0087] In addition to the methods and devices described in the above embodiments, this application also provides a computer-readable storage medium storing computer code that, when executed, performs the method described in any of the preceding embodiments.
[0088] This application also provides a computer program product that, when executed by a computer device, performs the method described in any of the preceding claims.
[0089] This application also provides a computer device, the computer device comprising:
[0090] One or more processors;
[0091] Memory, used to store one or more computer programs;
[0092] When the one or more computer programs are executed by the one or more processors, the one or more processors cause the one or more processors to perform the method as described in any of the preceding methods.
[0093] It should be noted that this application can be implemented in software and / or a combination of software and hardware, for example, using an application-specific integrated circuit (ASIC), a general-purpose computer, or any other similar hardware device. In one embodiment, the software program of this application can be executed by a processor to implement the steps or functions described above. Similarly, the software program of this application (including related data structures) can be stored in a computer-readable recording medium, such as RAM memory, a magnetic or optical drive, a floppy disk, or similar devices. Furthermore, some steps or functions of this application can be implemented in hardware, for example, as circuitry that cooperates with a processor to perform the various steps or functions.
[0094] Furthermore, a portion of this application can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can invoke or provide the methods and / or technical solutions according to this application through the operation of the computer. Those skilled in the art will understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executing the instructions, or the computer compiling the instructions and then executing the corresponding compiled program, or the computer reading and executing the instructions, or the computer reading and installing the instructions and then executing the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to a computer.
[0095] Communication media include media through which communication signals containing, for example, computer-readable instructions, data structures, program modules, or other data are transmitted from one system to another. Communication media can include guided transmission media (such as cables and wires (e.g., optical fibers, coaxial cables, etc.)) and wireless (unguided transmission) media capable of propagating energy waves, such as sound, electromagnetic, RF, microwave, and infrared. Computer-readable instructions, data structures, program modules, or other data can be embodied as modulated data signals in, for example, wireless media (such as carrier waves or similar mechanisms embodied as part of spread spectrum technology). The term "modulated data signal" refers to a signal whose one or more characteristics are altered or set in a manner that encodes information in the signal. Modulation can be analog, digital, or a hybrid modulation technique.
[0096] By way of example and not limitation, computer-readable storage media may include volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules or other data. For example, computer-readable storage media include, but are not limited to, volatile memories such as random access memory (RAM, DRAM, SRAM); and non-volatile memories such as flash memory, various read-only memories (ROM, PROM, EPROM, EEPROM), magnetic and ferromagnetic / ferroelectric memories (MRAM, FeRAM); and magnetic and optical storage devices (hard disks, magnetic tapes, CDs, DVDs); or other media now known or hereafter developed capable of storing computer-readable information / data for use by a computer system.
[0097] Herein, one embodiment of this application includes an apparatus comprising a memory for storing computer program instructions and a processor for executing the program instructions, wherein when the computer program instructions are executed by the processor, the apparatus is triggered to run a method and / or technical solution based on the foregoing embodiments of this application.
[0098] It will be apparent to those skilled in the art that this application is not limited to the details of the exemplary embodiments described above, and that this application can be implemented in other specific forms without departing from the spirit or essential characteristics of this application. Therefore, the embodiments should be considered exemplary and non-limiting in all respects, and the scope of this application is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be embraced within this application. No reference numerals in the claims should be construed as limiting the scope of the claims. Furthermore, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices recited in the apparatus claims may also be implemented by a single unit or device in software or hardware. The terms "first," "second," etc., are used to indicate names and do not indicate any particular order.
Claims
1. A method for determining information about acid-base sensitive binding proteins, wherein, The method includes: Based on target protein information, the target amino acid corresponding to the amino acid in the binding region of the target protein, and the noise information of the binding protein corresponding to the target protein in the initial time step, the binding protein information corresponding to the target protein is generated through the trained first binding protein generation model. The target protein information includes target protein sequence information and target protein structure information. The target amino acid and the amino acid in the binding region of the target protein match the preset acid-base sensitivity information. The binding protein information includes binding protein sequence information and binding protein structure information. The binding protein sequence information contains the target amino acid. The first binding protein generation model includes a first-order matching prediction model, and the generation of binding protein information corresponding to the target protein through the trained first binding protein generation model includes: Based on the target protein information and the noise information of the binding protein corresponding to the target protein in the initial time step, the first velocity field corresponding to the initial time step is generated by the first flow matching prediction model. Based on the first velocity field, determine the noise information corresponding to the binding protein in the next time step; Based on the noise information of the binding protein in the next time step, and in conjunction with the target amino acid, update the noise information of the binding protein in the next time step; determine whether a preset first stopping condition is met; If the first stopping condition is not met, based on the noise information of the binding protein and the target protein information in the updated next time step, the first velocity field corresponding to the next time step is generated again through the first flow matching prediction model, and steps i and j are repeated. If the first stopping condition is met, the binding protein information is determined.
2. The method according to claim 1, wherein, The first stopping condition includes at least one of the following: The running time steps have reached the preset time steps; The difference between the determined first velocity field and the first velocity field determined in the previous time step is less than the first threshold. The difference between the noise information of the updated binding protein and the noise information of the updated binding protein in the previous time step is less than the second threshold.
3. The method according to claim 1 or 2, wherein, The method further includes: Based on the corresponding pH information and the binding protein structure information, the trained second binding protein generation model generates the binding protein structure information and the corresponding quality score under the pH information. The second binding protein generation model includes a second flow matching prediction model. The process of generating the binding protein structure information and corresponding quality score under the pH information using the trained second binding protein generation model includes: Based on the pH information and the binding protein structure information in the current time step, the second velocity field corresponding to the current time step is generated by the second flow matching prediction model. Based on the second velocity field, k determines the binding protein structure information corresponding to the next time step; and determines whether the preset second stopping condition is met. If the second stopping condition is not met, based on the binding protein structure information corresponding to the next time step and the pH information, continue to generate the second velocity field corresponding to the next time step through the second flow matching prediction model; repeat step k. If the second stopping condition is met, determine the binding protein structure information and the corresponding quality score under the pH information.
4. The method according to claim 3, wherein, The method of generating the binding protein information corresponding to the target protein based on target protein information, the target amino acid corresponding to the binding region of the target protein, and the noise information of the binding protein corresponding to the target protein in the initial time step, through the trained first binding protein generation model, includes: Based on the target protein information, determine the length of one or more sequences corresponding to the binding protein; For each sequence length, based on the target protein information, the target amino acid corresponding to the binding region of the target protein, and the noise information of the binding protein corresponding to the target protein in the initial time step, the binding protein information corresponding to that sequence length is generated by the trained first binding protein generation model. The process of generating binding protein structure information and corresponding quality scores based on the corresponding pH information and the binding protein structure information using a trained second binding protein generation model includes: Based on the corresponding pH information and the binding protein structure information corresponding to each sequence length, the trained second binding protein generation model generates the binding protein structure information and the corresponding quality score for each sequence length under the pH information. The method further includes: Based on the quality score, the corresponding binding protein sequence information and the corresponding binding protein structure information are preferably determined from the binding protein sequence information corresponding to the length of the one or more sequences and the binding protein structure information corresponding to the pH information.
5. The method according to any one of claims 1 to 4, wherein, The method further includes: Based on the noise information of sample-binding proteins in the sampling time step and the corresponding sample target protein information, the first binding protein generation model generates the corresponding sample-binding protein prediction information. Based on a preset first loss function, a first loss is determined between the sample binding protein information corresponding to the sample binding protein and the sample binding protein prediction information. Based on the first loss, the first binding protein generation model is updated using the backpropagation algorithm to obtain a trained first binding protein generation model.
6. The method according to claim 3 or 4, wherein, The method further includes: Based on the sample binding protein structure information and the corresponding sample pH information in the sample binding protein information, the second binding protein generation model generates sample binding protein structure prediction information and corresponding prediction quality score under the sample pH information. Based on the sample binding protein structure information and the sample pH information, the sample binding protein structure information under the sample pH information is generated using a first protein structure prediction model; based on the sample binding protein structure information under the sample pH information and the sample binding protein structure prediction information under the sample pH information, the corresponding quality score is determined. Based on a preset second loss function, a second loss is determined between the sample binding protein structure prediction information under the sample pH information and the sample binding protein structure information under the sample pH information, as well as between the predicted quality score and the quality score. Based on the second loss, the second binding protein generation model is updated using the backpropagation algorithm to obtain a trained second binding protein generation model.
7. The method according to claim 5 or 6, wherein, The method further includes: Obtain the complex information corresponding to the sample protein complex, wherein the sample protein complex contains multiple chains; Each chain in the sample protein complex is taken as the sample binding protein, and the remaining chain in the sample protein complex is taken as the sample target protein corresponding to the sample binding protein. The corresponding sample binding protein information and the corresponding sample target protein information are determined.
8. The method according to claim 7, wherein, The complex information includes complex sequence information and complex structure information. The complex information corresponding to the obtained sample protein complex includes: Based on the complex sequence information, the corresponding complex structure information is determined using a second protein structure prediction model.
9. The method according to any one of claims 1 to 8, wherein, The method further includes: Based on the preset acid-base sensitivity information, an amino acid pair set is determined, wherein the amino acid pair set includes one or more amino acid pairs that match the acid-base sensitivity information; Based on the amino acids in the binding region of the target protein, the corresponding target amino acid is determined, wherein the amino acid in the binding region of the target protein and the target amino acid belong to an amino acid pair in the amino acid pair set.
10. The method according to claim 9, wherein, The step of determining an amino acid pair set based on the preset acid-base sensitivity information, wherein the amino acid pair set includes one or more amino acid pairs that match the acid-base sensitivity information, including: Based on the acidity coefficient corresponding to each amino acid, one or more first amino acids whose acidity coefficient matches the preset acid-base sensitivity information are determined. Based on the pH sensitivity information, one or more second amino acids that match each first amino acid are determined, thereby determining an amino acid pair set, wherein the amino acid pair set includes amino acid pairs consisting of each first amino acid and a matching second amino acid.
11. A computer device for determining information about acid-base sensitive binding proteins, comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the steps of the method as described in any one of claims 1 to 10.
12. A computer-readable storage medium having a computer program / instructions stored thereon, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method as described in any one of claims 1 to 10.
13. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the steps of the method as described in any one of claims 1 to 10.