Target determination method and device, computer device, and computer program

A computer-assisted method predicts and screens protein properties efficiently, addressing the laborious nature of traditional directed evolution by identifying target proteins quickly and accurately.

JP7746656B2Active Publication Date: 2025-10-01TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2024532391
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-05-09
Filing Date
2023-03-29
Publication Date
2025-10-01
Estimated Expiration
2043-03-29

AI Technical Summary

Technical Problem

Current directed evolution techniques for proteins are laborious and time-consuming, requiring multiple cycles of mutation, recombination, and screening to achieve desired protein properties.

Method used

A method utilizing a computer device and program that employs a pre-trained model to predict index values for objects in a set, screens based on conditions, determines a mapping relationship between indices and features, and identifies target objects efficiently, reducing the need for extensive experimental screening.

Benefits of technology

Accelerates the process of obtaining desired protein properties by improving efficiency and reducing time costs through computational methods, enhancing the speed and accuracy of protein evolution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007746656000018
    Figure 0007746656000018
  • Figure 0007746656000019
    Figure 0007746656000019
  • Figure 0007746656000020
    Figure 0007746656000020
Patent Text Reader

Abstract

The present application relates to the technical field of artificial intelligence, in particular to a method and apparatus for determining an object, and a computer device and program. The method includes the steps of: obtaining an index predicted value for each predetermined index of each object in a first object set (202); screening an object whose index predicted value satisfies an index value screening condition from the first object set to obtain a second object set (204); determining a mapping relationship between the predetermined index and the object features based on the index experimental values ​​for the predetermined index of a plurality of objects in the first object set and the object features (206); and determining a target object that meets the index requirements of the predetermined index from the second object set based on the mapping relationship (208).
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This application claims priority from a Chinese patent application filed with the China Patent Office on May 9, 2022, bearing application number 202210498684.7 and entitled "Object Determination Method, Apparatus, Computer Device and Storage Medium," the entire contents of which are incorporated herein by reference.

[0002] The present application relates to the field of computer technology, and in particular to a target determination method and apparatus, a computer device, and a computer program. [Background technology]

[0003] With the development of computer technology, directed evolution technology has emerged. Directed evolution allows proteins with new functions and properties to be obtained in a relatively short period of time. By setting clear goals, molecules can be redesigned. Directed evolution has already become an important research tool in fields such as drug development and chemical engineering.

[0004] In traditional protein directed evolution, an initial protein is established for a target function, a library of mutants is constructed at one or more positions, the most common mutants are determined by screening, and these mutants are then randomly recombined and screened, and the screened mutants are used in another round of "mutation, recombination, and screening" cycles (loops), and this process is repeated until the desired protein properties are achieved.

[0005] However, currently, most directed evolution techniques are laborious, time-consuming, and time-intensive. Summary of the Invention [Problem to be solved by the invention]

[0006] The embodiments provided in the present application aim to provide a target determination method and apparatus, a computer device, and a computer program. [Means for solving the problem]

[0007] According to one aspect, the present application provides a method for target determination, which is implemented by a computer device, the method comprising: obtaining a predicted index value for each predetermined index for each subject in the first subject set, the predicted index value being a value for the subject's predetermined index predicted by a pre-trained model; Screening the first object set for objects whose predicted index values ​​satisfy the index value screening conditions to obtain a second object set; Determine a mapping relationship between the predetermined index and the target feature based on the index experimental values ​​and target features of a plurality of objects in the first object set, where the index experimental values ​​refer to the values ​​of the objects for the predetermined index obtained by experiment; and determining, based on the mapping relationship, target objects from the second object set that satisfy index requirements of the predetermined index;

[0008] According to another aspect, the present application further provides an object determination device, the device comprising: a prediction value obtaining module for obtaining a predicted index value for each predetermined index of each subject in the first subject set, the predicted index value being a value for the predetermined index of the subject predicted by a pre-trained model; an object set acquisition module for screening objects whose index prediction values ​​satisfy an index value screening condition from the first object set to acquire a second object set; a mapping relationship determination module for determining a mapping relationship between a predetermined index and an object feature based on an index experimental value for the predetermined index of a plurality of objects in the first object set and an object feature, wherein the index experimental value refers to an object value for the predetermined index obtained by an experiment; and A target object determining module is included for determining, based on the mapping relationship, a target object from the second object set that satisfies an index requirement of the predetermined index.

[0009] According to another aspect, the present application further provides a computing device including a memory and one or more processors, the memory having computer-readable instructions stored therein that, when executed by the processors, cause the one or more processors to perform the steps of the target determination method described above.

[0010] According to another aspect, the present application further provides one or more non-volatile readable storage media having computer-readable instructions stored thereon that, when executed by one or more processors, cause the one or more processors to perform the steps of the target determination method described above.

[0011] According to another aspect, the present application further provides a computer program product (computer program), the computer program product including computer-readable instructions that, when executed by a processor, implement the steps of the target determination method described above.

[0012] The details of one or more embodiments of the application are set forth in the drawings and description that follow. Other features, objects, and advantages of the application will become apparent from the description, drawings, and claims. [Brief explanation of the drawings]

[0013] In order to more clearly explain the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings that need to be used in the description of the embodiments. It is obvious that the drawings described below are only some embodiments of the present application, and those skilled in the art can obtain other drawings based on these drawings without any creative work. [Figure 1] FIG. 1 illustrates an application environment for a target determination method in some embodiments. [Figure 2] 1 is a flowchart of a targeting method according to some embodiments. [Figure 3A] FIG. 1 illustrates the application of enzymes in some examples. [Figure 3B] FIG. 1 illustrates the principles of machine learning assisted directed evolution in some embodiments. [Figure 4] FIG. 1 shows the average fitness of amino acids in some examples. [Figure 5] 1 is a flowchart of a targeting method according to some embodiments. [Figure 6] FIG. 1 illustrates the principle of the targeting method in some embodiments. [Figure 7] FIG. 1 illustrates the principle of the targeting method in some embodiments. [Figure 8] FIG. 1 illustrates an application environment for a target determination method in some embodiments. [Figure 9] FIG. 1 illustrates an application environment for a target determination method in some embodiments. [Figure 10] FIG. 1 illustrates fitness distributions for different datasets in some embodiments. [Figure 11] FIG. 1 shows the effect of different methods on four protein directed evolution datasets in some examples. [Figure 12] FIG. 10 illustrates the effect of different methods on a dataset in some embodiments. [Figure 13] FIG. 1 is a block diagram illustrating a configuration of an object determination device according to some embodiments. [Figure 14] FIG. 1 illustrates the internal configuration of a computer device according to some embodiments. [Figure 15] FIG. 1 illustrates the internal configuration of a computer device according to some embodiments. DETAILED DESCRIPTION OF THE INVENTION

[0014] In order to make the purpose, technical solution and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the drawings and examples. It should be understood that the specific examples described herein are only for the purpose of interpreting the present application, and do not limit the present application.

[0015] The object determination method provided in the embodiments of the present application can be used in an application environment such as that shown in Figure 1, in which a terminal 102 communicates with a server 104 via a network. A data storage system can store data that the server 104 needs to process. The data storage system can be integrated into the server 104 or can be located in the cloud or on another server.

[0016] Specifically, the server 104 obtains the predicted index values ​​for each predetermined index of each object in the first object set, screens out objects from the first object set whose predicted index values ​​satisfy the index value screening conditions, obtains a second object set, and determines a mapping relationship between the predetermined indexes and the object features based on the experimental index values ​​for the predetermined indexes of the multiple objects in the first object set and the object features, and then determines target objects that meet the index requirements of the predetermined indexes from the second object set based on the mapping relationship. After determining the target objects, the server 104 can store the target objects and further send the target objects to the terminal 102, and the terminal 102 can display related information of the target objects.

[0017] The terminal 102 may be, but is not limited to, various desktop computers, laptops, smartphones, tablet computers, IoT devices, and portable wearable devices. The IoT devices may be smart speakers, smart TVs, smart air conditioners, intelligent vehicle-mounted devices, etc. The portable wearable devices may be smart watches, smart bracelets, head-mounted devices, etc. The server 104 may be implemented as an independent server or a server cluster consisting of multiple servers.

[0018] In some embodiments, the indicator prediction value may be predicted by a trained indicator detection model, which may be based on artificial intelligence and machine learning, such as a neural network model.

[0019] The schemes (technical solutions) provided in the embodiments of this application relate to technologies such as artificial intelligence neural networks, and are specifically illustrated by the following embodiments.

[0020] In some embodiments, a target determination method is provided, as shown in Figure 2, which may be executed by a terminal or a server, or may be jointly executed by a terminal and a server. The method is described by taking the server 104 in Figure 1 as an example, and specifically includes the following steps:

[0021] Step 202: Obtain an index prediction value for each predetermined index for each subject in the first subject set.

[0022] The first object set includes a plurality of objects, and the objects may be real substances, including but not limited to at least one of proteins, materials, batteries, etc. The objects may also be abstract concrete concepts, for example, the object may be a fast charging protocol for a battery.

[0023] An object may have a plurality of corresponding indices, and a predetermined indices may be any one of the plurality of indices of the object. For example, if the object is a protein, the indices of the object may include, but are not limited to, at least one of fitness, enrichment score, activity, brightness, etc.; if the object is a material, the indices of the object may include, but are not limited to, at least one of material composition or composition ratio, etc.; if the object is a battery fast charging protocol, the indices of the object may include, but are not limited to, each parameter in the battery fast charging protocol.

[0024] The index predicted value is a value corresponding to the predetermined index predicted for the subject. The index predetermined value is a value of the subject corresponding to the predetermined index predicted by a pre-trained model, which has the function of predicting the value of the subject corresponding to the predetermined index. The index predicted value can be predicted by a trained index detection model. The index detection model may be a neural network model.

[0025] Each object in the first object set may belong to an object category, which may include, but is not limited to, at least one of substances such as proteins or materials, and may also include abstract concepts such as battery charging protocols. For example, each object in the first object set may belong to a certain type of protein, e.g., each object may be a mutant protein obtained after mutating the same type of protein. A mutant protein is a protein that has not undergone mutation, and a mutant protein is a protein obtained after mutations are introduced into the wild-type protein. Protein-directed evolution can obtain a desired protein through mutation. Protein-directed evolution includes two mutation scenarios: one is a k-site saturation mutagenesis scenario and the other is a non-saturation mutagenesis scenario.

[0026] The k-site saturation mutagenesis scenario is used to mutate amino acids at k designated (predetermined) mutation sites, and in the mutant protein generated by this scenario, an amino acid at at least one of the k designated mutation sites is obtained after mutation. For example, when k=4, an amino acid at at least one of the four designated mutation sites in the resulting mutant protein is obtained by mutation. That is, the positions and number of mutation sites in the k-site saturation mutagenesis scenario are fixed, and mutations can only occur at the k designated mutation sites. A mutation site refers to a position in a protein where a mutation can occur. Therefore, a mutation site may also be referred to as a mutation position, and each position in a protein has one amino acid.

[0027] In the non-saturation mutagenesis scenario, the mutation sites are not fixed, but the number of amino acids at which mutations occur is fixed. For example, if each mutant protein obtained in the non-saturation mutagenesis scenario has amino acids at two positions that can be mutated, the positions at which the mutations occur can be the same or different, e.g., one mutation occurs at positions 1 and 2, and the other mutation occurs at positions 3 and 4.

[0028] When the object category is proteins, each object in the first object set may be a mutant protein generated in a k-site saturation mutagenesis scenario or a mutant protein generated in a non-saturation mutagenesis scenario. A mutant protein may also be referred to as a variant. A protein may be represented by an amino acid sequence. For example, if the first object set includes n mutant proteins, the first object set may be:

[0029]

number

[0030] Specifically, for each object in the first object set, the server can predict an index prediction value for a predetermined index for each object, and obtain an index prediction value for each object, for example, by using a trained index detection model to predict and obtain the index prediction value.

[0031] In some embodiments, the server may screen a plurality of objects from a first object set to obtain a reference object set, and determine object features for each object in the reference object set, where the object features refer to the features of the objects. The server may also experimentally obtain values ​​corresponding to a predetermined index for each object in the reference object set and obtain experimental index values ​​for each object in the reference object set, where the experimental index values ​​refer to the object values ​​corresponding to the predetermined index obtained experimentally, i.e., the experimental index values ​​for the objects are true values ​​for the predetermined index for the objects. For example, if the objects are proteins and the predetermined index is fitness, the experimental index values ​​refer to values ​​related to the protein's fitness obtained experimentally. The server may train an index detection model using the object features and the experimental index values ​​for each object in the reference object set to obtain a trained index detection model, determine object features for each object in the first object set, input the object features for each object in the first object set into the trained index detection model, and use the trained index detection model to predict and obtain predicted values ​​for each corresponding index for each object in the first object set.

[0032] In some embodiments, the target is a mutant protein, the target feature is a protein feature, and the protein feature may be a feature obtained by encoding based on amino acids at mutation sites in the mutant protein. For example, amino acids can be encoded based on experimental index values ​​for a predetermined index of the mutant protein to obtain amino acid features corresponding to the amino acids, and the protein feature of the mutant protein can be obtained based on the amino acid features corresponding to the amino acids at each mutation position. For example, for a mutant protein generated in a k-site saturation mutagenesis scenario, the protein feature of the mutant protein can be obtained using the amino acid feature of the k-site amino acid. For a mutant protein generated in a non-saturation mutagenesis scenario, if mutations occur at two positions, a vector consisting of the amino acid features of the amino acids at these two positions can be determined as the protein feature of the mutant protein.

[0033] Step 204: From the first object set, objects whose predicted index values ​​satisfy the index value screening conditions are screened to obtain a second object set.

[0034] The index value screening condition includes that the predicted index value is greater than a first index threshold, which may be preset or may be set according to needs, and the second object set is a set of objects screened from the first object set, and the predicted index value of the objects in the second object set satisfies the index value screening condition.

[0035] Specifically, the server compares the predicted index value of each object in the first object set with a first index threshold, and can form a second object set from each object whose predicted index value is greater than the first index threshold. For example, if the object is a mutant protein, the predetermined index is affinity, and the first index threshold is an affinity threshold, the second object set can be formed from each object in the first object set whose affinity is greater than the affinity threshold, and the affinity threshold can be set in advance or according to needs.

[0036] Step 206: Determine a mapping relationship between the predetermined index and the object feature based on the index experimental values ​​for the predetermined index of the plurality of objects in the first object set and the object feature.

[0037] Wherein, the plurality of objects in the first object set may refer to each object in the above-mentioned reference object set. The mapping relationship between the predetermined index and the object feature is used to reflect the change of the value of the predetermined index with the change of the object feature, and the mapping relationship between the predetermined index and the object feature can be expressed by a curve, for example, the mapping relationship can be expressed by a curve y1=f1(x), where y1 represents the predetermined index and x represents the object feature.

[0038] Specifically, after obtaining the reference object set, the server can not only use the reference object set to train an index detection model, but also use the index experimental values ​​and object features for the predetermined index of each object in the reference object set to determine the mapping relationship between the predetermined index and the object features.

[0039] In some embodiments, after obtaining the index experimental value of each object in the reference object set, for each object in the reference object set, the server sets the object features and index experimental value of the object as points on the curve y1=f1(x), and uses multiple points on multiple curves to perform fitting to the obtained multiple points, thereby obtaining the curve y1=f1(x) representing the mapping relationship.

[0040] Step 208: Determine the target object that satisfies the index requirement of the predetermined index from the second object set based on the mapping relationship.

[0041] The index requirement of the predetermined index may be, for example, at least one of the following: the index experimental value is as large as possible, or the index experimental value is greater than a second index threshold value; and the target object is an object in the second object set that meets the index requirement of the predetermined index.

[0042] Specifically, the mapping relationship between the predetermined index and the target feature is a first mapping relationship, and the server performs statistical calculations based on the first mapping relationship to obtain a second mapping relationship between the target statistical index and the target feature. Based on the second mapping relationship, the server determines a statistical index value for the target statistical index of each object in the second object set. Based on the statistical index values ​​of each object, the server selects objects from the second object set. Based on the selected objects, the server obtains objects that meet the index requirements of the predetermined index. The first mapping relationship represents a rule by which the value of the predetermined index changes with changes in the target feature, and the second mapping relationship represents a rule by which the value of the target statistical index changes with changes in the target feature. For example, the second mapping relationship can be represented by a curve y2=f2(x), where y2 represents the target statistical index and x represents the target feature.

[0043] The target statistical indicator may be one or more. For example, the target statistical indicator may include, but is not limited to, at least one of expected improvement (EI), probability of improvement (PI), upper confidence bound (UCB), or Thompson sampling (TS). The first mapping relationship may be further referred to as a probabilistic surrogate model, and the second mapping relationship may be further referred to as an acquisition function. The acquisition function is constructed using the posterior probability distribution obtained by the probabilistic surrogate model, and the most "potential" next experimental point is selected by maximizing the acquisition function. The acquisition function is responsible for testing the proposed new point based on the trade-off between exploration and exploitation. Exploration refers to selecting a point as far away from known points as possible for the next experiment, i.e., exploring as much unknown territory as possible. Exploitation refers to selecting a point as close to known points as possible for the next experiment, i.e., mining as many points around known points as possible.

[0044] In some embodiments, the server can determine a statistical index value for a target statistical index of each object in the second object set based on the second mapping relationship, and determine a selection target from each object in the second object set based on the statistical index value of each object. Specifically, there can be one or more selection targets, and the object corresponding to the largest statistical index value can be determined as the selection target, or the object whose statistical index value is greater than a third index threshold can be determined as the selection target, and the third index threshold can be set according to needs. The server can obtain an object that meets the index requirements of a predetermined index based on the selection target. For example, the server can determine the selection target as an object that meets the index requirements of a predetermined index.

[0045] In some embodiments, after obtaining the selection target, the server obtains the experimental index value of the selection target, compares the experimental index value of the selection target with a second index threshold, and determines that the experimental index value of the selection target reaches the second index threshold, and determines that the selection target is the target target; where, after obtaining the selection target, the server can experimentally select the corresponding experimental index value; if it determines that the experimental index value of the selection target does not reach the second index threshold, the server adds the selection target to the reference target set, and again uses the experimental index value and object features for the predetermined index of each object in the reference target set to determine a first mapping relationship between the predetermined index and the object features, and again determines a target target that meets the index requirements of the predetermined index from the second target set based on the first mapping relationship; in this way, the server continuously loops, and when it finds a selection target whose experimental index value reaches the second index threshold, it terminates the loop and determines the selection target whose experimental index value reaches the second index threshold as the target target; or when the number of loops reaches a count threshold, it determines the screened selection target as the target target.

[0046] The above-mentioned target determination method obtains predicted index values ​​for each predetermined index of each target in a first target set, selects from the first target set targets whose predicted index values ​​satisfy index value screening conditions, obtains a second target set, determines a mapping relationship between the predetermined indexes and the target features based on experimental index values ​​for the predetermined indexes of multiple targets in the first target set and the target features, and determines target targets that meet the index requirements of the predetermined indexes from the second target set based on the mapping relationship. Because the second target set is screened from the first target set, the efficiency of determining target targets that meet the index requirements of the predetermined indexes from the second target set is higher than the efficiency of screening target targets from the first target set. This reduces the time cost of determining target targets.

[0047] In real design and application scenarios, for example, when an environmental scientist obtains environmental conditions by designing the placement positions of sensors, a chemist obtains new substances by designing experiments, or a pharmaceutical company fights diseases by designing new drugs, these design problems are usually considered and solved as optimization problems such as the following (only maximization problems are considered, and minimization problems can be converted into maximization problems by simply taking negative signs):

[0048]

number

[0049] In contrast, the target determination method provided in this application can accelerate the process of obtaining the desired object and improve efficiency, thereby reducing time costs. For example, the target determination method provided in this application can be used in computationally assisted protein evolution to obtain the desired protein. As shown in Figure 3A, proteins play an important role in people's lives. For example, enzymes play an important role in human society, from daily life to industry. Some everyday detergents contain enzymes to promote the decomposition of grease and other contaminants. Enzymes are essential for processes such as fermentation and degradation in the food industry. In pharmaceuticals and fine chemicals, enzymes serve as environmentally friendly, efficient catalysts, replacing several traditional chemical production processes that require heavy metals and consume a lot of energy. Enzymes also play a key role in bioenergy generation. Directed evolution can obtain proteins with novel functions and properties in a relatively short period of time. By artificially setting clear goals, scientists can redesign molecules, and directed evolution has already become an important research tool in fields such as drug development and chemical engineering. Machine learning-assisted directed evolution can be employed. For example, as shown in Figure 3B, the process of machine learning-assisted directed evolution can include four steps: 1) establish an initial protein for the target function and construct a mutant library at k positions; 2) train a model using existing data; 3) predict other mutants in the mutant library using the trained model; and 4) select the optimal mutant for experimental testing and add it to the training set for the next round of model training. Computation-assisted directed evolution can accelerate optimization and reduce the experimental burden.

[0050] In some embodiments, the objects in the first object set are mutant proteins, and the method further includes obtaining a reference object set by screening based on the first object set, where the reference object set satisfies a condition that each type of amino acid occurs at least a target number of times at each mutation position. The step of obtaining a predicted index value for each predetermined index for each object in the first object set includes training an index detection model based on the object features and experimental index values ​​of each object in the reference object set; and predicting the predicted index value for each object in the first object set using the trained index detection model.

[0051] Each target in the reference target set may be a mutant protein, or the reference target set may include wild-type and mutant proteins. The target in the first target set is a mutant protein, and the target number of times may be preset or set according to needs, for example, two times. The mutation position is a mutation site. The reference target set satisfies the condition that each amino acid occurs at each mutation position at least the target number of times. For example, if there are 20 amino acids and the target number is 2, then in each protein in the reference target set, these 20 amino acids will occur at least twice at each mutation site. Taking a mutant protein generated in a k-site saturation mutagenesis scenario as an example, if there are four mutation sites and each of the 20 amino acids occurs twice at each mutation site, 40 samples can be selected as initial samples from the sample space. In this way, the coverage of the coding information of the amino acids contained in the initial samples can be maximized while minimizing the number of required experiments, thereby reducing experimental costs. The sample space may include mutant proteins and may also include wild-type proteins, the sample refers to the protein, the reference set may change over time, and the initial sample refers to the initially determined reference set.

[0052] The indicator detection model is used to determine the target value corresponding to a predetermined indicator based on the target features, i.e., to determine the indicator predicted value for the predetermined indicator of the target. The indicator detection model can be a neural network model, such as XGBOD (Improving Supervised Outlier Detection with Unsupervised Representation Learning), or of course, other models, but this is not limited thereto. The basic process of XGBOD is to use various unsupervised models to learn the original data to obtain outlier scores for each sample, and then use the outlier scores as a new data representation format; then, merge them with the original features to generate a new feature space; finally, train an XGBoost classifier in the new feature space, and regard the output as the prediction result.

[0053] Specifically, the server obtains an experimental index value for a predetermined index for each object in the reference object set (referred to as the experimental index value corresponding to the object), and obtains an object feature for each object in the reference object set based on the experimental index value for the predetermined index for each object in the reference object set. The server inputs the object features of the object into an index detection model waiting to be trained to make a prediction, thereby obtaining a predicted index value for the predetermined index for the object (referred to as the predicted index value corresponding to the object), and then adjusts the model parameters of the index detection model based on the difference between the experimental index value corresponding to the object and the corresponding predicted index value until the model converges, thereby obtaining a trained index detection model. The server inputs the object features of each object in the first object set into the trained index detection model to make a prediction, thereby obtaining a predicted index value corresponding to each object in the first object set.

[0054] In some embodiments, the server determines the amino acid signature of each type of amino acid based on the experimental index value of each subject in the reference subject set, and when determining the subject signature (i.e., protein signature) of the subject in the first subject set, the server can use the amino acid signature of each type of amino acid determined based on the reference subject set to determine the protein signature of the subject in the first subject set. For example, in the scenario of k-site saturation mutagenesis, the server can determine the amino acid signature at each mutation position of each type of amino acid based on the experimental index value of each subject in the reference subject set. For the subject in the first subject set, the server determines each amino acid at the mutation position in the subject, and determines the amino acid signature corresponding to the amino acid at each mutation position in the subject from the determined "amino acid signature at each mutation position of each type of amino acid", and then determines the vector consisting of the determined amino acid signature as the subject signature (i.e., protein signature) of the subject.

[0055] In some embodiments, obtaining a reference target set by screening based on a first target set is a sample selection strategy for determining the initial sample in Bayesian optimization, which can improve optimization efficiency and reduce the time cost of Bayesian optimization. This can maximize the coverage of the amino acid coding information contained in the initial sample and minimize the number of required experiments, thereby reducing experimental costs. Among these, Bayesian optimization typically uses Gaussian process (GP) regression based on Gaussian distribution as the a priori probabilistic surrogate model. GP is flexible and scalable, and in theory, can surrogate any linear or nonlinear function. Of course, Gaussian process regression based on Student-t priors can also be used as the a priori probabilistic surrogate model, combining robust regression (Gaussian process based on Student-t distribution) with outlier detection to separate data points into outliers and inliers, thereby eliminating the impact of outliers on model fitting. Gaussian processes based on student-t priors can be abbreviated as "Robust GP", and Gaussian processes based on Gaussian distributions can be abbreviated as "GP".

[0056] In this embodiment, the reference target set satisfies the condition that "each type of amino acid appears at least the target number of times at each mutation position," so the number of each amino acid in the obtained reference target set can be balanced. As a result, by training the index detection model based on the reference target set, the accuracy of the training can be improved, and the accuracy of the index prediction value predicted by the trained index detection model can be improved.

[0057] In some embodiments, the step of determining a mapping relationship between the predetermined index and the target feature based on the index experimental values ​​for the predetermined index and the target feature of a plurality of objects in the first object set includes the steps of: determining the target feature of each object in the reference object set based on the index experimental values ​​for the predetermined index of each object in the reference object set; and determining a mapping relationship between the predetermined index and the target feature based on the index experimental values ​​for the predetermined index and the target feature of each object in the reference object set.

[0058] Specifically, the server can obtain the object characteristics of each object in the reference object by performing statistical calculations on the index experimental values ​​corresponding to each object in the reference object set, where the index experimental values ​​corresponding to the object refer to the index experimental values ​​related to a predetermined index of the object, and the statistical calculations include, but are not limited to, calculating at least one of the average value, maximum value, or minimum value.

[0059] In some embodiments, the mapping relationship between the predetermined index and the object feature is a first mapping relationship, and the mapping relationship may be represented by a curve, for example, the first mapping relationship is represented by a curve y1=f1(x), and for each object in the reference object set, the server can set the object feature and the index experimental value of the object as points on the curve y1=f1(x), and use multiple points on multiple curves to perform fitting to the obtained multiple points to generate a curve y1=f1(x) representing the mapping relationship.

[0060] In this embodiment, since the reference target set satisfies the condition that "each type of amino acid appears at least the target number of times at each mutation position," each amino acid in the obtained reference target set can be balanced. Therefore, determining the target features of each target in the reference target set based on the experimental index values ​​for the specified indexes of each target in the reference target set can make the range of coverage of the encoded information of amino acids relatively larger, that is, the range of information covered by the target features can be broadened.

[0061] In some embodiments, the step of obtaining a reference object set by screening based on the first object set may include the steps of: obtaining a current score set, where the current score set includes a current score corresponding to each type of amino acid; obtaining a second protein set based on the first object set, and selecting a target protein from the second protein set based on the current score set; reducing the current scores corresponding to each amino acid at each mutation position in the target protein in the current score set, and transferring the target protein from the second protein set to the first protein set. (Transfer) and if the current score set indicates that the first protein set does not satisfy the condition that "each type of amino acid occurs at least the target number of times at each mutation position," return to the step of screening (selecting) target proteins from the second protein set based on the current score set, and repeat this process until the current score set indicates that the first protein set satisfies the condition that "each type of amino acid occurs at least the target number of times at each mutation position," and then determine the first protein set as the reference set.

[0062] The current score set includes a current score corresponding to each type of amino acid. The current score is an integer, for example, 2. The current scores corresponding to each type of amino acid may be different or the same, and each type of amino acid may correspond to one current score or multiple current scores, for example, amino acids of the same type have corresponding current scores at different mutation positions. The current score set is constantly changing. A target protein is selected from the second protein set based on the current score set, and there may be one or more target proteins.

[0063] Specifically, the server can determine the first object set as the second protein set, i.e., the objects in the second protein set match the objects in the first object set, or the server can obtain objects from the first object set other than the objects in the reference object set to form the second object set.

[0064] In some embodiments, the current score set is determined as the current protein score of each mutant protein in the second protein set, and the mutant protein corresponding to the largest current protein score is determined as the target protein. The server can obtain a first protein sequence by sorting each mutant protein in the second protein set according to the order of the current protein scores from highest to lowest, and determine the mutant protein arranged before the sorting threshold in the first protein sequence as the target protein. The sorting threshold may be preset or may be set according to needs, for example, may be any one of first place, second place, etc.

[0065] In some embodiments, the server can decrement the current scores corresponding to the amino acids at each mutation position in the target protein in the current score set and transfer the target protein from the second protein set to the first protein set.

[0066] For example, if the target in the first target set is a mutant protein generated in a k-site saturation mutagenesis scenario, the current score set will include the current scores of each type of amino acid at each type of mutation position, taking four mutation positions as an example, and the current score set may be expressed in the form of a matrix, which may also be referred to as the current score matrix. In the matrix corresponding to the current score set, the uth row and wth column represent the current score corresponding to the uth amino acid at the wth mutation position, where 1≦u≦m and 1≦w≦k, where m represents the number of types of amino acids, e.g., 20, and k represents the number of mutation positions, e.g., 4. For example, if the amino acid at the first mutation position in the mutant protein is the first type of amino acid, the corresponding current score is the score in the first row and first column of the matrix.

[0067] When the object in the first object set is a mutant protein generated in a non-saturation mutagenesis scenario, the current score set includes a current score corresponding to each type of amino acid, and the current score set may be represented by a vector, which may further be referred to as a current score vector, where the score at the uth position in the vector represents the current score of the uth type of amino acid. For example, when the amino acid at the first mutation position in a mutant protein is the first type of amino acid, the corresponding current score is the first score in the vector.

[0068] In some embodiments, if the current score set indicates that the first protein set does not satisfy the condition that "each type of amino acid occurs at least the target number of times at each mutation position," the server returns to performing the step of selecting target proteins from the second protein set based on the current score set, and continues this process until the current score set indicates that the first protein set satisfies the condition that "each type of amino acid occurs at least the target number of times at each mutation position," and then determines the first protein set as the reference set. Specifically, the first protein set is constantly changing, and if no proteins are included in the initial first protein set, the scores in the initial current score set are all equal to the target number of times; for example, if the target number of times is 2, each current score in the initial current score set is equal to 2. After the target protein is determined, the target protein is transferred from the second protein set to the first protein set, and the current score corresponding to each amino acid at each mutation position in the target protein in the current score matrix is ​​decremented by 1 each time, for example, from 2 to 1 or from 1 to 0, thereby constantly updating the current score matrix. If there is no score greater than 0 in the current score matrix, it is determined that the first protein set satisfies the condition that "each type of amino acid appears at least the target number of times at each mutation position," and the first protein set is determined as the reference set.

[0069] In this embodiment, a second protein set is obtained based on the first target set, the target protein value is screened from the second protein set, the current score set is updated based on the target protein, and the target protein is transferred from the second protein set to the first protein set, thereby constantly selecting proteins and updating the current score. If the current score set indicates that the first protein set meets the condition that "each type of amino acid appears at least the target number of times at each mutation position," the first protein set is determined to be the reference target set. This allows for quick selection of a reference target set from the first target set that meets the condition that each type of amino acid appears at least the target number of times at each mutation position, and reduces the number of experiments required to determine the reference target set, thereby reducing experimental costs and time costs.

[0070] In some embodiments, the step of obtaining a current score set includes the following steps: obtaining an initial score set, and the initial scores corresponding to each type of amino acid in the initial score set are the target number of times; and reducing the initial scores corresponding to each amino acid at each mutation position in the wild-type protein in the initial score set to obtain a current score set, and determining a first protein set based on the wild-type protein, where the wild-type protein is a protein without mutations.

[0071] Wherein, the first score corresponding to each type of amino acid in the initial score set is the target number of times, and if the target number is an integer, for example, 2, the first score is 2. For example, if the target in the first target set is a mutant protein generated in a scenario of saturation mutagenesis at k sites, the initial score set will include the first scores at each type of mutation position for each type of amino acid. Taking the example of four mutation positions, the initially obtained set can be represented in the form of a matrix, and each element in the matrix can all be 2. In the matrix corresponding to the initially obtained set, the uth row and wth column represent the first score corresponding to the uth type of amino acid at the wth mutation position, where 1≦u≦m and 1≦w≦k, m represents the number of types of amino acids, for example, 20, and k represents the number of mutation positions, for example, 4. For example, if the amino acid at the first mutation position in the mutant protein is the first type of amino acid, the corresponding first score is the first score in the first row and first column of the matrix.

[0072] When the target in the first target set is a mutant protein generated in a non-saturation mutagenesis scenario, the first score set includes the first scores corresponding to each type of amino acid, the first score set is represented by a vector, and all elements in the vector may be 2, and the score at the uth position in the vector represents the first score of the uth type of amino acid. When the amino acid at the first mutation position in the mutant protein is the first type of amino acid, the corresponding first score is the first score in the vector.

[0073] Specifically, in the case of a k-site saturation mutagenesis scenario, there are k mutation positions, and the server determines the amino acid corresponding to each mutation position from the wild-type protein, and gradually decrements the current score corresponding to the amino acid in the initial score set, subtracting 1 each time to obtain the current score set. Taking a four-site saturation mutagenesis scenario as an example, if there are four mutation positions and the corresponding amino acids in the wild-type protein at these four mutation positions are the first amino acid A1, the second amino acid A2, the third amino acid A3, and the fourth amino acid A4, respectively, the scores in the first row, first column, the second row, second column, the third row, third column, and the fourth row, fourth column in the matrix corresponding to the initial score set are each decremented by 1, and the initial score set after decrementing by 1 is determined as the current score set.

[0074] In the case of a non-saturation mutagenesis scenario, each type of mutant protein has a corresponding target number of mutation positions, for example, 2, and the server can obtain a mutation position set by counting (statistics) the mutation positions corresponding to each of the multiple mutant proteins in the non-saturation mutagenesis scenario, and the server can determine the amino acids corresponding to each mutation position in the mutation position set from the wild-type protein, and the server can determine the amino acids corresponding to each mutation position in the mutation position set from the wild-type protein, and gradually reduce the current scores corresponding to each type of amino acid determined in the initial score set, subtracting 1 each time, to obtain the current score set.

[0075] In some embodiments, the server may determine a set of wild-type proteins as the first protein set, i.e., the initial first protein set includes one wild-type protein.

[0076] In this embodiment, the initial score set is updated based on the wild-type protein without mutations, and the current score set is obtained, thereby accelerating the speed of score reduction, improving the efficiency of obtaining the reference target set, and reducing the time cost.

[0077] In some embodiments, the step of selecting a target protein from the second protein set based on the current score set includes the following steps: for each mutant protein in the second protein set, determine a current score from the current score set corresponding to each amino acid at each mutation position in the mutant protein; determine a current protein score for the mutant protein based on the obtained current scores; and select and obtain a target protein from the second protein set based on the current protein score.

[0078] Specifically, for each mutant protein in the second protein set, a current score corresponding to each amino acid at each mutation position in the mutant protein is determined from the current score set, the obtained current scores are summed, and the result of the sum is used as the current protein score of the mutant protein.

[0079] In some embodiments, the server can sort each mutant protein in the second protein set according to the order of the current protein score from highest to lowest, obtain a first protein sequence, and determine the mutant protein arranged before the sorting threshold in the first protein sequence as the target protein, where the sorting threshold may be preset or may be set according to needs, for example, any one of first place, second place, etc.

[0080] In this embodiment, mutant proteins whose current protein scores satisfy the condition that the score is relatively large are screened from the second protein set to obtain the target protein. The higher the current protein score, the stronger the update intensity for the current score set. This accelerates the speed at which the current score set indicates that the first protein set satisfies the condition that "each type of amino acid appears at least the target number of times at each mutation position," thereby improving the efficiency of obtaining the reference target set.

[0081] In some embodiments, each type of amino acid has a corresponding amino acid, and the scores in the current score set are uniquely labeled by the amino acid and the mutation position, and the step of determining from the current score set a current score corresponding to each amino acid at each mutation position in the mutant protein includes the following steps: for an amino acid at each mutation position, determine from the current score set a current score corresponding to the amino acid at the mutation position based on the amino acid corresponding to the amino acid and the mutation position.

[0082] Specifically, if the target in the first target set is a mutant protein generated in a k-site saturation mutagenesis scenario, the current score set includes the current scores at each mutation position for each type of amino acid, i.e., the scores in the current score matrix are uniquely identified by the amino acid and the mutation position. Take the example of four mutation positions, and the current score set can be expressed in the form of a matrix, where the uth row and wth column in the matrix corresponding to the current score set represent the current score corresponding to the uth amino acid at the wth mutation position, where 1≦u≦m and 1≦w≦k, where m represents the number of amino acid types, e.g., 20, and k represents the number of mutation positions, e.g., 4. For example, if the amino acid at the first mutation position in the mutant protein is the first type of amino acid, the corresponding current score is the score in the first row and first column in the matrix.

[0083] In the case of a non-saturation mutagenesis scenario, when the current score set is a current score vector, the uth element in the current score vector is the score of the uth amino acid; when the mutant protein has two mutation positions, the amino acids at the two mutation positions are the third and tenth amino acids, respectively; in this case, the score corresponding to the third amino acid is the score at the third position in the current score vector, and the score corresponding to the tenth amino acid is the score at the tenth position in the current score vector.

[0084] In some embodiments, in the case of a k-site saturation mutagenesis scenario, the server can determine the reference set to be screened and obtained by the following algorithm:

[0085] The input data of the algorithm is p, a set D train ={(S0, y0)}, matrix M.

[0086] where p refers to the number of occurrences of each type of amino acid at each mutation site, i.e., the target number, e.g., 2. train refers to the first protein set, where S0 in (S0, y0) represents the wild-type protein and y0 represents the experimental index value of the wild-type protein. M refers to the current score matrix.

[0087]

number

[0088] The step of initializing the current score matrix is ​​to determine whether the uth amino acid appears at the wth mutation position in S0. uw = p-1, otherwise M uw = p and M uw is the element in row u and column w in M, where 1≦w≦k.

[0089] The output data of the algorithm is the updated set D train The algorithm outputs D train is the initial referent set.

[0090] The steps of the algorithm are as follows:

[0091] Step 1: while ∃M uw>0 do, which means that if there is an element in matrix M that is smaller than 0, steps 2 to 5 are executed.

[0092] Step 2: Score each variant

[0093]

number

[0094] Step 3: Maximum score mutant protein i*=argmaxScore i and this step is used to determine the mutant protein with the highest current protein score.

[0095] Wherein, i* indicates that the mutant protein with the maximum current protein score is the i*th mutant protein in the first subject set.

[0096] Step 4: Update the set, D train ←(S i *,y i *), and the meaning of this step is to add the i*th mutant protein in the first subject set to the first protein set.

[0097] Among them, (S i *,y i *), Si* represents the i*th mutant protein in the first subject set, and yi* represents the index experimental value of the i*th mutant protein.

[0098] Step 5: Update the score matrix M, and set the uth amino acid to S i * appears at the wth mutation position, M uw =M uw -1, otherwise M uw =M uw is.

[0099] Step 6: end while, if there is no element in M ​​greater than 0, execute step 7.

[0100] Step 7:D train Output.

[0101] In some embodiments, in the case of a non-saturation mutagenesis scenario, the server can screen and obtain a reference population by the following algorithm:

[0102] The input data of the algorithm is p, and the set D train ={(S0,y0)}, vector Q.

[0103] where p refers to the number of occurrences of each type of amino acid at each mutation site, i.e., the target number, e.g., 2. train refers to the first protein set, and in (S0, y0), S0 represents the wild-type protein and y0 represents the experimental index value of the wild-type protein. Vector Q refers to the current score vector,

[0104]

number

[0105] The step of initializing the current score vector is to determine whether the uth amino acid appears at the mutation position of S0. u = p-1, otherwise Qu = p and Q u is the u-th element in vector Q.

[0106] The output data of the algorithm is the updated set D train The algorithm outputs D train is the initial referent set.

[0107] Step 1: while ∃Qu>0 do, which means that if there is an element in matrix Q that is smaller than 0, steps 2 to 5 are executed.

[0108] Step 2: Score each variant

[0109]

number

[0110] Step 3: Maximum score variant i*=argmaxScore i is selected, and this step is used to determine the mutant protein with the highest current protein score.

[0111] Wherein, i* indicates that the mutant protein with the maximum current protein score is the i*th mutant protein in the first subject set.

[0112] Step 4: Update the set, D train ←(S i *,y i*), which means that the i*th mutant protein in the first subject set is added to the first protein set.

[0113] Step 5: Update the score matrix Q so that the uth amino acid is S i * appears at the mutation position of Q u = p-1, otherwise Q u = p, and Qu is the u-th element in vector Q.

[0114] Step 6: end while, if there is no element in Q greater than 0, execute step 7.

[0115] Step 7:D train Output.

[0116] In this embodiment, each type of amino acid has a corresponding amino acid, and the scores in the current score set are uniquely labeled by the amino acid and the mutation position. For the amino acid at each mutation position, the current score corresponding to the amino acid at the mutation position is determined from the current score set based on the amino acid and mutation position corresponding to the amino acid, thereby allowing the current score at each mutation position of each amino acid to be determined quickly and accurately.

[0117] The protein encoding method (i.e., protein feature determination method) provided in this application can assist protein evolution in conjunction with Bayesian optimization. The process of encoding proteins to obtain protein features can be referred to as the protein feature representation process, and effective protein feature representation is crucial for finding optimal protein variants through Bayesian optimization. To construct accurate and information-rich low-dimensional feature representations in conjunction with relevant domain knowledge, this application proposes a new low-dimensional encoding scheme to represent each amino acid at each site. Specifically, two methods for calculating the amino acid representation at each site are provided for two experimental scenarios in protein directed evolution: a k-site saturation mutagenesis scenario and a non-saturation mutagenesis scenario.

[0118] In some embodiments, the target feature is a protein feature, and the step of determining the target feature of each object in the reference target set based on the experimental index value for a predetermined index of each object in the reference target set includes the steps of: for each mutation position, dividing the reference target set according to the type of amino acid at the mutation position, and obtaining a first sub-target set corresponding to each type of amino acid; for each type of amino acid at each mutation position, determining the amino acid feature at the mutation position of the amino acid based on the experimental index value of each object in the first sub-target set corresponding to the amino acid; and obtaining the protein feature of the object based on the amino acid feature of the amino acid at each mutation position in the object.

[0119] Wherein, the first sub-object set is a subset of the reference object set, and the first sub-object set includes at least some of the objects in the reference object set. The first sub-object set is uniquely determined by the mutation position and the type of amino acid. For example, if there are four mutation positions, which are the first mutation position, the second mutation position, the third mutation position, and the fourth mutation position, and there are 20 types of amino acids, which are type ii amino acids, and 1≦ii≦20, 80 first sub-object sets are generated, and at least one of the mutation positions and amino acids corresponding to different first sub-object sets is different. For example, first sub-object set 1 is the first sub-object set corresponding to the first mutation position and the first type of amino acid, and first sub-object set 2 is the first sub-object set corresponding to the first mutation position and the second type of amino acid.

[0120] Specifically, for each mutation position, the server divides the reference object set according to the type of amino acid at the mutation position, thereby obtaining a first subobject set corresponding to each type of amino acid. For example, for the kkth mutation position, the server can obtain the amino acid at the kkth mutation position from each object in the reference object set to construct an amino acid set corresponding to the kkth mutation position. For example, if the parameter object set includes 40 objects, the amino acid set will include 40 amino acids, and the type of amino acid at the kkth mutation position in different objects may be the same or different. After obtaining the amino acid set corresponding to the kkth mutation position, the server can divide the amino acid set according to the type of amino acid to obtain multiple subsets, classifying amino acids of the same type into the same subset and different amino acids into different subsets, each subset containing only one type of amino acid. Each subset obtained by dividing is a first subobject set corresponding to each type of amino acid at the kkth mutation position. For example, if the jth position in a protein is a mutation position, the first subobject set corresponding to the mutation position is V j (a)={i|S ij = a}, where j represents the mutation position, i is the ordinal number of the object in the reference object set, and S ij represents the amino acid at mutation position j in the i-th object Si in the reference object set, and a represents any one type of amino acid. For example, if there are 20 types of amino acids, a represents any one type of amino acid among these 20 types of amino acids. If a is the first type of amino acid (denoted as A1), the first sub-object set corresponding to amino acid A1 at mutation position j is V j (A1)={i|S ij =A1}.

[0121] In some embodiments, for each type of amino acid at each mutation position, the server can determine the amino acid characteristics at the mutation position of the amino acid based on the index experimental value of each object in the first sub-object set corresponding to the amino acid. For example, if the first sub-object set corresponding to amino acid A1 at mutation position j is V j (A1)={i|S ij =A1}, when calculating the amino acid feature at mutation position j of amino acid A1, V j (A1)={i|S ij =A1}, the experimental index values ​​of the subjects corresponding to the ordinal numbers of each subject can be used to calculate the amino acid characteristics at mutation position j of amino acid A1.

[0122] In this embodiment, for each type of amino acid at each mutation position, the amino acid characteristics at the mutation position of the amino acid are determined based on the experimental index values ​​of each subject in the first sub-subject set corresponding to the amino acid, so that the amino acid characteristics at different mutation positions of the same type of amino acid are related to the mutation position, i.e., the same amino acid has different characteristic expressions at different positions, for example, the characteristics of the same type of amino acid located at different mutation positions may be different from each other, thereby improving the accuracy of encoding for amino acids. The protein characteristic determination method provided in this embodiment can be used to encode mutant proteins generated in a k-site saturation mutagenesis scenario and obtain the protein characteristics of the mutant proteins.

[0123] The targeting method provided in this application can be applied to Bayesian optimization, which can be used to assist protein directed evolution. Bayesian optimization can effectively explore combinatorial space and find optimal solutions in the sample space by balancing exploration and exploitation with a small number of experimental samples and as few experiments as possible. However, in the application of Bayesian optimization, encoding strategies currently face several problems. On the one hand, high-dimensional encoding strategies are challenging for Bayesian optimization because a successful global optimization search requires an accurate and information-rich low-dimensional representation. On the other hand, classification labels (e.g., one-hot encoding) may result in missing knowledge about death variants from the available experimental data for a particular protein. This can be seen in Figure 4, where each letter corresponding to the abscissa represents one type of amino acid, e.g., V represents one type of amino acid. Figure 4 calculates the average fitness of 20 types of amino acids (AA, Amino Acid) at four mutation sites in 384 experimental samples (GB1 variants) selected from the GB1 dataset. The average fitness of each type of amino acid at each mutation site is obtained by calculating the average of the affinity measurements at that mutation site, and the corresponding standard deviation is shown as an error bar (i.e., a vertical line in Figure 4). As can be seen from Figure 4, the presence of some death variants at a particular mutation site can directly result in low or zero fitness, regardless of the choice of amino acids at other positions. Therefore, applying traditional protein encoding schemes to Bayesian optimization methods to support protein directed evolution is usually ineffective.

[0124] In contrast, according to the protein encoding method (i.e., the protein feature determination method) provided in the present application, the protein features obtained by encoding are accurate and information-rich low-dimensional features, so by applying the target determination method provided in the present application to Bayesian optimization, it is possible to rapidly assist protein directed evolution using Bayesian optimization.

[0125] In some embodiments, the step of determining an amino acid characteristic at a mutation position of an amino acid based on an index experimental value of each subject in the first sub-subject set corresponding to the amino acid includes the steps of: calculating statistics for the index experimental value of each subject in the first sub-subject set corresponding to the amino acid to obtain at least one index experimental statistical value; and determining an amino acid characteristic at a mutation position of the amino acid based on the at least one index experimental statistical value.

[0126] The index experimental statistics may be one or more, where multiple refers to at least two. The calculation of statistics may include, but is not limited to, at least one of the following: average value, minimum value, or minimum value calculation.

[0127] Specifically, the server can perform statistical calculations on the indicator experimental values ​​of each object in the first sub-object set corresponding to the amino acid to obtain at least one indicator experimental statistical value, and determine the amino acid characteristics at the mutation position of the amino acid based on the at least one indicator experimental statistical value.

[0128] In some embodiments, the step of calculating statistics for the experimental index values ​​of each subject in the first sub-subject set corresponding to the amino acid to obtain at least one experimental index value comprises the steps of: calculating an average value for the experimental index values ​​of each subject in the first sub-subject set corresponding to the amino acid to obtain a first average index value; and determining a maximum experimental index value from the experimental index values ​​of each subject in the first sub-subject set corresponding to the amino acid to obtain a first maximum index value, wherein the at least one experimental index value comprises at least one of the first average index value or the first maximum index value. j (A1)={i|S ij =A1}, when calculating the amino acid feature at mutation position j of amino acid A1, V j (A1)={i|S ij=A1}, the experimental index values ​​of the target corresponding to each sequence number are obtained, the average of each obtained experimental index value is calculated to obtain the first average index value, and the maximum value is determined from each experimental index value to obtain the first maximum index value, the first average index value is an experimental index statistical value, and the maximum value is also an experimental index statistical value, and the amino acid characteristics at mutation position j of amino acid A1 are determined based on at least one of the first average index value or the first maximum index value. In this embodiment, the first average index value and the first maximum index value are obtained statistically, so that the characteristics of the amino acid can be reflected, and thus, by using the first average index value or the first maximum index value as the experimental index statistical value, the accuracy of the experimental index statistical value can be improved.

[0129] In some embodiments, the step of determining the amino acid feature at the mutation position of the amino acid based on at least one index experimental statistic includes the step of configuring the amino acid feature at the mutation position of the amino acid by a first index average value and a first index maximum value. For example, the amino acid feature may be configured using the first index average value and the first index maximum value as feature values, that is, the amino acid feature includes the first index average value and the first index maximum value. For example, the first index average value may be expressed as formula (1), and the first index maximum value may be expressed as formula (2). y in formula (1) and formula (2) i represents the experimental index value for a predetermined index of the i-th object in the reference object set. In formulas (1) and (2), a represents an amino acid. In this embodiment, the amino acid features at the mutation position of the amino acid are configured using the first index average value and the first index maximum value, so that the amino acid features can be related to the experimental index value, thereby improving the accuracy of the target objects screened based on the amino acid features.

[0130]

number

[0131] In this embodiment, statistics are calculated for the experimental indicator values ​​of each mutant protein in the first sub-subject set corresponding to the amino acid, and the amino acid features at the mutation position of the amino acid are determined based on the statistical experimental indicator values, and data statistics are performed to improve the accuracy of the amino acid features obtained by encoding.The protein feature determination method provided in this embodiment can be applied to encoding the mutant protein generated in the k-site saturation mutagenesis scenario to obtain the protein features of the mutant protein.

[0132] In some embodiments, the target feature is a protein feature, and the step of determining the target feature of each target in the reference target set based on the index experimental value for a predetermined index of each target in the reference target set includes the steps of: for each type of amino acid, determining from the reference target set the target of amino acids contained in the amino acid at the mutation position, and obtaining a second sub-target set corresponding to the amino acid; for each type of amino acid, determining the amino acid feature of the amino acid based on the index experimental value of each target in the second sub-target set corresponding to the amino acid; and obtaining the protein feature of the target based on the amino acid feature of the amino acid at each mutation position in the target.

[0133] The second sub-object sets are subsets of the reference object set, each of which includes at least some of the objects in the reference object set, and each second sub-object set corresponds to one type of amino acid, with different second sub-object sets differing by one amino acid.

[0134] Specifically, the objects in the reference object set are proteins, and the reference object set may include mutant proteins and wild-type proteins. For each type of amino acid, the server can determine the objects of the amino acid contained in the amino acid at the mutation position from the reference object set and construct a second subobject set corresponding to the amino acid. For example, for a first type of amino acid A1, for each object, the server determines the amino acid at each mutation position of the object to construct an amino acid set corresponding to the object. After obtaining the amino acid sets corresponding to each object in the reference object set, it determines an amino acid set containing the first type of amino acid A1 from the amino acid sets corresponding to each object, and constructs a second subobject set corresponding to the amino acid A1 using the objects corresponding to the amino acid set containing A1.

[0135] In some embodiments, for each object in the reference object set, for each type of amino acid, the server can determine a mutation position corresponding to each type of amino acid from the object, where when the amino acid at mutation position 1 is A1, the mutation position corresponding to amino acid A1 is mutation position 1. Each type of amino acid may correspond to 0, 1, or multiple mutation positions, where multiple refers to at least two, for example, for the i-th object S in the reference object set, ij For each type of amino acid, a set N i(a) may be expressed as formula (3), where j is the mutation position. For each type of amino acid, a second sub-object set corresponding to each type of amino acid can be determined based on the set of mutation positions corresponding to each type of amino acid in each amino acid. For example, the second sub-object set V(a) may be expressed as formula (4), where for amino acid a, the number of mutation positions corresponding to amino acid a in the i-th object (i.e., |N i If (a)|) is not 0, the i-th object is the object in the second sub-object set corresponding to amino acid a. i (a) | is the set N i Represents the number of elements contained in (a).

[0136]

number

[0137]

number

[0138] In this embodiment, since the amino acid at the mutation position of each mutant protein in the second sub-subset corresponding to a certain amino acid contains that amino acid, for each type of amino acid, the amino acid characteristics of the amino acid are obtained based on the experimental index values ​​of each mutant protein in the second sub-subset corresponding to that amino acid, and the amino acid is then encoded based on the experimental index values ​​of proteins containing that amino acid, thereby improving the accuracy of amino acid encoding.The protein characteristic determination method provided in this embodiment can be applied to encoding mutant proteins generated in a non-saturation mutagenesis scenario to obtain the protein characteristics of the mutant protein.

[0139] In some embodiments, the step of determining a target object that meets the index requirements of the specified index from the second object set based on the mapping relationship includes the following steps: determine a statistical index value for the target statistical index of each object in the second object set based on the mapping relationship, and determine a selected object from the second object set based on the statistical index value; if the iteration stopping (termination) condition is not satisfied, add the selected object to the reference object set; return to the step of determining the object features of each object in the reference object set based on the index experimental value for the specified index of each object in the reference object set, and perform this process until the iteration stopping condition is satisfied; and the selected object obtained when the iteration stopping condition is satisfied is the target object that meets the index requirements of the specified index.

[0140] Wherein, the mapping relationship between the predetermined index and the target feature is the first mapping relationship, and the iteration stopping condition includes at least one of the following: the number of iterations (i.e., the number of loops) reaches a number threshold, and the experimental value of the index of the selected object reaches a second index threshold, and the selected object can change from time to time, and the selected object determined in different loops is different.

[0141] Specifically, the server can perform statistical calculations based on the first mapping relationship to obtain a second mapping relationship between the target statistical indicators and the object features, and determine a statistical indicator value for the target statistical indicator of each object in the second object set based on the second mapping relationship, where the statistical indicator value refers to the value for the target statistical indicator of the object. For example, the second mapping relationship is represented by a curve y2=f2(x), and to determine the index statistical value for the target statistical indicator of an object, the value of y2 can be calculated when x is the object feature of the object in the curve y2=f2(x), and the value of y2 is the index statistical value for the target statistical indicator of the object.

[0142] In some embodiments, there may be one or more selection targets, and the target corresponding to the largest statistical index value may be determined as the selection target, or the target whose statistical index value is greater than a third index threshold may be determined as the selection target, and the third index threshold may be set according to needs. The server may obtain targets that satisfy index requirements of predetermined indexes based on the selection targets, for example, the server may determine the selection target as an object that satisfies index requirements of predetermined indexes.

[0143] In some embodiments, if the iteration stopping condition is not satisfied, the server may add the selected object to the reference object set and return to the step of determining the object characteristics of each object in the reference object set based on the experimental index value for the predetermined index of each object in the reference object set, and perform this process until the iteration stopping condition is satisfied, and the selected object obtained when the iteration stopping condition is satisfied may be determined as the target object that meets the index requirement of the predetermined index. For example, the iteration stopping condition may be that the number of iterations (i.e., the number of loops) reaches a threshold number of times, and in this case, the selected object is determined as the target object when the number of iterations (i.e., the number of loops) reaches the threshold number of times.

[0144] In this embodiment, if the iteration stopping condition is not satisfied, a new selection target is determined again, so that a target target that meets the index requirements of the specified index can be gradually found. In addition, by adding a selection target to the reference target set, the number of targets in the reference target set increases. Therefore, by performing a step of determining the target features of each target in the reference target set in the process of determining a new selection target each time, the accuracy of the target features obtained by encoding can be gradually improved, and the accuracy of the target target that is finally selected can be improved.

[0145] In some embodiments, as shown in FIG. 5, a target determination method is provided, in which the target of the method is a mutant protein, and the method may be executed by a terminal or a server, or may be jointly executed by a terminal and a server. The method is described as being applied to a server as an example, and the method includes the following steps:

[0146] Step 502: An initial score set is obtained, and the initial scores corresponding to each type of amino acid in the initial score set are the target number of times.

[0147] Step 504: In the initial score set, the initial scores corresponding to the amino acids at each mutation position in the wild-type protein are subtracted to obtain a current score set, and a first protein set is determined based on the wild-type protein, where the wild-type protein is a protein without mutations, and a second protein set is obtained based on the first target set.

[0148] Step 506: For each mutant protein in the second protein set, determine a current score corresponding to each amino acid at each mutation position in the mutant protein from the current score set, determine a current protein score for the mutant protein based on each obtained current score, and select and obtain a target protein from the second protein set based on the current protein score.

[0149] Step 508: Decrement the current scores corresponding to the amino acids at each mutation position in the target protein in the current score set, and transfer the target protein from the second protein set to the first protein set.

[0150] Step 510: Determine whether there is a score greater than 0 in the current score set; if yes, return to step 506; if no, execute step 512.

[0151] Step 512: Determine the first protein set as the reference set.

[0152] Please refer to Figure 6, which shows the principle of the target determination method, which is used to determine mutant proteins with relatively high affinity. The candidate sample space includes wild-type proteins and multiple mutant proteins. The reference target set in step 508 is the initial reference target set, which can be changed later. The initial reference target set is, for example, the initial sample set in Figure 6. According to the method for determining the initial reference target set provided in the present application, the initial sample set is obtained by screening from the candidate sample space, and the samples in the initial sample set are any one of wild-type proteins or mutant proteins.

[0153] Step 514: Train an index detection model based on the object features and index experimental values ​​of each object in the reference object set, and use the trained index detection model to predict the index predicted value of each object in the first object set, and screen out objects from the first object set whose index predicted value satisfies the index value screening conditions to obtain a second object set.

[0154] As shown in Figure 6, after obtaining the initial sample set, the "affinity" step obtains the affinity of the samples in the initial sample set (obtained by experimental measurement); after obtaining the affinity, the "protein feature representation" step encodes the proteins in the initial sample set, obtains the protein features of each protein in the initial sample set, and uses the protein affinity (obtained by experimental measurement) and protein features to train the indicator detection model; after training, the protein features corresponding to the proteins in the candidate sample space are input into the indicator detection waiting model, and prediction is made to obtain the affinity prediction value of the protein; and proteins with affinity prediction values ​​greater than the affinity threshold are selected from the candidate sample space to form the second target set.

[0155] By screening the initial sample set from the candidate sample space (i.e., by screening the second target set based on the first target set), a search space pre-screening strategy is provided. Because many variants in the candidate sample space have relatively low affinity values, adopting a sample search space pre-screening strategy can eliminate samples with low affinity values ​​in advance, reducing the sample space that needs to be searched in Bayesian optimization and improving computational efficiency. For example, by adopting XGBOD, low-affinity variants can be removed from the sample search space. Specifically, in each Bayesian optimization iteration, samples in the candidate sample space are first pre-screened. By setting a threshold, samples below the threshold can be determined as low fitness (removal point, postive class), and samples above the threshold can be determined as high fitness (non-removal point, negative class). By training XGBOD using existing experimental samples (i.e., proteins with experimentally determined affinities) as a training set, and screening samples in the candidate sample space to eliminate potential sample points with very low affinities in advance, the number of samples in the sampling space can be reduced, improving the efficiency of the model.

[0156] The process of determining one selection target from the second target set each time can be considered as one Bayesian optimization. For example, in Figure 6, the process of using the acquisition function each time to screen samples from the remaining samples after samples with relatively low affinity have been removed is one Bayesian optimization.

[0157] Step 516: Based on the index experimental value for the predetermined index of each object in the reference object set, determine the object feature of each object in the reference object set, and based on the index experimental value for the predetermined index of each object in the reference object set and the object feature, determine the mapping relationship between the predetermined index and the object feature.

[0158] Among them, the mapping relationship between the predetermined index and the target feature is, for example, the probabilistic surrogate model in FIG. 6, and the probabilistic surrogate model may be any one of a Gaussian process regression model based on Gaussian distribution, a Gaussian process regression model based on Student-t distribution, etc.

[0159] Step 518: Determine a statistical index value for the target statistical index of each object in the second object set based on the mapping relationship, and determine a selected object from the second object set based on the statistical index value.

[0160] As shown in FIG. 6, the target statistical index is, for example, the acquisition function in FIG. 6, and the mapping relationship between the predetermined index and the target feature is used to determine the acquisition function, and a sample is selected from the remaining samples based on the acquisition function.

[0161] Step 520: Determine whether the iteration stopping condition is satisfied; if no, execute step 522; if yes, execute step 524.

[0162] Step 522: Add the selected object to the reference object set, and return to step 514.

[0163] Step 524: The selected object obtained when the iteration stopping condition is satisfied is determined as the target object that satisfies the index requirement of the predetermined index.

[0164] For example, the server may use an algorithm such as Bayesian optimization with prescreened search space via outlier detection (ODBO), which implements the process of determining a target object from a second object set based on a reference object set.

[0165] The input data for the algorithm is the initial sample set D t and the number of experiments is T (i.e., the number of loops is T).

[0166] Among them, the first sample set D t is the initial referent set, i.e., the referent set employed in the first iteration.

[0167] The output data of the algorithm is the optimum (s*, y*) in the sample space, where s* represents the optimum variant and y* represents the affinity of s* (obtained by experimental measurement).

[0168] The process of the algorithm is as follows:

[0169] T←1; / / Let 1 be t, where t represents the number of the current iteration; while t≦T do; / / Continue executing if t≦T, where T is the number of experiments (i.e., the number of loops); if Robust GP then; / / Run if Robust GP is used as the probabilistic surrogate model; D t Train a Gaussian process regression model based on the Student-t distribution using D t represents the initial sample set, i.e., the initial referent set (the referent set adopted at the first iteration), based on the fact that is equal to D1; D tin ={(S i ,y i )│|f t (S i )-y i |≦α}, and remove outliers based on the rejection threshold α, leaving normal points; / / D tin is D t represents the set of samples remaining after removing outliers (samples) from D tin Train a Gaussian process regression model based on Gaussian distribution using; else if GP then; / / Continue if Robust GP is used as the probabilistic surrogate model; D tUtilizing to train a Gaussian process regression model based on Gaussian distribution; if Naivo BO then; / / Run if Naive Bayes Optimization is used, where Naivo BO refers to Naive Bayes Optimization; Acquisition function based on current posterior probability

[0170]

number

[0171]

number

[0172] Among them, TuRBO is a global optimization method that constructs a series of local GP surrogate models to avoid over-searching the height uncertainty region in the search space from a global perspective, and makes full use of the quadratic convergence of the local reliable method to achieve a highly efficient solution.

[0173] An example will be used to explain the basic flow of the target determination method provided in this application. As shown in Figure 7, it mainly includes four steps: 1) obtaining initial experimental data; 2) obtaining a feature representation of the data; 3) performing pre-screening of the search space; and 4) using a Bayesian optimization algorithm to train a probabilistic surrogate model based on the initial experimental data for the screened search space. After training the surrogate model, the acquisition function is optimized to select experimental samples for the next round from the search space. Validation is performed for the given experimental design, the experimental results are added to the training set, and the posterior probability of the surrogate model is updated. This process is repeated until the design is maximized, resources are exhausted, or no conditions for space exploration improvement can be found. Figure 7(a) shows the initial experimental data obtained, i.e., the initial reference target set screened from the first target set. Eight mutants are shown in (a), each with a corresponding score, which represents fitness. For example, the mutant "H76L,K78R" represents one mutant, and the score of "H76L,K78R" is 0.18. Figure 7(b) shows the feature representation of the data (i.e., the process of determining amino acid features by encoding). The bar graph in Figure 7(b) shows the average fitness at the i-th mutation site of 20 amino acids, and the table in Figure 7(b) shows the average fitness at the i-th mutation site of five amino acids. For example, the average fitness of the amino acid "V" is 1.12. Figure 7(c) shows the process of search space pre-screening, in which the characteristics of mutants in the initial experimental data are determined, and "P1 P2 A1 A2" in Figure 7(c) are the mutant characteristics, where P is an abbreviation for position, A is an abbreviation for Amino Acid, P1 and P2 respectively represent the positions of amino acids in the mutant, and A1 and A2 respectively represent the amino acid characteristics. The determined mutant characteristics are used to perform pre-screening on the search space, and in Figure 7(c), solid circles represent outliers (i.e., mutants with relatively low fitness) and solid triangles represent normal points (i.e., mutants with relatively high fitness).Figure 7(d) shows the process of the Bayesian optimization algorithm. The four processes in Figure 7 allow us to determine highly fit mutants.

[0174] Regarding computationally assisted experimental design, the technical solution of this application provides a framework for highly efficient experimental design. A search space prescreening strategy pre-screens samples in the candidate sample space (i.e., screens from a first target set to obtain a second target set) and, combined with a Bayesian optimization algorithm, balances exploration and exploitation, effectively searches the sample space, and finds the optimal experimental design scheme in as few steps as possible. For the practical application scenario of protein directed evolution, the technical solution of this application designs an average fitness-based amino acid encoding strategy (i.e., a method for obtaining amino acid features) to accurately and effectively realize feature representation. To better assist experimenters in experimental design, a first sample screening strategy (i.e., a method for determining a reference target set) is also provided to help experimenters select the initial experimental samples, thereby ensuring that the coverage of amino acid encoding information contained in the initial number of samples is maximized and the number of initial experiments required is minimized. The Bayesian optimization algorithm for search space prescreening can reduce experimental costs and time costs. The technical proposal of this application realizes a framework for highly efficient experiment design, called ODBO (Bayesian optimization with prescreened search space via outlier detection). This method assists experiment design through search space screening plus Bayesian optimization, helping experimenters reduce experimental costs and time costs. For the practical adaptation scenario of protein directed evolution, an amino acid encoding strategy based on average fitness is provided to achieve accurate and effective feature representation.In order to better assist experimenters in experimental design, the present technical solution also provides an initial sample screening strategy to assist experimenters in selecting the initial experimental samples, which can ensure that the coverage of amino acid coding information contained in the initial number of samples is maximized and the number of required experiments is minimized, thereby reducing experimental costs.

[0175] The present invention can be used to perform computationally assisted protein directed evolution experimental design, while the proposed Bayesian optimization combined with search space prescreening can also be applied to other fields of automated experimental design, such as new materials development and fast charging protocols for batteries.

[0176] In materials science, discovering and creating materials with specific properties is expensive and time-consuming. With each new composition or material parameter, the space of candidate experiments grows exponentially. For example, to study the effect of one new parameter (e.g., the effect of introducing a dopant), approximately 10 experiments must be performed within a range of parameters, so that N parameters are 10 N Many experiments are required. With the emergence of each new parameter, the number of candidate experiments quickly exceeds the feasibility of exhaustive search. The diversity and complexity of material composition-structure-property (CSP) relationships, as well as the chaos of materials-processing parameters and atoms, can further confound research. Combined with the lack of optimal materials, these challenges threaten innovation and industrial progress. A Bayesian optimization-based materials discovery support method can guide laboratory experimenters in experimental design, accelerating the speed of material discovery and expending fewer resources in materials exploration experiments by balancing the exploration of unknown functions using experiments with the identification of extreme values ​​using prior knowledge.

[0177] Lithium-ion batteries are one of the most commonly used energy storage devices in electric vehicles. With advances in battery chemistry, a key issue is how to effectively determine charging protocols that balance the need for fast charging with maximizing battery life. However, determining an appropriate charging protocol is not easy. On the one hand, evaluating a battery's cycle life requires months to years. On the other hand, the vast parameter tuning space and the diversity of samples make experiments even more challenging. Narrowing the parameter range and shortening experimental time are crucial for the development of lithium-ion batteries. Computational experimental design can be used to reduce the cost of experimental optimization, and feedback from completed experiments can be used to inform subsequent experimental decisions, balancing experimental results and needs. That is, experimental parameter space with high uncertainty can be tested and explored, and promising parameters can be predicted based on completed experimental results. Ultimately, this can reduce the number and time required for experiments, reduce costs, and achieve the discovery of an effective charging protocol.

[0178] This application also provides one application scenario, which is a new material development scenario. The above-described target determination method is used in this application scenario. Specifically, the target determination method is applied in this application scenario as follows: the server obtains predicted index values ​​for each predetermined index of each material in a first material set; selects materials from the first material set whose predicted index values ​​satisfy the index value screening conditions to obtain a second material set; determines a mapping relationship between the predetermined index and material characteristics based on the experimental index values ​​and material characteristics of the predetermined index of the materials in the first material set; and determines a target material from the second material set that meets the index requirements of the predetermined index based on the mapping relationship. This allows for rapid identification of materials with specified properties. Among these, the materials in the first material set differ in at least one composition or the content of at least one composition. As shown in Figure 8, closed-loop optimization of perovskite electrolytes based on machine learning is shown. Bayesian optimization can be used to discover an effective experimental search for high-speed lithium-ion conductors from perovskite solid electrolytes.

[0179] This application also provides another application scenario, which is a battery fast charging protocol scenario, in which the above-described target determination method is used. Specifically, the target determination method is applied in this application scenario as follows: the server obtains predicted indicator values ​​for each predetermined indicator of each battery charging protocol in a first battery charging protocol set; selects battery charging protocols from the first battery charging protocol set whose predicted indicator values ​​satisfy the indicator value screening conditions to obtain a second battery charging protocol set; obtains a mapping relationship between the predetermined indicators and battery charging protocol characteristics based on the experimental indicator values ​​and battery charging protocol characteristics for the predetermined indicators of the multiple battery charging protocols in the first battery charging protocol set; and determines a target battery charging protocol from the second battery charging protocol set that meets the indicator requirements of the predetermined indicator based on the mapping relationship. This allows for rapid determination of a battery charging protocol with specified performance. Among these, at least one parameter of each battery charging protocol in the first battery charging protocol set is different, or at least one parameter value is different. As shown in Figure 9, closed-loop optimization of a battery fast charging protocol based on machine learning is illustrated. Machine learning methods can effectively optimize the parameter space and identify the current and voltage setting parameters of the fast charging protocol to maximize battery life.

[0180] The target determination method provided in this application is based on the Python language and the Botorch library and can be configured on a server equipped with a Linux operating system or a Windows operating system and CPU / GPU computing resources.

[0181] To verify the effectiveness of the target determination method provided in this application in supporting protein directed evolution, tests were conducted on four protein directed evolution datasets: 1) GB1 dataset (containing 55 mutation sites); 2) GB1 dataset (containing 4 mutation sites); 3) BRCA1 dataset; and 4) green fluorescent protein dataset.

[0182] Among them, GB1 refers to the B1 structural domain of protein G. Protein G is an immunoglobulin-binding protein expressed in C and G streptococci. The B1 structural domain of protein G (GB1) interacts with the Fc structural domain of immunoglobulins. Experiments were conducted on the generated GB1 dataset. Saturation mutagenesis was performed on four carefully selected residue sites, 39, 40, 41, and 51, from GB1. 149,361 mutants had experimentally measured fitness values. The fitness standard was the binding affinity with IgG-Fc. One or two amino acids were mutated across the entire 55-codon random region of the GB1 protein, and a total of 536,944 mutant data were collected.

[0183] BRCA1 is a multidomain protein belonging to the tumor suppressor gene family, with three domains most commonly mutated: the N-terminal RING domain, exons 11-13, and the BRCT domain. The BRCA1 RING structural domain is responsible for the E3 ubiquitin ligase activity of BRCA1 and mediates interactions between BRCA1 and other proteins. The effects of single or multiple point mutations in BRCA1 residues on the function of E3 ubiquitin ligase activity have been investigated. The dataset contains a total of 98,300 variants with E3 scores.

[0184] Green fluorescent protein (GFP), also known as green fluorescent protein, was first discovered in the jellyfish Aequorea victoria (avGFP) and emits green fluorescence when exposed to light. The local fitness landscape of avGFP was analyzed by estimating the fluorescence levels of genotypes obtained by random mutagenesis of the avGFP sequence. The dataset contains 54,025 distinct protein sequences. Table 1 provides detailed information on the four datasets used.

[0185] Table 1 provides detailed information on the protein directed evolution dataset.

[0186] [Table 1] Figure 10 shows the fitness distributions of different datasets. In Figure 10, the abscissa is the metric value, and the ordinate is density. Figure 10(a) shows the fitness distribution of dataset GB1(4), Figure 10(b) shows the fitness distribution of dataset GB1(55), Figure 10(c) shows the fitness distribution of dataset BRCA1, and Figure 10(d) shows the fitness distribution of dataset avGFP.

[0187] An initial sample screening strategy was employed to generate the initial sample set. For the GB1(4) dataset, which conforms to a saturation mutation scenario, each amino acid was set to occur at least twice at each position, resulting in 40 initial training samples. For the GB1(55), Ube4b, and avGFP datasets, which conform to a non-saturation mutation scenario, each amino acid was set to occur at least once at every position, resulting in 136, 217, and 142 initial training samples, respectively. For the ODBO algorithm, the filtering threshold for search space prescreening was set to 0.05. For each method, 10 different random seeds were used for each experiment. For the GB1(55), Ube4b, and avGFP datasets, each method screened one sample from the sample space each time, repeating 50 times. For the GB1(55) dataset, each method screened one sample from the sample space each time, repeating 100 times. Expected improvement (EI) was employed as the acquisition function. Ube4b and BRCA1 refer to the same protein.

[0188] Figure 11 summarizes the performance of four protein-directed evolution datasets for different methods. Dataset 1 refers to dataset GB1(4), dataset 2 refers to dataset GB1(55), dataset 3 refers to dataset Ube4, and dataset 4 refers to avGFP. Method 1 refers to the random screening method, method 2 refers to the combination of TuRBO and GP, method 3 refers to the combination of ODBO, uRBO, and GP, method 4 refers to the combination of ODBO, TuRBO, and RobustGP, and method 5 refers to the combination of ODBO, TuRBO, and RobustGP. (outside 1) JPEG0007746656000013.jpg16121 refers to the combination of ODBO with GP, Method 6 refers to the combination of ODBO with BO and GP, and Method 7 refers to the combination of ODBO with BO and RobustGP. Each of the four graphs in Figure 11 contains a line, which indicates true maximum fitness.

[0189] Figure 12 summarizes the comparison of four protein-directed evolution datasets using different methods, with each curve representing the average value obtained by each method using 10 different random seeds. F1 represents the combination of ODBO with TuRBO and GP, with q = 1; F2 represents the combination of ODBO with TuRBO and GP, with q = 5; F3 represents the combination of ODBO with TuRBO and GP, with q = 10; F4 represents the combination of ODBO with TuRBO and RobustGP, with q = 1; F5 represents the combination of ODBO with TuRBO and RobustGP, with q = 5; and F6 represents the combination of ODBO with TuRBO and RobustGP, with q = 10. G1 represents the combination of ODBO with TuRBO and GP, and the acquisition function is the expected improvement; G2 represents the combination of ODBO with TuRBO and GP, and the acquisition function is the upper confidence limit; G3 represents the combination of ODBO with TuRBO and GP, and the acquisition function is Thompson sampling. Q represents the number of samples selected at each iteration to perform the next round of experiments.

[0190] Therefore, we can find that ODBO achieves the best performance in all datasets. The search space pre-screening step can provide more effective sample collection and is advantageous for finding mutants with optimal properties faster. For example, for the saturated mutation scenario (i.e., GB1(4) dataset), the combined method of ODBO with TuRBO and RobustGP can obtain the best variant (fitness = 8.76) in one large sample space (204 = 16,000) with less than 50 evaluations. However, in methods that do not employ a pre-screening strategy, Bayesian optimization algorithms (e.g., (outside 2) TuRBO (TuRBO, TuRBO) can usually converge to a single bad local optimum, resulting in poor average performance. This demonstrates the importance of search space prescreening. For the case of non-saturating mutations, (Outside 3) Besides the BO-GP combination method, almost all Bayesian optimization methods use a given low-dimensional protein encoding strategy to find optimal variants. While all methods can only find near-optimal variants in the GB1(55) and avGFP datasets, the method proposed in this technical proposal outperforms other methods.

[0191] Table 2 shows the proportion of samples belonging to the previous 1%, 2%, and 5% affinity values ​​in the sample space screened by different calculation methods in the GB1(4) dataset in 50 rounds of recommendation screening. As can be seen, the adoption of sample space pre-screening is advantageous in selecting better samples from each round of sample screening for the next round of experimental testing.

[0192] [Table 2] The performance of Bayesian optimization algorithms with different acquisition functions and probabilistic surrogate models in protein-directed evolution was also tested. Figure 10 shows the performance of Bayesian optimization algorithms with different acquisition functions and probabilistic surrogate models on the GB1(4) dataset. Figure 10(a) shows the performance of the "ODBO,TuRBO+GP" method when using EI, UCB, PI, and TS as acquisition functions, respectively. Figure 10(b) shows the performance of the "ODBO,TuRBO+RobustGP" method when using EI, UCB, PI, and TS as acquisition functions, respectively.

[0193] Additionally, the computational resources consumed when each method was run on the GB1 dataset were calculated, as shown in Table 3. When using a conventional encoding scheme (shown here as feature Georgiev, which encodes using physicochemical properties), there are 76 feature dimensions, and using TuRBO requires significant computational resources and time. In contrast, when using the amino acid encoding scheme presented here, the feature dimension of amino acids can be reduced to four dimensions, significantly reducing computational time and resource consumption. Furthermore, by employing a search space prescreening strategy, the computational time and resources consumed can be significantly reduced. Furthermore, ODBO can find the optimum in the sample space with the fewest experimental steps, which is advantageous for reducing experimental and time costs.

[0194] [Table 3] It should be understood that although the steps in the flowcharts according to the above-described embodiments may be displayed sequentially according to the direction of the arrows, these steps are not necessarily executed sequentially in the order indicated by the arrows. Unless otherwise specified, the execution of these steps is not restricted to the order, and these steps may be executed in other orders. Furthermore, at least some of the steps in the flowcharts according to the above-described embodiments may include multiple steps or multiple stages, and these steps or stages may be executed at the same time or at different times. Furthermore, these steps or stages may be executed sequentially or alternately with other steps or at least some of the steps or stages in other steps.

[0195] Furthermore, based on the same inventive concept, an embodiment of the present application further provides an object determination device for realizing the above-mentioned object determination method. The problem-solving implementation scheme provided by the device is similar to the implementation scheme described in the above-mentioned method, so that the specific limitations of one or more object determination device embodiments provided below can refer to the limitations of the above-mentioned object determination method, and detailed descriptions thereof will be omitted here.

[0196] In some embodiments, as shown in FIG. 13, an object determination device is provided, which includes a predicted value obtaining module 1302, an object set obtaining module 1304, a mapping relationship determining module 1306, and a target object determining module 1308.

[0197] The prediction value obtaining module 1302 is used to obtain an index prediction value for each predetermined index for each object in the first object set.

[0198] The object set acquisition module 1304 is used to screen objects whose index prediction values ​​satisfy the index value screening conditions from the first object set, and acquire a second object set.

[0199] The mapping relationship determination module 1306 is used to determine a mapping relationship between the predetermined index and the target feature based on the index experimental values ​​for the predetermined index of a plurality of objects in the first object set and the target feature.

[0200] The target object determination module 1308 is used to determine target objects that satisfy the index requirements of the predetermined index from the second object set based on the mapping relationship.

[0201] In some embodiments, the objects in the first object set are mutant proteins, and the apparatus further includes a reference object set screening module for obtaining a reference object set by screening based on the first object set, wherein the reference object set satisfies the condition that each type of amino acid appears at least a target number of times at each mutation position, and the predicted value acquisition module further trains an index detection model based on the object features and the index experimental value of each object in the reference object set; and uses the trained index detection model to predict the index predicted value of each object in the first object set.

[0202] In some embodiments, the mapping relationship determination module is further used to obtain object features of each object in the reference object set based on the index experimental value for the predetermined index of each object in the reference object set; and determine a mapping relationship between the predetermined index and the object feature based on the index experimental value for the predetermined index of each object in the reference object set and the object feature.

[0203] In some embodiments, the reference target set screening module further obtains a current score set, where the current score set includes a current score corresponding to each type of amino acid; obtains a second protein set based on the first target set, and selects a target protein from the second protein set based on the current score set; reduces the current scores corresponding to each amino acid at each mutation position in the target protein in the current score set, and transfers the target protein from the second protein set to the first protein set; if the current score set indicates that the first protein set does not satisfy the condition that "each type of amino acid occurs at least the target number of times at each mutation position", returns to the step of selecting a target protein from the second protein set based on the current score set, and repeats this process until the current score set indicates that the first protein set satisfies the condition that "each type of amino acid occurs at least the target number of times at each mutation position", and then uses the first protein set to determine the reference target set.

[0204] In some embodiments, the reference set screening module further obtains an initial score set, and the initial scores corresponding to each type of amino acid in the initial score set are target numbers; and reduces the initial scores corresponding to each amino acid at each mutation position in the wild-type protein in the initial score set to obtain a current score set, which is used to determine a first protein set based on the wild-type protein, where the wild-type protein is a protein without mutations.

[0205] In some embodiments, the mapping relationship determination module is further used to determine, for each mutant protein in the second protein set, a current score from the current score set, corresponding to each amino acid at each mutation position in the mutant protein; determine a current protein score for the mutant protein based on each obtained current score; and select and obtain a target protein from the second protein set based on the current protein score.

[0206] In some embodiments, each type of amino acid has a corresponding amino acid, and the scores in the current score set are uniquely labeled by the amino acid and the mutation position, and the mapping relationship determination module is further used to determine, for the amino acid at each mutation position, a current score corresponding to the amino acid at the mutation position from the current score set based on the amino acid corresponding to the amino acid and the mutation position.

[0207] In some embodiments, the target feature is a protein feature, and the mapping relationship determination module further divides the reference target set according to the type of amino acid at each mutation position, and obtains a first sub-target set corresponding to each type of amino acid; for each type of amino acid at each mutation position, determine the amino acid characteristics at the mutation position of the amino acid based on the index experimental value of each target in the first sub-target set corresponding to the amino acid; and obtains the target protein feature based on the amino acid characteristics of the amino acid at each mutation position in the target.

[0208] In some embodiments, the mapping relationship determination module further calculates statistics for the index experimental values ​​of each object in the first sub-object set corresponding to the amino acid to obtain at least one index experimental statistical value; and is used to determine an amino acid characteristic at the mutation position of the amino acid based on the at least one index experimental statistical value.

[0209] In some embodiments, the mapping relationship determination module is further used to calculate an average value for the index experimental values ​​of each object in the first sub-object set corresponding to the amino acid to obtain a first index average value; and to determine the maximum index experimental value from the index experimental values ​​of each object in the first sub-object set corresponding to the amino acid to obtain a first index maximum value, and the at least one index experimental statistical value includes at least one of the first index average value or the first index maximum value.

[0210] In some embodiments, the mapping relationship determination module is further used to configure an amino acid signature at the mutation position of the amino acid according to the first index average value and the first index maximum value.

[0211] In some embodiments, the target feature is a protein feature, and the mapping relationship determination module further determines, for each type of amino acid, an amino acid target included in the amino acid at the mutation position from the reference target set, to obtain a second sub-target set corresponding to the amino acid; for each type of amino acid, determines an amino acid characteristic of the amino acid based on the index experimental value of each target in the second sub-target set corresponding to the amino acid; and is used to obtain a target protein feature based on the amino acid characteristic of the amino acid at each mutation position in the target.

[0212] In some embodiments, the target object determination module further determines, based on the mapping relationship, statistical index values ​​for target statistical indexes of each object in the second object set, and determines a selected object from the second object set based on the statistical index values; if the iteration stopping condition is not satisfied, adds the selected object to the reference object set, and returns to the step of determining the object features of each object in the reference object set based on the index experimental values ​​for the predetermined indexes of each object in the reference object set, and performs this process until the iteration stopping condition is satisfied; and is used to determine the selected object obtained when the iteration stopping condition is satisfied as a target object that meets the index requirements of the predetermined index.

[0213] Each module in the above-mentioned object determination device may be realized in whole or in part by software, hardware, or a combination thereof. Each module may be embedded in the form of hardware, or may be separate from the processor in a computer device, or may be stored in the form of software in a memory device in a computer device, so that the processor can call and execute the operations (steps) corresponding to each module.

[0214] In some embodiments, a computer device is provided, which may be a server and whose internal configuration may be as shown in FIG. 14 . The computer device may include a processor, a memory, an input / output interface (I / O), and a communication interface. The processor, memory, and I / O interface are connected via a system bus, and the communication interface is connected to the system bus via the I / O interface. The processor of the computer device is used to provide calculation and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, computer-readable instructions, and a database. The internal memory is used to provide an environment for executing the operating system and the computer-readable instructions in the non-volatile storage medium. The database of the computer device is used to store data related to the object determination method. The I / O interface of the computer device is used for the processor to exchange information with external devices. The communication interface of the computer device is used for communicating with external terminals via a network connection. The computer-readable instructions, when executed by the processor, are used to implement the object determination method.

[0215] In some embodiments, a computing device is provided, which may be a terminal and whose internal configuration may be as shown in FIG. 15 . The computing device may include a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. The processor, the memory, and the input / output interface are connected via a system bus, and the communication interface, the display unit, and the input device are connected to the system bus via the input / output interface. The processor of the computing device is used to provide calculation and control capabilities. The memory of the computing device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer-readable instructions. The internal memory is used to provide an environment for executing the operating system and computer-readable instructions in the non-volatile storage medium. The input / output interface of the computing device is used by the processor to exchange information with an external device. The communication interface of the computing device is used to communicate with an external terminal via a wired or wireless method, and the wireless method may be realized by Wi-Fi, a mobile cellular network, NFC, or other technologies. The computer-readable instructions, when executed by a processor, are used to realize the object determination method. The display unit of the computer device is used to generate a visible image, and may be a display screen, a projection device, or a virtual reality imaging device, the display screen may be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device may be a touch layer (film) covering the display screen, or may be a key, trackball, or touch panel installed on the outer shell of the computer device, or may even be an external keyboard, touch panel, or mouse, etc.

[0216] As will be understood by those skilled in the art, the configurations shown in Figures 14 and 15 are only some of the configurations related to the technical solution of the present application, and do not limit the computer equipment to which the technical solution of the present application is applied; a specific computer equipment may include more or fewer components than those shown, or may combine some components, or may have a different component layout.

[0217] In some embodiments, a computing device is provided that includes a memory and one or more processors, the memory having computer-readable instructions stored therein that, when executed by the processors, cause the one or more processors to perform steps in the target determination method described above.

[0218] In some embodiments, one or more non-volatile readable storage media are provided having computer-readable instructions stored therein that, when executed by one or more processors, cause the one or more processors to perform the steps in the target determination method described above.

[0219] In some embodiments, a computer program product is provided that includes computer readable instructions that, when executed by a processor, can implement the steps in the target determination method described above.

[0220] All user information (including, but not limited to, user device information, user personal information, etc.) and data (including, but not limited to, data for analysis, data for storage, data for display, etc.) related to this application are information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0221] As will be understood by those skilled in the art, all or part of the steps in the methods described above can be performed by instructing associated hardware using computer-readable instructions. The computer-readable instructions can be stored in a non-volatile computer-readable storage medium, and the computer-readable instructions, when executed, can include the steps of the method embodiments described above. Any references to storage devices, databases, or other media used in the embodiments provided herein may include at least one of non-volatile and volatile storage devices. Non-volatile storage devices may include read-only memory (ROM), magnetic tape, floppy disks, flash memory, optical memory, high-density embedded non-volatile memory, ReRAM, magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile storage devices may include random access memory (RAM) or external high-speed cache memory, etc. For example, and not by way of limitation, RAM may be in multiple forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database according to each embodiment provided herein may include at least one of a relational database and a non-relational database. The non-relational database may be, but is not limited to, a blockchain-based distributed database. The processor according to each embodiment provided herein may be, but is not limited to, a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a quantum computing-based data processing logic device, or the like.

[0222] Although the preferred embodiment of the present application has been described above, the present application is not limited to this embodiment, and any modification to the present application falls within the technical scope of the present application as long as it does not depart from the spirit of the present application.

Claims

1. 1. A computer-implemented method for determining an object, comprising: obtaining a predicted index value for each predetermined index for each subject in the first subject set, the predicted index value being a value for the predetermined index for the subject predicted by a pre-trained model; screening the objects whose predicted index values ​​satisfy the index value screening conditions from the first object set to obtain a second object set; determining a mapping relationship between the predetermined index and the target feature based on the index experimental values ​​of the predetermined index and the target feature of a plurality of objects in the first target set, wherein the index experimental values ​​refer to the values ​​of the objects for the predetermined index obtained by experiment; and determining, based on the mapping relationship, a target object from the second object set that satisfies the index requirement of the predetermined index; The method, wherein the objects in said first object set are mutant proteins.

2. 10. The method of claim 1, further comprising: obtaining a reference target set by screening based on the first target set, wherein the reference target set satisfies the condition that "each type of amino acid appears at least a target number of times at each mutation position"; obtaining a predicted index value for each predetermined index for each subject in the first subject set, training an indicator detection model based on the object features and indicator experimental values ​​of each object in the reference object set; and utilizing the trained indicator detection model to predict an indicator prediction value for each object in the first object set.

3. 3. The method of claim 2, The step of determining a mapping relationship between the predetermined index and the target feature based on the index experimental values ​​for the predetermined index of the plurality of objects in the first target set and the target feature includes: determining a target characteristic for each of the targets in the reference target set based on an index experimental value for the predetermined index for each of the targets in the reference target set; and determining a mapping relationship between a predetermined index and a target feature based on an index experimental value for the predetermined index and the target feature for each subject in the reference subject set.

4. 3. The method of claim 2, obtaining a reference object set by screening based on the first object set, obtaining a current score set, the current score set including current scores corresponding to each type of amino acid; obtaining a second protein set based on the first object set, and selecting a target protein from the second protein set based on a current score set; decrementing the current scores corresponding to the amino acids at each mutation position in the target protein in the current score set, and transferring the target protein from the second protein set to the first protein set; and The method includes a step of returning to and continuing to select a target protein from the second protein set based on the current score set when the current score set indicates that the first protein set does not satisfy the condition that "each type of amino acid occurs at least the target number of times at each mutation position," and a step of determining the first protein set as a reference target set when the current score set indicates that the first protein set satisfies the condition that "each type of amino acid occurs at least the target number of times at each mutation position."

5. 5. The method of claim 4, The step of obtaining the current score set includes: obtaining an initial score set, wherein the initial scores corresponding to each type of amino acid in the initial score set are target numbers; and The method includes a step of reducing the initial scores corresponding to the amino acids at each mutation position in the wild-type protein in the initial score set to obtain a current score set, and determining a first protein set based on the wild-type protein, wherein the wild-type protein is a protein in which no mutations occur.

6. 5. The method of claim 4, selecting a target protein from the second protein set based on the current score set, determining, for each mutant protein in said second set of proteins, a current score from a set of current scores corresponding to the amino acid at each mutation position in said mutant protein; determining a current protein score for the mutant protein based on each of the obtained current scores; and obtaining a target protein by screening from said second protein set based on the current protein score.

7. 7. The method of claim 6, each type of amino acid has a corresponding amino acid, and scores in the current score set are uniquely labeled by amino acid and mutation position; determining a current score corresponding to each amino acid at each mutation position in the mutant protein from the current score set, For each amino acid at the mutation position, determining a current score corresponding to the amino acid at the mutation position from a current score set based on the amino acid corresponding to the amino acid and the mutation position.

8. 4. The method of claim 3, the target feature is a protein feature; The step of determining the object characteristics of each object in the reference object set based on the index experimental value for the predetermined index of each object in the reference object set includes: for each mutation position, dividing the reference object set according to the type of amino acid at the mutation position, and obtaining a first sub-object set respectively corresponding to each type of amino acid; For each type of amino acid at each of the mutation positions, determining an amino acid signature at the mutation position of the amino acid based on the experimental index value of each subject in the first sub-subject set corresponding to the amino acid; and obtaining a protein signature for the subject based on an amino acid signature of an amino acid at each mutation position in the subject.

9. 9. The method of claim 8, determining an amino acid signature at the mutation position of the amino acid based on an experimental index value of each subject in the first sub-subject set corresponding to the amino acid, performing a statistical calculation on the index experimental value of each subject in the first sub-subject set corresponding to the amino acid to obtain at least one index experimental statistical value; and determining an amino acid characteristic at the mutation position of the amino acid based on the at least one indicative experimental statistic.

10. 10. The method of claim 9, The step of calculating statistics for the index experimental value of each subject in the first sub-subject set corresponding to the amino acid to obtain at least one index experimental statistical value includes: calculating an average value for the index experimental value of each subject in the first sub-subject set corresponding to the amino acid to obtain a first index average value; and The method includes a step of determining a maximum index experimental value from the index experimental values ​​of each subject in the first sub-subject set corresponding to the amino acid and obtaining it as a first index maximum value, wherein the at least one index experimental statistical value includes at least one of the first index average value or the first index maximum value.

11. 11. The method of claim 10, determining an amino acid signature at the mutation position of the amino acid based on the at least one indicative experimental statistic, The method includes a step of constructing an amino acid signature at the mutation position of the amino acid by the first index average value and the first index maximum value.

12. 4. The method of claim 3, the target feature is a protein feature; The step of determining the object characteristics of each object in the reference object set based on the index experimental value for the predetermined index of each object in the reference object set includes: For each type of amino acid, determining a target of the amino acid from the reference target set that includes the amino acid at the mutation position, and obtaining a second sub-target set corresponding to the amino acid; For each type of amino acid, determining an amino acid signature for the amino acid based on the index experimental value of each subject in the second sub-subject set corresponding to the amino acid; and obtaining a protein signature for the subject based on an amino acid signature of an amino acid at each mutation position in the subject.

13. 4. The method of claim 3, The step of determining, based on the mapping relationship, a target object from the second object set that satisfies the index requirement of the predetermined index, includes: determining a statistical index value for a target statistical index of each object in the second object set based on the mapping relationship, and determining a selection object from the second object set based on the statistical index value; adding the selected object to a reference object set if the iteration stopping condition is not satisfied; Returning to the step of determining the object features of each object in the reference object set based on the index experimental value for the predetermined index of each object in the reference object set, and continuing to perform the step until an iteration stopping condition is satisfied; and The method includes a step of determining a selected object obtained when an iteration stopping condition is satisfied as a target object that satisfies an index requirement of the predetermined index.

14. A device for determining an object, comprising: a prediction value obtaining module for obtaining an index prediction value for each predetermined index of each subject in the first subject set, the index prediction value being a value for the predetermined index of the subject predicted by a pre-trained model; an object set acquisition module for screening objects whose index prediction values ​​satisfy an index value screening condition from the first object set to acquire a second object set; a mapping relationship determination module for determining a mapping relationship between a predetermined index and an object feature based on an index experimental value for the predetermined index of a plurality of objects in the first object set and an object feature, wherein the index experimental value refers to an object value for the predetermined index obtained by an experiment; and a target object determination module for determining a target object that satisfies an index requirement of the predetermined index from the second object set based on the mapping relationship; The objects in the first object set are mutant proteins.

15. A computing device including a memory and a processor coupled to the memory, The storage device stores a computer program, A computing device, wherein the processor is configured to execute the computer program to implement the method of any one of claims 1 to 13.

16. A program for causing a computer to execute the method according to any one of claims 1 to 13.

Citation Information

Patent Citations

  • Novel protein molecular orientation evolvement method

    CN101353372A

  • Method of acquiring proteins with high affinity by computer aided design

    CN102511045A

  • Protein prediction model generation method, device and equipment and storage medium

    CN111048145A

  • Cyclic peptide design method, compound structure generation method, device and electronic equipment

    CN114333985A

  • Method and electronic system for predicting at least one fitness value of a protein, related computer program product

    EP3082056A1