Object Determination Method, Apparatus, Computer Device, and Storage Medium
By obtaining the index prediction value and experimental value of the object set, the mapping relationship between preset indicators and object characteristics is determined, which solves the time-consuming and labor-intensive problem in protein orientation evolution, and achieves rapid and efficient target object determination.
Patent Information
- Application Number
- CN202210498684.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-09
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2042-05-09
AI Technical Summary
The existing protein-oriented evolution technology is time-consuming and labor-intensive, and has a high time cost.
By obtaining the predicted values of indicators in the object set, filtering out the object set that meets the indicator requirements, determining the mapping relationship between the preset indicators and the object features, thereby quickly determining the target object.
Reduces the time cost of determining the target object and improves efficiency.
Smart Images

Figure CN115116539B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular, to an object determination method, apparatus, computer device, and storage medium. Background Art
[0002] With the development of computer technology, directed evolution technology has emerged. Directed evolution can obtain proteins with new functions and characteristics in a relatively short time. By clearly setting goals, molecules can be redesigned, and directed evolution has become an important research tool in fields such as new drug research and development and chemical engineering.
[0003] In traditional protein directed evolution, an initial protein is established for the target function, a variant library is constructed at one or more positions, the most common mutants are determined by screening, these mutants are randomly recombined and screened, and the screened mutants are used for the next round of "mutation, recombination, screening" cycle until the expected protein performance is achieved.
[0004] However, most of the current directed evolution technologies are laborious and time-consuming, with a relatively large time cost. Summary of the Invention
[0005] Based on this, in view of the above technical problems, it is necessary to provide an object determination method, apparatus, computer device, computer-readable storage medium, and computer program product that can reduce the time cost.
[0006] On the one hand, the present application provides an object determination method. The method includes: obtaining the index prediction values of each object in the first object set on a preset index; selecting the objects whose index prediction values meet the index value screening condition from the first object set to obtain a second object set; determining the mapping relationship between the preset index and the object characteristics based on the index experimental values and object characteristics of multiple objects in the first object set on the preset index; and determining the target object that meets the index requirements of the preset index from the second object set based on the mapping relationship.
[0007] On the other hand, the present application also provides an object determination apparatus. The apparatus includes: a prediction value acquisition module, configured to obtain the index prediction values of each object in the first object set on a preset index; an object set acquisition module, configured to select the objects whose index prediction values meet the index value screening condition from the first object set to obtain a second object set; a mapping relationship determination module, configured to determine the mapping relationship between the preset index and the object characteristics based on the index experimental values and object characteristics of multiple objects in the first object set on the preset index; and a target object determination module, configured to determine the target object that meets the index requirements of the preset index from the second object set based on the mapping relationship.
[0008] In some embodiments, the objects in the first object set are mutant proteins, and the device further includes a reference object set screening module for screening a reference object set from the first object set; the reference object set satisfies the condition that each amino acid appears at least a target number of times at each mutation position; the predicted value acquisition module is further configured to train an index detection model based on the object features and index experimental values of each object in the reference object set; and use the trained index detection model to predict the index predicted values of each object in the first object set.
[0009] In some embodiments, the mapping relationship determination module is further configured to determine the object features of each object in the reference object set based on the index experimental values of each object in the reference object set on the preset index; and determine the mapping relationship between the preset index and the object features based on the index experimental values and object features of each object in the reference object set on the preset index.
[0010] In some embodiments, the reference object set screening module is further configured to obtain a current score set; the current score set includes the current scores corresponding to each amino acid respectively; obtain a second protein set based on the first object set, and select a target protein from the second protein set based on the current score set; decrease the current scores corresponding to the amino acids at each mutation position in the target protein in the current score set, and move the target protein from the second protein set to the first protein set; in the case where the current score set indicates that the first protein set does not satisfy the condition that each amino acid appears at least a target number of times at each mutation position, return to the step of selecting a target protein from the second protein set based on the current score set until the current score set indicates that the first protein set satisfies the condition that each amino acid appears at least a target number of times at each mutation position, and determine the first protein set as the reference object set.
[0011] In some embodiments, the reference object set screening module is further configured to obtain an initial score set; the initial score corresponding to each amino acid in the initial score set is the target number of times; decrease the initial scores corresponding to the amino acids at each mutation position in the wild-type protein in the initial score set to obtain a current score set, and determine a first protein set based on the wild-type protein; the wild-type protein is a protein that has not undergone mutation.
[0012] In some embodiments, the mapping relationship determination module is further configured to, for each mutant protein in the second protein set, determine the current scores respectively corresponding to the amino acids at each mutation position in the mutant protein from the current score set; determine the current protein score of the mutant protein based on the obtained current scores; and select a target protein from the second protein set based on the current protein score.
[0013] In some embodiments, each amino acid corresponds to an amino acid, and the scores in the current score set are uniquely identified by the amino acid and the mutation position; the mapping relationship determination module is further configured to, for the amino acid at each mutation position, determine the current score corresponding to the amino acid at the mutation position from the current score set according to the amino acid corresponding to the amino acid and the mutation position.
[0014] In some embodiments, the object feature is a protein feature, and the mapping relationship determination module is further configured to, for each mutation position, divide the reference object set according to the type of the amino acid at the mutation position to obtain a first sub-object set corresponding to each amino acid; for each amino acid at each mutation position, determine the amino acid feature of the amino acid at the mutation position based on the index experimental values of each object in the first sub-object set corresponding to the amino acid; and obtain the protein feature of the object based on the amino acid features of the amino acids at each mutation position in the object.
[0015] In some embodiments, the mapping relationship determination module is further configured to perform statistical calculations on the index experimental values of each object in the first sub-object set corresponding to the amino acid to obtain at least one index experimental statistical value; and determine the amino acid feature of the amino acid at the mutation position based on the at least one index experimental statistical value.
[0016] In some embodiments, the object feature is a protein feature; the mapping relationship determination module is further configured to: for each amino acid, determine the objects in which the amino acid at the mutation position includes the amino acid from the reference object set to obtain a second sub-object set corresponding to the amino acid; for each amino acid, determine the amino acid feature of the amino acid based on the index experimental values of the respective objects in the second sub-object set corresponding to the amino acid; and obtain the protein feature of the object based on the amino acid features of the amino acids at each mutation position in the object.
[0017] In some embodiments, the target object determination module is further configured to determine the statistical index values of each object in the second object set on the target statistical index based on the mapping relationship, and determine the selected objects from the second object set based on the statistical index values; in the case where the iteration stop condition is not satisfied, add the selected objects to the reference object set; return the step of determining the object characteristics of each object in the reference object set based on the index experiment values of each object in the reference object set on the preset index until the iteration stop condition is satisfied; and determine the selected objects obtained when the iteration stop condition is satisfied as the target objects that meet the index requirements of the preset index.
[0018] On the other hand, the present application also provides a computer device. The computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps in the above object determination method are implemented.
[0019] On the other hand, the present application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the above object determination method are implemented.
[0020] On the other hand, the present application also provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the steps in the above object determination method are implemented.
[0021] The above object determination method, device, computer device, storage medium, and computer program product obtain the index prediction values of each object in the first object set on the preset index respectively, select the objects whose index prediction values meet the index value screening conditions from the first object set to obtain a second object set, determine the mapping relationship between the preset index and the object characteristics based on the index experiment values and object characteristics of multiple objects in the first object set on the preset index, and determine the target objects that meet the index requirements of the preset index from the second object set based on the mapping relationship. Since the second object set is screened from the first object set, determining the target objects that meet the index requirements of the preset index from the second object set is more efficient than screening the target objects from the first object set, thereby reducing the time cost of determining the target objects. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 It is an application environment diagram of the object determination method in some embodiments;
[0023] Figure 2 It is a flowchart of the object determination method in some embodiments;
[0024] Figure 3A Application diagram of an enzyme in some embodiments;
[0025] Figure 3B Schematic diagram of machine learning-assisted directed evolution in some embodiments;
[0026] Figure 4 Schematic diagram of the average fitness of amino acids in some embodiments;
[0027] Figure 5 Flow schematic diagram of an object determination method in some embodiments;
[0028] Figure 6 Schematic diagram of an object determination method in some embodiments;
[0029] Figure 7 Schematic diagram of an object determination method in some embodiments;
[0030] Figure 8 Application environment diagram of an object determination method in some embodiments;
[0031] Figure 9 Application environment diagram of an object determination method in some embodiments;
[0032] Figure 10 Fitness distribution diagram of different data sets in some embodiments;
[0033] Figure 11 Effect diagram of different methods on four protein directed evolution data sets in some embodiments;
[0034] Figure 12 Effect diagram of different methods on a data set in some embodiments;
[0035] Figure 13 Structural block diagram of an object determination device in some embodiments;
[0036] Figure 14 E Internal structure diagram of a computer device in some embodiments;
[0037] Figure 15 Internal structure diagram of a computer device in some embodiments. Detailed implementation manners
[0038] In order to make the objectives, technical solutions, and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0039] The object determination method provided by the embodiments of the present application can be applied to, for example Figure 1In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or placed on the cloud or other servers.
[0040] Specifically, the server 104 can obtain the index prediction values of each object in the first object set on a preset index, select the objects whose index prediction values meet the index value screening conditions from the first object set to obtain a second object set, determine the mapping relationship between the preset index and the object features based on the index experimental values and object features of multiple objects in the first object set, and determine the target objects that meet the index requirements of the preset index from the second object set based on the mapping relationship. After the server 104 determines the target objects, it can store the target objects and can also send the target objects to the terminal 102, and the terminal 102 can display the relevant information of the target objects.
[0041] Among them, the terminal 102 can be, but is not limited to, various desktop computers, laptop computers, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, etc. The portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The server 104 can be implemented by an independent server or a server cluster composed of multiple servers.
[0042] In some embodiments, the index prediction value can be predicted by a trained index detection model. The index detection model can be based on artificial intelligence and machine learning, such as a neural network model. Among them, Artificial Intelligence (AI) is to use a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results of theory, methods, technologies, and application systems. In other words, artificial intelligence is a comprehensive technology of computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence is also to study the design principles and implementation methods of various intelligent machines to make the machines have the functions of perception, reasoning, and decision-making.
[0043] Artificial intelligence technology is an interdisciplinary subject, involving a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0044] Machine learning (ML) is an interdisciplinary field that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills, and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.
[0045] With the research and progress of artificial intelligence technology, artificial intelligence technology has been studied and applied in multiple fields. For example, common ones include smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, driverless, autonomous driving, drones, robots, smart healthcare, smart customer service, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.
[0046] The solution provided by the embodiments of this application relates to technologies such as neural networks of artificial intelligence, and is specifically described through the following embodiments:
[0047] In some embodiments, as Figure 2 shown, a method for object determination is provided. This method can be executed by a terminal or a server, or jointly executed by a terminal and a server. Taking the server 104 in Figure 1 as an example, the method includes the following steps:
[0048] Step 202, obtain the index prediction values of each object in the first object set on a preset index.
[0049] Among them, the first object set includes multiple objects. An object can be a real substance, including but not limited to at least one of proteins, materials, or batteries, etc. An object can also be an abstract specific concept. For example, the object is a battery fast charging protocol.
[0050] An object can correspond to multiple indexes, and the preset index can be any one of the multiple indexes of the object. For example, if the object is a protein, the indexes of the object include but are not limited to at least one of fitness, concentration fraction, activity, or brightness, etc. If the object is a material, the indexes of the object include but are not limited to at least one of the composition of the material or the proportion of the composition, etc. If the object is a battery fast charging protocol, the indexes of the object include but are not limited to each parameter in the battery fast charging protocol.
[0051] The predicted value of an indicator is the value corresponding to the preset indicator predicted for an object. The predicted value of an indicator can be predicted by a trained indicator detection model. The indicator detection model can be a neural network model.
[0052] Each object in the first object set can belong to the same object category. The object category includes, but is not limited to, at least one of substances such as proteins or materials, and can also include abstract concepts such as battery charging protocols. For example, each object in the first object set belongs to a certain type of protein. For example, each object is a mutant protein obtained by mutating the same protein. Among them, the mutant protein is relative to the wild-type protein. The wild-type protein is a protein that has not undergone mutation, and the mutant protein is a protein obtained by mutating on the basis of the wild-type protein. In protein directed evolution, the desired protein can be obtained through mutation. Protein directed evolution includes two mutation scenarios, one is the saturation mutagenesis scenario at the k position, and the other is the non-saturation mutagenesis scenario.
[0053] The saturation mutagenesis scenario at the k position is used to mutate the amino acids at k specified mutation sites. In the mutant proteins generated in this scenario, the amino acid at at least one of the k specified mutation sites is obtained by mutation. For example, if k = 4, then in the obtained mutant protein, the amino acid at at least one of the specified 4 mutation sites is obtained by mutation. That is, the positions and numbers of the mutation sites in the saturation mutagenesis scenario at the k position are fixed, and the mutation only occurs at the specified k mutation sites. The mutation site refers to the position where mutation may occur in a protein. Therefore, the mutation site can also be called the mutation position, and there is one amino acid at each position in the protein.
[0054] In the non-saturation mutagenesis scenario, the mutation sites are not fixed but the number of mutated amino acids is fixed. For example, in each mutant protein obtained in the non-saturation mutagenesis scenario, there are 2 positions of amino acids obtained by mutation, but the positions where the mutation occurs can be the same or different. For example, one mutation occurs at positions 1 and 2, and one mutation occurs at positions 3 and 4.
[0055] If the object category is a protein, each object in the first object set can be a mutant protein generated in the saturation mutagenesis scenario at the k position or a mutant protein generated in the non-saturation mutagenesis scenario. The mutant protein can also be called a mutant. A protein can be represented by an amino acid sequence. Suppose the first object set includes n mutant proteins, then the first object set can be represented as , where n represents the number of mutants, S i represents a mutant, S i =(S i1,S i2 ,…,S iL ), S i represents the i-th amino acid sequence having L amino acids, S ij represents an amino acid, 1 ≤ j ≤ L, y i represents the fitness of the i-th protein, and the fitness is obtained by experimental measurement. The fitness of the protein characterizes the properties of the protein, and the fitness can be, for example, affinity.
[0056] Specifically, for each object in the first object set, the server can predict the metric prediction value of each object on a preset metric to obtain the metric prediction value of each object. For example, the metric prediction value can be predicted using a trained metric detection model.
[0057] In some embodiments, the server can screen multiple objects from the first object set to obtain a reference object set, determine the object features of each object in the reference object set, where the object features refer to the features of the object, and obtain the metric experimental value of each object in the reference object set on a preset metric through experimental means. The metric experimental value refers to the value of the object on the preset metric obtained through experimental means, that is, the metric experimental value of the object is the true value of the object on the preset metric. The server can use the object features of each object in the reference object set and the metric experimental values of each object to train the metric detection model to obtain a trained metric detection model, determine the object features of each object in the first object set, input the object features of each object in the first object set into the trained metric detection model, and use the trained metric detection model to predict the metric prediction value corresponding to each object in the first object set.
[0058] In some embodiments, the object is a mutant protein, and the object feature is a protein feature. The protein feature can be a feature encoded based on the amino acids at the mutation sites in the mutant protein. For example, the amino acids can be encoded based on the metric experimental value of the mutant protein on a preset metric to obtain the amino acid features corresponding to the amino acids, and the protein feature of the mutant protein can be obtained based on the amino acid features corresponding to the amino acids at each mutation position. For example, for a mutant protein generated in a saturation mutagenesis scenario at the k-th position, the protein feature of the mutant protein can be obtained using the amino acid feature of the amino acid at the k-th position. For a mutant protein generated in a non-saturation mutagenesis scenario, if mutations occur at 2 positions, the vector composed of the amino acid features of the amino acids at these 2 positions is determined as the protein feature of the mutant protein.
[0059] Step 204, select the objects from the first object set whose metric prediction values meet the metric value screening condition to obtain a second object set.
[0060] Among them, the index value screening condition includes that the index prediction value is greater than the first index threshold, and the first index threshold can be preset or set as needed. The second object set is a set composed of objects selected from the first object set, and the index prediction values of the objects in the second object set meet the index value screening condition.
[0061] Specifically, the server can compare the index prediction value of each object in the first object set with the first index threshold, and form the second object set by combining the objects whose index prediction values are greater than the first index threshold. For example, if the object is a mutant protein, the preset index is affinity, and the first index threshold is the affinity threshold, then the second object set is formed by combining the objects in the first object set whose affinity is greater than the affinity threshold, and the affinity threshold can be preset or set as needed.
[0062] Step 206: Based on the index experimental values and object characteristics of multiple objects in the first object set on the preset index, determine the mapping relationship between the preset index and the object characteristics.
[0063] Among them, the multiple objects in the first object set can refer to each object in the above-mentioned reference object set. The mapping relationship between the preset index and the object characteristics is used to reflect the change of the value of the preset index as the object characteristics change. The mapping relationship between the preset index and the object characteristics can be represented by a curve. For example, this mapping relationship can be represented by the curve y1 = f1(x), where y1 represents the preset index and x represents the object characteristics.
[0064] Specifically, after obtaining the reference object set, in addition to using the reference object set to train the index detection model, the server can also use the index experimental values and object characteristics of each object in the reference object set to determine the mapping relationship between the preset index and the object characteristics.
[0065] In some embodiments, after obtaining the index experimental values of each object in the reference object set, for each object in the reference object set, the server can use the object characteristics and index experimental value of the object as a point on the curve y1 = f1(x). For multiple points on multiple curves, fitting the obtained multiple points to generate the curve y1 = f1(x) representing this mapping relationship.
[0066] Step 208: Based on the mapping relationship, determine the target object that meets the index requirements of the preset index from the second object set.
[0067] Among them, the index requirements of the preset index can be, for example, that the index experimental value is as large as possible or at least one of the index experimental values is greater than the second index threshold. The target object is the object in the second object set that meets the index requirements of the preset index.
[0068] Specifically, the mapping relationship between the preset index and the object characteristics is the first mapping relationship. The server can perform statistical operations based on the first mapping relationship to obtain the second mapping relationship between the target statistical index and the object characteristics. Based on the second mapping relationship, determine the statistical index values of each object in the second object set on the target statistical index. Based on the statistical index values of each object, determine the selected objects from each object in the second object set. Based on the selected objects, obtain the objects that meet the index requirements of the preset index. The first mapping relationship represents the rule that the value of the preset index changes with the change of the object characteristics, and the second mapping relationship represents the rule that the value of the target statistical index changes with the change of the object characteristics. For example, the second mapping relationship can be represented by the curve y2 = f2(x), where y2 represents the target statistical index and x represents the object characteristics.
[0069] The target statistical index can be one or more. For example, the target statistical index includes but is not limited to at least one of Expected Improvement (EI), Probability of Improvement (PI), Upper Confidence Bound (UCB), or Thompson Sampling (TS), etc. The first mapping relationship can also be called a probability surrogate model, and the second mapping relationship can also be called an acquisition function. Among them, the acquisition function is constructed through the posterior probability distribution obtained by the probability surrogate model, and the next most "promising" experimental point is selected by maximizing the acquisition function. The acquisition function is responsible for testing the proposed new points based on the exploration and exploitation trade-off. Exploration means trying to select points far from the known points for the next experiment, that is, trying to explore unknown areas; exploitation means trying to select points close to the known points for the next experiment, that is, trying to mine the points around the known points.
[0070] In some embodiments, the server can determine the statistical index values of each object in the second object set on the target statistical index based on the second mapping relationship, and determine the selected objects from each object in the second object set based on the statistical index values of each object. Specifically, the selected objects can be one or more. The object corresponding to the largest statistical index value can be determined as the selected object, or the objects with statistical index values greater than the third index threshold can be determined as the selected objects. The third index threshold can be set as needed. The server can obtain the objects that meet the index requirements of the preset index based on the selected objects. For example, the server can determine the selected objects as the objects that meet the index requirements of the preset index.
[0071] In some embodiments, after obtaining a selected object, the server can obtain the index experimental value of the selected object, compare the index experimental value of the selected object with a second index threshold, and when it is determined that the index experimental value of the selected object reaches the second index threshold, determine the selected object as the target object. Among them, after obtaining the selected object, the corresponding index experimental value can be determined through experiments. If it is determined that the index experimental value of the selected object does not reach the second index threshold, the selected object can be added to the reference object set, and then the index experimental values and object characteristics of each object in the reference object set on the preset index are used to determine the first mapping relationship between the preset index and the object characteristics. Then, based on the first mapping relationship, a target object that meets the index requirements of the preset index is determined from the second object set, and the loop continues. When a selected object whose index experimental value reaches the second index threshold is found, the loop ends, and the selected object whose index experimental value reaches the second index threshold is determined as the target object, or when the number of loops reaches the number threshold, the selected object screened out is determined as the target object.
[0072] In the above object determination method, the index prediction values of each object in the first object set on the preset index are obtained, objects whose index prediction values meet the index value screening conditions are selected from the first object set to obtain a second object set, and based on the index experimental values and object characteristics of multiple objects in the first object set on the preset index, the mapping relationship between the preset index and the object characteristics is determined, and based on the mapping relationship, a target object that meets the index requirements of the preset index is determined from the second object set. Since the second object set is screened from the first object set, determining a target object that meets the index requirements of the preset index from the second object set is more efficient than screening the target object from the first object set, thus reducing the time cost of determining the target object.
[0073] In actual design application scenarios, such as: environmental scientists obtain environmental conditions by designing sensor deployment locations; chemists obtain new substances by designing experiments; pharmaceutical manufacturers develop new drugs to resist diseases, etc. Usually, these design problems are considered as the following optimization problems for solution (only considering maximization problems, minimization problems can be simply converted into minimization problems by taking the negative sign operation):
[0074]
[0075] Among them, \(x\) represents a \(d\)-dimensional decision vector, \(X\) represents the decision space, and \(f(x)\) represents the objective function. Corresponding to the above example, \(x\) can represent the sensor deployment location, experimental configuration, drug formulation, etc., and \(f(x)\) can represent the measure of the performance of the environment, experiment, formulation, etc. In these actual design application scenarios, there are many complex design decisions, and their optimization objectives usually have the following characteristics: high computational cost: ideally, the function can be executed multiple times to determine its optimal solution, but in actual optimization problems, it is unrealistic to calculate too many samples, and the computational cost is very high; black-box function: in actual problems, the structure of the objective function is difficult to describe mathematically, without first-order or higher-order derivatives, and cannot be solved by gradient descent or Newton-related algorithms; to find the global minimum / maximum value: a certain mechanism is needed to avoid falling into the local minimum / maximum value. Therefore, in order to obtain the required substance, a high time cost needs to be paid.
[0076] The object determination method proposed in this application can accelerate the process of obtaining the required object, improve the efficiency, and thus reduce the time cost. For example, the object determination method provided in this application can be applied to computational methods to assist protein evolution, so as to obtain the required protein. Proteins play an important role in people's lives. For example, enzymes are used in human society, from daily life to industry, as Figure 3A shown. In some daily-use washing powders, enzymes are contained to promote the decomposition of oil stains and other stains; enzymes are essential in the fermentation, degradation and other processes in the food industry; in drugs and fine chemicals, enzymes, as green and efficient catalysts, have replaced some production processes in traditional chemistry that require heavy metals and are energy-consuming; in addition, enzymes are the most important role in the development of bioenergy. Directed evolution can obtain proteins with new functions and characteristics in a relatively short time. By clearly setting artificial goals, scientists can redesign molecules, and it has become an important research tool in the fields of new drug research and development, chemical engineering, etc. Machine learning can be used to assist directed evolution, as Figure 3B shown. For example, the process of machine learning methods assisting directed evolution can include four steps: 1) establish an initial protein for the target function and construct a variant library at \(k\) positions; 2) train the model using existing data; 3) use the trained model to predict other mutants in the variant library; 4) select the optimal mutant for experimental testing and add it to the training set for the next round of model training. Computational methods can assist directed evolution to accelerate optimization and reduce the experimental burden.
[0077] In some embodiments, the objects in the first object set are mutant proteins, and the method further includes: screening a reference object set from the first object set; the reference object set satisfies the condition that each amino acid appears at least a target number of times at each mutation position; obtaining the index prediction values of each object in the first object set on a preset index, including: training an index detection model based on the object features and index experimental values of each object in the reference object set; using the trained index detection model to predict the index prediction values of each object in the first object set.
[0078] Among them, each object in the reference object set can be a mutant protein, or the reference object set can include wild-type proteins and mutant proteins. The objects in the first object set are mutant proteins, and the target number can be preset or set as needed, for example, it can be 2 times. The mutation position is the mutation site. The reference object set satisfies the condition that each amino acid appears at least a target number of times at each mutation position. For example, if there are 20 kinds of amino acids and the target number is 2, then in each protein in the reference object set, these 20 amino acids appear at least 2 times at each mutation site. Taking the mutant proteins generated by the saturation mutagenesis scenario at the k position as an example, there are 4 mutation sites, and each of the 20 kinds of amino acids appears 2 times at each mutation site, then 40 samples can be selected from the sample space as the initial samples. This not only ensures that the amino acid coding information covered in the initial sample quantity has the largest coverage range, but also requires the fewest number of experiments, reducing the experimental cost. Among them, the sample space can include mutant proteins and can also include wild-type proteins. The sample refers to a protein, the reference object set can be constantly changing, and the initial sample refers to the initially determined reference object set.
[0079] The index detection model is used to determine the value of the object on a preset index according to the object features, that is, to determine the index prediction value of the object on the preset index. The index detection model can be a neural network model. For example, it can be XGBOD (Improving Supervised Outlier Detection with Unsupervised Representation Learning), and of course, it can also be other models, which are not limited here. Among them, the basic process of XGBOD is to learn the original data by using a variety of unsupervised models, obtain the outlier scores of each sample, and use the outlier scores as a new data representation form. Subsequently, the original features are combined to generate a new feature space. Finally, an XGBoost classifier is trained on the new feature space, and its output is regarded as the prediction result.
[0080] Specifically, the server can obtain the metric experimental values of each object in the reference object set on a preset metric (denoted as the metric experimental value corresponding to the object), and determine the object characteristics of each object in the reference object set based on the metric experimental values of each object in the reference object set on the preset metric. The server can input the object characteristics of the object into the metric detection model to be trained for prediction, obtain the metric prediction value of the object on the preset metric (denoted as the metric prediction value corresponding to the object), and adjust the model parameters of the metric detection model based on the difference between the metric experimental value corresponding to the object and the corresponding metric prediction value until the model converges, obtaining the trained metric detection model. The server can input the object characteristics of each object in the first object set into the trained metric detection model for prediction, obtaining the metric prediction value corresponding to each object in the first object set.
[0081] In some embodiments, the server can determine the amino acid characteristics of each amino acid based on the metric experimental values of each object in the reference object set. Thus, when determining the object characteristics (i.e., protein characteristics) of the objects in the first object set, the server uses the amino acid characteristics of each amino acid determined by the reference object set to determine the protein characteristics of the objects in the first object set. For example, in the k-site saturation mutagenesis scenario, based on the metric experimental values of each object in the reference object set, the amino acid characteristics of each amino acid at each mutation position can be determined. For the objects in the first object set, the amino acids at the mutation positions in the object can be determined, and the amino acid characteristics corresponding to the amino acids at each mutation position in the object can be determined from the already determined "amino acid characteristics of each amino acid at each mutation position", and the vector composed of the determined amino acid characteristics is determined as the object characteristics (i.e., protein characteristics) of the object.
[0082] In some embodiments, the method of screening the reference object set from the first object set can be used as a sample selection strategy for determining the initial samples in Bayesian optimization, thereby improving the optimization efficiency, reducing the time cost of Bayesian optimization, ensuring that the amino acid coding information covered in the initial sample size has the largest coverage range, while minimizing the number of experiments required and reducing the experimental cost. Among them, Bayesian optimization usually uses Gaussian process (GP) regression based on the Gaussian distribution as the prior probability surrogate model. GP is flexible and scalable, and can theoretically surrogate any linear / non-linear function. Of course, Gaussian process regression based on the student-t prior can also be used as the prior probability surrogate model, combining robust regression (Gaussian process based on the student-t distribution) with outlier detection, classifying data points into outliers and inliers, thereby eliminating the influence of outliers on model fitting. The Gaussian process based on the student-t prior can be abbreviated as "Robust GP", and the Gaussian process based on the Gaussian distribution can be abbreviated as "GP".
[0083] In this embodiment, since the reference object set satisfies the condition that each amino acid appears at least the target number of times at each mutation position, the number of each amino acid in the obtained reference object set is balanced. Therefore, training the index detection model based on the reference object set improves the training accuracy, and thus improves the accuracy of the index prediction value predicted by the trained index detection model.
[0084] In some embodiments, determining the mapping relationship between the preset index and the object feature based on the index experimental values and object features of multiple objects in the first object set includes: determining the object features of each object in the reference object set based on the index experimental values of each object in the reference object set on the preset index; determining the mapping relationship between the preset index and the object feature based on the index experimental values and object features of each object in the reference object set on the preset index.
[0085] Specifically, the server can perform statistical calculations on the index experimental values corresponding to each object in the reference object set to obtain the object features of each object in the reference object. Among them, the index experimental value corresponding to the object refers to the index experimental value of the object on the preset index. Statistical calculations include but are not limited to calculating at least one of the mean, maximum value or minimum value.
[0086] In some embodiments, the mapping relationship between the preset index and the object feature is the first mapping relationship. The mapping relationship can be represented by a curve. For example, the first mapping relationship is represented by the curve y1 = f1(x). For each object in the reference object set, the server can use the object feature and the index experimental value of the object as points on the curve y1 = f1(x). Multiple points on multiple curves are used to fit the obtained multiple points to generate the curve y1 = f1(x) representing the mapping relationship.
[0087] In this embodiment, since the reference object set satisfies the condition that each amino acid appears at least the target number of times at each mutation position, the amino acids in the obtained reference object set are balanced. Therefore, based on the index experimental values of each object in the reference object set for the preset index, determining the object features of each object in the reference object set can make the coverage range of the amino acid coding information larger, that is, it improves the information range covered by the object features.
[0088] In some embodiments, screening the reference object set from the first object set includes: obtaining the current score set; the current score set includes the current scores corresponding to each amino acid respectively; obtaining the second protein set based on the first object set, and selecting the target protein from the second protein set based on the current score set; decreasing the current scores corresponding to the amino acids at each mutation position in the target protein in the current score set, and moving the target protein from the second protein set to the first protein set; in the case where the current score set indicates that the first protein set does not satisfy the condition that each amino acid appears at least the target number of times at each mutation position, returning to the step of selecting the target protein from the second protein set based on the current score set until the current score set indicates that the first protein set satisfies the condition that each amino acid appears at least the target number of times at each mutation position, and determining the first protein set as the reference object set.
[0089] Among them, the current score set includes the current scores corresponding to each amino acid respectively. The current score is an integer, for example, 2. The current scores corresponding to each amino acid can be different or the same. Each amino acid can correspond to one current score or multiple current scores. For example, the same amino acid may have different current scores at different mutation positions. The current score set is constantly changing. The target protein is selected from the second protein set based on the current score set, and the target protein can be one or more.
[0090] Specifically, the server can determine the first object set as the second protein set, that is, the objects in the second protein set are the same as the objects in the first object set. Or, the server can obtain the objects other than the objects in the reference object set from the first object set to form the second object set.
[0091] In some embodiments, for the current score set, the current protein scores of each mutant protein in the second protein set are determined, and the mutant protein corresponding to the maximum current protein score is determined as the target protein. The server may rank the mutant proteins in the second protein set in descending order of the current protein scores to obtain a first protein sequence, and determine the mutant proteins in the first protein sequence that are before the sorting threshold as the target proteins. The sorting threshold may be preset or set as needed, for example, any one of the 1st position or the 2nd position, etc.
[0092] In some embodiments, the server may decrease the current scores corresponding to each amino acid at each mutation position in the target protein in the current score set, and move the target protein from the second protein set to the first protein set.
[0093] For example, if the objects in the first object set are mutant proteins generated in a k-site saturation mutagenesis scenario, the current score set includes the current scores of each amino acid at each mutation position. Taking 4 mutation positions as an example, the current score set may be represented in matrix form. The current score set may also be referred to as the current score matrix. In the matrix corresponding to the current score set, the score in the u-th row and the w-th column represents the current score of the u-th amino acid at the w-th mutation position, where 1 ≤ u ≤ m, 1 ≤ w ≤ k, m is the number of types of amino acids, for example, 20, and k represents the number of mutation positions, for example, 4. For example, if the amino acid at the 1st mutation position in the mutant protein is the 1st amino acid, the corresponding current score is the score in the 1st row and the 1st column of the matrix.
[0094] If the objects in the first object set are mutant proteins generated in a non-saturation mutagenesis scenario, the current score set includes the current scores corresponding to each amino acid respectively. The current score set may be represented by a vector. The current score set may also be referred to as the current score vector. The score arranged at the u-th position in the vector represents the current score of the u-th amino acid. For example, if the amino acid at the 1st mutation position in the mutant protein is the 1st amino acid, the corresponding current score is the score arranged at the 1st position in the vector.
[0095] In some embodiments, when the current score set indicates that the first protein set does not meet the condition that each amino acid appears at least the target number of times at each mutation position, the server returns to perform the step of selecting a target protein from the second protein set based on the current score set until the current score set indicates that the first protein set meets the condition that each amino acid appears at least the target number of times at each mutation position, and then determines the first protein set as the reference object set. Specifically, the first protein set is constantly changing. If the initial first protein set does not include any proteins, the scores in the initial current score set are all equal to the target number. For example, if the target number is 2, each current score in the current score set is equal to 2, which is in the initial current score set. After determining the target protein, the target protein can be moved from the second protein set to the first protein set, and the current scores corresponding to the amino acids at each mutation position in the target protein are decreased by 1 each time, for example, from 2 to 1 or from 1 to 0, so as to continuously update the current score matrix. When there is no score greater than 0 in the current score matrix, it is determined that the first protein set meets the condition that each amino acid appears at least the target number of times at each mutation position, and thus the first protein set is determined as the reference object set.
[0096] In this embodiment, the second protein set is obtained based on the first object set, the target protein value is selected from the second protein set, the current score set is updated based on the target protein, and the target protein is moved from the second protein set to the first protein set, so as to continuously select proteins to update the current score. When the current score set indicates that the first protein set meets the condition that each amino acid appears at least the target number of times at each mutation position, the first protein set is determined as the reference object set, so as to quickly select the reference object set that meets the condition that each amino acid appears at least the target number of times at each mutation position from the first object set, reduce the number of experiments for determining the reference object set, and thus reduce the experimental cost and time cost.
[0097] In some embodiments, obtaining the current score set includes: obtaining an initial score set; the initial scores corresponding to each amino acid in the initial score set are the target number; decreasing the initial scores corresponding to the amino acids at each mutation position in the wild-type protein in the initial score set to obtain the current score set, and determining the first protein set based on the wild-type protein; the wild-type protein is a protein that has not undergone mutation.
[0098] Among them, the initial score corresponding to each amino acid in the initial score set is the target number of times, and the target number of times is an integer. For example, if it is 2, the initial score is 2. For example, if the objects in the first object set are mutant proteins generated in the k-site saturation mutagenesis scenario, the initial score set includes the initial scores of each amino acid at each mutation position. Taking 4 mutation positions as an example, the initial score set can be represented in matrix form, and each element in the matrix is 2. In the matrix corresponding to the initial score set, the element in the u-th row and w-th column represents the initial score of the u-th amino acid at the w-th mutation position, where 1 ≤ u ≤ m, 1 ≤ w ≤ k, m is the number of types of amino acids, for example 20, and k represents the number of mutation positions, for example 4. For example, if the amino acid at the 1st mutation position in the mutant protein is the 1st amino acid, the corresponding initial score is the initial score in the 1st row and 1st column of the matrix.
[0099] If the objects in the first object set are mutant proteins generated in a non-saturation mutagenesis scenario, the initial score set includes the initial scores corresponding to each amino acid respectively. The initial score set can be represented by a vector, and each element in the vector is 2. The score arranged at the u-th position in the vector represents the initial score of the u-th amino acid. If the amino acid at the 1st mutation position in the mutant protein is the 1st amino acid, the corresponding initial score is the score arranged at the 1st position in the vector.
[0100] Specifically, in the k-site saturation mutagenesis scenario with k mutation positions, the server can determine the amino acid corresponding to each mutation position from the wild-type protein, and decrement the current score corresponding to this amino acid in the initial score set by 1 each time to obtain the current score set. Taking the 4-site saturation mutagenesis scenario as an example, there are 4 mutation positions. If the amino acids corresponding to these 4 mutation positions in the wild-type protein are the 1st amino acid A1, the 2nd amino acid A2, the 3rd amino acid A3, and the 4th amino acid A4 respectively, then decrement the scores in the 1st row and 1st column, 2nd row and 2nd column, 3rd row and 3rd column, and 4th row and 4th column of the matrix corresponding to the initial score set by 1 each, and determine the initial score set after the decrement as the current score set.
[0101] In the non-saturation mutagenesis scenario, each mutant protein corresponds to a target number of mutation positions, and the target number is, for example, 2. The server can count the mutation positions corresponding to multiple mutant proteins in the non-saturation mutagenesis scenario to obtain a mutation position set. The server can determine the amino acid corresponding to each mutation position in the mutation position set from the wild-type protein, and the server can determine the amino acids corresponding to each mutation position in the mutation position set in the wild-type protein respectively. Decrement the current score corresponding to each determined amino acid in the initial score set by 1 each time to obtain the current score set.
[0102] In some embodiments, the server may determine the set of wild-type proteins as the first protein set, that is, the initial first protein set includes one wild-type protein.
[0103] In this embodiment, the initial score set is updated based on the wild-type protein without mutation to obtain the current score set, thereby accelerating the speed of score decrease, improving the efficiency of obtaining the reference object set, and reducing the time cost.
[0104] In some embodiments, selecting a target protein from the second protein set based on the current score set includes: for each mutant protein in the second protein set, determining the current scores corresponding to the amino acids at each mutation position in the mutant protein from the current score set; determining the current protein score of the mutant protein based on the obtained current scores; and selecting the target protein from the second protein set based on the current protein score.
[0105] Specifically, for each mutant protein in the second protein set, determine the current scores corresponding to the amino acids at each mutation position in the mutant protein from the current score set, perform a summation calculation on the obtained current scores, and determine the result of the summation calculation as the current protein score of the mutant protein.
[0106] In some embodiments, the server may arrange the mutant proteins in the second protein set in descending order according to the current protein score to obtain a first protein sequence, and determine the mutant proteins arranged before the sorting threshold in the first protein sequence as the target proteins. The sorting threshold may be preset or set as needed, for example, any one of the 1st or 2nd, etc.
[0107] In this embodiment, mutant proteins with current protein scores meeting the condition of relatively large scores are selected from the second protein set to obtain the target proteins. Since the larger the current protein score, the greater the impact on updating the current score set, the speed of making the current score set represent the condition that the first protein set satisfies that each amino acid appears at least the target number of times at each mutation position is accelerated, and the efficiency of obtaining the reference object set is accelerated.
[0108] In some embodiments, each amino acid corresponds to an amino acid, and the scores in the current score set are uniquely identified by the amino acid and the mutation position; determining the current scores corresponding to the amino acids at each mutation position in the mutant protein from the current score set includes: for the amino acid at each mutation position, determining the current score corresponding to the amino acid at the mutation position from the current score set according to the amino acid corresponding to the amino acid and the mutation position.
[0109] Specifically, in the case where the objects in the first object set are mutant proteins generated in a k-site saturation mutagenesis scenario, the current score set includes the current scores of each amino acid at each mutation position. That is, the scores in the current score matrix are uniquely identified by the amino acid and the mutation position. Taking 4 mutation positions as an example, the current score set can be represented in matrix form. In the matrix corresponding to the current score set, the element in the u-th row and w-th column represents the current score corresponding to the u-th amino acid at the w-th mutation position, where 1 ≤ u ≤ m, 1 ≤ w ≤ k, m is the number of types of amino acids, for example 20, and k represents the number of mutation positions, for example 4. For example, if the amino acid at the first mutation position in the mutant protein is the first amino acid, the corresponding current score is the score in the first row and first column of the matrix.
[0110] For a non-saturation mutagenesis scenario, if the current score set is a current score vector, and the u-th element in the current score vector is the score of the u-th amino acid. If the mutant protein has 2 mutation positions, and the amino acids at these two mutation positions are the 3rd amino acid and the 10th amino acid respectively, then the score corresponding to the 3rd amino acid is the score at the 3rd position in the current score vector, and the score corresponding to the 10th amino acid is the score at the 10th position in the current score vector.
[0111] In some embodiments, for a k-site saturation mutagenesis scenario, the server can determine the screened reference object set through the following algorithm:
[0112] Input data of this algorithm: p, set D train ={(S0, y0)}, matrix M.
[0113] Among them, p refers to the number of occurrences of each amino acid at each mutation site, that is, the target number, for example 2. D train refers to the first protein set, S0 in (S0, y0) represents the wild-type protein, and y0 represents the index experimental value of the wild-type protein. M refers to the current score matrix, where m is the number of types of amino acids, for example 20, AAINDEX(a) represents the coordinates of amino acid a in matrix M, that is, the position of the score corresponding to amino acid a in M.
[0114] The step of initializing the current score matrix is: if the u-th amino acid appears at the w-th mutation position of S0, then M uw = p - 1, otherwise M uw = p, M uw is the element in the u-th row and w-th column of M, 1 ≤ w ≤ k. [[ID=3@]]
[0115] Output data of this algorithm: the updated set D trainThe D output by the algorithm train is the initial set of reference objects.
[0116] Steps of this algorithm:
[0117] Step 1: while M uw > 0 do. This step means that if there are elements less than 0 in matrix M, then execute Steps 2 to 5.
[0118] Step 2: Calculate the score of each mutant:
[0119] Among them, Score i is the current protein score of the i-th mutant protein S i in the first object set, V i ={(u, w)|AAINDEX(S ij )} represents the coordinates (u, w) in matrix M of the current score corresponding to each amino acid S i at each mutation position in S ij .
[0120] Step 3: Select the mutant protein with the maximum score i* = argmaxScore i . This step is used to determine the mutant protein with the maximum current protein score.
[0121] Among them, i* indicates that the mutant protein with the maximum current protein score is the i*-th mutant protein in the first object set.
[0122] Step 4: Update set D train ←(S i *, y i *). This step means that add the i*-th mutant protein in the first object set to the first protein set.
[0123] Among them, S i * in (S i *, y i *) represents the i*-th mutant protein in the first object set, and y i * represents the experimental value of the index of the i*-th mutant protein.
[0124] Step 5: Update score matrix M. If the u-th amino acid appears at the w-th mutation position of S i *, then M uw = M uw - 1, otherwise M uw = M uw .
[0125] Step 6: end while. If there are no elements greater than 0 in M, then execute Step 7.
[0126] Step 7: Output D train .
[0127] In some embodiments, for the non-saturated mutagenesis scenario, the server can obtain the reference object set through the following algorithm:
[0128] Input data of this algorithm: p, set D train ={(S0, y0)}, vector Q.
[0129] Among them, p refers to the occurrence times of each amino acid at each mutation site, that is, the target times, for example, 2. D train refers to the first protein set. S0 in (S0, y0) represents the wild-type protein, and y0 represents the index experimental value of the wild-type protein. Vector Q refers to the current score vector, Q ∈ R [[ID=**19**]] m , m is the number of types of amino acids, for example, 20, and AAINDEX(a) represents the coordinate of amino acid a in vector Q, that is, the position of the score corresponding to amino acid a in Q.
[0130] The step of initializing the current score vector is: if the u-th amino acid appears at the mutation position of S0, then Q u = p - 1, otherwise Q u = p, Q u is the u-th element in vector Q.
[0131] The output data of this algorithm is: the updated set D train . The D output by the algorithm train is the initial reference object set.
[0132] Step 1: while Q u > 0 do. The meaning of this step is: if there are elements less than 0 in matrix Q, then execute Steps 2 to 5.
[0133] Step 2: Calculate the score of each mutant:
[0134] Among them, Score i is the current protein score of the i-th mutant protein S i in the first object set, B i ={u|AAINDEX(S ij )} represents the coordinate u of the current score corresponding to each amino acid S i at each mutation position in S ij in matrix Q.
[0135] Step 3: Select the mutant with the maximum score, \(i^*=\arg\max Score\). i This step is used to determine the mutant protein with the maximum score among the current proteins.
[0136] Among them, \(i^*\) indicates that the mutant protein with the maximum score among the current proteins is the \(i^*\)-th mutant protein in the first object set.
[0137] Step 4: Update set \(D\). train \(\leftarrow(S i ^*, y i ^*)\). This step means adding the \(i^*\)-th mutant protein in the first object set to the first protein set.
[0138] Step 5: Update the score matrix \(Q\). If the \(u\)-th amino acid appears at the mutation position of \(S i ^*\), then \(Q u = p - 1\); otherwise, \(Q u = p\). \(Q u \) is the \(u\)-th element in vector \(Q\).
[0139] Step 6: end while. If there is no element greater than 0 in \(Q\), then execute Step 7.
[0140] Step 7: Output \(D\). train
[0141] In this embodiment, since each amino acid corresponds to an amino acid, the scores in the current score set are uniquely identified by the amino acid and the mutation position. For the amino acid at each mutation position, according to the corresponding amino acid of the amino acid and the mutation position, the current score corresponding to the amino acid at the mutation position is determined from the current score set, so that the current score of each amino acid at each mutation position can be accurately and quickly determined.
[0142] The method for protein coding provided by this application (i.e., the method for determining protein characteristics) can combine Bayesian optimization to assist protein evolution. The process of encoding a protein to obtain protein characteristics can be called the process of protein feature representation. Effective protein feature representation is crucial for Bayesian optimization to find the best protein mutants. To better combine domain knowledge to construct an accurate and information-rich low-dimensional feature representation, a new low-dimensional coding strategy is proposed in this application to represent each amino acid at each site. Specifically, for two experimental scenarios in protein directed evolution: the saturation mutagenesis scenario and the non-saturation mutagenesis scenario at the \(k\)-th position, two methods are formulated to calculate the amino acid representation at each site.
[0143] In some embodiments, the object feature is a protein feature. Determining the object features of each object in the reference object set based on the index experimental values of each object in the reference object set on a preset index includes: for each mutation position, dividing the reference object set according to the type of amino acid at the mutation position to obtain a first sub-object set corresponding to each type of amino acid; for each type of amino acid at each mutation position, determining the amino acid feature of the amino acid at the mutation position based on the index experimental values of each object in the first sub-object set corresponding to the amino acid; and obtaining the protein feature of the object based on the amino acid features of the amino acids at each mutation position in the object.
[0144] Among them, the first sub-object set is uniquely determined by the mutation position and the type of amino acid. For example, if there are 4 mutation positions, namely the 1st mutation position, the 2nd mutation position, the 3rd mutation position, and the 4th mutation position, and there are 20 types of amino acids, namely the ii-th type of amino acid, 1 ≤ ii ≤ 20, then 80 first sub-object sets are generated. The mutation positions and at least one of the amino acids corresponding to different first sub-object sets are different. For example, the first sub-object set 1 is the first sub-object set corresponding to the 1st mutation position and the 1st type of amino acid, and the first sub-object set 2 is the first sub-object set corresponding to the 1st mutation position and the 2nd type of amino acid.
[0145] Specifically, for each mutation position, the server can divide the reference object set according to the type of amino acid at the mutation position to obtain a first sub-object set corresponding to each type of amino acid. For example, for the kk-th mutation position, the server can obtain the amino acid at the kk-th mutation position from each object in the reference object set to form an amino acid set corresponding to the kk-th mutation position. For example, if the parameter object set includes 40 objects, the amino acid set includes 40 amino acids, and the types of amino acids at the kk-th mutation position in different objects can be the same or different. After obtaining the amino acid set corresponding to the kk-th mutation position, the server can divide the amino acid set according to the type of amino acid to obtain multiple sub-sets, divide the same type of amino acid into the same sub-set, and divide different amino acids into different sub-sets. Each sub-set only includes one type of amino acid, and the obtained sub-sets are the first sub-object sets corresponding to each type of amino acid at the kk-th mutation position. For example, if the j-th position in the protein is the mutation position, the first sub-object set corresponding to this mutation position can be expressed as V j (a) = {i | S ij = a}, where j represents the mutation position, i is the number of the object in the reference object set, and S ij represents the i-th object S in the reference object set iThe amino acid at the mutation position j. a represents any amino acid. For example, if there are 20 amino acids, a represents any one of these 20 amino acids. If a is the first amino acid (denoted as A1), the first sub-object set corresponding to amino acid A1 at the mutation position j is V j (A1) = {i | S ij = A1}.
[0146] In some embodiments, for each amino acid at each mutation position, the server can determine the amino acid feature of the amino acid at the mutation position based on the index experimental values of each object in the first sub-object set corresponding to the amino acid. For example, the first sub-object set corresponding to amino acid A1 at the mutation position j is V j (A1) = {i | S ij = A1}, then when calculating the amino acid feature of amino acid A1 at the mutation position j, V j (A1) = {i | S ij = A1} can be used to determine the amino acid feature of amino acid A1 at the mutation position j by using the index experimental values of the objects corresponding to the numbers of each object in it.
[0147] In this embodiment, for each amino acid at each mutation position, the amino acid feature of the amino acid at the mutation position is determined based on the index experimental values of each object in the first sub-object set corresponding to the amino acid, so that the amino acid features of the same amino acid at different mutation positions are related to the mutation positions, that is, the same amino acid has different feature representations at different positions. For example, the features of the same type of amino acid at different mutation positions can be different, improving the accuracy of amino acid coding. The method for determining protein features provided in this embodiment can be applied to encoding the mutant proteins generated in the saturation mutagenesis scenario at the k position to obtain the protein features of the mutant proteins.
[0148] The object determination method provided in this application can be applied to Bayesian optimization to assist protein directed evolution using the Bayesian optimization method. Bayesian optimization can effectively explore the combinatorial space and find the optimal solution in the sample space with as few experimental times as possible among a small number of measurement samples by balancing exploration and exploitation. However, in the application of Bayesian optimization, the current coding strategies will inevitably encounter some problems. On the one hand, high-dimensional coding strategies are challenging for Bayesian optimization because successful global optimization searches require accurate and informative low-dimensional representations. On the other hand, classification labels (such as one-hot encoding) may cause the loss of knowledge about dead variants from the available experimental data of a specific protein. This can be seen from Figure 4 as shown in Figure 4 where each letter on the abscissa represents an amino acid. For example, V is an amino acid,Figure 4 The average fitness of 20 amino acids (AA, Amino Acid) at 4 mutation sites in 384 experimental samples (GB1 variants) selected from the GB1 dataset was calculated. The average fitness of each amino acid at each mutation site was obtained by calculating the average value of the affinity measurements at that mutation site, and the corresponding standard deviation is shown as error bars (i.e., Figure 4 the vertical lines in Figure 4 . It can be clearly seen from
[0149] that regardless of the amino acids selected at other positions, the presence of some dead variants at a specific mutation site will directly result in low fitness or zero fitness. Therefore, applying the existing protein coding method to the Bayesian optimization method to assist protein directed evolution usually has poor effects.
[0150] In some embodiments, determining the amino acid feature of an amino acid at a mutation position based on the index experimental values of each object in the first sub-object set corresponding to the amino acid includes: performing statistical calculations on the index experimental values of each object in the first sub-object set corresponding to the amino acid to obtain at least one index experimental statistical value; determining the amino acid feature of the amino acid at the mutation position based on the at least one index experimental statistical value.
[0151] Among them, the index experimental statistical value can be one or more, and multiple means at least two. Statistical calculations include but are not limited to at least one of calculating the mean, minimum value, or minimum value, etc.
[0152] Specifically, the server can perform statistical calculations on the index experimental values of each object in the first sub-object set corresponding to the amino acid to obtain at least one index experimental statistical value, and determine the amino acid feature of the amino acid at the mutation position based on the at least one index experimental statistical value. For example, the first sub-object set corresponding to amino acid A1 at mutation position j is V j (A1) = {i|S ij = A1}, then when calculating the amino acid feature of amino acid A1 at mutation position j, obtain V j (A1) = {i|S ij= A1}, calculate the mean value of each obtained index experimental value corresponding to the object with each number, obtain the first index mean value, determine the maximum value from each index experimental value, obtain the first index maximum value. The first index mean value is an index experimental statistical value, and the maximum value is also an index experimental statistical value. Based on at least one of the first index mean value or the first index maximum value, determine the amino acid feature of amino acid A1 at the mutation position j. For example, the first index mean value and the first index maximum value can be used as eigenvalue components of the amino acid feature respectively, that is, the amino acid feature includes the first index mean value and the first index maximum value. For example, the first index mean value can be expressed as formula (1), and the first index maximum value can be expressed as formula (2). y in formula (1) and formula (2) i represents the index experimental value of the i-th object in the reference object set on the preset index.
[0153]
[0154] Illustrate by way of example. Taking the mutant protein in the reference object set as the mutant protein generated in the saturation mutagenesis scenario at position k as an example, encode the amino acids by calculating the mean value or the maximum value of the measured affinity of each amino acid at each mutation site in the mutant protein. The feature of the corresponding mutant protein is represented by a feature vector composed of these amino acid encodings. This way allows the same amino acid to have different feature representations at different positions, creating a smoother local variable for regression.
[0155] In this embodiment, perform statistical calculations on the index experimental values of each mutant protein in the first sub-object set corresponding to the amino acid. Based on the statistically obtained index experimental statistical values, determine the amino acid feature of the amino acid at the mutation position, and improve the accuracy of the encoded amino acid feature through statistical data. The method for determining the protein feature provided in this embodiment can be applied to encoding the mutant protein generated in the saturation mutagenesis scenario at position k to obtain the protein feature of the mutant protein.
[0156] In some embodiments, the object feature is a protein feature; based on the index experimental values of each object in the reference object set on the preset index, determining the object features of each object in the reference object set includes: for each amino acid, determine the objects including the amino acid at the mutation position from the reference object set to obtain the second sub-object set corresponding to the amino acid; for each amino acid, based on the index experimental values of each object in the second sub-object set corresponding to the amino acid, determine the amino acid feature of the amino acid; based on the amino acid features of the amino acids at each mutation position in the object, obtain the protein feature of the object.
[0157] Among them, each second sub-object set corresponds to one kind of amino acid, and different second sub-object sets correspond to different amino acids.
[0158] Specifically, the objects in the reference object set are proteins, and the reference object set may include mutant proteins and wild-type proteins. For each kind of amino acid, the server can determine the objects in the reference object set that include this amino acid among the amino acids at the mutation positions, and form the second sub-object set corresponding to this amino acid. For example, for the first kind of amino acid A1, for each object, determine the amino acids at each mutation position of this object to form the amino acid set corresponding to this object. After obtaining the amino acid sets corresponding to each object in the reference object set, from the amino acid sets corresponding to each object respectively, determine the amino acid set that includes the first kind of amino acid A1, and combine the objects corresponding to the amino acid set including A1 into the second sub-object set corresponding to amino acid A1.
[0159] In some embodiments, for each object in the reference object set, for each kind of amino acid, the server can determine the mutation positions corresponding to each kind of amino acid respectively from this object. Among them, when the amino acid at mutation position 1 is A1, the mutation position corresponding to amino acid A1 is mutation position 1. Each kind of amino acid may correspond to 0, 1, or multiple mutation positions, and multiple means at least two. For example, for the i-th object S in the reference object set ij , the set N i (a) composed of the mutation positions corresponding to each kind of amino acid respectively can be expressed by formula (3), where j is the mutation position. For each kind of amino acid, based on the set composed of the mutation positions corresponding to each kind of amino acid in each amino acid, the second sub-object set corresponding to each kind of amino acid can be determined. For example, the second sub-object set V(a) can be expressed by formula (4). In formula (4), for amino acid a, if the number of mutation positions corresponding to amino acid a in the i-th object (i.e., |N i (a)|) is not 0, then the i-th object is used as the object in the second sub-object set corresponding to amino acid a. |N i (a)| represents the number of elements included in the set N i (a).
[0160] N i (a) = |{j|S ij = a}| (3) V(a) = {i||N i (a)| ≠ 0} (4)
[0161] In some embodiments, for each amino acid, the server may perform statistical calculations based on the measured values of the indicators of each object in the second sub-object set corresponding to the amino acid to obtain the amino acid feature of the amino acid. The statistical calculations include, but are not limited to, calculating at least one of the mean, maximum value, or minimum value. For example, the server may calculate the mean value of the measured values of the indicators of each object in the second sub-object set to obtain the second indicator mean value, obtain the maximum measured value of the indicators of each object in the second sub-object set to obtain the second indicator maximum value, and obtain the amino acid feature of the amino acid based on at least one of the second indicator mean value or the second indicator maximum value. For example, the second indicator mean value and the second indicator maximum value may be used as feature values to form the amino acid feature, that is, the amino acid feature includes the second indicator mean value and the second indicator maximum value. For example, the second indicator mean value may be expressed as formula (5), and the second indicator maximum value may be expressed as formula (6). The y in formula (5) and formula (6) i represents the measured value of the indicator of the i-th object in the reference object set on the preset indicator.
[0162] E max (a) = max i∈V y i (6)
[0163] In some embodiments, the server may encode the protein feature of the mutant protein based on each mutation position in the mutant protein and the amino acid feature of the amino acid at each mutation position. For example, for a mutant protein generated in a non-saturated mutagenesis scenario, if each mutant protein includes 2 mutation positions, a vector composed of these 2 mutation positions and the amino acid features of the amino acids corresponding to these 2 mutation positions respectively is determined as the protein feature of the mutant protein. Thus, for a mutant protein generated in a non-saturated mutagenesis scenario, the amino acids in the mutant protein can be encoded by calculating the average or maximum value of the fitness measurement values (measured values of indicators) of the proteins containing the amino acid at any position. The representation carrier of the mutant protein is composed of the mutation position and the corresponding mutated amino acid encoding. This encoding method is more in line with the biological significance of protein evolution and greatly reduces the dimension of the features.
[0164] In this embodiment, since the amino acid at the mutation position of each mutant protein in the second sub-object set corresponding to a certain amino acid includes this amino acid, for each amino acid, based on the index experimental values of the various mutant proteins in the second sub-object set corresponding to this amino acid, the amino acid feature of the amino acid is obtained, so as to encode this amino acid based on the index experimental value of the protein including this amino acid, improving the accuracy of amino acid encoding. The method for determining protein features provided in this embodiment can be applied to encoding the mutant proteins generated in the scenario of non-saturated mutagenesis to obtain the protein features of the mutant proteins.
[0165] In some embodiments, determining the target object that meets the index requirements of the preset index from the second object set based on the mapping relationship includes: based on the mapping relationship, determining the statistical index value of each object in the second object set on the target statistical index, and determining the selected object from the second object set based on the statistical index value; when the iteration stop condition is not met, adding the selected object to the reference object set; returning to the step of determining the object features of each object in the reference object set based on the index experimental values of each object in the reference object set on the preset index until the iteration stop condition is met; determining the selected object obtained when the iteration stop condition is met as the target object that meets the index requirements of the preset index.
[0166] Among them, the mapping relationship between the preset index and the object feature is the first mapping relationship. The iteration stop condition includes at least one of the number of iterations (i.e., the number of loops) reaching the number threshold and the index experimental value of the selected object reaching the second index threshold. The selected object can be constantly changing, and the selected objects determined in different loop numbers are different.
[0167] Specifically, the server can perform statistical calculations based on the first mapping relationship to obtain the second mapping relationship between the target statistical index and the object feature, and determine the statistical index value of each object in the second object set on the target statistical index based on the second mapping relationship. The statistical index value refers to the value of the object on the target statistical index. For example, if the second mapping relationship is represented by the curve y2 = f2(x), in order to determine the index statistical value of the object on the target statistical index, it is possible to calculate the value of y2 when x in the curve y2 = f2(x) is the object feature of this object, and determine the value of y2 as the index statistical value of this object on the target statistical index.
[0168] In some embodiments, the selected object(s) can be one or more. The object corresponding to the maximum statistical metric value can be determined as the selected object, or the object(s) with a statistical metric value greater than a third metric threshold can be determined as the selected object, and the third metric threshold can be set as needed. The server can obtain an object that meets the metric requirements of a preset metric based on the selected object. For example, the server can determine the selected object as an object that meets the metric requirements of the preset metric.
[0169] In some embodiments, when the iteration stop condition is not met, the server can add the selected object to the reference object set, and return the step of determining the object characteristics of each object in the reference object set based on the metric experimental values of each object in the reference object set on the preset metric, until the iteration stop condition is met, and determine the selected object obtained when the iteration stop condition is met as the target object that meets the metric requirements of the preset metric. For example, if the iteration stop condition is that the number of iterations (i.e., the number of loops) reaches a threshold number of times, then when the number of iterations (i.e., the number of loops) reaches the threshold number of times, the selected object is determined as the target object.
[0170] In this embodiment, when the iteration stop condition is not met, a new selected object is determined again, so as to gradually find the target object that meets the metric requirements of the preset metric. Since the selected object is added to the reference object set, the number of objects in the reference object set increases. Therefore, in the process of determining a new selected object each time, the step of determining the object characteristics of each object in the reference object set is executed, gradually improving the accuracy of the object characteristics obtained by coding and the accuracy of the finally selected target object.
[0171] In some embodiments, as Figure 5 shown, a method for object determination is provided. The object in this method is a mutant protein. This method can be executed by a terminal or a server, or can be jointly executed by a terminal and a server. Taking the application of this method to a server as an example, the method includes the following steps:
[0172] Step 502, obtain an initial score set; the initial score corresponding to each amino acid in the initial score set is the target number of times. Step 504, decrement the initial scores corresponding to the amino acids at each mutation position in the wild-type protein in the initial score set to obtain a current score set, and determine a first protein set based on the wild-type protein; the wild-type protein is a protein that has not mutated, and a second protein set is obtained based on the first object set.
[0173] Step 506: For each mutant protein in the second protein set, determine the current scores corresponding to the amino acids at each mutation position in the mutant protein from the current score set. Based on the obtained current scores, determine the current protein score of the mutant protein. Based on the current protein score, select a target protein from the second protein set.
[0174] Step 508: Decrease the current scores corresponding to the amino acids at each mutation position in the target protein in the current score set, and move the target protein from the second protein set to the first protein set.
[0175] Step 510: Determine whether there is a score greater than 0 in the current score set. If so, return to Step 506; if not, execute Step 512.
[0176] Step 512: Determine the first protein set as the reference object set.
[0177] As Figure 6 shown, a schematic diagram of an object determination method for determining mutant proteins with higher affinity is presented. The candidate sample space includes a wild-type protein and multiple mutant proteins. The reference object set in Step 508 is the initial reference object set, and the reference object set will change later. The initial reference object set is, for example, Figure 6 the initial sample set in. According to the method provided in this application for determining the initial reference object set, the initial sample set is screened from the candidate sample space. The samples in the initial sample set are either wild-type proteins or mutant proteins. Step 514: Based on the object characteristics and index experimental values of each object in the reference object set, train an index detection model. Using the trained index detection model, predict the index prediction values of each object in the first object set, and select the objects whose index prediction values meet the index value screening conditions from the first object set to obtain the second object set.
[0178] As Figure 6 shown, after obtaining the initial sample set, the "affinity" stage is used to obtain the affinity of the samples in the initial sample set (measured through experiments). After obtaining the affinity, in the "protein feature representation" stage, the proteins in the initial sample set are encoded to obtain the protein features of each protein in the initial sample set. Using the affinity of the protein (measured through experiments) and the protein features to train the index detection model. After training, input the protein features corresponding to the proteins in the candidate sample space into the index detection model to be tested, predict the affinity prediction value of the protein, and select the proteins with affinity prediction values greater than the affinity threshold from the candidate sample space to form the second object set.
[0179] By screening out an initial sample set from the candidate sample space (i.e., screening out a second object set based on the first object set), a search space pre-screening strategy is provided. The affinity values of many mutants in the candidate sample space are low. By adopting the sample search space pre-screening strategy, samples with low affinity values are eliminated in advance, reducing the sample space to be searched in Bayesian optimization and improving the calculation efficiency. For example, XGBOD can be used to eliminate mutants with low affinity from the sample search space. Specifically, in each iteration process of Bayesian optimization, the samples in the candidate sample space can be pre-screened first. Through threshold setting, samples below the threshold are judged as low fitness (elimination points, Positive class), and samples above the threshold are judged as high fitness (non-elimination points, Negative class). Using samples with existing experimental values (i.e., proteins whose affinity has been determined through experiments) as the training set to train XGBOD, screening the samples in the candidate sample space, and filtering out potential sample points with very low affinity in advance to reduce the sample volume in the sampling space and improve the model efficiency.
[0180] The process of determining a selected object from the second object set each time can be regarded as a Bayesian optimization. For example, Figure 6 in, the process of screening out samples from the remaining samples after filtering out samples with lower affinity using the acquisition function each time is a Bayesian optimization.
[0181] Step 516: Based on the index experimental values of each object in the reference object set on the preset index, determine the object characteristics of each object in the reference object set. Based on the index experimental values and object characteristics of each object in the reference object set on the preset index, determine the mapping relationship between the preset index and the object characteristics.
[0182] Among them, the mapping relationship between the preset index and the object characteristics is, for example, Figure 6 the probability surrogate model in. The probability surrogate model can be any one of the Gaussian process regression model based on Gaussian distribution or the Gaussian process regression model based on student-t distribution, etc.
[0183] Step 518: Based on the mapping relationship, determine the statistical index values of each object in the second object set on the target statistical index. Based on the statistical index values, determine the selected object from the second object set.
[0184] Such as Figure 6 in, the target statistical index is, for example, Figure 6 the acquisition function in. Determine the acquisition function using the mapping relationship between the preset index and the object characteristics, and select samples from the remaining samples based on the acquisition function.
[0185] Step 520: Determine whether the iteration stop condition is satisfied. If not, execute Step 522; if so, execute Step 524.
[0186] Step 522: Add the selected object to the reference object set, and return to Step 514.
[0187] Step 524: Determine the selected object obtained when the iteration stop condition is satisfied as the target object that meets the index requirements of the preset index.
[0188] For example, the server can use the following algorithm, which is Bayesian optimization with prescreened search space via outlier detection (ODBO), to implement the process of determining the target object from the second object set based on the reference object set.
[0189] Input data of this algorithm: Initial sample set D t , number of experiments T (i.e., the number of loops is T).
[0190] Among them, the initial sample set D t is the initial reference object set, that is, the reference object set used in the first iteration.
[0191] Output data of this algorithm: Optimal value (s*, y*) in the sample space. s* represents the optimal mutant, and y* represents the affinity of s* (measured through experiments).
[0192] Process of this algorithm:
[0193] t ← 1; / / Assign 1 to t, where t represents the current iteration number;
[0194] while t ≤ T do; / / When t ≤ T, execute downward, where T is the number of experiments (i.e., the number of loops);
[0195] if Robust GP then; / / If Robust GP is used as the probability surrogate model, execute downward; [[ID=~36]]
[0196] Use D t to train a Gaussian process regression model based on the student - t distribution; / / When t is 1, D t According to being equal to D1, it represents the initial sample set, that is, the initial reference object set (the reference object set used in the first iteration);
[0197] D tin ={(S i ,yi )||f t (S i )-y i |≤α} Filter out outliers based on the rejection threshold α and retain normal points; / / D tin Denote the set composed of the remaining samples after filtering out outliers (samples) from D t ;
[0198] Use D tin to train a Gaussian process regression model based on a Gaussian distribution;
[0199] else if GP then; / / If Robust GP is used as the probabilistic surrogate model, execute downward;
[0200] Use D t to train a Gaussian process regression model based on a Gaussian distribution;
[0201] if Naivo BO then; / / If Naive Bayesian Optimization is used, execute downward. Naivo BO refers to Naive Bayesian Optimization;
[0202] Maximize the acquisition function according to the current posterior probability / / D_search refers to the second object set, that is, the set composed of the remaining samples after filtering out samples with low affinity from the candidate sample space; S t+1 The candidate is the sample selected from the second object set (i.e., the above-mentioned selected object);
[0203] Experimentally evaluate the data point S t+1 and update the experimental result to the observed sample set D t+1 ←(S t+1 ,y t+1 ) / / y t+1 To determine the affinity of S t+1 (measured experimentally);
[0204] else if TuRBO then; / / TuRBO (Trust region Bayesian optimization);
[0205] Set the confidence threshold interval Ω according to the current confidence threshold TR;
[0206] Within the trust region Ω, randomly sample several points and maximize the acquisition function according to the current posterior probability
[0207] Experimentally evaluate the data point S t+1 and update the experimental result to the observed sample set D t+1←(S t+1 , y t+1 );
[0208] end if; / / End this loop;
[0209] Update the probabilistic surrogate model;
[0210] t ← t + 1 / / Increment t by 1;
[0211] end while;
[0212] Output (s*, y*).
[0213] Among them, TuRBO is a global optimization method. By constructing a series of local GPs surrogate models, it can avoid over-exploring highly uncertain regions in the search space from a global perspective. At the same time, it can make full use of the second-order convergence of the trust-region method locally for efficient solution.
[0214] Illustrate the basic process of the object determination method provided by this application. For example, Figure 7 As shown, it mainly includes four steps: 1) Obtain initial experimental data; 2) Characterize the data; 3) Pre-screen the search space; 4) For the screened search space, the Bayesian optimization algorithm trains a probabilistic surrogate model through the initial experimental data. After training the surrogate model, by optimizing the acquisition function, the next-round experimental samples are selected in the search space. Verify the proposed experimental design, add the experimental results to the training set, and update the posterior of the surrogate model. This process is repeated continuously until the design is maximized, the resources are exhausted, or the space is explored to conditions where it is unlikely to find improvements. Figure 7 In (a), the obtained initial experimental data is the initial reference object set screened from the first object set. (a) shows 8 mutants, and each mutant corresponds to a score. This score represents fitness. For example, the mutant "H76L, K78R" represents a mutant, and the score of "H76L, K78R" is 0.18. Figure � In (b), the data is characterized (i.e., the process of encoding to determine amino acid features), Figure 7 The bar chart in (b) shows the average fitness of 20 amino acids at the i-th mutation site, Figure 7 The table in (b) shows the average fitness of 5 amino acids at the i-th mutation site. For example, the average fitness of the amino acid "V" is 1.12. Figure 7 In (c), it is the process of pre-screening the search space. In this process, the characteristics of the mutants in the initial experimental data are determined, Figure 7"P1 P2 A1 A2" in (c) are the characteristics of the mutant, where P is the abbreviation of position, A is the abbreviation of Amino Acid, P1 and P2 respectively represent the positions of the amino acids in the mutant, and A1 and A2 respectively represent the characteristics of the amino acids. The characteristics of the determined mutant are used to pre-screen the search space. Figure 7 In (c), the solid circles represent outliers (i.e., mutants with lower fitness), and the solid triangles represent normal points (i.e., mutants with higher fitness). Figure 7 In (d) is the process of the Bayesian optimization algorithm. Through Figure 7 The four processes can determine mutants with high fitness.
[0215] For computational method-assisted experimental design, the proposed solution of this application presents an efficient, experimental design-oriented framework. By using a search space pre-screening strategy to pre-screen the samples in the candidate sample space in advance (i.e., screening the second object set from the first object set), combined with the Bayesian optimization algorithm, it balances exploration and exploitation, effectively explores the sample space, and finds the optimal experimental design scheme in as few steps as possible. In the proposed solution of this application, for the actual application scenario of protein directed evolution, an amino acid coding strategy based on average fitness is designed to accurately and effectively perform feature representation (i.e., the method of obtaining amino acid features). To better assist experimentalists in experimental design, an initial sample selection strategy is also proposed to assist experimentalists in selecting initial experimental samples (i.e., the method of determining the reference object set), so as to ensure that the amino acid coding information covered in the initial sample size has the largest coverage range and the fewest required initial experimental times. Through the Bayesian optimization algorithm with search space pre-screening, the experimental cost and time cost are reduced. An efficient, experimental design-oriented framework is implemented in the proposed solution of this application, which is called ODBO (Bayesian optimization with prescreened searchspace via outlier detection). This method assists experimental design by screening the search space and combining Bayesian optimization, helping experimentalists reduce experimental cost and time cost. For the actual application scenario of protein directed evolution, an amino acid coding strategy based on average fitness is proposed to accurately and effectively perform feature representation. To better assist experimentalists in experimental design, an initial sample selection strategy is also proposed in this solution to assist experimentalists in selecting initial experimental samples, so as to ensure that the amino acid coding information covered in the initial sample size has the largest coverage range, while the required number of experiments is the fewest, reducing the experimental cost.
[0216] The present invention can be used to solve the experimental design of computational methods for assisting protein directed evolution. The Bayesian optimization combined with search space pre-screening proposed herein can also be applied to the automated experimental design in other fields, such as new material development and battery fast charging protocols, etc.
[0217] In materials science, discovering materials with definite properties is both expensive and time-consuming. With the addition of each new component or material parameter, the space of candidate experiments grows exponentially. For example, if studying the effect of a new parameter (such as introducing doping) requires approximately 10 experiments within the parameter range, then N parameters will require 10 N possible experiments. With the emergence of each new parameter, the number of candidate experiments quickly exceeds the feasibility of exhaustive exploration. The diversity and complexity of the material composition-structure-property (CSP) relationship, including material-processing parameters and atomic disorder, make the research more chaotic. Coupled with the scarcity of the best materials, these challenges threaten innovation and industrial progress. The method of assisted material discovery based on Bayesian optimization can guide laboratory experimenters in experimental design. It balances the experiment of exploring unknown functions using experiments and identifying extreme values using prior knowledge, and can accelerate the speed of material discovery in the experiment of material exploration while consuming fewer resources.
[0218] Lithium-ion batteries are one of the most commonly used energy storage devices for electric vehicles. With the continuous progress of battery chemistry technology, an important issue is how to effectively determine the charging protocol to best balance the need for fast charging while maximizing the battery service life. However, it is not easy to determine a suitable charging protocol. On the one hand, estimating the cycle life of a battery takes from several months to several years. On the other hand, the huge parameter adjustment space and the diversity of samples make the experiment even more difficult. How to further reduce the parameter range and shorten the experimental time is crucial for the development of lithium-ion batteries. The method of computational-assisted experimental design can be used to reduce the cost of experimental optimization, use the feedback of completed experiments to provide information for subsequent experimental decisions, balance the relationship between experimental results and requirements, that is, test the experimental parameter space with high uncertainty and explore it, and predict promising parameters based on the completed experimental results. Ultimately, it can reduce the number and time of required experiments, reduce costs, and find an effective charging protocol.
[0219] The present application also provides an application scenario, which is a new material development scenario, and this application scenario applies the above object determination method. Specifically, the application of the object determination method in this application scenario is as follows: The server can obtain the index prediction values of each material in the first material set on a preset index, select the materials whose index prediction values meet the index value screening conditions from the first material set to obtain a second material set, determine the mapping relationship between the preset index and the material characteristics based on the index experimental values and material characteristics of multiple materials in the first material set, and determine the target materials that meet the index requirements of the preset index from the second material set based on the mapping relationship. Thus, materials with specified properties can be quickly determined. Among them, at least one component of each material in the first material set is different or the content of at least one component is different. As Figure 8 shown, it demonstrates the closed-loop optimization of perovskite electrolytes based on machine learning, and uses Bayesian optimization to achieve an effective experimental search for fast lithium-ion conductors from perovskite solid electrolytes.
[0220] The present application also provides an application scenario, which is a battery fast charging protocol scenario, and this application scenario applies the above object determination method. Specifically, the application of the object determination method in this application scenario is as follows: The server can obtain the index prediction values of each battery charging protocol in the first battery charging protocol set on a preset index, select the battery charging protocols whose index prediction values meet the index value screening conditions from the first battery charging protocol set to obtain a second battery charging protocol set, determine the mapping relationship between the preset index and the battery charging protocol characteristics based on the index experimental values and battery charging protocol characteristics of multiple battery charging protocols in the first battery charging protocol set, and determine the target battery charging protocols that meet the index requirements of the preset index from the second battery charging protocol set based on the mapping relationship. Thus, battery charging protocols with specified properties can be quickly determined. Among them, at least one parameter of each battery charging protocol in the first battery charging protocol set is different or the value of at least one parameter is different. As Figure 9 shown, it demonstrates the closed-loop optimization of battery fast charging protocols based on machine learning. Through machine learning methods, the parameter space is effectively optimized, and the current and voltage configuration parameters of the fast charging protocol are specified to maximize the battery life.
[0221] The object determination method provided by the present application can be deployed on a server equipped with a Linux operating system or a Windows operating system and CPU / GPU computing resources based on the Python language and the Botorch library.
[0222] To verify the effectiveness of the object determination method provided in this application in assisting protein directed evolution, tests were conducted on four protein directed evolution datasets: 1) GB1 dataset (with 55 mutant segments); 2) GB1 dataset (with 4 mutant segments); 3) BRCA1 dataset; 4) green fluorescent protein dataset.
[0223] Among them, GB1 refers to the B1 domain of protein G. Protein G is an immunoglobulin binding protein expressed in group C and G streptococci. The B1 domain of protein G (GB1) interacts with the Fc domain of immunoglobulin. We conducted experiments on the generated GB1 datasets respectively. Saturation mutagenesis was performed at four carefully selected residue sites 39, 40, 41, and 51 in GB1. There are experimentally measured fitness values among 149,361 variants. The fitness criterion is the binding affinity with IgG-Fc. One or two amino acids were mutated in the entire 55-codon random region of the GB1 protein, and a total of 536,944 mutant data were collected.
[0224] BRCA1 is a multi-domain protein belonging to the tumor suppressor gene family and most commonly mutated in three domains: the N-terminal RING domain, exons 11-13, and the BRCT domain. The BRCA1 RING domain is responsible for the E3 ubiquitin ligase activity of BRCA1 and mediates the interaction between BRCA1 and other proteins. The functional effects of single or multiple point mutations of BRCA1 residues on E3 ubiquitin ligase activity were studied. This dataset contains a total of 98,300 mutants with E3 scores.
[0225] Green fluorescent protein (GFP), also known as green fluorescent protein, was first discovered in a jellyfish with the scientific name Aequorea victoria (avGFP) and exhibits green fluorescence when exposed to light. The local fitness landscape of avGFP was analyzed by estimating the fluorescence levels of genotypes obtained through random mutagenesis of the avGFP sequence. This dataset includes 54,025 different protein sequences. Table 1 details the information of the four datasets used.
[0226] Table 1 Details of protein directed evolution datasets.
[0227]
[0228] Figure 10 The fitness distributions of different datasets are shown. Figure 10 In it, the abscissa is the metric value and the ordinate is Density (concentration or density), Figure 10 In (a), it is the fitness distribution of the GB1(4) dataset,Figure 10 In (b), it is the fitness distribution of the dataset GB1(55). Figure 10 In (c), it is the fitness distribution of the dataset BRCA1. Figure 10 In (d), it is the fitness distribution of the dataset avGFP.
[0229] An initial sample selection strategy is adopted to generate an initial sample set. For the GB1(4) dataset that conforms to the saturated mutation scenario, it is set that each amino acid appears at least 2 times at each position, and 40 initial training samples are obtained. For the GB1(55), Ube4b, and avGFP datasets that conform to the non - saturated mutation scenario, it is set that each amino acid appears at least once at all positions, and 136, 217, and 142 initial training samples are obtained respectively. For the ODBO algorithm, the filtering threshold for pre - screening the search space is set to 0.05. For each method, we conduct each experiment with 10 different random seeds. Each method selects one sample from the sample space each time in the GB1(55), Ube4b, and avGFP datasets and runs 50 iterations. For the GB1(55) dataset, one sample is selected from the sample space each time and 100 iterations are run. The expected improvement (EI) is adopted as the acquisition function. Ube4b and BRCA1 refer to the same protein.
[0230] Figure 11 Summarizes the performance of different methods on four protein directed evolution datasets. Among them, dataset 1 refers to the dataset GB1(4), dataset 2 refers to the dataset GB1(55), dataset 3 refers to the dataset Ube4, and dataset 4 refers to avGFP. Method 1 refers to the method of random selection (Random), method 2 refers to the method of TuRBO combined with GP, method 3 refers to the method of ODBO combined with TuRBO and GP, method 4 refers to the method of ODBO combined with TuRBO and RobustGP, method 5 refers to the method of BO combined with GP, method 6 refers to the method of ODBO combined with BO and GP, method 7 refers to the method of ODBO combined with BO and RobustGP. Figure 11 Each of the four figures in it includes a straight line, and this straight line refers to the true maximum fitness.
[0231] Figure 12Summarizes the comparison of different methods for four protein directed evolution datasets. Each curve represents the average value obtained by each method over 10 different random seeds. Among them, F1 represents the method of ODBO combined with TuRBO and GP with q = 1, F2 represents the method of ODBO combined with TuRBO and GP with q = 5, F3 represents the method of ODBO combined with TuRBO and GP with q = 10, F4 represents the method of ODBO combined with TuRBO and RobustGP with q = 1, F5 represents the method of ODBO combined with TuRBO and RobustGP with q = 5, and F6 represents the method of ODBO combined with TuRBO and RobustGP with q = 10. G1 represents ODBO combined with TuRBO and GP with the acquisition function of expected improvement, G2 represents ODBO combined with TuRBO and GP with the acquisition function of confidence bound strategy, and G3 represents ODBO combined with TuRBO and GP with the acquisition function of Thompson sampling. q represents the number of samples selected for the next round of experiments in each iteration.
[0232] It can be found that ODBO achieves the best performance on all datasets. The search space pre-screening step can perform more effective sample acquisition, which helps to find mutants with optimal properties faster. For example, for the saturation mutagenesis scenario (i.e., the GB1(4) dataset), the method of ODBO combined with TuRBO and RobustGP can find the optimal variable (fitness = 8.76) with fewer than 50 evaluations in a large sample space (204 = 16000). However, among the methods without using the pre-screening strategy, Bayesian optimization algorithms (such as BO, TuRBO) usually converge to a poor local optimum, thus reducing the average performance. This illustrates the importance of search space pre-screening. For the non-saturation mutagenesis case, except for the method of BO combined with GP, almost all Bayesian optimization methods adopt the proposed low-dimensional protein coding strategy to find the optimal mutants. Although all methods can only find near-optimal mutants in the GB1(55) and avGFP datasets, the method proposed in this scheme is also superior to other methods.
[0233] Table 2 shows the proportions of samples belonging to the top 1%, 2%, and 5% affinity values in the sample space screened by different computational methods in the 50-round recommended selection on the GB1(4) dataset. It can be seen that adopting sample space pre-screening is more conducive to selecting better samples from the sample selection in each round for the next round of experimental tests.
[0234] Table 2
[0235] Method Top 1% Top 2% Top 5% Random 1.8 3.6 6.4 Naive BO+GP 14 20.6 31.2 TuRBO+GP 20.8 32.2 45 ODBO,BO+GP 29.6 41 62.2 ODBO,TuRBO+GP 31.6 44.6 67.2 ODBO,BO+RobustGP 35.6 50 65.8 ODBO,TuRBO+RobustGP 41.2 58.2 71.2
[0236] In addition, the performance of the Bayesian optimization algorithm with different acquisition functions and probabilistic surrogate models in protein directed evolution was also tested. As Figure 10 shown, the performance of the Bayesian optimization algorithm with different acquisition functions and probabilistic surrogate models in the GB1(4) dataset is presented. Figure 10 In (a) of [reference], the batch size at each iteration is different, Figure 10 and (b) of [reference] shows the performance when using EI, UCB, PI, and TS as acquisition functions in the "ODBO, TuRBO+GP" method, Figure 10 and (c) of [reference] shows the performance when using EI, UCB, PI, and TS as acquisition functions in the "ODBO, TuRBO+RobustGP" method.
[0237] We also calculated the computational resources consumed by each method when running on the GB1 dataset, as shown in Table 3. Using the traditional coding method (here, the feature georgiev encoded using physicochemical properties) has 76-dimensional features, and using TuRBO requires a large amount of computational resources and time. When using the amino acid coding rule proposed by us, the feature dimension of amino acids can be reduced to 4 dimensions, and the computational time and resource consumption can be significantly reduced. In addition, by adopting the search pre-screening strategy, the computational time and resources consumed can be greatly reduced. Moreover, ODBO can find the optimal value in the sample space with the fewest experimental steps, which helps to reduce the experimental cost and time cost.
[0238] Table 3
[0239]
[0240] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same moment, but can be executed at different moments. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.
[0241] Based on the same inventive concept, an embodiment of this application further provides an object determination device for implementing the object determination method involved above. The implementation solutions provided by this device for solving problems are similar to the implementation solutions described in the above method. Therefore, the specific limitations in one or more embodiments of the object determination device provided below can refer to the limitations on the object determination method in the above text, and will not be elaborated here.
[0242] In some embodiments, as Figure 13 shown, an object determination device is provided, including: a predicted value acquisition module 1302, an object set obtaining module 1304, a mapping relationship determination module 1306, and a target object determination module 1308, where:
[0243] The predicted value acquisition module 1302 is configured to acquire the predicted index values of each object in the first object set on a preset index.
[0244] The object set obtaining module 1304 is configured to select the objects in the first object set whose predicted index values meet the index value screening condition to obtain a second object set.
[0245] The mapping relationship determination module 1306 is configured to determine the mapping relationship between the preset index and the object feature based on the experimental index values and object features of multiple objects in the first object set on the preset index.
[0246] The target object determination module 1308 is configured to determine the target object that meets the index requirements of the preset index from the second object set based on the mapping relationship.
[0247] In some embodiments, the objects in the first object set are mutant proteins, and the device further includes a reference object set screening module, configured to screen and obtain a reference object set based on the first object set; the reference object set satisfies the condition that each amino acid appears at least a target number of times at each mutation position; the predicted value acquisition module is further configured to train an index detection model based on the object features and experimental index values of each object in the reference object set; and use the trained index detection model to predict the predicted index values of each object in the first object set.
[0248] In some embodiments, the mapping relationship determination module is further configured to determine the object features of each object in the reference object set based on the experimental index values of each object in the reference object set on the preset index; and determine the mapping relationship between the preset index and the object feature based on the experimental index values and object features of each object in the reference object set on the preset index.
[0249] In some embodiments, the reference object set screening module is further configured to obtain a current score set; the current score set includes the current scores respectively corresponding to each amino acid; obtain a second protein set based on the first object set, and select a target protein from the second protein set based on the current score set; decrease the current scores respectively corresponding to each amino acid at each mutation position in the target protein in the current score set, and move the target protein from the second protein set to the first protein set; in the case that the current score set indicates that the first protein set does not meet the condition that each amino acid appears at least the target number of times at each mutation position, return to the step of selecting a target protein from the second protein set based on the current score set until the current score set indicates that the first protein set meets the condition that each amino acid appears at least the target number of times at each mutation position, and determine the first protein set as the reference object set.
[0250] In some embodiments, the reference object set screening module is further configured to obtain an initial score set; the initial score corresponding to each amino acid in the initial score set is the target number of times; decrease the initial scores respectively corresponding to each amino acid at each mutation position in the wild-type protein in the initial score set to obtain a current score set, and determine a first protein set based on the wild-type protein; the wild-type protein is a protein that has not undergone mutation.
[0251] In some embodiments, the mapping relationship determination module is further configured to, for each mutant protein in the second protein set, determine the current scores respectively corresponding to each amino acid at each mutation position in the mutant protein from the current score set; determine the current protein score of the mutant protein based on the obtained respective current scores; and select a target protein from the second protein set based on the current protein score.
[0252] In some embodiments, each amino acid corresponds to an amino acid, and the scores in the current score set are uniquely identified by the amino acid and the mutation position; the mapping relationship determination module is further configured to, for each amino acid at each mutation position, determine the current score corresponding to the amino acid at the mutation position from the current score set according to the amino acid corresponding to the amino acid and the mutation position.
[0253] In some embodiments, the object feature is a protein feature, and the mapping relationship determination module is further configured to, for each mutation position, divide the reference object set according to the type of the amino acid at the mutation position to obtain a first sub-object set respectively corresponding to each amino acid; for each amino acid at each mutation position, determine the amino acid feature of the amino acid at the mutation position based on the index experimental values of each object in the first sub-object set corresponding to the amino acid; and obtain the protein feature of the object based on the amino acid features of the amino acids at each mutation position in the object.
[0254] In some embodiments, the mapping relationship determination module is further configured to perform statistical calculations on the index experimental values of each object in the first sub-object set corresponding to the amino acid to obtain at least one index experimental statistical value; and determine the amino acid characteristics of the amino acid at the mutation position based on the at least one index experimental statistical value.
[0255] In some embodiments, the object feature is a protein feature; the mapping relationship determination module is further configured to: for each type of amino acid, determine the objects in the reference object set that include the amino acid at the mutation position to obtain a second sub-object set corresponding to the amino acid; for each type of amino acid, determine the amino acid characteristics of the amino acid based on the index experimental values of the respective objects in the second sub-object set corresponding to the amino acid; and obtain the protein feature of the object based on the amino acid characteristics of the amino acids at each mutation position in the object.
[0256] In some embodiments, the target object determination module is further configured to determine the statistical index values of each object in the second object set on the target statistical index based on the mapping relationship, and determine the selected object from the second object set based on the statistical index values; in the case where the iteration stop condition is not satisfied, add the selected object to the reference object set; return to the step of determining the object features of each object in the reference object set based on the index experimental values of each object in the reference object set on the preset index until the iteration stop condition is satisfied; and determine the selected object obtained when the iteration stop condition is satisfied as the target object that meets the index requirements of the preset index.
[0257] Each module in the above object determination device can be implemented in whole or in part by software, hardware, and their combination. Each of the above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form so that the processor can call and execute the operations corresponding to each of the above modules.
[0258] In some embodiments, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 14As shown in the figure. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store the data involved in the object determination method. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it implements an object determination method.
[0259] In some embodiments, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as Figure 15 shown in the figure. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface, the display unit, and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner. The wireless manner can be achieved through WIFI, a mobile cellular network, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements an object determination method. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the computer device housing, or an external keyboard, touchpad, or mouse, etc.
[0260] Those skilled in the art can understand that Figure 14 and Figure 15The structure shown is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0261] In some embodiments, a computer device is provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above object determination method are implemented.
[0262] In some embodiments, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above object determination method are implemented.
[0263] In some embodiments, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the steps in the above object determination method are implemented.
[0264] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data that have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.
[0265] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memories can include read-only memory (ROM), magnetic tapes, floppy disks, flash memories, optical memories, high-density embedded non-volatile memories, resistive random access memories (ReRAM), magnetoresistive random access memories (MRAM), ferroelectric random access memories (FRAM), phase change memories (PCM), graphene memories, etc. Volatile memories can include random access memory (RAM) or external cache memories, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.
[0266] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0267] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.
Claims
1. An object determination method, characterized in that, The method includes: Filtering a reference object set based on a first object set; the objects in the first object set are mutant proteins; the reference object set satisfies the condition that each amino acid appears at least a target number of times at each mutation position; Training an index detection model based on the object characteristics of each object in the reference object set and the index experimental values on a preset index; Using the trained index detection model to predict the index prediction values of each object in the first object set on the preset index; Selecting the objects in the first object set whose index prediction values meet the index value screening conditions to obtain a second object set; Determining the mapping relationship between the preset index and the object characteristics based on the index experimental values and object characteristics of multiple objects in the first object set on the preset index; Determining target objects that meet the index requirements of the preset index from the second object set based on the mapping relationship.
2. The method according to claim 1, wherein The filtering the reference object set based on the first object set includes: Obtaining a second protein set based on the first object set, selecting a target protein from the second protein set, updating the current score set based on the target protein, and moving the target protein from the second protein set to the first protein set, continuously selecting proteins to update the current score, and characterizing the situation where the first protein set satisfies the condition that each amino acid appears at least a target number of times at each mutation position in the current score set, and determining the first protein set as the reference object set.
3. The method according to claim 2, wherein The determining the mapping relationship between the preset index and the object characteristics based on the index experimental values and object characteristics of multiple objects in the first object set on the preset index includes: Determining the object characteristics of each object in the reference object set based on the index experimental values of each object in the reference object set on the preset index; Determining the mapping relationship between the preset index and the object characteristics based on the index experimental values and object characteristics of each object in the reference object set on the preset index.
4. The method according to claim 2, wherein The obtaining a second protein set based on the first object set, selecting a target protein from the second protein set, updating the current score set based on the target protein, and moving the target protein from the second protein set to the first protein set, continuously selecting proteins to update the current score, and characterizing the situation where the first protein set satisfies the condition that each amino acid appears at least a target number of times at each mutation position in the current score set, and determining the first protein set as the reference object set includes: Obtaining the current score set; the current score set includes the current scores corresponding to each amino acid respectively; Obtaining a second protein set based on the first object set, and selecting a target protein from the second protein set based on the current score set; Decreasing the current scores corresponding to the amino acids at each mutation position in the target protein in the current score set, and moving the target protein from the second protein set to the first protein set; In the case where the current score set indicates that the first protein set does not meet the condition that each amino acid appears at least the target number of times at each mutation position, return the step of selecting a target protein from the second protein set based on the current score set until the current score set indicates that the first protein set meets the condition that each amino acid appears at least the target number of times at each mutation position, and determine the first protein set as the reference object set.
5. The method according to claim 4, characterized in that, The obtaining of the current score set includes: Obtain an initial score set; the initial score corresponding to each amino acid in the initial score set is the target number of times. Decrease the initial scores corresponding to the amino acids at each mutation position in the wild-type protein in the initial score set to obtain the current score set, and determine the first protein set based on the wild-type protein; the wild-type protein is a protein that has not undergone mutation.
6. The method according to claim 4, characterized in that, The selecting of the target protein from the second protein set based on the current score set includes: For each mutant protein in the second protein set, determine the current scores corresponding to the amino acids at each mutation position in the mutant protein from the current score set. Determine the current protein score of the mutant protein based on the obtained current scores. Select the target protein from the second protein set based on the current protein score.
7. The method according to claim 6, wherein Each amino acid corresponds to an amino acid score, and the scores in the current score set are uniquely identified by the amino acid and the mutation position. The determining of the current scores corresponding to the amino acids at each mutation position in the mutant protein from the current score set includes: For the amino acid at each mutation position, determine the current score corresponding to the amino acid at the mutation position from the current score set according to the amino acid score corresponding to the amino acid and the mutation position.
8. The method according to claim 3, wherein The object feature is a protein feature, and the determining of the object features of the objects in the reference object set based on the index experimental values of each object in the reference object set on the preset index includes: For each mutation position, divide the reference object set according to the type of the amino acid at the mutation position to obtain the first sub-object set corresponding to each amino acid. For each amino acid at each mutation position, determine the amino acid feature of the amino acid at the mutation position based on the index experimental values of each object in the first sub-object set corresponding to the amino acid. Obtain the protein feature of the object based on the amino acid features of the amino acids at each mutation position in the object.
9. The method according to claim 8, wherein The determining of the amino acid feature of the amino acid at the mutation position based on the index experimental values of each object in the first sub-object set corresponding to the amino acid includes: Perform statistical calculations on the index experimental values of each object in the first sub-object set corresponding to the amino acid to obtain at least one index experimental statistical value. Determine the amino acid feature of the amino acid at the mutation position based on the at least one index experimental statistical value.
10. The method according to claim 3, wherein The object feature is a protein feature; determining the object features of each object in the reference object set based on the index experimental values of each object in the reference object set on the preset index includes: For each amino acid, determine, from the reference object set, the objects in which the amino acid at the mutation position includes the amino acid, to obtain a second sub-object set corresponding to the amino acid; For each amino acid, based on the index experimental values of each object in the second sub-object set corresponding to the amino acid, determine the amino acid feature of the amino acid; Based on the amino acid features of the amino acids at each mutation position in the object, obtain the protein feature of the object.
11. The method according to claim 3, wherein Determining the target object that meets the index requirements of the preset index from the second object set based on the mapping relationship includes: Based on the mapping relationship, determine the statistical index values of each object in the second object set on the target statistical index, and determine the selected objects from the second object set based on the statistical index values; In the case where the iteration stop condition is not satisfied, add the selected objects to the reference object set; Return the step of determining the object features of each object in the reference object set based on the index experimental values of each object in the reference object set on the preset index until the iteration stop condition is satisfied; Determine the selected objects obtained when the iteration stop condition is satisfied as the target objects that meet the index requirements of the preset index.
12. An object determination device, characterized in that, The device includes: A reference object set screening module, configured to screen a reference object set based on a first object set; the objects in the first object set are mutant proteins; the reference object set satisfies the condition that each amino acid appears at least a target number of times at each mutation position; A predicted value acquisition module, configured to train an index detection model based on the object features of each object in the reference object set and the index experimental values on a preset index; use the trained index detection model to predict the index predicted values of each object in the first object set on the preset index; An object set obtaining module, configured to select, from the first object set, the objects whose index predicted values meet the index value screening condition, to obtain a second object set; A mapping relationship determination module, configured to determine the mapping relationship between the preset index and the object features based on the index experimental values and object features of multiple objects in the first object set on the preset index; A target object determination module, configured to determine the target objects that meet the index requirements of the preset index from the second object set based on the mapping relationship.
13. The object determination device according to claim 12, characterized in that, The reference object set screening module is further configured to obtain a second protein set based on the first object set, select a target protein from the second protein set, update the current score set based on the target protein, and move the target protein from the second protein set to the first protein set. Continuously select proteins to update the current score. When the current score set represents the condition that the first protein set satisfies that each amino acid appears at least the target number of times at each mutation position, the first protein set is determined as the reference object set.
14. The object determination device according to claim 13, characterized in that, The mapping relationship determination module determines the object characteristics of each object in the reference object set based on the index experimental values of each object in the reference object set on the preset index; and determines the mapping relationship between the preset index and the object characteristics based on the index experimental values and object characteristics of each object in the reference object set on the preset index.
15. The object determination device according to claim 13, characterized in that, The reference object set screening module obtains a current score set; the current score set includes the current scores corresponding to each amino acid respectively. Obtain a second protein set based on the first object set, and select a target protein from the second protein set based on the current score set; decrease the current scores corresponding to the amino acids at each mutation position in the target protein in the current score set, and move the target protein from the second protein set to the first protein set; when the current score set represents the condition that the first protein set does not satisfy that each amino acid appears at least the target number of times at each mutation position, return to the step of selecting a target protein from the second protein set based on the current score set until the current score set represents the condition that the first protein set satisfies that each amino acid appears at least the target number of times at each mutation position, and determine the first protein set as the reference object set.
16. The object determination device according to claim 15, characterized in that, The reference object set screening module obtains an initial score set; the initial score corresponding to each amino acid in the initial score set is the target number; decrease the initial scores corresponding to the amino acids at each mutation position in the wild-type protein in the initial score set to obtain the current score set, and determine the first protein set based on the wild-type protein. The wild-type protein is a protein that has not undergone mutation.
17. The object determination device according to claim 15, characterized in that, The reference object set screening module determines, for each mutant protein in the second protein set, the current scores corresponding to the amino acids at each mutation position in the mutant protein from the current score set. Determine the current protein score of the mutant protein based on the obtained current scores. Select a target protein from the second protein set based on the current protein score.
18. The object determination device according to claim 17, characterized in that Each amino acid corresponds to an amino acid score, and the scores in the current score set are uniquely identified by the amino acid and the mutation position. For each amino acid at a mutation position, the reference object set screening module determines the current score corresponding to the amino acid at the mutation position from the current score set according to the amino acid score corresponding to the amino acid and the mutation position.
19. The object determination device according to claim 14, characterized in that, The object feature is a protein feature. For each mutation position, the mapping relationship determination module divides the reference object set according to the type of amino acid at the mutation position to obtain a first sub-object set corresponding to each amino acid; for each amino acid at each mutation position, based on the index experimental values of each object in the first sub-object set corresponding to the amino acid, the amino acid feature of the amino acid at the mutation position is determined. Based on the amino acid features of the amino acids at each mutation position in the object, the protein feature of the object is obtained.
20. The object determination device according to claim 19, characterized in that, The mapping relationship determination module performs statistical calculations on the index experimental values of each object in the first sub-object set corresponding to the amino acid to obtain at least one index experimental statistical value; based on the at least one index experimental statistical value, the amino acid feature of the amino acid at the mutation position is determined.
21. The object determination device according to claim 14, characterized in that The object feature is a protein feature; for each amino acid, the mapping relationship determination module determines the objects in the reference object set whose amino acids at the mutation position include the amino acid to obtain a second sub-object set corresponding to the amino acid; for each amino acid, based on the index experimental values of each object in the second sub-object set corresponding to the amino acid, the amino acid feature of the amino acid is determined; based on the amino acid features of the amino acids at each mutation position in the object, the protein feature of the object is obtained.
22. The object determination device according to claim 14, characterized in that, The target object determination module, based on the mapping relationship, determines the statistical index values of each object in the second object set on the target statistical index, and determines the selected objects from the second object set based on the statistical index values; when the iteration stop condition is not satisfied, the selected objects are added to the reference object set; return the step of determining the object features of each object in the reference object set based on the index experimental values of each object in the reference object set on the preset index until the iteration stop condition is satisfied; the selected objects obtained when the iteration stop condition is satisfied are determined as the target objects that meet the index requirements of the preset index.
23. A computer device includes a memory and a processor, the memory storing a computer program, wherein, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 11 are implemented.
24. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 11 are implemented.
25. A computer program product comprising a computer program, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 11 are implemented.