Gene design and evolution path inference method based on adaptive constrained depth model

The genetic design and evolutionary path inference method are constructed through adaptive band constraint depth model, which solves the problem of difficult analysis of high-dimensional genotype space in traditional methods, realizes the direct correlation between genotype mutations and phenotypic changes, and improves the efficiency of synthetic biological design regulatory sequences.

CN120388610AActive Publication Date: 2025-07-29GUANGDONG UNIV OF TECH
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510503439.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-07-29
Estimated Expiration
2045-04-22

AI Technical Summary

Technical Problem

The existing technology is difficult to globally characterize the relationship between gene sequence-expression-fitness, and traditional methods are difficult to effectively visualize and analyze the evolutionary paths of high-dimensional genotype spaces, and the efficiency of synthetic biological design regulation sequences is inefficient.

Method used

Adaptive band-constrained depth model is adopted, and evolutionary vectors are generated by constructing the autoencoder model and the Transformer model, and small-batch iterative training is performed. Combined with Monte Carlo sampling and fitness landscape analysis, evolutionary sequences that meet the conditions are extracted.

Benefits of technology

The direct association between genotype mutations and phenotypic changes is achieved, the efficiency of synthetic biological design regulated sequences is improved, and the target sequence is quickly found through low-dimensional fitness landscape visualization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388610A_ABST
    Figure CN120388610A_ABST
Patent Text Reader

Abstract

The invention relates to a gene design and evolution path inference method based on a self-adaptive constrained depth model, which comprises the following steps: generating a random gene sequence set through a preset algorithm rule, converting the random gene sequence set into an evolutionary sequence containing evolutionary information, namely an evolutionary vector, and calculating mutation robustness; constructing a deep self-encoding network model, dynamically determining the optimal prototype number of the network through small-batch iterative training by adopting a prototype clustering analysis method, and performing model parameter optimization based on training data; inputting the evolutionary vector into an encoder module of an auto-encoder to obtain a low-dimensional representation of the evolutionary vector, namely an evolutionary space; establishing a deep regression prediction model, and accurately calculating the protein expression level corresponding to each regulatory gene sequence; establishing a low-dimensional tensor according to the obtained low-dimensional representation of the evolutionary vector and the corresponding expression value, performing Monte Carlo sampling in the low-dimensional tensor, mapping the protein expression value to the fitness, and generating a fitness contour map containing an evolution trajectory, namely obtaining a fitness terrain containing evolution information; and dyeing the low-dimensional tensor by using mutation robustness, and selecting a gene sequence with required properties on the graph. According to the technical scheme, the efficiency problem of existing gene sequence design is effectively solved, and the problem that the fitness terrain cannot represent evolution information of each genotype is effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to technical fields such as computers, genetic big data and gene sequence design, and in particular to a gene design and evolutionary path inference method based on an adaptive constrained deep model. Background Art

[0002] The evolution of cis-regulatory elements (CREs) plays a leading role in the evolution of gene expression. Mutations in these regulatory sequences can disrupt their binding to transcription factors, thereby altering the temporal patterns, spatial distribution, and expression levels of gene expression, ultimately affecting the phenotypic traits and adaptive capacity of an organism. Notably, transcription factors evolve more slowly due to their need to coordinate the regulation of multiple target genes, while CREs exhibit faster evolutionary dynamics. This difference has led to cis-regulatory elements being considered an important molecular basis for driving the evolution of phenotypic diversity.

[0003] Currently, the regulatory sequence space grows exponentially with the number of base pairs. For example, an 80-base pair sequence has 4 80 There are many possibilities. Traditional studies are mostly based on local neighborhood sampling of natural sequences, covering only a very small part of the sequence space, and are unable to globally characterize the relationship between sequence-expression-fitness. Traditional adaptive landscapes usually group genotypes based on sequence similarity, but genotypes with similar sequences may have very different functions. The high dimensionality of the genotype space makes it difficult for traditional methods to effectively visualize and analyze global evolutionary paths, and the rational design of regulatory sequences in synthetic biology relies on trial-and-error experiments, which is inefficient. Deep learning technology, with its powerful nonlinear fitting ability and feature abstraction potential, provides a new approach for constructing sequence-expression global mapping and designing sequences. However, in the field of regulatory sequence evolution modeling, existing models ignore evolutionary constraints, resulting in unstable performance of synthetically designed sequences in real biological environments. Therefore, there is an urgent need to develop a new technical solution to comprehensively solve the problems existing in existing technologies. Summary of the Invention

[0004] The purpose of the present invention is to provide a gene design and evolutionary path inference method based on an adaptive constrained deep model, which can globally characterize the relationship between sequence-expression-fitness and design gene sequences through the evolvability space.

[0005] In order to solve the above technical problems, the present invention adopts the following technical solutions:

[0006] A method for gene design and evolutionary path inference based on an adaptive constrained deep model, comprising the following steps:

[0007] R1: With the constraint of fixing the front and back base pairs, insert a random sequence with the number of base pairs L in the middle region, that is, a sequence that may not exist in nature, transform it into an evolvability vector E containing evolutionary information, and calculate the corresponding mutation robustness;

[0008] R2: Construct a prototype-regularized autoencoder model AANet;

[0009] R3: Perform mini-batch iterative training on AANet based on the partial evolvability vector E output by R1 to dynamically determine the number of prototype parameters of AANet; complete the full-scale training of AANet using all evolvability vectors E, and input the evolvability vector E into the encoder of AANet to output the projection of the evolvability vector in the low-dimensional space, that is, construct an evolvability space;

[0010] R4: Construct a Transformer depth prediction model, and generate gene sequences to train the Transformer model according to the method of R1. Input the gene sequences obtained in R1 into the trained depth prediction model to obtain the corresponding expression level for each vector;

[0011] R5: Perform tensor fusion on the low-dimensional space projection of the evolvability vector obtained in R3 and the protein expression level values obtained in R4 to generate a low-dimensional tensor; perform Monte Carlo sampling on the low-dimensional tensor, calculate the fitness values of the sampling points through a preset protein expression level - fitness mapping function, and generate an isocontour map of the fitness landscape based on the spatial distribution of the fitness values;

[0012] R6: Extract evolution sequences that meet the conditions from the low-dimensional fitness landscape tensor according to the mutation robustness.

[0013] Among them, step R1 is specifically:

[0014] R1-1: Generate gene sequences in the way of fixing the front and back base pairs and having L random base pairs in the middle;

[0015] R1-2: For each base sequence s0 with length L, only considering single-base mutations, there are 3 kinds of mutation sequences corresponding to each base, and there are a total of 3L kinds of mutation possibilities. For each mutation sequence s i in its 3L mutation neighborhood, calculate the difference d i = f(s i ) - f(s0), where f(s) represents the expression value predicted by the deep model;

[0016] R1-3: Sort all d i values so that they satisfy d i ≤ d i+1, an increasing 240 - dimensional evolvability vector E = [d1, d2,..., d 3L is obtained.

[0017] R1 - 4: Calculate the mutation robustness: For any given base sequence s i , first construct its 3L mutation neighborhood, and then count the number of variants with expression changes less than the threshold ε in this neighborhood. The threshold is set to twice the standard deviation of the genome - wide protein expression (ε = 2σ). The formula for calculating the mutation robustness is:

[0018]

[0019] In the formula represents the number of variants with expression difference d ij less than the threshold. Evolutionary analysis shows that gene regulatory elements under stabilizing selection usually exhibit significantly enhanced mutation robustness, and this property can effectively buffer the perturbation of random mutations to key phenotypic traits.

[0020] Among them, the specific steps of R2 are as follows:

[0021] R2 - 1: Construct a network model with an encoder EN and a decoder DN. The encoder EN consists of multiple fully - connected layers, and the input is the L network nodes of the sequence; through multiple fully - connected layers, it is connected to k - 1 network nodes, that is, the output of the encoder EN; where L is the length of the input sequence s i , and k is the number of prototype quantities. The network accepts a sequence s i , and generates a prototype space of a specific dimension k through the encoder EN.

[0022] The decoder DN is a mirror image of the encoder EN. The input of DN is the output of EN, and the output of DN is the L network nodes. Try to reconstruct the points in the prototype space through DN.

[0023] R2 - 2: Prototype analysis is a high - dimensional data reduction method based on the extreme value principle. Its mathematical model represents the observed data set E as a convex combination of a finite set of prototypes. These prototypes are essentially pure types representing extreme trade - off relationships in the feature space. For example, in evolutionary biology, different prototypes may correspond to the optimal phenotypic configurations under specific ecological niches. The formula for prototype analysis is as follows:

[0024]

[0025] Among them, there is a prototype regularization constraint:

[0026]

[0027] α ij ≥0, i = 1,…, n, j = 1,…, k

[0028] c1, …, c k is the prototype, and the dimension of the prototype is the same as that of s i , and α ij is the constraint of the convex combination, and F is the encoder EN in R2-1. According to the idea of prototype analysis, any sequence E i can be linearly represented by the convex combination of prototypes:

[0029] Because of the constraint Therefore, the number of prototypes set in the network can be reduced to k - 1, and the remaining one can be reproduced through this constraint.

[0030] R2-3: The reconstruction loss of the network is as follows:

[0031]

[0032] Among them, the specific steps of R3 are as follows:

[0033] R3-1: Use the partial evolvability vector E obtained in step R1 to train AANet in mini-batches: Adopt different numbers of prototype numbers k, train the same batch of data, calculate the reconstruction losses of the network corresponding to different prototype numbers using the MSE formula in R2-2, and select the points with smaller reconstruction losses of the training data and fewer prototype numbers.

[0034] R3-2: Input all the evolvability vectors E obtained in step R1 into the autoencoder network constructed in R2. The specific steps are as follows:

[0035] Determine the number of prototypes through step R3-1, set the parameters of the network, and input all the evolvability vectors to train the autoencoder network.

[0036] Extract all the low-dimensional output vectors of the encoder EN, and all the vectors form the evolvability space.

[0037] Among them, the specific steps of R4 are as follows:

[0038] R4-1: The DNA-protein expression prediction model based on Transformer adjusts the structure, converts the DNA sequence into a high-dimensional vector through the embedding layer, and superimposes the position encoding to retain the sequence order. Adopt a pure encoder architecture and stack N layers, with each layer containing a multi-head self-attention module and a feed-forward network: The multi-head attention calculates the multi-level global dependencies between bases in parallel, the feed-forward network strengthens the local feature expression, and the residual connection and layer normalization ensure the training stability. Finally, the global features output by the encoder are pooled and mapped to the continuous protein expression prediction value through the fully connected layer.

[0039] R4-2: Regenerate a random dataset. In the training dataset, the ATCG values of the sequences are converted into one-hot encoding. Each gene sequence has its corresponding gene expression level, which is used as the output. Use this dataset to train the Transformer model.

[0040] R4-3: Convert the ATCG values of all the sequences generated by R1 into one-hot encoding and input them into the deep prediction model trained in R4-2. Let the deep prediction model network be a function f to obtain the expression levels of all evolvability vectors as f(E).

[0041] Among them, step R5 is specifically as follows:

[0042] R5-1: Generate a low-dimensional tensor using the evolvability vectors obtained by R3.

[0043] R5-2: Conduct Monte Carlo simulation on the low-dimensional tensor obtained in R5-1, randomly sample 1,000 points, perform cubic spline interpolation on the fitness axis through a curve from protein expression value to fitness, and map the corresponding protein expression values obtained by R4 to fitness, and scale the fitness to between 0 and 1.

[0044] R5-3: At this time, the fitness values of each point in the low-dimensional tensor are between 0 and 1. Then, draw a contour plot through these points to obtain a low-dimensional fitness landscape diagram. A series of studies such as the study of evolutionary paths or the quantification of epistasis can be carried out on this diagram.

[0045] Among them, step R6 is specifically as follows:

[0046] R6-1: Stain the low-dimensional tensor obtained in R5-1 with the mutation robustness MR obtained by R1-4.

[0047] R6-2: Since the regulatory sequences of genes whose expression levels are regulated by stabilizing selection are often more robust to mutation effects, analyze this diagram, analyze the tensor distribution, and extract the evolutionary sequence E that meets the conditions. For example: if genes with expression levels regulated by stabilizing selection are required, then select the sequence E with better mutation robustness. i 。

[0048] R6-3: Since the evolvability vector E i and the gene sequence s i are in one-to-one correspondence, therefore through E i the gene sequence s i can be deduced. Compared with the prior art, the gene design and evolutionary path inference method based on the adaptive constrained deep model provided by the above technical solution has the following beneficial effects:

[0049] 1. The present invention utilizes the evolvability vector to collect the evolutionary information of each gene sequence. By traversing the predicted effects of all single-base mutations in the sequence on the phenotype, that is, traversing the evolutionary possibilities of the gene, the evolvability vector of each genotype is constructed to comprehensively characterize its potential to generate phenotypic variations. The constructed fitness landscape directly associates genotype mutations with phenotypic changes, avoiding reliance on sequence similarity compared with the traditional method of grouping based on gene sequence similarity.

[0050] 2. The high dimensionality of the genotype space makes it difficult for traditional methods to effectively visualize and analyze the global evolutionary path. The present invention realizes the effective low-dimensional visualization of the fitness landscape through the prototype-regularized autoencoder network. The efficiency of the trial-and-error experiments for synthetic biology design and regulation sequences is low. The present invention uses deep learning technology to construct a sequence-expression model and an autoencoder model, and can quickly find the target sequence through the evolvability space, improving the efficiency of the trial-and-error experiments for designing regulation sequences. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 The autoencoder model in the embodiment of the present invention;

[0052] Figure 2 The construction method of the evolvability vector of the present invention;

[0053] Figure 3 The Transformer model in the embodiment of the present invention;

[0054] Figure 4 The evolvability space in the embodiment of the present invention;

[0055] Figure 5 The specific implementation flowchart of the gene design and evolutionary path inference method based on the adaptive constrained depth model in the implementation of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0056] In order to make the objectives and advantages of the present invention clearer, the present invention will be specifically described below in conjunction with embodiments. It should be understood that the following text only describes one or several specific implementation manners of the present invention, and does not strictly limit the specific protection scope claimed by the present invention.

[0057] This embodiment takes the generated dataset of the regulation sequence and the protein expression level of the controlled gene as an example, and combines the attached Figures 1 to 4 , to further analyze and explain the present invention.

[0058] The gene design and evolutionary path inference method based on the adaptive constrained depth model includes the following steps:

[0059] Step 1. Generate a random gene sequence in the way of fixing the first 13 and the last 17 bases and making the middle 80 bases random, and transform the middle 80-base sequence into an evolvability vector E containing evolutionary information (refer to Figure 1 ):

[0060] 1.1 Generate 100,000 gene sequences in the way of fixing the first 13 bases and the last 17 bases and making the middle 80 bases random;

[0061] 1.2 For each base sequence s0 with length L, only considering single-base mutations, there are 3 kinds of mutant sequences corresponding to each base, and there are a total of 3L kinds of mutation possibilities. For each mutant sequence s i in its 3L mutation neighborhood, calculate the difference d i = f(s i ) - f(s0), where f(s) represents the expression value predicted by the deep model;

[0062] 1.3 Sort all d i values to make it satisfy d i ≤ d i+1 , and obtain a monotonically increasing 240-dimensional evolvability vector D = [d1, d2,..., d 3L .

[0063] 1.4 Calculate the mutation robustness using the formula , set the threshold ε as twice the standard deviation of the whole-genome protein expression level (ε = 2σ), represents the number of variants with expression difference less than the threshold (refer to Figure 1 ).

[0064] Step 2. Build a prototype-regularized autoencoder model AANet and build a Transformer deep prediction model, specifically:

[0065] 2.1 Build a network model with an encoder EN and a decoder DN. The encoder EN consists of multiple fully connected layers, and the input is L network nodes of the sequence; through multiple fully connected layers, it is connected to k - 1 network nodes, that is, the output of the encoder EN; where L is the length of the input sequence s i , and k is the number of prototype quantities. The network accepts a sequence s i , and generates a prototype space with a specific dimension k through the encoder EN.

[0066] The decoder DN is a mirror image of the encoder EN. The input of DN is the output of EN, and the output of DN is L network nodes. Try to reconstruct the points in the prototype space through DN.

[0067] 2.2 The formula for prototype analysis is as follows:

[0068]

[0069] Among them, there is a prototype regularization constraint:

[0070]

[0071] α ij ≥0, i = 1, …, n, j = 1, …, k

[0072] c1, …, c k are prototypes, and the dimension of the prototypes is the same as that of s i . α ij is the constraint of the convex combination, and F is the encoder EN in R2-1. According to the idea of prototype analysis, any sequence E i can be linearly represented by the convex combination of prototypes:

[0073] Because of the constraint Therefore, the number of prototypes set in the network can be reduced to k - 1, and the remaining one can be reproduced through this constraint.

[0074] 2.3 The reconstruction loss of the network is as follows:

[0075]

[0076] Step 3: Use the partial evolvability vector E obtained in Step 1 to train AANet in small batches to determine the number of prototypes required for AANet obtained in Step 2; use the entire evolvability vector E to train AANet and input it into the encoder part of the autoencoder to obtain the two-dimensional representation of the evolvability vector, that is, the evolvability space.

[0077] 3.1 Use different numbers of prototype quantities to train the same batch of data, and calculate the reconstruction loss of the network corresponding to different prototype quantities using the MSE formula in Step 2.2.

[0078] 3.2 Draw an elbow graph with the prototype quantity as the horizontal axis and the reconstruction loss as the vertical axis, and obtain that the prototype quantity corresponding to the elbow inflection point is 3. According to the constraint, the number of nodes in the middle layer of the network can be obtained as 2.

[0079] 3.3 Determine the prototype quantity as 3 through Step 3.1, set the parameters of the network, and set the number of nodes in the middle layer to 2 (refer to Figure 2 ), and train the autoencoder network with the input of the entire evolvability vector.

[0080] 3.4 Since the determined prototype quantity is 3, extract all 2D output vectors of the encoder EN, and all the vectors form the evolvability space, and the vectors are all wrapped by the prototype vectors.

[0081] Step 4: Input all the evolvability vectors obtained in Step 1 into the trained deep prediction model to obtain the corresponding expression levels for each vector.

[0082] 4.1 The DNA-protein expression level prediction model based on Transformer adjusts the structure to convert the DNA sequence into a high-dimensional vector through an embedding layer and superimposes positional encoding to retain the sequence order (refer to Figure 3 ). A pure encoder architecture is stacked with 3 layers, and each layer contains a multi-head self-attention module and a feed-forward network: the multi-head attention calculates the multi-level global dependencies between bases in parallel, the feed-forward network strengthens the local feature expression, and the residual connection and layer normalization ensure the training stability. Finally, after the global features output by the encoder are pooled, they are mapped to the continuous protein expression level prediction value through a fully connected layer, and the attention weights simultaneously provide the interpretability basis for the key regulatory base regions.

[0083] 4.2 Regenerate 100,000 random datasets according to the method in Step 1. In the training dataset, the ATCG values of the sequences are converted into one-hot encodings, and each gene sequence has its corresponding gene expression level, which is used as the output. Use this dataset to train the Transformer model.

[0084] 4.3 Convert the ATCG values of all the sequences generated in Step 1 into one-hot encodings and input them into the deep prediction model trained in 4.1. Let the deep prediction model network be a function f, and obtain the expression levels of all evolvability vectors as f(E).

[0085] Step 5: The optimal number of prototypes determined in Step 3 is 3, and a 2D representation of the evolvability vectors can be obtained. Draw a 2D scatter plot. Sample in the scatter plot, map it to fitness through a curve from a protein expression value to fitness, and then draw a contour plot to obtain a low-dimensional fitness landscape.

[0086] R5-1: Generate a low-dimensional tensor using the evolvability vectors obtained in Step 3.

[0087] R5-2: Conduct Monte Carlo simulation in the low-dimensional tensor obtained in Step 5.1, randomly sample 1000 points, perform cubic spline interpolation on the fitness axis through a curve from a protein expression value to fitness, and map the corresponding protein expression values obtained in Step 4 to fitness, and scale the fitness to between 0 and 1.

[0088] R5-3: At this time, the fitness values of each point in the low-dimensional tensor are between 0 and 1. Then, draw a contour plot through these points to obtain a low-dimensional fitness landscape map. A series of studies such as the evolutionary path or the quantification of epistasis can be conducted on this map.

[0089] Step 6. Extract the evolution sequences that meet the conditions from the low-dimensional fitness landscape tensor according to the mutational robustness.

[0090] 6.1 Stain the low-dimensional tensor obtained in Step 5.1 with the mutational robustness MR obtained in Step 1.4.

[0091] 6.2 Since the gene regulatory elements under the action of stabilizing selection usually exhibit significantly enhanced mutational robustness, by analyzing the low-dimensional tensor, it can be known that the sequences strongly affected by stabilizing selection in gene expression are often far from Prototype 1.

[0092] 6.3 According to the required properties, if sequences under strong stabilizing selection need to be found, search for sequences as far as possible from Prototype 1.

[0093] In the gene sequence exploration and design method provided in this embodiment, an evolvability vector is used to reflect the evolutionary possibility of genes, comprehensively characterize the potential of generating phenotypic variations, and avoid relying on sequence similarity; at the same time, a Transformer deep model is used to realize the prediction from gene sequences to gene expression levels, achieving an accurate sequence-expression level correspondence.

[0094] At the same time, the present invention uses a prototype-regularized autoencoder network to realize the effective low-dimensional visualization of the fitness landscape. Through the fitness landscape, the target sequence can be quickly found, improving the efficiency of the trial-and-error experiment for designing regulatory sequences. In summary, the method of the present invention effectively solves the problem that the existing methods group based on gene sequence similarity, resulting in a large change in the fitness landscape; improves the efficiency of the trial-and-error experiment for designing regulatory sequences, and also has a high calculation speed, so it is suitable for use in the real world.

[0095] The embodiments of the present invention have been described in detail above with reference to the accompanying drawings. However, the present invention is not limited to the above embodiments. For those of ordinary skill in the art in this technical field, after learning the content recorded in the present invention, without departing from the principle of the present invention, several equivalent transformations and substitutions can still be made, and these equivalent transformations and substitutions should also be regarded as belonging to the protection scope of the present invention.

Claims

1. A method for gene design and evolutionary path inference based on an adaptive depth model with constraints, characterized in that Including the following steps: R1: With fixed front and back base pairs as constraints, insert a random sequence with the number of base pairs L in the middle region, that is, a sequence that may not exist in nature, transform it into an evolvability vector E containing evolutionary information, and calculate the corresponding mutation robustness; R2: Construct a prototype-regularized autoencoder model AANet; R3: Based on the partial evolvability vector E output by R1, perform mini-batch iterative training on AANet to dynamically determine the number of prototype parameters of AANet; Complete the full-scale training of AANet using all evolvability vectors E, and input the evolvability vector E into the encoder of AANet to output the projection of the evolvability vector in the low-dimensional space, that is, construct an evolvability space; R4: Construct a Transformer deep prediction model, and generate gene sequences to train the Transformer model according to the method of R1. Input the gene sequences obtained in R1 into the trained deep prediction model to obtain the corresponding expression level for each vector; R5: Perform tensor fusion on the low-dimensional space projection of the evolvability vector obtained in R3 and the protein expression level values obtained in R4 to generate a low-dimensional tensor; perform Monte Carlo sampling on the low-dimensional tensor, calculate the fitness values of the sampling points through a preset protein expression level - fitness mapping function, and generate an isocontour map of the fitness landscape based on the spatial distribution of the fitness values; R6: Extract evolution sequences that meet the conditions in the low-dimensional fitness landscape tensor according to the mutation robustness.

2. The method for gene design and evolution path inference based on an adaptive band-constrained depth model according to claim 1, wherein The specific steps of step R1 are as follows: R1-1: Generate gene sequences in the way of fixing the front and back base pairs and having L random base pairs in the middle; R1-2: For each base sequence s of length L i , considering only single-base mutations, there are 3 types of mutant sequences corresponding to each base, and there are a total of 3L possible mutations. For each mutant sequence s ij in its 3L mutant neighborhood, calculate the difference d ij = f(s ij ) - f(s i ), where f(s) represents the expression value predicted by the deep model; R1-3: Sort all d values in the base sequence s i so that they satisfy d ij ≤ d ij to obtain a monotonically increasing 3L-dimensional evolvability vector E i(j+1) = [d i , d i1 ,..., d i2 ,..., d i(3L) . R1-4: Calculate Mutation Robustness: For any given base sequence s i , first construct its 3L mutation neighborhood, and then count the number of variants with an expression change less than the threshold ε in this neighborhood. The threshold is set to twice the standard deviation of the protein expression across the whole genome (ε = 2σ). The formula for calculating mutation robustness is: In the formula represents the expression difference d ij The number of variants less than the threshold. Evolutionary analysis shows that gene regulatory elements under stabilizing selection usually exhibit significantly enhanced mutational robustness, a property that can effectively buffer the perturbation of key phenotypic traits by random mutations.

3. The method for gene design and evolution path inference based on an adaptive band-constrained depth model according to claim 1, wherein The specific steps of step R2 are as follows: R2-1: Construct a network model with an encoder EN and a decoder DN. The encoder EN consists of multiple fully connected layers, and the input is 3L network nodes of the sequence; through multiple fully connected layers, it is connected to k - 1 network nodes, which is the output of the encoder EN; where 3L is the length of the input sequence E i , and k is the number of prototype quantities. The network accepts a sequence E i , and generates a prototype space of a specific dimension k through the encoder EN. The decoder DN is a mirror image of the encoder EN. The input of DN is the output of EN, and the output of DN is 3L network nodes. Reconstruct the points in the prototype space as much as possible through DN. R2-2: Prototype analysis is a high-dimensional data dimensionality reduction method based on the extreme value principle. Its mathematical model expresses the observed data set E as a convex combination of a finite set of prototypes. These prototypes are essentially pure types representing extreme trade-off relationships in the feature space. For example, in evolutionary biology, different prototypes may correspond to the optimal phenotype configurations under specific ecological niches. The formula of prototype analysis is as follows: Among them, there is a prototype regularization constraint: α ij ≥ 0, i = 1, …, n, j = 1, …, k c1,…,c k is a prototype, and the dimension of the prototype is the same as that of s i , and α ij is the constraint of the convex combination, and F is the encoder EN in R2-1. According to the idea of prototype analysis, any sequence E i can be linearly represented by the convex combination of prototypes: Because of the constraint Therefore, the number of prototypes set in the network can be reduced to k - 1, and the remaining one can be reproduced through this constraint. The reconstruction loss of the network is as follows:

4. The method for gene design and evolutionary path inference based on an adaptive band-constrained depth model according to claim 1, wherein The specific steps of step R3 are as follows: R3-1: Use the partial evolvability vector E obtained in step R1 to perform mini-batch training on AANet. The specific steps are as follows: Adopt different numbers of prototype numbers k to train the same batch of data, calculate the reconstruction losses of the networks corresponding to different prototype numbers using the MSE formula in R2-2, and select the points with smaller reconstruction losses of the training data and fewer prototype numbers. R3-2: Input all the evolvability vectors E obtained in step R1 into the autoencoder network constructed in R2. The specific steps are as follows: Determine the number of prototypes through step R3-1, set the parameters of the network, and train the autoencoder network with the input of all evolvability vectors. Extract all the low-dimensional output vectors of the encoder EN, and all the vectors form an evolvability space.

5. The method for gene design and evolution path inference based on an adaptive band-constrained depth model according to claim 1, wherein Input all the gene sequences obtained in R1 into the trained deep prediction model. Specifically, step R4 is as follows: R4-1: After adjustment, the Transformer model can be used for the task of predicting protein expression levels from DNA sequences. Its input is a DNA sequence, such as a base sequence of length L. Each base needs to be converted into a high-dimensional vector through an embedding layer first, and position encoding is added to retain the sequence order features. Since the task is regression prediction rather than sequence generation, the model only retains the encoder part and stacks N layers, combines multiple self-attention mechanisms to capture long-range dependencies in the DNA sequence, and finally maps the global features output by the encoder to a continuous value through a fully connected layer. The output is the predicted value of the protein expression level for a single sample. R4-2: The encoder of the Transformer is stacked by multiple layers, and each layer contains two core modules: the self-attention mechanism and the feed-forward neural network. First, calculate the correlation weights between all bases in the sequence through the self-attention mechanism, and use multi-head attention to capture different levels of regulatory patterns in parallel; then the feed-forward network performs a non-linear transformation on the vector at each position to enhance the local feature expression ability. Each layer alleviates the problem of gradient disappearance through residual connections and normalization. Finally, the encoder outputs a global feature vector of the same length as the input sequence, which integrates the global dependencies and local structural information of the DNA sequence, and is directly mapped to the predicted value of the protein expression level through a fully connected layer. R4-3: Regenerate a random dataset according to the method of R1. In the training dataset, the values of ATCG in the sequence are converted into one-hot encoding, and each gene sequence has its corresponding gene expression level, which is used as the output. Use this dataset to train the Transformer model. R4-4: Convert the ATCG values of all the sequences generated in R1 into one-hot encoding, input them into the deep prediction model trained in R4-3, and let the deep prediction model network be a function f to obtain the expression levels of all evolvability vectors as f(E).

6. The method for gene design and evolutionary path inference based on an adaptive band-constrained depth model according to claim 1, wherein Specifically, step R5 is as follows: R5-1: Generate a low-dimensional tensor using the evolvability vectors obtained in R3. R5-2: Conduct Monte Carlo simulation on the low-dimensional tensor obtained in R5-1, randomly sample 1000 points, perform cubic spline interpolation on the fitness axis through a curve from protein expression value to fitness, and map the corresponding protein expression value obtained in R4 to the fitness, and scale the fitness to between 0 and 1. R5-3: At this time, the fitness value of each point in the low-dimensional tensor is between 0 and 1. Then draw a contour map through these points to obtain a low-dimensional fitness landscape map. A series of studies such as the study of evolutionary paths or the quantification of epistasis can be carried out on this map.

7. The method for gene design and evolution path inference based on an adaptive band-constrained depth model according to claim 1, wherein Specifically, step R6 is as follows: R6-1: Stain the low-dimensional tensor obtained in R5-1 with the mutation robustness MR obtained in R1-4. R6-2: Since the regulatory sequences of genes whose expression levels are subject to stabilizing selection tend to be more robust to mutational effects, analyze this figure, analyze the tensor distribution, and extract the evolutionary sequence E that meets the conditions. For example, if a gene whose expression level is subject to stabilizing selection is required, then select the sequence E with better mutational robustness. i 。 R6-3: Due to the evolvability vector E i and the gene sequence s i being in one-to-one correspondence, the gene sequence s i can be deduced inversely through E i .

Citation Information

Patent Citations

  • Remote sensing image content description method based on variational self-attention reinforcement learning

    CN111126282A

  • Traditional Chinese medicine prescription recommendation method and system based on new crown protein heterogeneous network clustering

    CN112863634A

  • Techniques for early detection of variants of interest

    CN118284941A

  • Adaptive method and device during testing based on dual-path adversarial lifting denoising

    CN119152255A