Prediction model training method, guide nucleotide sequence screening method, and related device
By training an AI model to predict gene expression data based on nucleotide sequences, the problem of low sgRNA screening efficiency in existing technologies has been solved, enabling rapid and accurate screening of suitable sgRNAs and improving the application efficiency of the CRISPR-Cas9 gene editing system.
Patent Information
- Application Number
- CN202511264316.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-05
- Publication Date
- 2025-12-30
- Estimated Expiration
- 2045-09-05
AI Technical Summary
Existing technologies are labor-intensive and time-consuming in finding suitable sgRNAs, and it is difficult to quickly and accurately screen for efficient sgRNAs, resulting in low application efficiency of the CRISPR-Cas9 gene editing system.
Nucleotide sequences can be screened by training a prediction model, using AI models to predict nucleotide sequences, and combining this with gene expression data for screening, thus reducing the need for experimental verification.
This improved the efficiency of sgRNA screening, reduced manpower and time costs, and enhanced the application efficiency of the CRISPR-Cas9 gene editing system.
Smart Images

Figure CN120748487B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of bioinformatics processing, specifically to a predictive model training method, a guided nucleotide sequence screening method, and related apparatus. Background Technology
[0002] Gene editing systems, exemplified by CRISPR-Cas9, are revolutionary technologies that allow scientists to modify DNA with unprecedented precision. This system consists of endonucleases and guide RNA (sgRNA). The core technology is finding the key sgRNA, which guides the endonuclease to precisely locate the cleavage sites on nucleic acids, allowing the gene editing task to be completed according to the biologist's design.
[0003] Currently, the approach to finding suitable sgRNAs involves searching for potential sgRNAs based on target genes. This typically involves comprehensively considering information such as the GC content (the ratio of guanine nucleotide G to cytosine nucleotide C in the gene), sequence length, and chemical modifications of the sgRNA sequence to identify a suitable candidate. However, the number of candidate sgRNAs identified based on this information often reaches hundreds, and determining which one is optimal remains unknown. This necessitates researchers conducting experiments on each candidate sgRNA and comparing the results to find the optimal sgRNA. However, this method is extremely labor-intensive and time-consuming, making it inefficient and slow in finding the optimal sgRNA. Summary of the Invention
[0004] This application provides a prediction model training method, a guided nucleotide sequence screening method, and related apparatus. By training a prediction model, the model can be used to predict the corresponding gene expression data for nucleotide sequences, thereby improving the efficiency of guided nucleotide sequence optimization.
[0005] The first aspect of this application provides a method for training a gene expression level prediction model, the method comprising:
[0006] Obtain the initial prediction model;
[0007] Multiple sets of training samples are obtained. Each set of training samples includes a nucleotide sequence and the actual gene expression level data of the edited gene corresponding to the nucleotide sequence. The nucleotide sequence includes a guide nucleotide sequence and a gene sequence associated with the guide nucleotide sequence. The edited gene is obtained by gene editing of the gene sequence based on the guide nucleotide sequence.
[0008] Multiple sets of training samples are input into the initial prediction model, so that the initial prediction model predicts the gene expression level data corresponding to the nucleotide sequence based on the nucleotide sequence in each set of training samples, and adjusts the model parameters based on the prediction results and the actual gene expression level data until convergence, thereby obtaining the target prediction model.
[0009] The target prediction model is used to predict gene expression levels based on the input gene sequence and its corresponding guide nucleotide sequence.
[0010] A second aspect of this application provides a guided nucleotide sequence screening method, the method comprising:
[0011] Obtain the target prediction model corresponding to the gene sequence of the target species; the target prediction model is obtained by training the gene expression level prediction model training method of the target species on the gene sequence of the target species.
[0012] Obtain multiple candidate guide nucleotide sequences determined based on the gene sequences of the target species;
[0013] The gene sequence of the target species and each of the candidate guide nucleotide sequences are input into the target prediction model, so that the target prediction model predicts the gene expression level data corresponding to each candidate guide nucleotide sequence based on the gene sequence of the target species and the candidate guide nucleotide sequence.
[0014] Based on the gene expression data corresponding to each of the candidate guide nucleotide sequences predicted by the target prediction model, the optimal guide nucleotide sequence is determined among the plurality of candidate guide nucleotide sequences.
[0015] A third aspect of this application provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the methods of the first or second aspect described above.
[0016] A fourth aspect of this application provides a computer storage medium storing instructions that, when executed on a computer, cause the computer to perform the methods described in the first or second aspect.
[0017] As can be seen from the above technical solutions, the embodiments of this application have the following advantages:
[0018] By training a prediction model to predict the corresponding gene expression data for nucleotide sequences, and then finding a suitable guide nucleotide sequence based on the gene expression level, it is possible to reduce the manpower and time costs of determining a suitable guide nucleotide sequence without having to conduct experiments. Compared with simple bioinformatics analysis, this method can greatly improve the efficiency of guide nucleotide sequence optimization. Attached Figure Description
[0019] Figure 1 This is a flowchart illustrating the gene expression level prediction model training method in the embodiments of this application;
[0020] Figure 2 This is a flowchart illustrating the method for obtaining the initial prediction model in an embodiment of this application.
[0021] Figure 3 This is an exemplary schematic diagram illustrating the changing trend of the loss function during model training in the genome learning phase of this application embodiment;
[0022] Figure 4 This is another flowchart illustrating the method for obtaining the initial prediction model in this application embodiment;
[0023] Figure 5 This is an exemplary schematic diagram illustrating the changing trend of the loss function during the model training phase in the genome and expression level learning stage of this application.
[0024] Figure 6 This is an exemplary schematic diagram illustrating the multiple stages of learning the genome, learning the transcriptome, and learning from experimental data in the embodiments of this application.
[0025] Figure 7 This is an exemplary schematic diagram illustrating the changing trend of the loss function during the experimental data learning phase in this application embodiment;
[0026] Figure 8 This is an exemplary schematic diagram illustrating the performance of the model during the experimental data learning phase in the embodiments of this application;
[0027] Figure 9 This is a flowchart illustrating the guided nucleotide sequence screening method in the embodiments of this application;
[0028] Figure 10 This is a schematic diagram of a computer device in an embodiment of this application. Detailed Implementation
[0029] This application provides a prediction model training method, a guided nucleotide sequence screening method, and related apparatus. By training a prediction model, the model can be used to predict the corresponding gene expression data for nucleotide sequences, thereby improving the efficiency of guided nucleotide sequence optimization.
[0030] In recent years, gene editing technology derived from the CRISPR-Cas9 system has developed into a rapid technique for editing, silencing (making genes lose their ability to express themselves), or activating genes. Due to the targeted gene editing capabilities of this type of system, dozens of therapies that replace defective genes by introducing exogenous genes have been approved for treating human diseases. Furthermore, the activation or silencing of gene expression based on modified Cas9 proteins has also been effectively applied in agriculture. These technologies have effectively enhanced the cold resistance and disease resistance of crops, and significantly improved the nutritional value and shelf life of agricultural products.
[0031] However, CRISPR-Cas9-type gene editing systems suffer from significant off-target effects (sgRNA binding to non-target DNA). Studies have shown that off-target effects occur when there are three mismatches between the target sequence and the sgRNA (mismatch: the gene and RNA do not bind together according to the complementary pairing principle). Off-target effects can lead to harmful events such as sequence mutations, deletions, rearrangements, immune responses, and activation of oncogenes, thus limiting the application of gene editing systems. To address this issue, sgRNA optimization is necessary.
[0032] Engineered bacteria, such as *Escherichia coli*, *Bacillus subtilis*, *Lactobacillus*, *Acetobacter*, *Corynebacterium glutamicum*, and *Pseudomonas*, have demonstrated enormous application potential in fields such as industrial biotechnology, environmental remediation, waste treatment, industrial production, food processing, and agricultural production due to their outstanding metabolic diversity, environmental robustness, and adaptability. Currently, these bacteria are widely used in various biotechnology fields, including biopharmaceuticals, aquaculture, pollutant degradation, functional food development, food fermentation, plant growth promotion, and the production of industrially valuable compounds.
[0033] However, despite the wide range of applications for these bacteria, current technologies still struggle to rapidly, accurately, and cost-effectively screen for efficient sgRNAs from gene regulatory regions. Developing a new technology that can quickly and accurately assist researchers in identifying potential sgRNAs would significantly improve the efficiency of these bacteria in industrial applications and help researchers explore their functions more deeply.
[0034] While various techniques exist for identifying efficient sgRNAs in gene regulatory regions, each has its own distinct advantages and disadvantages. Bioinformatics prediction tools can rapidly screen for potential sgRNAs, but their accuracy is limited by the sophistication of algorithms and databases, potentially leading to false positives or false negatives. Experimental validation techniques provide the most reliable data through direct biological experiments, but these are time-consuming, costly, and involve complex data analysis. Analysis methods based on gene expression data, combining gene expression levels with single-cell techniques, can reveal the nuanced effects of sgRNAs on gene activity, but the data is complex, analysis is challenging, and single-cell culture contributes to high experimental costs. Therefore, despite progress in sgRNA screening, each method has limitations, necessitating further technological innovation and optimization to improve screening efficiency and accuracy.
[0035] To address the aforementioned technical problems, this application proposes a gene expression level prediction model training method and a guided nucleotide sequence screening method. Compared with related technical solutions, the underlying algorithm used in this application differs. Related technical solutions often employ sequence alignment as the underlying algorithm for bioinformatics analysis, requiring massive computational resources to obtain a result. In contrast, this application uses an AI model as the underlying algorithm; once the model tool is trained, it can quickly output results with minimal computational resources. Furthermore, regarding the result generation process, this application considers nucleotide sequence information such as the organism's genome and transcriptome, achieving higher accuracy, while related technical solutions neglect this information.
[0036] The following describes the training method for the gene expression level prediction model in the embodiments of this application:
[0037] Please see Figure 1 One embodiment of the gene expression level prediction model training method in this application includes:
[0038] 101. Obtain the initial prediction model;
[0039] The method in this embodiment can be executed by a computer device, which can be a terminal, server, or other device with data processing and computing capabilities. The computer device can obtain an initial prediction model to be trained, which can be, for example, any AI model capable of processing nucleotide sequence data.
[0040] 102. Obtain multiple sets of training samples, each set of training samples including nucleotide sequences and the actual gene expression data of the edited gene corresponding to the nucleotide sequence. The nucleotide sequence includes a guide nucleotide sequence, a gene sequence associated with the guide nucleotide sequence, and a regulatory element sequence corresponding to the associated gene. The edited gene is obtained by gene editing of the gene sequence based on the guide nucleotide sequence.
[0041] The computer device can also acquire multiple sets of training samples, each set of training samples including a nucleotide sequence and the actual gene expression data of the edited gene corresponding to the nucleotide sequence. The nucleotide sequence includes a guide nucleotide sequence and a gene sequence associated with the guide nucleotide sequence. The edited gene is obtained by gene editing based on the guide nucleotide sequence of the gene sequence.
[0042] For example, the guide nucleotide sequence can be a guide nucleotide sequence such as sgRNA that can guide effector proteins (such as cleavage enzymes / activating proteins) to a specific site of a gene to facilitate gene editing. For example, it can guide endonucleases to a specific site of a gene for cleavage to facilitate gene editing, or guide effector proteins to a specific site of a gene to bind to regulate genome transcription, and so on.
[0043] Guide RNA is an RNA molecule with a specific sequence and structure that recognizes and binds to target DNA or RNA, thereby guiding related enzymes or protein complexes to perform precise cutting, modification, or transcription. Based on their function and mechanism of action, guide RNAs can be classified into various types, including but not limited to crRNA (CRISPR RNA) and tracrRNA (trans-activating CRISPR RNA) in gene editing systems, as well as artificially designed sgRNA (single-stranded guide RNA). In addition, there are guide RNAs involved in the RNA editing process, and rRNA that plays a role in ribosome synthesis and function.
[0044] The actual gene expression level data of the edited gene can be a specific numerical value of the actual expression level of the edited gene, or it can be data that represents the change in the actual expression level of the edited gene, such as the rate of change, ratio, or difference between the expression level of the edited gene before and after editing.
[0045] The gene sequence associated with the guide nucleotide sequence can be a nucleotide sequence of a gene from any species, such as a gene sequence from a bacterial genome, or a gene sequence from the genome of another plant or animal. This gene sequence may include gene regulatory region sequences, gene coding region sequences, and other gene-related nucleotide sequences, without limitation here.
[0046] 103. Input multiple sets of training samples into the initial prediction model, so that the initial prediction model predicts the gene expression level data corresponding to the nucleotide sequence based on the nucleotide sequence in each set of training samples, and adjusts the model parameters based on the prediction results and the actual gene expression level data until convergence, to obtain the target prediction model.
[0047] The initial prediction model can be trained based on multiple sets of training samples. During training, it predicts the gene expression level corresponding to the nucleotide sequence in each training sample. The model parameters are then adjusted based on the prediction results and the actual gene expression data until convergence. In other words, the model learns the relationship between nucleotide sequences and gene expression levels during training, continuously adjusting the learning effect based on the actual gene expression data of the edited gene. When the convergence condition is met, the model possesses the ability to accurately predict the gene expression level corresponding to the nucleotide sequence. Therefore, the trained target prediction model can be used to predict the corresponding gene expression level for an input gene sequence and its corresponding guide nucleotide sequence.
[0048] Among them, the gene expression data predicted by the model can be used to determine the appropriate guide nucleotide sequence. For example, if gene editing is for the purpose of enhancing gene expression, a guide nucleotide sequence corresponding to a higher gene expression level can be found as the appropriate guide nucleotide sequence; if gene editing is for the purpose of suppressing gene expression, a guide nucleotide sequence corresponding to a lower gene expression level can be found as the appropriate guide nucleotide sequence.
[0049] Similarly, the gene expression data output by the model prediction can be a specific numerical value of gene expression, or it can be a representation of the change in gene expression, such as the rate of change, ratio, or difference between the gene expression level in the predicted nucleotide sequence before it was edited and the gene expression level after it was edited based on the guide nucleotide sequence.
[0050] It should be noted that inputting multiple sets of training samples into the initial prediction model can be done by inputting all the training samples into the model at once, or by inputting the training samples into the model in batches so that the model can learn and optimize step by step, thereby improving the training efficiency of the model.
[0051] In this embodiment, a prediction model is trained to predict the corresponding gene expression data for nucleotide sequences, and then a suitable guide nucleotide sequence is found based on the gene expression level. This eliminates the need to determine the suitable guide nucleotide sequence through experiments, reducing manpower and time costs. Compared with simple bioinformatics analysis, this method can greatly improve the efficiency of guide nucleotide sequence optimization.
[0052] The initial prediction model mentioned above can be an untrained model or a pre-trained model. Therefore, based on Figure 1 In one optional implementation of the illustrated embodiment, the initial prediction model can be a model pre-trained based on gene nucleotide sequences. Specifically, the method for obtaining the initial prediction model may include... Figure 2 The following steps are shown:
[0053] 201. Obtain the first initial prediction model;
[0054] 202. Obtain multiple sets of first training samples, each set of first training samples including the nucleotide sequence of the first gene;
[0055] 203. Input multiple sets of the first training samples into the first initial prediction model, so that the first initial prediction model predicts the i to i+n nucleotides of the first gene nucleotide sequence based on the first j nucleotides of the first gene nucleotide sequence in each set of the first training samples, and adjusts the model parameters of the first initial prediction model according to the prediction results and the true values of the i to i+n nucleotides until convergence, thereby obtaining the initial prediction model.
[0056] A computer device can acquire a first initial prediction model and multiple sets of first training samples. Each set of first training samples includes a first gene nucleotide sequence. The multiple sets of first training samples are input into the first initial prediction model to allow the model to learn the genome. The first initial prediction model predicts the i-th to i+n-th nucleotides of the first gene nucleotide sequence based on the first j nucleotides of the first gene nucleotide sequence in each set of first training samples. The model parameters of the first initial prediction model are adjusted based on the prediction results and the actual values of the i-th to i+n-th nucleotides until convergence, thus obtaining the initial prediction model. Here, j, i, and n are all integers, and i > j ≥ 1.
[0057] It should be noted that inputting multiple sets of first training samples into the first initial prediction model can be done by inputting these multiple sets of first training samples into the model all at once for training, or by inputting these multiple sets of first training samples into the model in batches for training, so that the model can learn and optimize step by step, thereby improving the training efficiency of the model.
[0058] For example, the GPT model (a model framework proposed by OpenAI) is chosen as the underlying computational method, while the *E. coli* genome is used as the training dataset for the model. At this stage, the *E. coli* genome can be cut into sequences containing nucleotides of fixed or variable lengths (including A, T, C, and G nucleotides). These nucleotide sequences are then fed into the GPT model for learning, with the model's task being to predict the next one or more nucleotides.
[0059] It should be noted that using the *E. coli* genome as the training dataset for the model is merely an illustrative example; of course, genomes of other species or multiple species can also be used. Similarly, choosing the GPT model as the basis for the computational method is only an example; any other AI model can be used for this purpose.
[0060] When the model is learning, it is only shown a portion of the nucleotide information of the first gene nucleotide sequence. Based on this information, the model predicts one or more subsequent nucleotides. For example, when n=0 and ij=1, that is, when the model is predicting the i-th nucleotide, it is only shown the first i-1 nucleotides (where i ranges from [2, N], and N is the length of the nucleotide sequence input to the model). After learning a gene segment, the model will predict the nucleotides from position 2 to N. Alternatively, when n≥1, the model learns the feature information of the first j nucleotides of the first gene nucleotide sequence and predicts the i-th to i+n nucleotides. For example, learning the feature information of the first 5 nucleotides of the first gene nucleotide sequence and predicting the 7th to 9th nucleotides.
[0061] Next, the difference between the model's predictions and the actual results is calculated based on the nucleotide sequence of the input gene and the nucleotide sequence predicted by the model. The model is then optimized based on this difference. At this stage, machine learning algorithms can be used to calculate the difference between the actual and predicted values and optimize the model. During this stage, the model can learn the entire genome of the species being used.
[0062] During the genome learning phase, this initial prediction model continuously learns and optimizes its parameters by predicting nucleotides at specific positions in the nucleotide sequence, enabling it to capture patterns and features within the gene sequence. Once the model training converges, it gains the ability to predict subsequent nucleotide sequences based on the input gene nucleotide sequence, providing a foundation for subsequent gene expression level prediction.
[0063] like Figure 3The graph shows the trend of the loss function during model training in the genome learning phase. "LogStep" in the graph represents the sequence number of the loss record; the figure shows 600 loss records. In some alternative approaches, since the model needs to calculate the loss for each learning iteration, storing and visualizing the results of each calculation would result in very dense data points. Therefore, it is preferable to record the loss only once every n learning iterations (n ≥ 1). Figure 3 It can be seen that the loss value tends to stabilize in the later stage of model training, which confirms that the model training has reached convergence.
[0064] Therefore, through this stage of genome learning, the model can learn the patterns in the nucleotide sequences of genes in the genome, which lays a solid foundation for subsequent gene expression level prediction. After obtaining an initial prediction model with genome learning capabilities, applying it to guide the prediction of the relationship between nucleotide sequences and gene expression levels can more accurately capture the complex relationship between the two, thereby improving the accuracy and efficiency of prediction.
[0065] based on Figure 1 In another optional implementation of the illustrated embodiment, the initial prediction model can be a model pre-trained based on gene nucleotide sequences and their expression levels. Specifically, the method for obtaining the initial prediction model may include... Figure 4 The following steps are shown:
[0066] 401. Obtain the second initial prediction model;
[0067] 402. Obtain multiple sets of second training samples, each set of second training samples including the nucleotide sequence of the second gene and the gene expression level data measured by gene expression of the second gene nucleotide sequence;
[0068] 403. Input multiple sets of the second training samples into the second initial prediction model, so that the second initial prediction model predicts the gene expression level data corresponding to the second gene nucleotide sequence based on the second gene nucleotide sequence in each set of the second training samples, and adjusts the model parameters based on the prediction results and the gene expression level data in the second training samples until convergence, to obtain the initial prediction model.
[0069] The computer device can acquire a second initial prediction model and multiple sets of second training samples. Each set of second training samples includes a second gene nucleotide sequence and gene expression level data measured from the gene expression of the second gene nucleotide sequence. Multiple sets of second training samples can be input into the second initial prediction model, allowing the model to learn the relationship between gene nucleotide sequences and their expression levels. The second initial prediction model then predicts the gene expression level data corresponding to the second gene nucleotide sequence based on the second gene nucleotide sequence in each set of second training samples, and adjusts the model parameters based on the prediction results and the gene expression level data in the second training samples until convergence, thus obtaining the initial prediction model.
[0070] It should be noted that inputting multiple sets of second training samples into the second initial prediction model can be done by inputting these multiple sets of second training samples into the model all at once for training, or by inputting these multiple sets of second training samples into the model in batches for training, so that the model can learn and optimize step by step, thereby improving the training efficiency of the model.
[0071] For example, during the model training phase, where the model learns the relationship between gene nucleotide sequences and expression levels, it is desirable for the model to grasp certain patterns between genes and their expression levels (in this phase, expression level refers to the amount of DNA translated into messenger RNA, which is RNA that can be translated into protein based on its genetic information). Therefore, pre-sequential transcriptomes of a single species or multiple species can be used to allow the model to continue learning. Transcriptome data includes genes in the genome that can be effectively transcribed, and the expression levels of those genes in specific culture environments.
[0072] In this stage, gene sequences with nucleotide lengths less than or equal to a specific length can be selected, or all gene sequences can be chosen as the dataset for model learning. This data is then fed into the second initial prediction model for training and learning. The goal of model training at this stage is for the model to learn the relationship between genes and their expression levels within the genome. Therefore, the model's learning task can be adjusted to predict gene expression levels; that is, for the genes input into the model, the model will generate gene expression levels based on the gene sequence.
[0073] For the gene expression levels generated by the model, algorithms related to machine learning are used to calculate the difference between the model's predictions and the actual results, and an appropriate optimizer is selected to further enhance the model's performance. In this stage, to ensure the model fully captures the relationship between genes and expression levels, it can be allowed to learn multiple times across the entire transcriptome until the difference between the model's predictions and the gene expression levels in the transcriptome is minimized.
[0074] like Figure 5The table shows the trend of the loss function during model training in the genome and expression level learning stages. It includes 1000 loss records for model training. To control the fluctuation of the loss during visualization, the loss can be logarithmically processed, and the logarithmic result is used as the ordinate. Alternatively, one could choose to record the loss once for each n training iterations (n ≥ 1). Figure 5 It can be seen that, except for a few samples that deviate significantly from the equilibrium value, the loss values of other samples during model training are basically maintained within the equilibrium value range, indicating that the change in the loss value tends to be stable, and it can be determined that the model training has reached convergence.
[0075] Therefore, pre-training the initial prediction model to learn the relationship between genes and expression levels allows the model to better understand the intrinsic connection between gene sequences and gene expression levels. In this way, when predicting the relationship between guide nucleotide sequences and gene expression levels, the model can more accurately capture the complex relationship between the two, thereby improving the accuracy and reliability of the predictions. Furthermore, this pre-training method can accelerate the convergence speed of the model in subsequent tasks, improving the overall training efficiency of the model.
[0076] The aforementioned second initial prediction model can be an untrained model or a pre-trained model. Therefore, in some optional embodiments, the second initial prediction model can be obtained by training the model based on the aforementioned genome learning training phase. Specifically, the computer device can acquire a first initial prediction model and multiple sets of first training samples, each set of first training samples including a first gene nucleotide sequence. Multiple sets of first training samples can be input into the first initial prediction model to learn the nucleotide information in the gene nucleotide sequence. The first initial prediction model predicts the i-th to i+n-th nucleotides of the first gene nucleotide sequence based on the first j nucleotides of the first gene nucleotide sequence in each set of first training samples, and adjusts the model parameters of the first initial prediction model according to the prediction results and the actual values of the i-th to i+n-th nucleotides until convergence, thus obtaining the second initial prediction model. Here, j, i, and n are all integers and i > j ≥ 1.
[0077] The model training in this stage is similar to that described above. Figure 2 The model training for the first initial prediction model shown in the figure is similar to the stage of genome learning, and will not be described again here.
[0078] like Figure 6 The figure illustrates the multiple stages of the model's learning from genomes, transcriptomes, and experimental data. As shown, the computer device can acquire the AI model for stage one, genome learning. The training process for this stage can be as follows: Figure 2The method flow is shown below. After completing Phase One training, the model can proceed to Phase Two for transcriptome learning. The training process in this phase can be as follows: Figure 4 The method flow is shown below. After completing Phase Two training, the model can enter Phase Three to learn from the experimental data. The training process in this phase can be as follows: Figure 1 The method flow is shown below. Detailed explanations and descriptions of each stage of the training process can be found in the previous text and will not be repeated here.
[0079] Therefore, combining the above embodiments and their various optional implementation methods, the final target prediction model can be obtained based on training in stages one and three, stages two and three, stages one, two and three, or even only stage three. Among these, training the model in stages one, two, and three yields the best training effect and higher accuracy in predicting gene expression levels. This is because the model not only learns the patterns of gene nucleotide sequences in the genome but also the correspondence between gene nucleotide sequences and their expression levels. Furthermore, it can be fine-tuned using a large amount of experimental data, thereby further improving the model's prediction accuracy.
[0080] In the experimental data learning phase, gene sequences, regulatory element sequences, guide nucleotide sequences, and their corresponding gene expression levels can be used as training samples. These data are input into a model that has already completed genome and transcriptome learning, allowing the model to further learn the relationship between guide nucleotide sequences and gene expression levels. The model predicts the corresponding gene expression levels based on the input guide nucleotide sequences and adjusts its parameters based on the prediction results and the actual gene expression data from the experimental data until the model converges. Through this training phase, the model can more accurately capture the complex relationship between guide nucleotide sequences and gene expression levels, thereby improving the prediction accuracy and reliability in practical applications. The final target prediction model can be combined and optimized based on the training results from the above multiple phases to achieve the best prediction performance.
[0081] During the experimental data learning phase, the model's learning objective focuses more on the fine-grained relationship between guide nucleotide sequences and gene expression levels. Since experimental data typically originates from specific experimental conditions and samples, it often contains richer and more specific biological information. Therefore, incorporating experimental data into model training not only helps the model gain a deeper understanding of the mechanism by which guide nucleotide sequences affect gene expression levels, but also improves the model's predictive ability under specific experimental conditions.
[0082] Based on the above embodiments and their multiple optional implementations, during the training process of the initial prediction model, the model parameters of the initial prediction model are adjusted based on the prediction results output by the model and the real gene expression data until convergence. This can be achieved by constructing a loss function based on the prediction results and the real gene expression data, and adjusting the model parameters of the initial prediction model according to the loss value of the loss function until convergence.
[0083] Based on the above embodiments and their various alternative implementations, the guide nucleotide sequence may include an sgRNA sequence, and the gene sequence may include sequences of genes associated with the sgRNA sequence and sequences of gene regulatory elements. The sgRNA sequence may include the sgRNA itself and a PAM sequence (Protospacer Adjacent Motif) located on the genome following the sgRNA itself.
[0084] For example, during Phase 3 training, experimentally collected data can include sgRNA, genes associated with sgRNA, regulatory element sequences associated with the genes, and gene expression levels observed after gene editing based on the current sgRNA. The sgRNA can be divided into two parts: the sgRNA itself, containing several nucleotides, and a PAM sequence (containing a small number of nucleotides) located after the sgRNA on the genome. The regulatory element sequences can be promoter sequences, enhancer sequences, terminator sequences, etc., that can affect the expression of the currently edited gene. The sgRNA can find corresponding nucleotide sequences with base pairing within or on the complementary strand of these regulatory element sequences.
[0085] Therefore, at this stage, there is a gene regulatory element sequence and a sgRNA sequence of several nucleotides (including the sgRNA sequence body and the PAM sequence). The sgRNA body is responsible for allowing nucleotide cutting enzymes to find the target gene, while the PAM sequence provides the Cas9 protein with the site for nucleic acid cleavage. This site is usually the space between the third and fourth nucleotides upstream of the PAM sequence.
[0086] The PAM sequence is typically located downstream (3' end) of the target DNA sequence, adjacent to the sgRNA recognition region. The Cas9 protein recognizes the PAM sequence; only when the correct PAM is detected will Cas9 be activated and cleave the DNA. This ensures that the Cas protein accurately locates the target DNA and avoids off-target effects.
[0087] Next, regulatory element sequences, sgRNA sequences, gene sequences, etc., can be assembled together to form data suitable for input into the model for learning. Optionally, some N nucleotides (N indicates that this nucleotide does not carry any biological information) can be added to the end of the sequences to make each sequence of uniform length, for example, 512 nucleotides. Then, these sequence data are divided into training sets and test sets (the training set is used for model learning, and the test set is used to test the model's performance) for the model to learn.
[0088] Throughout the learning process, the model receives synthetic nucleotide sequence data containing information from multiple genomic fragments and then predicts gene expression levels. The model is optimized by calculating the difference between the predicted and actual gene expression levels until it achieves its best performance on the test set. Figure 7 The table shown in the figure illustrates the trend of the loss function during the model training phase under the experimental data learning stage. It includes 1000 loss records of the model training. As can be seen from the figure, the loss value continuously decreases in the later stage of model training. When the loss value is less than the preset threshold, it can be determined that the model training has reached convergence.
[0089] For example, R can be used 2 Metrics are used to measure the model's performance on the test set. R 2 R is a commonly used metric in statistics to measure how well a predicted result fits the true value. 2 The closer the value is to 1, the better the model performs. For example... Figure 8 The figure shows the model's performance during the experimental data learning phase, and the R-value of the model on the test set is also shown. 2 The value gradually approaches 1, indicating that the training effect and performance of the model are gradually optimized.
[0090] Based on the above embodiments and their various alternative implementations, the gene sequence, the first gene nucleotide sequence, and the second gene nucleotide sequence can all originate from the genome of the same species, such as from the genome of *E. coli*. Therefore, by training a model based on gene nucleotide sequences from the genome of the same species, the trained model can predict the corresponding gene expression levels for the genome of that species and its guide nucleotide sequence, thus helping to improve the accuracy of predicting genome expression levels for the same species.
[0091] In other embodiments, the gene sequence, the first gene nucleotide sequence, and the second gene nucleotide sequence may also be derived from the genomes of different species. This approach can expand the application scope of the model in predicting expression levels across genomes of different species, thereby enabling the prediction of gene expression levels in different species.
[0092] The following will be discussed in the preceding text. Figure 1Based on the illustrated embodiments and their multiple alternative implementations, embodiments of this application are described in further detail. Please refer to... Figure 9 One embodiment of the guided nucleotide sequence screening method in this application includes:
[0093] 901. Obtain the target prediction model corresponding to the gene sequence of the target species;
[0094] 902. Obtain multiple candidate guide nucleotide sequences determined based on the gene sequence of the target species;
[0095] In this embodiment, the target prediction model can be derived from the aforementioned Figure 1 The gene expression prediction model training method of the illustrated embodiment and its multiple optional embodiments is obtained by training the gene sequence of the target species. Therefore, the trained target prediction model can be used to predict the gene expression levels of the target species and its corresponding multiple candidate guide nucleotide sequences.
[0096] The gene sequence of the target species may include nucleotide sequences related to the gene, such as regulatory element sequences and coding region sequences, and is not limited here. The method for determining multiple candidate guide nucleotide sequences based on the gene sequence of the target species can be as follows: multiple candidate guide nucleotide sequences corresponding to the gene sequence of the target species are designed based on the base complementarity principle, taking into account factors such as target specificity, editing efficiency, and off-target risk. Specifically, bioinformatics tools such as Benchling, CRISPRscan, CRISPRitz, and sgRNA Scorer 2.0 can be used to design the candidate guide nucleotide sequences corresponding to the gene sequence.
[0097] 903. Input the gene sequence, regulatory element sequence, and each candidate guide nucleotide sequence of the target species into the target prediction model so that the model can predict the corresponding gene expression data for each candidate guide nucleotide sequence based on the gene sequence of the target species and the corresponding candidate guide nucleotide sequence.
[0098] Computer equipment can input the gene sequence, regulatory element sequence, and each candidate guide nucleotide sequence of the target species into a target prediction model. Based on the performance obtained through pre-training, the model processes the gene sequence of the target species and each candidate guide nucleotide sequence to predict the gene expression level data corresponding to each candidate guide nucleotide sequence.
[0099] The gene expression data output by the model prediction can be a specific numerical value of gene expression, or it can be a representation of the change in gene expression, such as the rate of change, ratio, or difference between the expression level of the gene sequence of the target species before it was edited and the expression level of the gene sequence after it was edited based on the candidate guide nucleotide sequence.
[0100] 904. Based on the gene expression data corresponding to each of the candidate guide nucleotide sequences predicted by the target prediction model, determine the optimal guide nucleotide sequence among the plurality of candidate guide nucleotide sequences;
[0101] The gene expression data predicted by the model can be used to determine the appropriate guide nucleotide sequence. For example, if the purpose of gene editing is to enhance gene expression, a candidate guide nucleotide sequence with a higher gene expression level can be selected from multiple candidate guide nucleotide sequences as the appropriate guide nucleotide sequence. If the purpose of gene editing is to suppress gene expression, a candidate guide nucleotide sequence with a lower gene expression level can be selected from multiple candidate guide nucleotide sequences as the appropriate guide nucleotide sequence.
[0102] In this embodiment, a prediction model is trained to predict the corresponding gene expression data for nucleotide sequences, and then a suitable guide nucleotide sequence is found based on the gene expression level. This eliminates the need to determine the suitable guide nucleotide sequence through experiments, reducing manpower and time costs. Compared with simple bioinformatics analysis, this method can greatly improve the efficiency of guide nucleotide sequence optimization.
[0103] The following describes in further detail the above-described embodiments and their respective optional implementations, using computing resources with a Linux system as an example.
[0104] First, the computer device needs to have the corresponding drivers installed according to the specific model of the CPU, GPU, and other hardware. To ensure the smooth operation of the AI model, users also need to create a corresponding virtual environment or instantiate the relevant container based on the selected model, and install tools and their dependencies such as Python, Java, C++, or R for model training.
[0105] Next, users need to integrate the dataset to be trained into a file that is easy to input into the model. For example, experimental data can be integrated into a CSV file, arranged column by column in the order of regulatory gene sequences, regulatory element sequences, sgRNA sequences, and experimental test results. During model training, users should select an appropriate data loading method according to the format of the data file to read the data into the model. For encoding the raw data, users can choose suitable methods such as one-hot encoding, BytePair Encoding, or nucleotide encoding. Regarding the statistics of the model prediction results, users can select indicators such as R², Pearson correlation coefficient, Spearman correlation coefficient, or accuracy to evaluate the model performance based on the distribution characteristics of the input data.
[0106] Of course, the above application examples are merely illustrative of one applicable scenario for the embodiments of this application, and the specific configurations or data formats involved do not constitute a limitation on the embodiments of this application. Users can adjust the data in the input file according to actual needs, such as modifying regulatory element sequences, sgRNA sequences, etc., to explore the effects of different sequences on gene expression levels. In addition, users can further analyze the effects of different guide nucleotide sequences on gene editing effects based on the output gene expression data, thereby screening out the optimal guide nucleotide sequence.
[0107] It is worth noting that the gene expression level prediction model training method and guide nucleotide sequence screening method provided in this application are not only applicable to gene editing scenarios based on gene editing systems such as CRISPR / Cas9, but can also be applied to other bioinformatics research that requires prediction of gene expression levels. By utilizing the prediction model provided in this application, users can more efficiently and accurately screen suitable guide nucleotide sequences, providing strong support for gene editing and other bioinformatics research.
[0108] The computer device in the embodiments of this application is described below. Please refer to [link / reference]. Figure 10 One embodiment of the computer device in this application includes:
[0109] The computer device 1000 may include one or more central processing units (CPUs) 1001 and a memory 1005, wherein the memory 1005 stores one or more applications or data.
[0110] The memory 1005 can be volatile or persistent storage. The program stored in the memory 1005 can include one or more modules, each module including a series of instruction operations on the computer device. Furthermore, the central processing unit 1001 can be configured to communicate with the memory 1005 and execute the series of instruction operations in the memory 1005 on the computer device 1000.
[0111] The computer device 1000 may also include one or more power supplies 1002, one or more wired or wireless network interfaces 1003, one or more input / output interfaces 1004, and / or one or more operating systems, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.
[0112] The central processing unit 1001 can perform the aforementioned... Figure 1 or Figure 9 The specific operations performed by the computer device in the illustrated embodiment will not be described in detail here.
[0113] This application also provides a computer storage medium, one embodiment of which includes: the computer storage medium storing instructions, which, when executed on a computer, cause the computer to perform the aforementioned... Figure 1 or Figure 9 The operations performed by the computer device in the illustrated embodiment.
[0114] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0115] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between apparatuses or units through some interfaces, and may be electrical, mechanical, or other forms.
[0116] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0117] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0118] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
Claims
1. A method for training a gene expression level prediction model, characterized in that, The method comprises: obtaining an initial prediction model; obtaining a plurality of groups of training samples, each group of the training samples comprising a nucleotide sequence and real gene expression data of an edited gene corresponding to the nucleotide sequence, the nucleotide sequence comprising a guide nucleotide sequence and a gene sequence associated with the guide nucleotide sequence, the edited gene being obtained by gene editing of the gene sequence based on the guide of the guide nucleotide sequence; inputting the plurality of groups of the training samples into the initial prediction model, so that the initial prediction model predicts gene expression data corresponding to the nucleotide sequence according to the nucleotide sequence in each group of the training samples, and adjusts model parameters based on the prediction result and the real gene expression data until convergence, to obtain a target prediction model; the target prediction model is used to predict gene expression data of an input gene sequence and its corresponding guide nucleotide sequence; the obtaining of the initial prediction model comprises: obtaining a first initial prediction model; obtaining a plurality of groups of first training samples, each group of the first training samples comprising a first gene nucleotide sequence; inputting the plurality of groups of the first training samples into the first initial prediction model, so that the first initial prediction model predicts i-th to i+n-th nucleotides of the first gene nucleotide sequence according to the first j nucleotides of the first gene nucleotide sequence in each group of the first training samples, and adjusts model parameters of the first initial prediction model according to the prediction result and real values of the i-th to i+n-th nucleotides until convergence, to obtain the initial prediction model; wherein j, i and n are all integers and i > j ≥ 1; alternatively, the obtaining of the initial prediction model comprises: obtaining a second initial prediction model; obtaining a plurality of groups of second training samples, each group of the second training samples comprising a second gene nucleotide sequence and gene expression data measured by gene expression of the second gene nucleotide sequence; inputting the plurality of groups of the second training samples into the second initial prediction model, so that the second initial prediction model predicts gene expression data corresponding to the second gene nucleotide sequence according to the second gene nucleotide sequence in each group of the second training samples, and adjusts model parameters based on the prediction result and gene expression data in the second training sample until convergence, to obtain the initial prediction model.
2. The method of claim 1, wherein, the obtaining of the second initial prediction model comprises: obtaining a first initial prediction model; obtaining a plurality of groups of first training samples, each group of the first training samples comprising a first gene nucleotide sequence; inputting the plurality of groups of the first training samples into the first initial prediction model, so that the first initial prediction model predicts i-th to i+n-th nucleotides of the first gene nucleotide sequence according to the first j nucleotides of the first gene nucleotide sequence in each group of the first training samples, and adjusts model parameters according to the prediction result and real values of the i-th to i+n-th nucleotides until convergence, to obtain the second initial prediction model; wherein j, i and n are all integers and i > j ≥ 1.
3. The method of claim 2, wherein, The gene sequence, the first gene nucleotide sequence, and the second gene nucleotide sequence are from a same species of genome.
4. The method according to any one of claims 1 to 3, characterized in that, The adjusting the model parameters based on the prediction result and the real gene expression data until convergence comprises: The adjusting the model parameters based on the prediction result and the real gene expression data until convergence comprises:
5. The method according to any one of claims 1 to 3, characterized in that, The guide nucleotide sequence comprises an sgRNA sequence, and the gene sequence comprises a gene regulatory element sequence associated with the sgRNA sequence. The sgRNA sequence comprises an sgRNA body and a PAM sequence located behind the sgRNA body on a gene.
6. A method of guided nucleotide sequence screening, comprising, The method comprises: obtaining a target prediction model corresponding to a gene sequence of a target species; the target prediction model is trained by the gene expression prediction model training method according to any one of claims 1 to 5 on the gene sequence of the target species; obtaining a plurality of candidate guide nucleotide sequences determined based on the gene sequence of the target species; inputting the gene sequence of the target species and each of the candidate guide nucleotide sequences into the target prediction model, so that the target prediction model predicts gene expression data corresponding to each of the candidate guide nucleotide sequences according to the gene sequence of the target species and the candidate guide nucleotide sequence; determining an optimal guide nucleotide sequence from the plurality of candidate guide nucleotide sequences according to the gene expression data corresponding to each of the candidate guide nucleotide sequences predicted by the target prediction model. 7.A computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the computer device is configured to perform the method according to any one of claims 1-6 when the computer program is executed by the processor. The processor implements the method according to any one of claims 1 to 6 when executing the computer program.
8. A computer storage medium, characterized in that The computer storage medium stores instructions, and the instructions make the computer execute the method according to any one of claims 1 to 6 when executed on the computer. The computer storage medium stores instructions, and the instructions make the computer execute the method according to any one of claims 1 to 6 when executed on the computer.
Citation Information
Patent Citations
Gene expression prediction method and device
CN115019876A
Bovinized crispr / bocas9 gene editing system, method and use
WO2024230247A1