Prediction Method, Device, Electronic Device, Program and Medium for Gene Editing Results
By constructing gene editing data, combining the characteristics of gene methylation and guide RNA, the prediction accuracy of gene editing results is improved, and the problem of single prediction basis in the prior art is solved.
Patent Information
- Application Number
- CN202210467922.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-29
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2042-04-29
AI Technical Summary
In the prior art, the prediction basis for gene editing results is single, the accuracy cannot be guaranteed, and the influencing factors cannot be fully considered.
By obtaining the target gene methylation data, target gene sequence data and guide RNA sequence data of the target genome, the gene editing data is constructed, and the gene editing results prediction model is input for prediction, taking into account the effects of methylation and guide RNA.
It improves the prediction accuracy of gene editing results and can more comprehensively reflect the influence of various factors in the gene editing process.
Smart Images

Figure CN114783518B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure belongs to the technical field of gene data analysis, and particularly relates to a method, apparatus, electronic device, program, and medium for predicting gene editing results. Background Art
[0002] CRISPR-Cas is the third-generation gene editing technology following the introduction of gene editing technologies such as ZFN and TALENs. In just a few years, CRISPR-Cas technology has become popular worldwide and has become one of the most efficient, simplest, lowest-cost, and easiest-to-use gene editing and gene modification technologies, becoming the current mainstream gene editing system.
[0003] However, for the detection of gene editing results, usually only the influence of the used sgRNA (guide RNA) on the editing results can be considered. However, there are other factors that affect gene editing results. Therefore, in the related technologies, the basis for predicting gene editing results is single, and the accuracy of predicting gene editing results cannot be guaranteed. Summary of the Invention
[0004] The present disclosure provides a method, apparatus, electronic device, program, and medium for predicting gene editing results.
[0005] Some embodiments of the present disclosure provide a method for predicting gene editing results, the method comprising:
[0006] Obtaining target gene methylation data, target gene sequence data of a target genome, and guide RNA sequence data corresponding to the target gene sequence data;
[0007] Constructing gene editing data according to the target gene methylation data, the target gene sequence data, and the guide RNA sequence data;
[0008] Inputting the gene editing data into a gene editing result prediction model for prediction.
[0009] Optionally, the gene editing data includes: a spliced gene feature composed of a methylation gene feature and a guide RNA feature; the constructing gene editing data according to the target gene methylation data, target gene sequence data, and guide RNA sequence data includes:
[0010] Performing aggregated feature extraction on the target gene sequence data towards the guide RNA sequence data to obtain a guide RNA feature, and performing aggregated feature extraction on the target gene sequence data towards the target gene methylation data to obtain a methylation gene feature;
[0011] Splicing the methylation gene feature and the guide RNA feature to obtain a spliced gene feature.
[0012] Optionally, the aggregating feature extraction of the target gene sequence data to the guide RNA sequence data to obtain guide RNA features, and the aggregating feature extraction of the target gene sequence data to the target gene methylation data to obtain methylated gene features include:
[0013] Calculating the correlation degree between each element in the target gene sequence data and each value in the guide RNA sequence data;
[0014] Performing weighted summation on the correlation degrees of each element in the target gene sequence data to obtain guide RNA features;
[0015] And, calculating the correlation degree between each element in the target gene sequence data and each value in the target gene methylation data;
[0016] Performing weighted summation on the correlation degrees of each element in the target gene methylation data to obtain methylated gene features.
[0017] Optionally, before performing weighted summation on the correlation degrees of each element in the target gene sequence data to obtain guide RNA features, the method further includes:
[0018] Performing normalization processing on the correlation degrees of each element in the target gene sequence data;
[0019] Before performing weighted summation on the correlation degrees of each element in the target gene methylation data to obtain methylated gene features, the method further includes:
[0020] Performing normalization processing on the correlation degrees of each element in the target gene methylation data.
[0021] Optionally, the splicing of the methylated gene features and the guide RNA features to obtain spliced gene features includes:
[0022] Performing a convolution operation and a pooling operation on the methylated gene features and the guide RNA features to obtain a methylated gene matrix feature and a guide RNA matrix feature with the same dimension;
[0023] Splicing the methylated gene matrix feature and the guide RNA matrix feature to obtain spliced gene features.
[0024] Optionally, the gene editing result prediction model is obtained through the following steps:
[0025] Obtaining sample target gene methylation data, sample target gene sequence data of a sample genome, and sample guide RNA sequence data corresponding to the sample target gene sequence data;
[0026] Performing aggregation feature extraction on the sample target gene sequence data with respect to the sample guide RNA sequence data to obtain sample guide RNA features, and performing aggregation feature extraction on the sample target gene methylation data with respect to the sample guide RNA sequence data to obtain sample methylated gene features;
[0027] Splicing the sample methylated gene features and the sample guide RNA features to obtain sample spliced gene features;
[0028] Training a gene editing result prediction model to be trained using the sample spliced gene features.
[0029] Optionally, the training of the gene editing result prediction model to be trained using the sample spliced gene features includes:
[0030] Inputting the sample spliced gene features into at least two different gene editing result prediction models to be trained respectively;
[0031] When the at least two different trained gene editing result prediction models all meet the corresponding training requirements, confirming that the at least two different gene editing result prediction models are all trained.
[0032] Optionally, the confirming that the at least two different gene editing result prediction models are all trained when the at least two different trained gene editing result prediction models all meet the corresponding training requirements includes:
[0033] Calculating the verification results corresponding to the at least two different trained gene editing result prediction models;
[0034] Combining at least two of the verification results to obtain a comprehensive verification result;
[0035] When the comprehensive verification result meets the training requirements, confirming that the at least two gene editing result prediction models are all trained.
[0036] Optionally, the verification result includes: a loss value; the combining at least two of the verification results to obtain a comprehensive verification result includes:
[0037] Combining the loss values of the at least two gene editing result prediction models to obtain a comprehensive loss value;
[0038] The confirming that the at least two gene editing result prediction models are all trained when the comprehensive verification result meets the training requirements includes:
[0039] When the comprehensive loss value is less than the loss value threshold, it is confirmed that all of the at least two gene editing result prediction models are trained.
[0040] Optionally, the gene editing result prediction model at least includes: an edited gene recognition model and a gene editing probability prediction model; the inputting the sample spliced gene features into at least two different gene editing result prediction models to be trained includes:
[0041] Inputting the sample spliced gene features into the edited gene recognition model to obtain a recognition result indicating whether the sample genome is edited, and inputting the sample spliced gene features into the gene editing probability prediction model to obtain a predicted probability value that the sample genome has been edited;
[0042] When all of the at least two different gene editing result prediction models after training meet the corresponding training requirements, the confirmation that all of the at least two different gene editing result prediction models are trained includes:
[0043] Comparing the sample label of the sample genome with the recognition result and the predicted probability value to respectively obtain a first loss value of the edited gene recognition model and a second loss value of the gene editing probability prediction model;
[0044] When the comprehensive loss value obtained by combining the first loss value and the second loss value is less than the preset loss value, it is confirmed that both the edited gene recognition model and the gene editing probability prediction model are trained.
[0045] Some embodiments of the present disclosure provide a prediction device for gene editing results, the device includes:
[0046] An acquisition module, configured to acquire target gene methylation data, target gene sequence data of a target genome, and guide RNA sequence data corresponding to the target gene sequence data;
[0047] A data processing module, configured to construct gene editing data according to the target gene methylation data, the target gene sequence data, and the guide RNA sequence data;
[0048] A prediction module, configured to input the gene editing data into a gene editing result prediction model for prediction.
[0049] Optionally, the data processing module is further configured to:
[0050] Performing aggregated feature extraction on the target gene sequence data with respect to the guide RNA sequence data to obtain guide RNA features, and performing aggregated feature extraction on the target gene sequence data with respect to the target gene methylation data to obtain methylated gene features;
[0051] Splicing the methylated gene features and the guide RNA features to obtain spliced gene features.
[0052] Optionally, the data processing module is further configured to:
[0053] Calculating the correlation between each element in the target gene sequence data and each value in the guide RNA sequence data;
[0054] Performing weighted summation on the correlations of each element in the target gene sequence data to obtain guide RNA features;
[0055] And calculating the correlation between each element in the target gene sequence data and each value in the target gene methylation data;
[0056] Performing weighted summation on the correlations of each element in the target gene methylation data to obtain methylated gene features.
[0057] Optionally, the data processing module is further configured to:
[0058] Performing normalization processing on the correlations of each element in the target gene sequence data;
[0059] Performing normalization processing on the correlations of each element in the target gene methylation data.
[0060] Optionally, the data processing module is further configured to:
[0061] Performing convolution operation and pooling operation on the methylated gene features and the guide RNA features to obtain methylated gene matrix features and guide RNA matrix features with the same dimension;
[0062] Splicing the methylated gene matrix features and the guide RNA matrix features to obtain spliced gene features.
[0063] Optionally, the apparatus further includes: a training module, configured to:
[0064] Obtaining sample target gene methylation data, sample target gene sequence data of a sample genome, and sample guide RNA sequence data corresponding to the sample target gene sequence data;
[0065] Performing aggregated feature extraction on the sample target gene sequence data with respect to the sample guide RNA sequence data to obtain sample guide RNA features, and performing aggregated feature extraction on the sample target gene methylation data with respect to the sample guide RNA sequence data to obtain sample methylated gene features;
[0066] Splicing the sample methylated gene features and the sample guide RNA features to obtain sample spliced gene features;
[0067] Training the gene editing result prediction model to be trained using the sample spliced gene features.
[0068] Optionally, the training module is further configured to:
[0069] Inputting the sample spliced gene features into at least two different gene editing result prediction models to be trained respectively;
[0070] When the trained at least two different gene editing result prediction models all meet the corresponding training requirements, confirming that the at least two different gene editing result prediction models are all trained.
[0071] Optionally, the training module is further configured to:
[0072] Calculating the verification results corresponding to the trained at least two different gene editing result prediction models;
[0073] Combining at least two of the verification results to obtain a comprehensive verification result;
[0074] When the comprehensive verification result meets the training requirements, confirming that the at least two gene editing result prediction models are all trained.
[0075] Optionally, the verification result includes: a loss value; the training module is further configured to:
[0076] Combining the loss values of the at least two gene editing result prediction models to obtain a comprehensive loss value;
[0077] The step of when the comprehensive verification result meets the training requirements, confirming that the at least two gene editing result prediction models are all trained, includes:
[0078] When the comprehensive loss value is less than the loss value threshold, confirming that the at least two gene editing result prediction models are all trained.
[0079] Optionally, the training module is further configured to:
[0080] Input the spliced gene features of the sample into the edited gene recognition model to obtain a recognition result indicating whether the genome of the sample is edited, and input the spliced gene features of the sample into the gene editing probability prediction model to obtain a predicted probability value that the genome of the sample has been edited;
[0081] Compare the sample label of the sample genome with the recognition result and the predicted probability value to respectively obtain a first loss value of the edited gene recognition model and a second loss value of the gene editing probability prediction model;
[0082] When the combined loss value obtained by combining the first loss value and the second loss value is less than the preset loss value, confirm that both the edited gene recognition model and the gene editing probability prediction model are trained.
[0083] Some embodiments of the present disclosure provide a computing device, including:
[0084] A memory storing computer-readable code;
[0085] One or more processors, when the computer-readable code is executed by the one or more processors, the computing device executes the method for predicting gene editing results as described above.
[0086] Some embodiments of the present disclosure provide a computer program, including computer-readable code, when the computer-readable code runs on a computing device, causing the computing device to execute the method for predicting gene editing results as described above.
[0087] Some embodiments of the present disclosure provide a non-transitory computer-readable medium storing the method for predicting gene editing results as described above.
[0088] A method, device, electronic device, program, and medium for predicting gene editing results provided by the present disclosure construct gene editing data for predicting gene editing results by using target gene sequence data, guide RNA, and target gene methylation data, so that the gene editing result prediction model can comprehensively consider the effects of gene methylation and guide RNA on gene editing during the prediction process, improving the accuracy of gene editing prediction.
[0089] The above description is only an overview of the technical solution of the present disclosure. In order to be able to understand the technical means of the present disclosure more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features, and advantages of the present disclosure more obvious and understandable, the specific embodiments of the present disclosure are specifically exemplified below. BRIEF DESCRIPTION OF THE DRAWINGS
[0090] To more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present disclosure. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0091] Figure 1 Schematically shows a flowchart of a method for predicting gene editing results provided by some embodiments of the present disclosure;
[0092] Figure 2 Schematically shows one of the flowcharts of another method for predicting gene editing results provided by some embodiments of the present disclosure;
[0093] Figure 3 Schematically shows another flowchart of a method for predicting gene editing results provided by some embodiments of the present disclosure;
[0094] Figure 4 Schematically shows a schematic diagram of the principle of a method for predicting gene editing results provided by some embodiments of the present disclosure;
[0095] Figure 5 Schematically shows a third flowchart of another method for predicting gene editing results provided by some embodiments of the present disclosure;
[0096] Figure 6 Schematically shows a flowchart of a method for training a gene editing result prediction model provided by some embodiments of the present disclosure;
[0097] Figure 7 Schematically shows one of the flowcharts of another method for training a gene editing result prediction model provided by some embodiments of the present disclosure;
[0098] Figure 8 Schematically shows a second flowchart of another method for training a gene editing result prediction model provided by some embodiments of the present disclosure;
[0099] Figure 9 Schematically shows a third flowchart of another method for training a gene editing result prediction model provided by some embodiments of the present disclosure;
[0100] Figure 10 Schematically shows a fourth flowchart of another method for training a gene editing result prediction model provided by some embodiments of the present disclosure;
[0101] Figure 11A logical schematic diagram of a method for training another gene editing result prediction model provided by some embodiments of the present disclosure is schematically shown;
[0102] Figure 12 A structural schematic diagram of a prediction device for gene editing results provided by some embodiments of the present disclosure is schematically shown;
[0103] Figure 13 A block diagram of a computing processing device for performing the method according to some embodiments of the present disclosure is schematically shown;
[0104] Figure 14 A storage unit for holding or carrying program code for implementing the method according to some embodiments of the present disclosure is schematically shown. Detailed implementation manners
[0105] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Apparently, the described embodiments are some, rather than all, of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present disclosure without creative efforts shall fall within the protection scope of the present disclosure.
[0106] It should be noted that the CRISPR / Cas system is an immune system of prokaryotes used to resist the invasion of foreign genetic material and provide acquired immunity for bacteria. When bacteria are invaded by viruses or foreign plasmids, they will generate corresponding "memories" so as to resist subsequent invasions. The CRISPR / Cas system can recognize foreign DNA and cut them off to silence the expression of foreign genes. This is similar to the principle of RNA interference (RNAi) in eukaryotes. Due to this precise targeting function, the CRISPR / Cas system has been developed into an efficient gene editing tool. In nature, there are multiple categories of the CRISPR / Cas system, among which the CRISPR / Cas9 system is the most deeply studied and most maturely applied category. CRISPR / Cas9 is the third generation of "genome site-directed editing technology" after "zinc finger endonuclease (ZFN)" and "transcription activator-like effector nuclease (TALEN)". With the advantages of low cost, convenient operation, high efficiency, etc., CRISPR / Cas9 has quickly become popular in laboratories around the world and has become a powerful helper for biological research. In the era of TALEN and ZFN, scientists often had to spend a large amount of money and hand over gene editing work to biological companies. Now, in the laboratory, people can easily achieve gene editing using CRISPR / Cas9 technology.
[0107] CRISPR stands for Clustered Regularly Interspersed Short Palindromic Repeats. The CRISPR sequence consists of numerous short and conserved repeat regions (repeat) and spacer regions (spacer). The repeat region contains palindromic sequences to form hairpin structures. The spacer region is the exogenous DNA sequence captured by bacteria and is the "blacklist" of the bacterial immune system. When these exogenous genetic materials invade again, the CRISPR / Cas system will give an accurate strike. The upstream leader region is considered the promoter of the CRISPR sequence. There is also a polymorphic family gene upstream, and the proteins encoded by this gene can all act together with the CRISPR sequence region and are named CRISPR associated genes (Cas). The Cas genes and the CRISPR sequence have co-evolved to form the highly conserved CRISPR / Cas system in bacteria.
[0108] When a virus invades, the proteins encoded by Cas1 and Cas2 will scan the exogenous DNA, and the DNA sequence of M is used as a candidate protospacer sequence. The Cas1 / 2 protein complex cuts the protospacer sequence from the exogenous DNA and, with the assistance of other enzymes, inserts the protospacer sequence downstream of the leader region adjacent to the CRISPR sequence. Then, the DNA is repaired to close the opened double-strand break, and a new spacer sequence is added to the CRISPR sequence in the genome.
[0109] Currently, there are three ways for the CRISPR / Cas system to synthesize crRNA, namely type I, type II, and type III. The CRISPR / Cas9 system belongs to type II and is the most mature and widely used type currently. The CRISPR sequence transcribes pre-CRISPR-derived RNA (pre-crRNA) and trans-acting crRNA (tracrRNA) under the regulation of the leader region. Among them, tracrRNA is an RNA with a hairpin structure transcribed from the repeat region, and pre-crRNA is a large RNA molecule transcribed from the entire CRISPR sequence. The pre-crRNA, tracrRNA, and the protein encoded by Cas9 will be assembled. According to the type of the invader, the corresponding spacer sequence RNA is selected and cut with the assistance of RNaseⅢ, and finally a short crRNA (containing a single type of spacer sequence RNA and part of the repeat region) is formed. The complex composed of crRNA, Cas9, and tracrRNA is the tool for the next step of cutting.
[0110] The Cas9 / tracrRNA / crRNA complex can precisely target the DNA of invaders. The complex scans the entire exogenous DNA sequence, locates the region of the PAM / protospacer, unwinds the DNA double strand, and forms an R-Loop. The crRNA hybridizes with the complementary strand, while the other strand remains single-stranded. Subsequently, the HNH nuclease activity of the Cas9 protein cleaves the DNA strand complementary to the crRNA, and its RuvC active site cleaves the non-complementary strand. Cas9 induces the formation of a double-strand break (DSB), silences the expression of exogenous DNA, and eliminates the invaders.
[0111] The application of CRISPR / Cas9 requires the presence of a relatively conserved PAM sequence (NGG) near the region to be edited, and the ability of the gRNA to complementarily pair with the base sequence upstream of the PAM. The CRISPR / Cas technology has been widely applied. Under the combined action of the guide RNA (gRNA) and the Cas9 protein, the genomic DNA of the cell to be edited is regarded as viral or exogenous DNA and is precisely cleaved. In addition to basic editing methods such as gene knockout and gene replacement, it can also be used for gene activation, disease model construction, and even gene therapy.
[0112] DNA methylation is a form of DNA chemical modification that can alter genetic expression without changing the DNA sequence. A large number of studies have shown that DNA methylation can cause changes in chromatin structure, DNA conformation, DNA stability, and the way DNA interacts with proteins, thereby controlling gene expression. DNA methylation generally occurs at CpG (a dinucleotide chain composed of cytosine (C) and guanine (G), where p refers to the phosphate between C and G) sites, while non-CpG methylation is more common in embryonic stem cells.
[0113] DNA methylation sequencing can be classified into three major categories according to its principle: bisulfite sequencing, restriction enzyme-based sequencing, and targeted enrichment of methylated sites sequencing. Among them, bisulfite treatment + sequencing was once considered the gold standard for DNA methylation analysis. The process is as follows: After bisulfite treatment, the target fragment is amplified by PCR, and the PCR product is sequenced. The sequence is compared with the untreated sequence to determine whether methylation has occurred at the CPG site. This method is reliable and highly accurate, and can clarify the methylation status of each CpG site in the target fragment.
[0114] When the CRISPR / Cas9 system performs gene editing, tracrRNA (trans-activating crRNA), pre-crRNA, and Cas9 protein are first transcribed and expressed. Then, tracrRNA activates ribonuclease III to modify pre-crRNA to form mature crRNA. Subsequently, crRNA, tracrRNA, and Cas9 protein form a complex. Through the recognition of crRNA, the complex is targeted to the target DNA. Meanwhile, tracrRNA activates Cas9 protein, and Cas9 acts as a nuclease to cleave the target DNA, causing double-strand breaks (DSBs). After the generation of DSBs, the cell can repair the DNA double-strand through different repair methods. In this system, crRNA and tracrRNA can be replaced by a single-stranded guide RNA (sgRNA), simplifying the recognition component RNA. Generally, the sgRNA is artificially designed, and the success of editing is related to the length and site selection of the sgRNA.
[0115] The CRISPR / Cas12a system belongs to type V system. Cas12a is a single-subunit protein. Compared with Cas9 protein, this protein has the following characteristics: ① It does not require tracrRNA to recognize and cleave DNA; ② The PAM sequence is rich in thymine and is usually TTTN; ③ The PAM sequence is located at the 5′ end of the recognized DNA; ④ The recognition point and the cleavage point are far from the PAM; ⑤ Sticky ends are generated after cleavage; ⑥ The protein molecular weight is smaller; ⑦ The required crRNA sequence is shorter. In addition to Cas9 and Cpf1 proteins, the Cas13 series of proteins have also been gradually taken seriously. A nuclease C2c2 (Cas13a) that can counter RNA viruses can bind to RNA targets and cleave RNA to achieve bacterial self-defense.
[0116] It has been found through research that the sites where the genome has been successfully edited can usually be methylated. Therefore, in the process of predicting the gene editing of the target gene in the genome, this disclosure incorporates the methylation data of the target gene in the genome to propose a method for predicting gene editing results to improve the accuracy of predicting gene editing results.
[0117] Figure 1 A schematic flow diagram of a method for predicting gene editing results provided by this disclosure is schematically shown. The method includes:
[0118] Step 101, obtaining the methylation data of the target gene, the target gene sequence data of the target genome, and the guide RNA sequence data corresponding to the target gene sequence data.
[0119] In the embodiments of the present disclosure, the target genome can be the genome of organisms with DNA and / or RNA structures such as humans, livestock, bacteria, viruses, etc. The target genome can be a genome that has been gene-edited or a genome that has not been gene-edited. It can be understood that although the gene editing process is for editing gene fragments in the genome, whether the gene editing is successful is uncertain, but this does not affect the feasibility of the gene editing result prediction method provided by some embodiments of the present disclosure. The target gene methylation data is data obtained by performing gene sequencing after methylation conversion of the target genome.
[0120] Currently, DNA methylation detection techniques based on sequencers can be divided into several methods according to different library construction methods, such as direct bisulfite sequencing, MeDIP sequencing, MBD sequencing, digestion-bisulfite sequencing, etc. For example:
[0121] The main steps of the bisulfite direct sequencing method include DNA fragmentation, end repair of DNA fragments, ligation of methylation sequencing adapters, bisulfite conversion, PCR amplification, sequencing, and sequence comparison. Specifically, after the fragmented DNA is end-modified and adenine is added to the 3' end, it is directly ligated to a methylated sequencing adapter (all sites on the adapter are modified to the methylated state). Under suitable reaction conditions, for single-stranded DNA molecules, bisulfite is used to remove the amino group of unmethylated cytosine and convert it into uracil, while methylated cytosine remains unchanged, that is, bisulfite conversion is performed. Then PCR amplification is carried out to convert all uracils into thymines. Finally, the PCR product is sequenced and compared with the untreated sequence to determine whether methylation occurs at the CpG site.
[0122] The MeDIP sequencing and MBD sequencing methods are based on the fact that in mammals, methylation generally occurs at the 5th carbon atom of cytosine in CpG. Therefore, proteins MBD or 5'-methylcytosine antibody MeDIP that specifically bind to methylated DNA can be used to enrich highly methylated DNA fragments. Combined with the second-generation high-throughput sequencing, the enriched DNA fragments are sequenced. Specifically, the method for separating methylated DNA fragments by the MDB method is called methylated CpG immunoprecipitation (MCIp). MeDIP can be used to immunoprecipitate and highly specifically enrich methylated DNA fragments through the 5-methylcytosine antibody. The 5-methylcytosine antibody can also bind to single methylated cytosines at non-CpG sites, so it has higher specificity than MBD. This technology is called methylated DNA immunoprecipitation. Combining with the new generation sequencing technology can screen abnormally methylated genes in a high-throughput manner, and this method avoids the limitations of the application of restriction enzymes at the enzyme cutting sites.
[0123] The enzymatic digestion-bisulfite sequencing method is based on bisulfite sequencing using an enzymatic digestion method. The purpose is to enrich the DNA fragments to be tested, reduce the size of the sequencing DNA library, and reduce the sequencing cost. This method can successfully enrich some CpG islands (8% of the measured data aligns to different CpG islands). Moreover, this method reduces the size of the sequencing DNA library to a certain extent, and after bisulfite conversion, there is no need to perform subsequent methylation site identification work.
[0124] Of course, the above gene methylation methods are only exemplary descriptions. The specific gene methylation sequencing method can be determined according to actual needs and is not limited here. The present disclosure can pre-measure the gene methylation data of the target gene fragment in the genome through experiments for the server to obtain and use when predicting the gene editing results.
[0125] Common methods for detecting methylation at specific sites include: 1. Methylation-specific PCR (MS-PCR): After bisulfite treatment, MS-PCR can be carried out. In the traditional MSP method, usually two pairs of primers are designed. One pair of MSP primers amplifies the DNA template after bisulfite treatment, while the other pair amplifies the unmethylated fragment. If the first pair of primers can amplify a fragment, it indicates that methylation exists at the detection site. If the second pair of primers can amplify a fragment, it indicates that methylation does not exist at the detection site; 2. Bisulfite treatment + sequencing: After bisulfite treatment, the target fragment is amplified by PCR, and the PCR product is sequenced. The sequence is compared with the untreated sequence to determine whether methylation occurs at the CpG site. There are also combined bisulfite restriction analysis (COBRA), fluorescence quantitative method (Methylight), methylation-sensitive high-resolution melting curve analysis, pyrosequencing. For details, reference can be made to the gene methylation sequencing technology in the relevant art and will not be elaborated here.
[0126] The target gene sequence data (target sequence) refers to the sequence data of the specified gene fragment in the target genome, usually the sequence data of the gene fragment that needs to be gene-edited. The sequence data of the guide RNA (sgRNA) is the sequence data of the guide RNA used for gene editing of the target gene fragment. For details, reference can be made to the above detailed introduction of sgRNA and will not be elaborated here. It should be noted that for different target gene fragments to be edited, the sequence data of the guide RNA may also be different. Even for the same target gene fragment, there can be multiple choices for the sequence data of the guide RNA, which can be pre-set according to actual needs and is not limited here.
[0127] The execution subject of the present disclosure can be an electronic device with functions such as data processing, data storage, and data transmission, can be a server that provides data support for a terminal, or can be a terminal with data processing and data display functions. In the following description, the server will be exemplarily described as the execution subject, but this does not mean that the execution subject of the present disclosure can only be the server, and the execution subject can also be replaced according to actual needs, which can be specifically set according to actual needs and will not be limited here.
[0128] In an embodiment of the present disclosure, before predicting the gene editing result of a target genome, the server needs to obtain the target gene sequence data of the target gene fragment that may be gene-edited in the target genome, and the guide RNA sequence data for gene-editing the target gene fragment. In particular, considering that gene fragments that can be gene-sequenced can usually also be methylated, the server in the present disclosure will also obtain the target gene methylation data of the target genome for subsequent model prediction.
[0129] Step 102: Construct gene editing data according to the target gene methylation data, the target gene sequence data, and the guide RNA sequence data.
[0130] In an embodiment of the present disclosure, the server constructs gene editing data by extracting the feature vectors input into the gene editing result prediction model from the target gene methylation data, the target gene sequence data, and the guide RNA sequence data, so that the methylation data features of the target gene fragment can be incorporated into the data participating in the prediction of the gene editing result.
[0131] Step 103: Input the gene editing data into the gene editing result prediction model for prediction.
[0132] In an embodiment of the present disclosure, the gene editing result prediction model is a machine learning model or a mathematical model for predicting the gene editing result of a target gene fragment. This gene editing result prediction model is also trained based on the sample splicing gene features combined with the methylated gene features. Therefore, this model not only learns the influence of the guide RNA on the gene editing process, but also learns the influence of gene methylation on the gene editing process, so as to more accurately predict the gene editing result of the genome by integrating the features of gene methylation and the guide RNA.
[0133] In practical applications, the server inputs the spliced gene features into the gene editing result prediction model for prediction to obtain the gene editing result of the target genome, and then the gene editing result can be output through the client or the display function it has. For example, the gene editing result of the target genome is displayed on the screen, or a gene editing report for the target genome is output after further analysis by the analysis system, etc. It can be specifically set according to actual needs and is not limited here.
[0134] By using the target gene sequence data, guide RNA, and target gene methylation data to construct gene editing data for predicting gene editing results, the gene editing result prediction model can comprehensively consider the impacts of gene methylation and guide RNA on gene editing during the prediction process, improving the accuracy of gene editing prediction.
[0135] Optionally, the gene editing data includes: a spliced gene feature composed of a methylated gene feature and a guide RNA feature. Refer to Figure 2 , the step 102 includes:
[0136] Step 1021, perform aggregated feature extraction on the target gene sequence data towards the guide RNA sequence data to obtain the guide RNA feature, and perform aggregated feature extraction on the target gene sequence data towards the target gene methylation data to obtain the methylated gene feature.
[0137] In the embodiments of the present disclosure, the aggregated feature extraction uses the Attention mechanism to selectively screen out a small amount of important information from a large amount of information and focus on these important information, ignoring most of the unimportant information. Its main focusing process is reflected in the calculation of the weight coefficient. The larger the weight, the more focused on its corresponding eigenvalue, that is, the weight represents the importance of the information, and the corresponding feature is the knowledge it needs to focus on learning. Applying this Attention model to the present disclosure is to regard the guide RNA sequence data and the target gene methylation data as important information respectively, and then aggregate the target-based sequence data to each element in the guide RNA sequence data and the target gene methylation data according to different weight coefficients, so as to obtain the guide RNA feature and the methylated gene feature.
[0138] Step 1022, splice the methylated gene feature and the guide RNA feature to obtain the spliced gene feature.
[0139] In the embodiments of the present disclosure, the server splices the methylation gene features and the guide RNA features, so that the data input into the gene editing result prediction model can not only characterize the guide RNA features, but also take into account the methylation features of the genome. The splicing method can be selected according to the vector dimensions of the methylation gene features and the guide RNA features. For example, after directly splicing two features with different dimensions, the missing values are filled with specific values such as 0 or infinitesimal, or directly splicing two different features with the same dimension into a matrix. The specific splicing method can be set according to actual needs and is not limited here.
[0140] In the embodiments of the present disclosure, the guide RNA features are extracted from the target gene sequence data and the guide RNA sequencing sequence respectively by using the aggregation feature extraction algorithm, and the methylation gene features are extracted from the target gene sequence data and the target gene methylation data. Then, the spliced gene features obtained by splicing the two gene features are used for prediction by the gene editing result prediction model, so that the gene editing result prediction model can comprehensively consider the influence of gene methylation and guide RNA on gene editing during the prediction process, improving the accuracy of gene editing prediction.
[0141] Optionally, referring to Figure 3 , step 1021 includes:
[0142] Step 10211, calculate the correlation between each element in the target gene sequence data and each value in the guide RNA sequence data.
[0143] In the embodiments of the present disclosure, referring to Figure 4 , the constituent elements of the guide RNA sequence data in Source (data source) are constructed into a <Key, Value> data pair structure. Given each element Query in Target (target gene sequence data), the similarity or correlation between Query and each Key is calculated.
[0144] The specific method for calculating similarity or correlation can adopt different functions and calculation mechanisms. For example:
[0145] The dot product calculation function shown in formula (1):
[0146] Similarity(Query, Key i ) = QueryKey i (1)
[0147] The Cosine similarity calculation function shown in formula (2):
[0148]
[0149] The MLP network calculation function as shown in formula (3):
[0150] similarity(Query, Key i ) = MLP(Query, Key i ) (3)
[0151] Of course, the above calculation function is only an exemplary description. The specific calculation method of similarity or correlation can be set according to actual needs and is not limited here.
[0152] Step 10212: Normalize the relevance of each element in the target gene sequence data.
[0153] In the embodiments of the present disclosure, the obtained result can be normalized with reference to the following formula (4):
[0154]
[0155] Where Lx represents the total number of elements, and Sim i represents the value of the element.
[0156] Step 10213: Perform weighted summation on the relevance of each element in the target gene sequence data to obtain the guide RNA feature.
[0157] In the embodiments of the present disclosure, the weight coefficient of each Key corresponding to Value is obtained, and then the Value is weighted and summed to obtain the final attention value. Specifically, it can be combined through different correlations by the following formula (5):
[0158]
[0159] Step 10214: Calculate the relevance between each element in the target gene sequence data and each value in the target gene methylation data.
[0160] This step is similar to the description in step 10211, and only the guide RNA sequence data needs to be replaced with the target gene methylation data, which will not be elaborated here.
[0161] Step 10215: Normalize the relevance of each element in the target gene methylation data.
[0162] This step is similar to the description in step 10212, and only the guide RNA sequence data needs to be replaced with the target gene methylation data, which will not be elaborated here.
[0163] Step 10216: Weighted sum the relevance of each element in the target gene methylation data to obtain the methylation gene feature.
[0164] This step is similar to the description of Step 10213. Only the guide RNA sequence data needs to be replaced with the target gene methylation data, which will not be elaborated here.
[0165] Optionally, referring to Figure 5 , Step 1022 includes:
[0166] Step 10221: Perform convolution operation and pooling operation on the methylation gene feature and the guide RNA feature to obtain methylation gene matrix features and guide RNA matrix features with the same dimension.
[0167] In the embodiments of the present disclosure, one-hot encoding is performed on sgRNA (guide RNA sequence data) and target sequence (target gene sequence data). For example, if the length of sgRNA is m and the length of target sequence is n, after one-hot encoding, sgRNA is represented as a vector of m * 4, and target sequence is represented as a guide RNA feature of n * 4 after one-hot encoding.
[0168] Perform Huffman encoding on the position information of the target gene methylation data of the target sequence. Using, for example, 5-bit encoded position information, a sequence of n * 5 and methylation data of n * 1 is obtained, where the AGT site is represented by infinitesimal ∞ (only C has a methylation state). The position encoding and the methylation data are spliced to obtain methylation data with position information, which is a methylation gene feature of n * 6.
[0169] Use the attention model to extract features from sgRNA and target sequence to obtain a guide RNA feature as a vector of m * n. Use attention to extract features from the target sequence and the methylation gene feature with position information. During the operation, the encoding matrix of the target sequence is filled with infinitesimal ∞ to make the representation matrix size of the target sequence, for example, a guide RNA feature of n * 6. After attention feature extraction, a methylation gene feature of n * n is obtained.
[0170] Step 10222: Splice the methylation gene matrix feature and the guide RNA matrix feature to obtain a spliced gene feature.
[0171] In the embodiments of the present disclosure, convolution and pooling operations are respectively performed on the methylation gene matrix features and the guide RNA matrix features to obtain the (c1, p, q) guide RNA matrix features and the (c2, p, q) methylation gene matrix features, and then the two matrices are concatenated to obtain the concatenated gene features of (c1 + c2, p, q).
[0172] Referring to Figure 6 , the schematic diagram shows the flowchart of a method for training a gene editing result prediction model provided by the present disclosure. The method includes:
[0173] Step 201, obtaining the sample target gene methylation data, the sample target gene sequence data of the sample genome, and the sample guide RNA sequence data corresponding to the sample target gene sequence data.
[0174] Step 202, performing aggregated feature extraction on the sample target gene sequence data to the sample guide RNA sequence data to obtain sample guide RNA features, and performing aggregated feature extraction on the sample target gene methylation data to the sample guide RNA sequence data to obtain sample methylated gene features.
[0175] Step 203, concatenating the sample methylated gene features and the sample guide RNA features to obtain sample concatenated gene features.
[0176] Step 204, using the sample concatenated gene features to train the gene editing result prediction model to be trained.
[0177] In the embodiments of the present disclosure, the feature extraction and data concatenation processes during training are similar to those in the above-mentioned gene editing result prediction method and will not be elaborated here. The difference is that the sample genome is labeled with tag information for describing the standard prediction results for the gene editing result prediction model to verify during training.
[0178] In the embodiments of the present disclosure, by using the aggregated feature extraction algorithm, guide RNA features are respectively extracted from the target gene sequence data and the guide RNA sequencing sequence, and methylated gene features are extracted from the target gene sequence data and the target gene methylation data. Then, the concatenated gene features obtained by concatenating the two gene features are used to train the gene editing result prediction model, enabling the gene editing result prediction model to learn the comprehensive influence of gene methylation and guide RNA on gene editing and improving the accuracy of gene editing prediction.
[0179] Optionally, referring to Figure 7 , step 204 includes:
[0180] Step 2041: Input the sample spliced gene features into at least two different gene editing result prediction models to be trained respectively.
[0181] In the embodiments of the present disclosure, the sample spliced gene features can be input into different gene editing result prediction models for training simultaneously, so as to achieve efficient processing of multiple different training tasks.
[0182] Step 2042: When the at least two different gene editing result prediction models after training all meet the corresponding training requirements, confirm that the at least two different gene editing result prediction models have all completed training.
[0183] In the embodiments of the present disclosure, if there are multiple different training tasks simultaneously, in order to ensure the training effect of each model, during the model training process, the training results of multiple different gene editing result prediction models can be comprehensively considered to determine whether to confirm that each gene editing result prediction model has completed training. Thus, not only the training efficiency of multiple different gene editing result prediction models is improved by performing multiple model training tasks simultaneously, but also the model performance of multiple different gene editing result prediction models is provided through the way of collaborative model verification.
[0184] Optionally, referring to Figure 8 , step 2042 includes:
[0185] Step 20421: Calculate the verification results corresponding to the at least two different gene editing result prediction models after training.
[0186] In the embodiments of the present disclosure, the verification results of different gene editing result prediction models can be verification metrics such as loss values, similarities, etc. Of course, the verification functions of different gene editing result prediction models can be the same or different, which can be specifically set according to actual needs and are not limited here.
[0187] Step 20422: Combine at least two of the verification results to obtain a comprehensive verification result.
[0188] In the embodiments of the present disclosure, when the calculation methods of the verification results are the same, the comprehensive verification result can be obtained by weighted summation of multiple verification results. If they are different, the different verification results can also be normalized by setting corresponding normalization functions and then combined to obtain the comprehensive verification result, which can be specifically set according to actual needs and is not limited here.
[0189] Step 20423: When the comprehensive verification result meets the training requirements, confirm that the at least two gene editing result prediction models have all completed training.
[0190] In an embodiment of the present disclosure, when the comprehensive verification result meets the characteristic numerical range or is less than or greater than the characteristic threshold, it can be confirmed that multiple different gene editing result prediction models have all been trained, so that the multiple different gene editing result prediction models can cooperate with each other for training, providing the training efficiency of the gene editing result prediction model.
[0191] Optionally, referring to Figure 9 , the step 20422 may include:
[0192] Step 20422A, combining the loss values of the at least two gene editing result prediction models to obtain a comprehensive loss value.
[0193] Optionally, referring to Figure 9 , the step 20423 may include:
[0194] Step 20423A, when the comprehensive loss value is less than the loss value threshold, confirm that the at least two gene editing result prediction models have all been trained.
[0195] In an embodiment of the present disclosure, when the verification result is a loss value, different gene editing results can select the same or different loss value calculation functions to calculate their corresponding loss values, and then the loss values are weighted and summed to obtain a comprehensive loss value. Considering that generally the smaller the loss value, the better the performance of the model, so it can be determined whether multiple gene editing result prediction models have all been trained by judging whether the comprehensive loss value is less than the loss value threshold, thereby improving the training efficiency of the gene editing result prediction model.
[0196] Optionally, the gene editing result prediction model at least includes: an edited gene recognition model and a gene editing probability prediction model. Referring to Figure 10 , the step 2041 includes:
[0197] Step S1, inputting the sample spliced gene feature into the edited gene recognition model to obtain a recognition result indicating whether the sample genome has been edited, and inputting the sample spliced gene feature into the gene editing probability prediction model to obtain a predicted probability value that the sample genome has been edited.
[0198] Referring to Figure 10 , the step 2042 includes:
[0199] Step S2, comparing the sample label of the sample genome with the recognition result and the predicted probability value to respectively obtain a first loss value of the edited gene recognition model and a second loss value of the gene editing probability prediction model.
[0200] Step S3: When the combined loss value obtained by combining the first loss value and the second loss value is less than the preset loss value, it is confirmed that both the edited gene recognition model and the gene editing probability prediction model are trained.
[0201] In the disclosed embodiment, with reference to Figure 11 , convolution, pooling, and softmax operations are respectively performed on the spliced gene features obtained after splicing, and an edited gene recognition model for predicting the training task of "whether to edit" and a gene editing probability prediction model for the training task of "probability of the site being edited" are set. The two training tasks of "whether to edit" and "probability of the site being edited" can both use the cross-entropy loss function as shown in the following formula (6) to calculate the loss:
[0202]
[0203] where S is the type of output. In the loss function L1 of "whether to edit", S = 2 (not edited, edited), and in the loss function L2 of "probability of the site being edited", S = n + 1, where n is the length of the target gene sequence data, and 1 represents the unedited state.
[0204] Then, the cross-entropy loss values of the two models are summed, that is, L = L1 + L2 is calculated to obtain the combined loss value.
[0205] In the disclosed embodiment of the present disclosure, the models for the two different training tasks of whether to edit and the edited probability are trained by using the loss value verification method, which improves the training efficiency of the models for the two different training tasks of whether to edit and the edited probability.
[0206] Figure 12 The structural schematic diagram of a prediction device 30 for gene editing results provided by the present disclosure is schematically shown. The device includes:
[0207] An acquisition module 301, configured to acquire the target gene methylation data of the target genome, the target gene sequence data, and the guide RNA sequence data corresponding to the target gene sequence data;
[0208] A data processing module 302, configured to construct gene editing data according to the target gene methylation data, the target gene sequence data, and the guide RNA sequence data;
[0209] A prediction module 303, configured to input the gene editing data into a gene editing result prediction model for prediction.
[0210] Optionally, the data processing module 302 is further configured to:
[0211] Performing aggregation feature extraction on the target gene sequence data towards the guide RNA sequence data to obtain guide RNA features, and performing aggregation feature extraction on the target gene sequence data towards the target gene methylation data to obtain methylated gene features;
[0212] Splicing the methylated gene features and the guide RNA features to obtain spliced gene features.
[0213] Optionally, the data processing module 302 is further configured to:
[0214] Calculating the correlation degree between each element in the target gene sequence data and each value in the guide RNA sequence data;
[0215] Performing weighted summation on the correlation degrees of each element in the target gene sequence data to obtain guide RNA features;
[0216] And calculating the correlation degree between each element in the target gene sequence data and each value in the target gene methylation data;
[0217] Performing weighted summation on the correlation degrees of each element in the target gene methylation data to obtain methylated gene features.
[0218] Optionally, the data processing module 302 is further configured to:
[0219] Performing normalization processing on the correlation degree of each element in the target gene sequence data;
[0220] Performing normalization processing on the correlation degree of each element in the target gene methylation data.
[0221] Optionally, the data processing module 302 is further configured to:
[0222] Performing a convolution operation and a pooling operation on the methylated gene features and the guide RNA features to obtain methylated gene matrix features and guide RNA matrix features with the same dimension;
[0223] Splicing the methylated gene matrix features and the guide RNA matrix features to obtain spliced gene features.
[0224] Optionally, the device further includes: a training module, configured to:
[0225] Obtaining sample target gene methylation data, sample target gene sequence data of a sample genome, and sample guide RNA sequence data corresponding to the sample target gene sequence data;
[0226] Performing aggregation feature extraction on the sample target gene sequence data with respect to the sample guide RNA sequence data to obtain sample guide RNA features, and performing aggregation feature extraction on the sample target gene methylation data with respect to the sample guide RNA sequence data to obtain sample methylated gene features;
[0227] Splicing the sample methylated gene features and the sample guide RNA features to obtain sample spliced gene features;
[0228] Using the sample spliced gene features to train a gene editing result prediction model to be trained.
[0229] Optionally, the training module is further configured to:
[0230] Inputting the sample spliced gene features into at least two different gene editing result prediction models to be trained respectively;
[0231] When the trained at least two different gene editing result prediction models all meet the corresponding training requirements, confirming that the at least two different gene editing result prediction models are all trained.
[0232] Optionally, the training module is further configured to:
[0233] Calculating the corresponding verification results of the trained at least two different gene editing result prediction models;
[0234] Combining at least two of the verification results to obtain a comprehensive verification result;
[0235] When the comprehensive verification result meets the training requirements, confirming that the at least two gene editing result prediction models are all trained.
[0236] Optionally, the verification result includes: a loss value; the training module is further configured to:
[0237] Combining the loss values of the at least two gene editing result prediction models to obtain a comprehensive loss value;
[0238] The step of when the comprehensive verification result meets the training requirements, confirming that the at least two gene editing result prediction models are all trained, includes:
[0239] When the comprehensive loss value is less than a loss value threshold, confirming that the at least two gene editing result prediction models are all trained.
[0240] Optionally, the training module is further configured to:
[0241] Input the spliced gene features of the sample into the edited gene recognition model to obtain a recognition result indicating whether the sample genome has been edited, and input the spliced gene features of the sample into the gene editing probability prediction model to obtain a predicted probability value that the sample genome has been edited;
[0242] Compare the sample label of the sample genome with the recognition result and the predicted probability value to respectively obtain a first loss value of the edited gene recognition model and a second loss value of the gene editing probability prediction model;
[0243] When the combined comprehensive loss value obtained from the first loss value and the second loss value is less than the preset loss value, confirm that both the edited gene recognition model and the gene editing probability prediction model are trained.
[0244] In the embodiments of the present disclosure, gene editing data for predicting gene editing results is constructed by using target gene sequence data, guide RNA sequence data, and target gene methylation data, enabling the gene editing result prediction model to comprehensively consider the influence of gene methylation and guide RNA sequence data on gene editing during the prediction process, thereby improving the accuracy of gene editing prediction.
[0245] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative effort.
[0246] Each component embodiment of the present disclosure can be implemented in hardware, or in software modules running on one or more processors, or in a combination thereof. Those skilled in the art should understand that a microprocessor or a digital signal processor (DSP) can be used in practice to implement some or all of the functions of some or all of the components in the computing processing device according to the embodiments of the present disclosure. The present disclosure can also be implemented as a device or device program (such as a computer program and a computer program product) for executing part or all of the methods described herein. Such a program implementing the present disclosure can be stored on a non-transitory computer-readable medium, or can be in the form of one or more signals. Such signals can be downloaded from an Internet website, provided on a carrier signal, or provided in any other form.
[0247] For example, Figure 13A computing processing device is shown that can implement the method according to the present disclosure. Traditionally, the computing processing device includes a processor 410 and a computer program product or a non-transitory computer-readable medium in the form of a memory 420. The memory 420 can be an electronic memory such as flash memory, EEPROM (Electrically Erasable Programmable Read-Only Memory), EPROM, hard disk, or ROM. The memory 420 has a storage space 430 for program code 431 for performing any method steps in the above-described method. For example, the storage space 430 for program code can include respective program codes 431 for implementing various steps in the above method. These program codes can be read from or written to one or more computer program products. These computer program products include program code carriers such as hard disks, compact discs (CDs), memory cards, or floppy disks. Such computer program products are typically portable or fixed storage units as described with reference to Figure 13 The storage unit may have a storage section, a storage space, etc. with an arrangement similar to that of the memory 420 in the Figure 12 computing processing device. The program code can be compressed in an appropriate form, for example. Generally, the storage unit includes computer-readable code 431’, that is, code that can be read by a processor such as 410, which, when run by the computing processing device, causes the computing processing device to execute each of the steps in the method described above.
[0248] It should be understood that although the steps in the flowchart of the drawings are shown sequentially in the direction of the arrows, these steps are not necessarily executed sequentially in the order indicated by the arrows. Unless otherwise clearly stated in this document, there is no strict order restriction for the execution of these steps, and they can be executed in other orders. Moreover, at least some of the steps in the flowchart of the drawings may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.
[0249] As used herein, "one embodiment", "an embodiment", or "one or more embodiments" means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the present disclosure. In addition, note that the examples of the phrase "in one embodiment" herein do not necessarily all refer to the same embodiment.
[0250] In the description provided herein, numerous specific details are set forth. However, it can be understood that embodiments of the present disclosure may be practiced without these specific details. In some instances, well-known methods, structures, and techniques have not been shown in detail so as not to obscure the understanding of this description.
[0251] In a claim, any reference sign between parentheses shall not be construed as limiting the claim. The word "comprising" does not exclude the presence of elements or steps not listed in a claim. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present disclosure may be implemented by means of hardware comprising several different elements and by means of a suitably programmed computer. In a unit claim enumerating several devices, several of these devices may be embodied by the same item of hardware. The use of the words first, second, and third, etc. does not denote any order. These words may be interpreted as names.
[0252] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present disclosure and are not intended to limit them. Although the present disclosure has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present disclosure.
Claims
1. A method for predicting the results of gene editing, characterized in that, The method includes: Obtaining target gene methylation data, target gene sequence data of a target genome, and guide RNA sequence data corresponding to the target gene sequence data; Constructing gene editing data according to the target gene methylation data, the target gene sequence data, and the guide RNA sequence data; Inputting the gene editing data into a gene editing result prediction model for prediction; The gene editing data includes: a spliced gene feature composed of a methylation gene feature and a guide RNA feature; the constructing of the gene editing data according to the target gene methylation data, the target gene sequence data, and the guide RNA sequence data includes: performing aggregated feature extraction on the target gene sequence data towards the guide RNA sequence data to obtain a guide RNA feature, and performing aggregated feature extraction on the target gene sequence data towards the target gene methylation data to obtain a methylation gene feature; splicing the methylation gene feature and the guide RNA feature to obtain a spliced gene feature.
2. The method according to claim 1, characterized in that, The performing of the aggregated feature extraction on the target gene sequence data towards the guide RNA sequence data to obtain a guide RNA feature, and the performing of the aggregated feature extraction on the target gene sequence data towards the target gene methylation data to obtain a methylation gene feature includes: Calculating the correlation between each element in the target gene sequence data and each value in the guide RNA sequence data; Performing weighted summation on the correlations of each element in the target gene sequence data to obtain a guide RNA feature; And calculating the correlation between each element in the target gene sequence data and each value in the target gene methylation data; Performing weighted summation on the correlations of each element in the target gene methylation data to obtain a methylation gene feature.
3. The method according to claim 1, wherein The splicing of the methylation gene feature and the guide RNA feature to obtain a spliced gene feature includes: Performing a convolution operation and a pooling operation on the methylation gene feature and the guide RNA feature to obtain a methylation gene matrix feature and a guide RNA matrix feature with the same dimension; Splicing the methylation gene matrix feature and the guide RNA matrix feature to obtain a spliced gene feature.
4. The method according to any one of claims 1 to 3, characterized in that, The gene editing result prediction model is obtained through the following steps: Obtaining sample target gene methylation data, sample target gene sequence data of a sample genome, and sample guide RNA sequence data corresponding to the sample target gene sequence data; Performing aggregated feature extraction on the sample target gene sequence data towards the sample guide RNA sequence data to obtain a sample guide RNA feature, and performing aggregated feature extraction on the sample target gene methylation data towards the sample guide RNA sequence data to obtain a sample methylation gene feature; Splicing the sample methylation gene feature and the sample guide RNA feature to obtain a sample spliced gene feature; Training a gene editing result prediction model to be trained using the sample spliced gene feature.
5. The method according to claim 4, wherein Training the gene editing result prediction model to be trained by using the sample splicing gene features includes: Inputting the sample splicing gene features into at least two different gene editing result prediction models to be trained respectively; When the at least two different trained gene editing result prediction models all meet the corresponding training requirements, it is confirmed that the at least two different gene editing result prediction models are all trained.
6. The method according to claim 5, wherein The step of, when the at least two different trained gene editing result prediction models all meet the corresponding training requirements, confirming that the at least two different gene editing result prediction models are all trained includes: Calculating the corresponding verification results of the at least two different trained gene editing result prediction models; Combining at least two of the verification results to obtain a comprehensive verification result; When the comprehensive verification result meets the training requirements, it is confirmed that the at least two gene editing result prediction models are all trained.
7. The method according to claim 6, characterized in that, The verification result includes: a loss value; the step of combining at least two of the verification results to obtain a comprehensive verification result includes: Combining the loss values of the at least two gene editing result prediction models to obtain a comprehensive loss value; The step of, when the comprehensive verification result meets the training requirements, confirming that the at least two gene editing result prediction models are all trained includes: When the comprehensive loss value is less than the loss value threshold, it is confirmed that the at least two gene editing result prediction models are all trained.
8. The method according to claim 5, characterized in that The gene editing result prediction model at least includes: an edited gene recognition model and a gene editing probability prediction model; the step of inputting the sample splicing gene features into at least two different gene editing result prediction models to be trained respectively includes: Inputting the sample splicing gene features into the edited gene recognition model to obtain a recognition result indicating whether the sample genome is edited, and inputting the sample splicing gene features into the gene editing probability prediction model to obtain a predicted probability value that the sample genome has been edited; The step of, when the at least two different trained gene editing result prediction models all meet the corresponding training requirements, confirming that the at least two different gene editing result prediction models are all trained includes: Comparing the sample label of the sample genome with the recognition result and the predicted probability value to obtain a first loss value of the edited gene recognition model and a second loss value of the gene editing probability prediction model respectively, where the sample label is the gene editing site and the standard gene editing probability of the sample genome; When the comprehensive loss value obtained by combining the first loss value and the second loss value is less than the preset loss value, it is confirmed that both the edited gene recognition model and the gene editing probability prediction model are trained.
9. A prediction device for gene editing results, characterized in that, The device includes: An acquisition module configured to acquire target gene methylation data, target gene sequence data of a target genome, and guide RNA sequence data corresponding to the target gene sequence data; A data processing module, configured to construct gene editing data according to the target gene methylation data, the target gene sequence data, and the guide RNA sequence data; A prediction module, configured to input the gene editing data into a gene editing result prediction model for prediction; Wherein, the gene editing data includes: a spliced gene feature composed of a methylated gene feature and a guide RNA feature; the constructing of the gene editing data according to the target gene methylation data, the target gene sequence data, and the guide RNA sequence data includes: performing aggregated feature extraction on the target gene sequence data towards the guide RNA sequence data to obtain a guide RNA feature, and performing aggregated feature extraction on the target gene sequence data towards the target gene methylation data to obtain a methylated gene feature; splicing the methylated gene feature and the guide RNA feature to obtain a spliced gene feature.
10. A computing processing device, characterized in that, Comprising: A memory, which stores computer-readable code; One or more processors, when the computer-readable code is executed by the one or more processors, the computing processing device executes the method for predicting gene editing results according to any one of claims 1-8.
11. A computer program product, characterized in that, Comprising computer-readable code, when the computer-readable code runs on a computing processing device, causing the computing processing device to execute the method for predicting gene editing results according to any one of claims 1-8.
12. A non-transitory computer-readable medium, characterized in that, A computer program for storing the method for predicting gene editing results according to any one of claims 1-8 is stored therein.
Citation Information
Patent Citations
CRISPR / Cas9 sgRNA activity prediction method based on deep learning
CN111613274A
Differential expression gene prediction system based on layered self-attention mechanism
CN114283888A