CRISPR-Cas13d gRNA design method and system based on intracellular dynamic environment and application

By integrating intracellular RNA secondary structure and RBP binding data through the SCALPEL model and using deep learning methods to predict gRNA efficiency, the problem of insufficient targeting activity and specificity of CRISPR-Cas13d gRNA in different cellular environments was solved, achieving more efficient and safer therapeutic effects.

CN121122401APending Publication Date: 2025-12-12SHANDONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511246156.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-02
Publication Date
2025-12-12

AI Technical Summary

Technical Problem

Existing CRISPR-Cas13d gRNA design models cannot accurately predict targeting activity and specificity under different cellular environments, leading to off-target effects and affecting treatment safety.

Method used

The SCALPEL model was constructed to integrate intracellular RNA secondary structure information and RBP binding data. Combined with deep learning methods, gRNA efficiency was predicted. RNA structure and binding sites were predicted through optimized icSHAPE experiments and PrismNet. Sequence features were processed using a Transformer network, and an efficiency difference threshold was introduced for evaluation.

Benefits of technology

Accurately predicting gRNA efficiency in different cellular environments can improve targeting activity and specificity, reduce off-target effects, and enhance the safety and efficacy of treatment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121122401A_ABST
    Figure CN121122401A_ABST
Patent Text Reader

Abstract

The invention discloses a CRISPR-Cas13d gRNA design method and system based on an intracellular dynamic environment and application, and relates to the technical field of RNA editing and deep learning. The method comprises the following steps: acquiring secondary structure information and RBP binding site distribution condition of target RNA in a cell line; constructing a gRNA design model based on a deep learning algorithm, and learning a gRNA recognition process, an RNA structure of a target site and a correlation among an RBP binding process by using the gRNA design model according to secondary structure information of the target RNA in the cell line and a distribution condition of the RBP binding site to obtain a gRNA efficiency prediction result; and introducing an efficiency difference threshold to evaluate a gRNA efficiency prediction result. According to the method, the gRNA efficiency can be efficiently predicted in different cell environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of RNA editing and deep learning technologies, and in particular to a CRISPR-Cas13d gRNA design method, system, and application based on the dynamic intracellular environment. Background Technology

[0002] The statements in this section are merely background information related to the present invention and do not necessarily constitute prior art.

[0003] The CRISPR system (Clustered Regularly Interspaced Short Palindromic Repeats) is an acquired immune system derived from bacteria and archaea, which has been transformed into a revolutionary gene-editing tool. The CRISPR-Cas13 system (Clustered Regularly Interspaced Short Palindromic Repeats–CRISPR-associated protein 13) is a gRNA-mediated targeted RNA editing technology. Its biggest difference from CRISPR-Cas9 is that it targets RNA, performing RNA knockdown, and can be used for single-base RNA editing and RNA detection. Compared to traditional RNA interference (RNAi) technology, CRISPR-Cas13 has stronger targeting specificity and higher knockdown efficiency, thus leading to its rapid development. Furthermore, because RNA editing does not require PAM sequences and does not cause permanent damage to the genome, CRISPR-Cas13 has broad prospects for medical applications. However, due to the complexity of the intracellular environment, especially the influence of RNA secondary structure and RNA-binding proteins (RBPs), designing efficient and highly specific CRISPR-Cas13 gRNAs remains a huge challenge.

[0004] In recent years, studies have attempted to improve the prediction accuracy of RfxCas13d gRNA (Ruminococcusflavefaciens Cas13d guide RNA) in terms of targeting activity by utilizing general information. Some studies have used random forest (RF) models to investigate the impact of basic design rules and mispairing on gRNA activity. Deep learning models have been used to improve gRNA targeting efficiency prediction by combining DeepCas13 with computationally predicted gRNA secondary structure folding as a separate input. Others have proposed the TIGER model, trained on large-scale datasets, which considers more secondary structure features, thus further elucidating the design principles of gRNAs. Specifically, TIGER systematically studies off-target effects by simultaneously inputting gRNA and target RNA sequences. Based on FACS sorting or CRISPR-Cas13d proliferation screening, in vivo assessment of the activity of a large number of gRNAs in different mammalian cell lines has been achieved.

[0005] However, these models lack a systematic assessment of cell-environment-specific factors, particularly RNA secondary structure and RNA-binding proteins (RBPs), making it difficult to accurately predict gRNA efficiency under different cellular conditions. In clinical applications, especially during high-dose gRNA delivery, off-target effects on unwanted cells, tissues, and organs are often unavoidable, severely impacting treatment safety. Therefore, designing gRNAs with cellular environment awareness capabilities, optimized for specific cell types, and combining in vivo information to distinguish between target and non-target cells has become crucial for improving gRNA targeting activity and specificity.

[0006] In conclusion, developing gRNA optimization strategies that can sense the cellular environment and adapt to specific cell types to more accurately assess and enhance the targeting activity and specificity of gRNA is of great significance for promoting the safe and efficient application of RNA editing technology in clinical practice. Summary of the Invention

[0007] To address the shortcomings of existing technologies, the present invention aims to provide a CRISPR-Cas13d gRNA design method, system, and application based on the dynamic intracellular environment. This method integrates RNA secondary structure information at the whole transcriptome level from different cell lines and binding data of RNA-binding proteins (RBPs) within those cell lines. It also incorporates data on the impact of CRISPR-Cas13d gRNA on cell growth or GFP expression obtained through existing high-throughput screening (i.e., gRNA editing efficiency data). Using deep learning methods, a relationship model between intracellular RNA secondary structure, RBP binding, and gRNA recognition is constructed and named SCALPEL (Specific CRISPR-Cas13 gRNA design through deep Learning Prediction using in vivo Experimental RNA structure and binding information). This deep learning model enables efficient prediction of gRNA efficiency under different cellular environments.

[0008] To achieve the above objectives, the present invention is implemented through the following technical solution: The first aspect of this invention provides a CRISPR-Cas13d gRNA design method based on the dynamic intracellular environment, comprising the following steps: To obtain secondary structure information of target RNA and distribution of RBP binding sites within cell lines; A gRNA design model was constructed based on deep learning algorithms. The gRNA design model was used to learn the gRNA recognition process, the RNA structure of the target site, and the relationship between the RBP binding process based on the secondary structure information of the target RNA in the cell line and the distribution of RBP binding sites, so as to obtain the gRNA efficiency prediction results. An efficiency difference threshold was introduced to evaluate the gRNA efficiency prediction results.

[0009] Furthermore, the RNA secondary structure information of the cell line was obtained through the optimized icSHAPE experiment, and the distribution of intracellular RBP binding sites was predicted using PrismNet.

[0010] Furthermore, the gRNA design model includes a Transformer-based sequence processing network, a feature fusion network for integrating additional features, and a regression network for predicting efficiency.

[0011] Furthermore, the specific steps for using a gRNA design model to learn the gRNA recognition process, the RNA structure of the target site, and the relationship between the RBP binding process based on the secondary structure information of the target RNA and the distribution of RBP binding sites within the cell line are as follows: A dataset was constructed using RNA secondary structure information from different cell lines, binding information of different RBPs within cells, and gRNA editing efficiency data. Sequence features were obtained by encoding and extracting gRNA and target RNA sequences using a sequence processing network. By fusing sequence features and data from the dataset, integrated features are obtained. Regression networks were used to predict gRNA efficiency based on the integrated features.

[0012] Furthermore, the target RNA sequence is derived from the sequence at the corresponding target location on the transcriptome, and is truncated to the same length as the gRNA sequence. The gRNA sequence comes from large-scale CRISPR screen data that has been collected and uniformly processed, and the gRNA sequence and the target RNA sequence are inversely complementary.

[0013] Furthermore, the specific steps for encoding and extracting sequence features from gRNA and target RNA sequences using sequence processing networks are as follows: First, the gRNA sequence is preprocessed; Next, predict the secondary structure information of the gRNA; Subsequently, the gRNA sequence, target RNA sequence, and secondary structure of the gRNA were one-hot encoded and converted into binary information; Finally, the BERT pre-trained model was used to process the gRNA sequence, target RNA sequence, and secondary structure of gRNA to obtain sequence features.

[0014] Furthermore, the specific steps for processing gRNA sequences, target RNA sequences, and gRNA secondary structures using the BERT pre-trained model are as follows: The target RNA sequence was 3-mer encoded using a BERT pre-trained model to obtain a deep feature representation, which was then dimensionality reduced. Subsequently, the one-hot encoded features of the gRNA sequence, target RNA sequence, and gRNA secondary structure were embedded and fused through a convolutional layer.

[0015] Furthermore, the specific steps for evaluating gRNA efficiency prediction results by introducing an efficiency difference threshold are as follows: Proliferation screening experiments were conducted based on gRNA efficiency prediction results; An efficiency difference threshold was introduced to evaluate the results of the proliferation screening experiment.

[0016] A second aspect of the present invention provides a CRISPR-Cas13d gRNA design system based on the dynamic intracellular environment, comprising: The data acquisition module is configured to acquire secondary structure information of target RNA and distribution of RBP binding sites within the cell line; The model prediction module is configured to build a gRNA design model based on a deep learning algorithm. The gRNA design model is used to learn the gRNA recognition process, the RNA structure of the target site, and the relationship between the RBP binding process based on the secondary structure information of the target RNA in the cell line and the distribution of RBP binding sites, so as to obtain the gRNA efficiency prediction results. The evaluation adjustment module is configured to introduce an efficiency difference threshold to evaluate the gRNA efficiency prediction results.

[0017] The third aspect of this invention provides an application of the CRISPR-Cas13dgRNA design system based on the dynamic intracellular environment in different cellular environments, as described in the second aspect.

[0018] The above one or more technical solutions have the following beneficial effects: This invention discloses a CRISPR-Cas13d gRNA design method, system, and application based on the dynamic intracellular environment, which can accurately predict the targeting effect of gRNA. Using deep learning methods, a relationship model between intracellular RNA secondary structure, RBP binding, and gRNA recognition was constructed and named SCALPEL. This invention utilizes this deep learning model to efficiently predict gRNA efficiency under different cellular environments, solving for the first time the problem that existing models cannot predict gRNA activity in different cell types, and enabling the design of highly specific gRNAs for different cell lines.

[0019] This invention constructs a gRNA design model based on deep learning algorithms. This model can more accurately and dynamically describe how RNA secondary structure interacts with protein binding and gRNA recognition. Furthermore, this invention integrates RNA secondary structure information at the whole transcriptome level from different cell lines and binding data of RNA-binding proteins within those cell lines. Combined with existing high-throughput screening data on the effects of CRISPR-Cas13d gRNA on cell growth or GFP expression (i.e., gRNA editing efficiency data), the gRNA design model was trained to obtain optimized model parameters.

[0020] This invention evaluates the predictive performance of gRNA design models by introducing an efficiency difference threshold, providing a double guarantee for the accurate prediction of gRNA design models.

[0021] Advantages of additional aspects of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0022] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0023] Figure 1 This is a flowchart of the optimizedSHAPE technique used in Embodiment 1 of the present invention to analyze the secondary structure of human cell line RNA; Figure 2 This is a schematic diagram illustrating the use of the PrismNet tool to obtain RBP binding information in Embodiment 1 of the present invention; Figure 3 This is a schematic diagram of the SCALPEL deep learning model architecture in Embodiment 1 of the present invention; Figure 4 This is a comparison of SCALPEL in Embodiment 1 of the present invention with existing models in terms of ROC curves and other performance evaluation metrics; Figure 5 This is a flowchart of the verification experiment conducted in Embodiment 1 of the present invention; Figure 6 The images show the flow cytometry analysis and fluorescence imaging results obtained in Example 1 of this invention. Figure 7 This is a graph showing the proportion of gRNAs that accurately determined the differences in efficacy among different cell lines using the SCALPEL model in Embodiment 1 of the present invention. Detailed Implementation

[0024] It should be noted that the following detailed descriptions are exemplary and intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0025] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of exemplary embodiments according to the invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof. The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0026] Example 1: Embodiment 1 of this invention provides a CRISPR-Cas13d gRNA design method based on the dynamic intracellular environment. It creatively utilizes in vivo dynamic information, namely target RNA secondary structure information and RBP binding information, to assist in gRNA design, including the following steps: Step 1: Obtain secondary structure information of target RNA and distribution of RBP binding sites in the cell line.

[0027] In this embodiment, the secondary structure information of the target RNA in the cell line was obtained through the optimized icSHAPE experiment, and the distribution of intracellular RBP binding sites was predicted using the existing protein-RNA interaction deep neural network (PrismNet) based on structure information.

[0028] Specifically, such as Figure 1 As shown, this embodiment first resolved the RNA secondary structure information in different cell lines. Target RNA structure is crucial for Cas13d gRNA activity because Cas13d preferentially binds to single-stranded RNA (ssRNA). This embodiment utilized an optimized version of icSHAPE technology to resolve the RNA secondary structure of three human cell lines: A375, HEK293FT, and HAP1. Specifically, cDNA libraries were constructed, and the RNA secondary structure was detected through high-throughput sequencing using optimized icSHAPE. Optimized icSHAPE removes spurious signals caused by premature termination of reverse transcription through RNase I digestion, and combined with randomized reverse transcription and cDNA library construction on magnetic beads, significantly reduces sample loss caused by multiple purification steps.

[0029] HEK293FT, A375, and HAP1 live cells were treated with 2-(azidomethyl)nicotinic acid imine (NAI-N3) at 37°C for five minutes. This chemical reagent modifies single-stranded regions of cellular RNA. RNA was then extracted, click-coated, and reverse transcribed using random primers to generate cDNA. The chemically modified nucleotides inhibited reverse transcription and generated a stop signal. Next, RNA was digested using RNase I, biotin bound the modified RNA molecules, and streptavidin magnetic beads were used to capture them. Sequencing libraries were constructed on the beads, and these signals were detected during sequencing. Finally, high-throughput sequencing was performed. After quality control, this embodiment obtained 200 million usable data points from each biological replicate library. The icSHAPE score (ranging from 0 to 1) was calculated using the bioinformatics tool icSHAPE-pipe, with higher scores indicating a greater probability of single-stranded RNA. This embodiment thus obtained a large amount of RNA structural data.

[0030] On the other hand, this embodiment uses PrismNet to predict the distribution of RBP binding sites in different cell lines.

[0031] PrismNet is an artificial intelligence model that can predict RBP binding in specific cellular environments using intracellular RNA structure information. After measuring intracellular RNA structure data from multiple cell lines, the existing PrismNet model can help predict the distribution of RBP binding in different cell lines in this embodiment. Figure 2 As shown, this embodiment uses in vivo RNA structure data (icSHAPE scores) and gRNA target sequences from different cell lines as input to obtain 172 RBP binding probabilities. This embodiment first uses PrismNet to obtain dynamic RBP binding information in HEK293FT, K562, HAP1, and A375 cells. Using in vivo RNA structure data and gRNA target sequences from different cell lines as input, this embodiment obtains 172 RBP binding probabilities. These probabilities reflect the overall binding conditions around the target site, and this information can be used as input for subsequent prediction of gRNA efficiency. They will be integrated into a deep learning model along with RNA structure information in the future.

[0032] Step 2: Construct a gRNA design model based on deep learning algorithms. Using the gRNA design model, learn the gRNA recognition process, the RNA structure of the target site, and the interrelationship between the RBP binding process based on the secondary structure information of the target RNA in the cell line and the distribution of RBP binding sites, and obtain the gRNA efficiency prediction results.

[0033] Step 2.1: Construct a gRNA design model based on deep learning algorithms.

[0034] The gRNA design model includes a Transformer-based sequence processing network, a feature fusion network for integrating additional features, and a regression network for predicting efficiency.

[0035] Step 2.2: Train the gRNA design model.

[0036] Among them, the gRNA design model is used to learn the relationship between gRNA recognition process, RNA structure of target site and RBP binding process based on the secondary structure information of target RNA in cell line and the distribution of RBP binding sites.

[0037] This embodiment constructs a neural network model to describe the relationship between gRNA recognition, the RNA structure at the target site, and RBP binding. This interaction model, by incorporating intracellular structural information and RBP binding information, can more accurately and dynamically describe how RNA secondary structure interacts with protein binding and gRNA recognition. In new cellular environments, given the RNA structural information and the RBP binding information predicted by PrismNet, this model can be used to predict the efficiency of gRNA in the new environment.

[0038] Step 2.2.1: Construct a dataset using secondary structure information of target RNA in different cell lines, binding information of different RBPs in cells, and gRNA editing efficiency data.

[0039] This embodiment also collects and integrates existing high-throughput gRNA editing efficiency data.

[0040] To date, four working systems have been used to screen the efficiency of different gRNAs in CRISPR-Cas13d and their effects on cells. These include: screening using green fluorescent protein (GFP) or stainable mRNAs such as CD46 and CD55; screening for gRNAs with different efficiencies using flow cytometry sorting; classifying cells based on intracellular GFP signals and then obtaining the sequences of gRNAs within different classes through sequencing; and designing gRNAs that target essential intracellular mRNAs and using cell viability to represent the efficiency of different gRNAs. This example collected a total of 290,000 gRNA editing efficiency data.

[0041] Step 2.2.2: Use a sequence processing network to encode and extract sequence features from the gRNA and target RNA sequences, such as... Figure 3As shown in the figure. The target RNA sequence is derived from the sequence at the corresponding target location on the transcriptome, and for ease of encoding, it is truncated to the same length as the gRNA sequence. In this embodiment, the gRNA sequence comes from large-scale CRISPR screen data that has been collected and uniformly processed; the gRNA sequence and the target RNA sequence are inversely complementary.

[0042] Step 2.2.2.1: First, preprocess the gRNA sequence. To process gRNAs of different lengths, target RNAs shorter than 30 nt are extended to 30 nt by extending the 5' flanking sequence and filling the gRNA with 'N'.

[0043] Step 2.2.2.2: Next, predict the secondary structure information of gRNA. In this embodiment, RNAfold software is used to predict the secondary structure of gRNA.

[0044] Step 2.2.2.3: Subsequently, the gRNA sequence, target RNA sequence, and secondary structure of gRNA were one-hot encoded into binary information, that is, the classification data were converted into binary vectors, where each category is represented by an independent bit, with only the bit corresponding to the category being 1 and the other bits being 0.

[0045] Step 2.2.2.4: Finally, the BERT pre-trained model is used to process the gRNA sequence, target RNA sequence, and secondary structure of gRNA to obtain sequence features.

[0046] Specifically, SCALPEL uses a BERT pre-trained model to process gRNA sequences, target RNA sequences, and gRNA secondary structures. BERT is a Transformer model with bidirectional encoder representation, widely used in natural language processing for pre-trained Transformer models. The target RNA sequence is 3-mer encoded using the BERT pre-trained model to obtain a deep feature representation with a feature tensor dimension of (N, 28, 768), which is then dimensionality-reduced. Subsequently, the one-hot encoded features of the gRNA sequence, target RNA sequence, and gRNA secondary structure are embedded and fused through convolutional layers to achieve feature integration. This embodiment integrates information from both the target and gRNA sequences by processing the target RNA sequence and gRNA sequence, facilitating the detection of off-target effects caused by base mismatches or insertions / deletions. Furthermore, the use of the BERT pre-trained model enhances SCALPEL's ability to capture implicit relationships between sequences, offering significant advantages over traditional CNN or RNN models.

[0047] Step 2.2.3: Merge sequence features and data from the dataset to obtain integrated features.

[0048] Specifically, the sequence features obtained by processing the BERT pre-trained model in step 2.2.2 are fused with additional features other than the gRNA sequence, target RNA sequence, and gRNA secondary structure to obtain integrated features. The additional features come from the dataset constructed in step 2.2.1, including cellular environment information (i.e., cell type-specific target RNA structure and protein binding probability) and target location, etc.

[0049] Cellular environment information includes icSHAPE scores and protein binding probabilities from different cell lines. Target location refers to the region and relative position of the gRNA targeting the transcript. A multi-source feature fusion network is used to concatenate all features along their feature channel dimensions. SEBlock and global attention mechanisms are employed to ensure the effectiveness of feature fusion and to perform dimensionality reduction.

[0050] Step 2.3.3: Use a regression network to predict gRNA efficiency of the integrated features.

[0051] This embodiment introduces a Transformer-based architecture that integrates in vivo cellular environment information, which has significant advantages over existing models that rely solely on traditional CNN or RNN architectures.

[0052] In vivo cellular environment information, including cell type-specific target RNA structures and protein binding probabilities, significantly improved model performance when analyzing the impact of different features. By integrating in vivo cellular environment information, SCALPEL accurately predicted the gRNA efficiency of EIF5 transcripts in the HAP1 cell line, even when trained solely on HEK293FT cell line data. In contrast, SCALPEL trained only on HEK293FT sequences struggled to capture cell line-specific efficiencies, highlighting the crucial role of in vivo RNA structure data in predicting dynamic efficiencies.

[0053] Step 3: Introduce an efficiency difference threshold to evaluate the gRNA efficiency prediction results.

[0054] Step 3.1: Conduct proliferation screening experiments based on the gRNA efficiency prediction results.

[0055] Existing models cannot predict gRNA activity in different cellular environments, while SCALPEL can accurately predict changes in gRNA efficiency across different cell lines. To validate the performance of SCALPEL in large-scale applications, this example demonstrates proliferation screening experiments conducted in two different cell lines (A375 and HeLa) using a library containing 1000 gRNAs that target 120 essential genes and 30 non-essential genes.

[0056] The specific experimental procedure includes the following steps: Step 3.1.1: Design of the screening experiment.

[0057] This embodiment selected 120 essential genes (DepMap 24Q4 public dataset; DepMap score <-1, in ≥500 cell lines) and 30 non-essential genes (DepMap 24Q4 public dataset; DepMap score between -0.1 and +0.1, in ≥500 cell lines) as controls. For each gene, this embodiment selected representative transcripts tagged as “MANE Select” or “Ensembl canonical” in the Ensembl database. Based on the corresponding in vivo data from the two cell lines, this embodiment designed six high-scoring gRNAs for the coding regions of each essential and non-essential gene. In addition, the library in this embodiment also included 100 non-target (NT) gRNAs, which had more than three mismatches with the hg38 human transcriptome.

[0058] Step 3.1.2: Construction of doxycycline-induced RfxCas13d stably expressing cells.

[0059] The doxycycline-induced RfxCas13d-NLS HeLa and A375 cells used in the screening assay were generated via lentiviral transduction under low infection coefficient (MOI < 0.1) conditions. Approximately 48 hours post-lentiviral infection, cells were screened using 5 μg / ml paclitaxel S (Thermo Fisher Scientific). Following screening, single cell clones were isolated, and RfxCas13d expression was validated by Western blotting over the next two weeks.

[0060] Step 3.1.3: Cloning of the gRNA library.

[0061] The gRNA library was synthesized as single-stranded oligonucleotides (Azenta / Genewiz) and subsequently amplified by PCR in a single reaction. PCR conditions were as follows: 98°C for 30 seconds, followed by 9 cycles of 98°C for 10 seconds, 63°C for 10 seconds, 72°C for 15 seconds, and a final extension at 72°C for 3 minutes. The PCR products were purified by agarose gel electroporation and cloned into the BsmBI-digested pLentiRfxGuide-Puro vector (Addgene 138151) using the Gibson assembly method. The resulting library was then electroporated into Endura-electroplated *E. coli* (Lucigen). Each gRNA in the library had over 1000 clones. Next-generation sequencing (NGS) was then performed to verify that the library contained the complete gRNA sequence. Step 3.1.4: Lentiviral production and screening experiments.

[0062] To generate lentiviruses expressing mixed libraries, LentiX-293T cells were injected with 4 × 10⁻⁶ cells per cell line 24 hours prior to transfection. 6 The cells were seeded at a density of 10 cells / 10 cm culture dish. In this example, lentivirus was produced by co-transfecting the library plasmid and packaging plasmid (psPAX2 and pCMV-VSVG) using the Perfect transfection reagent (Vazyme). 72 hours after transfection, the supernatant containing lentivirus was collected, filtered through a 0.45 μm membrane, and stored at -80°C for later use.

[0063] A375 and HeLa cells stably expressing RfxCas13d were transduced with a lentiviral library at low multiplicity of infection (MOI < 0.3). Twenty-four hours after transduction, cells were selected using 1 μg / ml puromycin (Sigma-Aldrich) to remove uninfected cells. Four biological replicates were performed in this embodiment. After puromycin selection, RfxCas13d expression was induced by adding 1 μg / ml doxycycline (Sigma-Aldrich). Selected cells were cultured continuously for 14 days and passaged every two days to maintain the initial cell number. At least 10 cells were collected on day 0 and every 7 days. 6 The cells were collected and stored at -80°C for subsequent extraction of genomic DNA.

[0064] Step 3.1.5: Library construction and high-throughput sequencing.

[0065] For each sample, genomic DNA was extracted from the frozen cell pellet using the FastPure Cell / TissueDNA Isolation Mini Kit (Vazyme) following the manufacturer's instructions. A two-step PCR strategy was then employed to amplify the gRNA sequence. PCR 1 amplified the crRNA-containing fragment, and the subsequent PCR 2 reaction was used to attach Illumina sequencing adapters, barcodes, and stagger sequences to prevent issues associated with single-template sequencing.

[0066] The PCR1 reaction conditions were as follows: 25 μl NEBNext Ultra II Q5 Master Mix (NEB), 0.5 μM each of forward and reverse primers, and 100 ng genomic DNA per μl. The PCR conditions were: 98 °C for 30 seconds, followed by a 24× reaction (98 °C, 10 seconds; 65 °C, 75 seconds), then at 65 °C for 5 minutes. The unpurified PCR1 product was used directly as a template for a second PCR.

[0067] The second PCR was performed in a 100 μl reaction volume, consisting of: 50 μl NEBNext Ultra II Q5Master Mix (NEB, M0544), 10 μl PCR1 product, and 0.5 μM of forward and reverse PCR2 primers. The conditions for the second PCR were: 98 °C for 30 seconds, 17× (98 °C, 10 seconds; 65 °C, 75 seconds), and 65 °C for 5 minutes.

[0068] These PCR products were purified by agarose gel electrophoresis and quantified using the Qubit dsDNA HS kit (ThermoFisher). The purified products were then mixed in equimolar ratios and sequenced on an Illumina NovaSeq 6000 platform.

[0069] Step 3.2: Introduce an efficiency difference threshold to evaluate the results of the proliferation screening experiment.

[0070] Step 3.2.1: Screening and analyzing experimental sequencing data.

[0071] Raw sequencing data from different samples were provided in FastQ format. First, Cutadapt was used to trim the data based on known anchor sequences, ensuring all data had the same length. Subsequently, the MAGeCK software package was used for data processing.

[0072] Step 3.2.2: Results Analysis.

[0073] To assess efficiency variations between the A375 and HeLa cell lines, this embodiment defines a threshold for efficiency differences. First, the results of all replicate experiments are summarized, and the L1 distance is calculated. The top 20% of L1 distances from these replicate experiments are defined as the threshold for efficiency differences (threshold = 0.6069). This embodiment found that gRNA counts and gRNA LFCs were highly reproducible across all cell lines and time points. Essential genes screened in this embodiment were significantly knocked down as expected. Furthermore, this embodiment noted that gRNAs targeting non-essential genes did not lead to significant gene knockdown in all screenings, demonstrating that cell fitness was not affected by non-specific side effects in the screening setup of this embodiment. This embodiment found that SCALPEL correctly predicts highly efficient gRNAs and outperforms existing models, highlighting the robustness of the model in this embodiment.

[0074] Existing models cannot predict gRNA activity in different cellular environments, while SCALPEL can accurately predict changes in gRNA efficiency across different cell lines. This embodiment evaluates SCALPEL's performance in predicting differences in gRNA efficiency. An efficiency difference threshold is defined for this purpose, determined by estimating random noise between repeated screenings. In this embodiment, the top 5% of the L1 distances from these cell lines is defined as the efficiency difference threshold for random noise. Figure 7 As shown, this embodiment observed that approximately 32% of gRNAs exhibited varying efficiency under different cellular conditions. Specifically, approximately 22% of gRNAs showed significantly higher efficiency in the HeLa cell line, while approximately 10% of gRNAs showed higher efficiency in the A375 cell line. This embodiment also found that SCALPEL performed exceptionally well in predicting gRNAs with significant dynamic efficiency changes. Specifically, SCALPEL accurately classified 77.3% of gRNAs, which showed significantly higher efficiency in the A375 cell line; simultaneously, 94.2% of gRNAs showed higher efficiency in the HeLa cell line.

[0075] This example demonstrates that, unlike previous models, SCALPEL can optimize gRNA design based on the specific environmental characteristics of a tissue, ensuring its efficient targeting in the target tissue or organ. Previous gRNA design models failed to adequately consider the specificity of the cellular environment, resulting in gRNAs that might be effective only in cell lines but less effective in specific tissues or organs. The technology in this embodiment, by incorporating cellular environmental information from the target tissue, optimizes gRNA design, ensuring its highly efficient targeting capability within specific tissues or organs, thereby improving the precision and efficacy of treatment.

[0076] For example, this embodiment demonstrates significant potential in cancer treatment. By designing specific gRNAs for different types of cancer cells, side effects can be reduced and the precision and efficiency of treatment can be improved. Specifically, this embodiment demonstrated its potential application value in cervical cancer treatment through experiments targeting HeLa cells, and its feasibility in melanoma treatment through screening experiments targeting A375 cells. Both experiments demonstrate the potential of this embodiment in targeted therapy for specific cancers. Furthermore, this embodiment has broad application prospects, especially in treating diseases of other vital organs. For example, the liver, lungs, and kidneys are prone to disease. By specifically targeting diseased tissues or organs, treatment efficacy can be greatly improved and damage to normal tissues reduced. For instance, in the treatment of liver cancer (such as hepatocellular carcinoma) and lung cancer, precisely targeting diseased cells in the liver or lungs can significantly improve treatment efficacy and reduce side effects without affecting healthy tissues. Further applications can be extended to organs such as the heart, intestines, and brain, especially those organs susceptible to disease due to environmental or genetic factors, providing broad prospects for precision medicine and personalized treatment.

[0077] This embodiment can accurately predict the targeting effect of gRNA, and its effect is significantly better than the current state-of-the-art gRNA design tools.

[0078] After model training, this embodiment compares SCALPEL with three state-of-the-art models, such as... Figure 4 As shown, the method calculates the correlation between model predictions and actual observations (including Pearson and Spearman correlations), the area under the operating characteristic curve (AUROC), and the area under the precision-recall curve (AUPRC), and evaluates all predictions of the model in this embodiment and other models. The results show that SCALPEL performs best on these four metrics. SCALPEL's AUROC is 0.883, significantly higher than the state-of-the-art models.

[0079] This embodiment solves for the first time the problem that existing models cannot predict the activity of gRNA in different cell types, and enables the design of highly specific gRNAs for different cell lines.

[0080] To verify this, this embodiment used SCAPLE to design two types of gRNAs in HeLa and HEK293T cell lines: (1) ss gRNA, targeting regions with high flexibility in both HeLa and HEK293T cells, which showed high editing efficiency in both cell lines; and (2) dysite gRNA, targeting regions with dynamic structural changes in both cell lines, which the model predicted would be highly efficient in one cell line and less efficient in the other. qRT-PCR results showed that the editing efficiency of ss gRNA was above 70% in both cell lines, while the editing efficiency of dysite gRNA showed a significant difference between the two cell lines. This indicates that SCAPLE fully learns the cell-specific environmental factors in different cell lines (including cell line-specific RNA secondary structures and RBP binding information), and can accurately design highly specific gRNAs for different cell lines.

[0081] This embodiment facilitates the design of gRNAs in animal models by integrating cellular environmental factors that are closely related to gRNA activity.

[0082] This embodiment attempts to design gRNA to target and knock down maternal genes in zebrafish embryos. Traditional DNA techniques for studying maternal genes in zebrafish are often very complex, involving cumbersome hybridization steps to obtain homozygous individuals, and the homozygosity of some genes can be lethal, hindering effective research. In contrast, RNA targeting technology can directly intervene in the expression of maternal genes, providing a more convenient method. However, current RNA targeting technologies have a high failure rate because the design rules for gRNAs in animal models are not yet clear. Past deep learning models have mostly been trained based on cell line data, and their performance in animal models remains unclear. The SCALPEL model has a unique advantage: it can precisely design gRNAs based on the specific environmental characteristics of cells and animal models, thereby effectively improving the success rate of RNA targeting technology in animal models. This innovative technology provides strong support for gene function research and precise intervention in disease models.

[0083] Using the SCALPEL model for gRNA design, combined with in vivo RNA structure, RBP binding information, and high-throughput screening data, this embodiment designed highly efficient gRNAs targeting the key maternal genes dnd1 and nanog in early zebrafish development. High-scoring and low-scoring gRNAs were injected into zebrafish 1-cell stages, and their efficiency was subsequently assessed by flow cytometry and qRT-PCR. The results showed that the highly efficient gRNAs significantly reduced the expression of dnd1 and nanog transcripts, inducing significant developmental defects. This indicates that the model in this embodiment effectively captures environmental factors affecting gRNA efficiency in animal models, highlighting its versatility.

[0084] Furthermore, the method of this embodiment not only demonstrates application potential in zebrafish, but also, considering the successful application of Cas13d technology in various animal models, can be extended to mice, monkeys, and other animal models. By integrating in vivo RNA structure data and RBP binding information, this embodiment can design more specific and efficient gRNAs for different animal models, thereby promoting the widespread application of this technology in more species, especially in gene function research and the establishment of disease models.

[0085] For example, Cas13d technology has been successfully applied in mice and non-human primate models (such as monkeys), providing a powerful tool for studying the role of RNA in embryonic development, neurological diseases (such as Alzheimer's and Parkinson's diseases), and immune responses (such as targeted therapy for RNA associated with viral infection).

[0086] However, existing gRNA design models fail to fully consider cell-specific environments, making it impossible to design highly efficient, specific gRNAs based on the specific cellular characteristics of animal models. This significantly impacts targeting efficacy. The SCALPEL model, by combining cellular environment data, RNA structural characteristics, and RBP binding information, can precisely design gRNAs for each animal model, ensuring they function only in specific target cells or tissues. This advantage not only improves its efficiency and convenience in various animal models but also provides a safer guarantee for disease treatment research.

[0087] In conclusion, SCALPEL has enormous application potential in animal models. This embodiment can not only advance gene research in model organisms such as zebrafish, but also be extended to animals closely related to humans, such as mice, rats, monkeys, and pigs, ultimately providing new directions and solutions for human gene therapy and disease intervention.

[0088] To better illustrate the superiority of the SCALPEL model in this embodiment, the following experiment was conducted: The 5' untranslated region (5' UTR) plays a crucial role in regulating key stages of the life cycle of infectious viruses, such as SARS-CoV-2 and other coronaviruses. The 5' UTR of these viruses is highly structurally specific and contains several conserved stem-loop elements essential for viral translation and replication. Furthermore, the 5' UTR has a lower mutation rate compared to protein-coding regions, making it a promising target for antiviral strategies. Figure 5 As shown, to evaluate whether the model in this embodiment can design efficient gRNAs for viral interference, this embodiment integrates the in vitro 5' UTR structures of five viruses. These structures were previously identified using the icSHAPE method, and their protein binding maps in the HEK293T host cell environment were predicted using PrismNet. Using the model in this embodiment, three high-scoring and three low-scoring gRNAs targeting the 5' UTRs of five viruses from five different coronavirus genera and lineages were designed.

[0089] The specific experimental procedure includes the following steps: S1: Construct an enhanced green fluorescent protein (destabilized EGFP) reporter system containing an unstable domain.

[0090] S1.1: Synthesize an expression vector containing viral UTR.

[0091] This embodiment selected five coronaviruses: SARS-CoV-2 MERS-CoV, BtCoV-HKU5, HCoV-NL63, HCoV-HKU1, and BtCoV-HKU9. The 5' UTR, 3' UTR, and extended regions of these viruses were constructed into expression vectors synthesized by AuGCT.

[0092] S1.2: Construct destabilized EGFP reporter plasmid.

[0093] Specifically, this embodiment first introduces an unstable domain coding sequence, namely the PEST sequence, after the enhanced green fluorescent protein (EGFP) sequence. This sequence is typically used to label proteins to give them a short half-life, so that when the EGFP RNA is degraded by the cas13d enzyme, the protein will also degrade rapidly, thus facilitating fluorescence observation. Subsequently, homologous recombination technology is used to construct the destabilized EGFP sequence into a viral UTR expression vector, so that the 5' UTR can control the expression of EGFP.

[0094] S2: Design gRNAs that target the UTR.

[0095] S2.1: Obtain RNA structure data and RBP binding data.

[0096] To design gRNAs targeting the 5' UTR of five viruses, this embodiment directly used in vitro RNA structure data detected by the icSHAPE method. RBP binding probabilities were predicted using PrismNet based on sequence and icSHAPE data.

[0097] For RNA structure data, this embodiment obtained icSHAPE data (icSHAPE score) of coronavirus UTR from previous research by Sun et al.

[0098] For RBP binding data, this embodiment first trained 172 RBP models using CLIP and icSHAPE data from the same cell line, provided by PrismNet. To predict the binding probability of the 172 RBPs, this embodiment divided the data into two parts (including structural data and excluding structural data) and used either "pu" mode (predicting using both sequence and structure) or "seq" mode (predicting using only sequence) respectively. For gRNAs targeting transcripts with icSHAPE scores, the target sequence and the corresponding icSHAPE score at each site were extended to the 5' and 3' ends, with a total length of 101 nt, as input to PrismNet. For gRNAs targeting transcripts without icSHAPE scores, the 101 nt target sequence was used for prediction.

[0099] S2.2 Design of gRNA targeting 5'UTR.

[0100] Using the RNA structure and RBP binding information obtained in S2.1 as input, and employing the SCALPEL model, this embodiment randomly selected 3 high-efficiency gRNAs and 3 low-efficiency gRNAs for each virus for subsequent experimental testing.

[0101] S3: Construction of gRNA expression plasmid.

[0102] S3.1: Annealing of gRNA complementary sequences.

[0103] The gRNA sequence was synthesized as a complementary single-stranded DNA (ssDNA) oligonucleotide. The complementary gRNA sequence was mixed with T4 DNA ligase buffer (final concentration 1×) and 0.5 μL of T4PNK, for a total reaction volume of 10 μL. The annealing procedure was as follows: reacted at 37°C for 30 min, then at 95°C for 5 min, followed by cooling to 4°C at a rate of 5°C / min.

[0104] S3.2: Enzyme digestion and dephosphorylation of the gRNA expression vector.

[0105] The gRNA expression vector was obtained from Addgene (Addgene 138151). The vector plasmid was digested with BsmBI restriction enzyme at 37°C for 30 minutes. The specific reaction system was as follows: 5 μg vector, 3 μl FastDigest BsmBI buffer, 3 μl FastAP, 6 μl 10×FastDigest buffer, 0.6 μl 100mM DTT, and the total reaction volume was 60 μl.

[0106] S3.3: Ligation of annealing products with enzyme digestion vectors and sequencing identification.

[0107] The annealing product was diluted 100-fold and then reacted with T4 DNA ligase at 16°C for 1 hour. The reaction mixture consisted of 1 μl of annealing product, 50 ng of the digested vector, 5 μl of 2× Quick Ligase buffer, and 1 μl of T4 DNA ligase (NEB), which was then inserted into the digested vector. The ligation product was then transformed into Stbl3 bacteria, plated, and cultured. After 12 hours, the colonies were identified by Sanger sequencing to obtain the correct ligation product.

[0108] S4: Cell transfection experiment to verify the effect of gRNA.

[0109] Twelve hours before transfection, HEK293T cells were introduced at a rate of 4 × 10⁶ cells per well. 5 Cells were seeded at a density of [number] cells per well in 6-well plates. RfxCas13d expression plasmid and gRNA expression plasmid were co-transfected with the 5' UTR-EGFP expression plasmid at a molar ratio of 2:2:1. EGFP expression was detected by flow cytometry (FACS) 48 hours after transfection using Ecfect (Vazyme) reagent. Flow cytometry data were analyzed using FlowJo software.

[0110] Representative results such as Figure 6 As shown in the figure. Through fluorescence microscopy and flow cytometry, this embodiment observed that, compared with inefficient gRNA, efficient gRNA significantly reduced GFP intensity in all viruses, demonstrating that the model of this embodiment can be designed with gRNAs with high targeting efficiency.

[0111] It is worth noting that targeting coronaviruses is only the beginning of this technology's application. With a deeper understanding of the RNA structures of various viruses, the method in this embodiment can be extended to the treatment of infections caused by a variety of viruses. For example, RNA viruses such as influenza, Ebola, yellow fever, Zika, and HIV pose a significant threat to human health. Traditional antiviral strategies for these viruses often face problems such as viral mutation and escape, and drug resistance. The method in this embodiment, by precisely and specifically targeting viral RNA, can effectively reduce this escape phenomenon, providing a new treatment approach. Compared with traditional methods, this method not only improves treatment efficiency but also enhances specificity, reduces side effects on host cells, and has higher safety.

[0112] Furthermore, this embodiment offers convenient and flexible design of the target region, allowing for optimization against different viruses. It can target not only coding regions but also conserved regions such as the 5' UTR and 3' UTR, which are crucial for viral translation and replication. By targeting these regions, the viral life cycle can be more precisely disrupted while minimizing impact on the host. This method is more specific than traditional antiviral approaches, enabling precise targeting of specific viral regions, resulting in more significant effects and fewer potential side effects and harms. Therefore, this embodiment not only provides a novel direction for viral infection treatment but also opens up broader prospects for early intervention and precision treatment of diseases. This is of great significance for precisely targeting key regulatory regions of the virus, reducing the possibility of viral mutation escape, and improving the specificity, efficiency, and safety of antiviral therapy.

[0113] Example 2: Embodiment 2 of this invention provides a CRISPR-Cas13d gRNA design system based on the dynamic intracellular environment. By incorporating intracellular environmental information features into the gRNA design model, the accuracy of model predictions is significantly improved, and cell specificity can be achieved for the first time. Specifically, it includes: The data acquisition module is configured to acquire secondary structure information of target RNA and distribution of RBP binding sites within the cell line; The model prediction module is configured to build a gRNA design model based on a deep learning algorithm. The gRNA design model is used to learn the gRNA recognition process, the RNA structure of the target site, and the relationship between the RBP binding process based on the secondary structure information of the target RNA in the cell line and the distribution of RBP binding sites, so as to obtain the gRNA efficiency prediction results. The evaluation adjustment module is configured to introduce an efficiency difference threshold to evaluate the gRNA efficiency prediction results.

[0114] Example 3: Embodiment 3 of the present invention provides the application of the CRISPR-Cas13d gRNA design system based on the dynamic intracellular environment of Embodiment 2 in different cellular environments.

[0115] The steps and methods involved in Examples 2 and 3 above correspond to those in Example 1. For specific implementation details, please refer to the relevant description section of Example 1.

[0116] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed in this application can be implemented in electronic hardware or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application. In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. A computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the flow or function according to the embodiments of this application is generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in or transmitted through a computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that a computer can access or a data processing device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk (SSD)). The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A CRISPR-Cas13d gRNA design method based on the dynamic intracellular environment, characterized in that, Includes the following steps: To obtain secondary structure information of target RNA and distribution of RBP binding sites within cell lines; A gRNA design model was constructed based on deep learning algorithms. The gRNA design model was used to learn the gRNA recognition process, the RNA structure of the target site, and the relationship between the RBP binding process based on the secondary structure information of the target RNA in the cell line and the distribution of RBP binding sites, so as to obtain the gRNA efficiency prediction results. An efficiency difference threshold was introduced to evaluate the gRNA efficiency prediction results.

2. The CRISPR-Cas13d gRNA design method based on the dynamic intracellular environment as described in claim 1, characterized in that, The secondary structure information of the target RNA in the cell line was obtained by the optimized icSHAPE experiment, and the distribution of intracellular RBP binding sites was predicted by PrismNet.

3. The CRISPR-Cas13d gRNA design method based on the dynamic intracellular environment as described in claim 1, characterized in that, The gRNA design model includes a Transformer-based sequence processing network, a feature fusion network for integrating additional features, and a regression network for predicting efficiency.

4. The CRISPR-Cas13d gRNA design method based on the intracellular dynamic environment as described in claim 3, characterized in that, The specific steps for using a gRNA design model to learn the gRNA recognition process, the RNA structure of the target site, and the relationship between the RBP binding process based on the secondary structure information of the target RNA and the distribution of RBP binding sites in the cell line are as follows: A dataset was constructed using RNA secondary structure information from different cell lines, binding information of different RBPs within cells, and gRNA editing efficiency data. Sequence features were obtained by encoding and extracting gRNA and target RNA sequences using a sequence processing network. By fusing sequence features and data from the dataset, integrated features are obtained. Regression networks were used to predict gRNA efficiency based on the integrated features.

5. The CRISPR-Cas13d gRNA design method based on the intracellular dynamic environment as described in claim 4, characterized in that, The target RNA sequence is derived from the sequence at the corresponding target location on the transcriptome and is truncated to the same length as the gRNA sequence. The gRNA sequence comes from large-scale CRISPR screen data that has been collected and uniformly processed. The gRNA sequence and the target RNA sequence are inversely complementary.

6. The CRISPR-Cas13d gRNA design method based on the intracellular dynamic environment as described in claim 4, characterized in that, The specific steps for encoding and extracting sequence features from gRNA and target RNA sequences using a sequence processing network are as follows: First, the gRNA sequence is preprocessed; Next, predict the secondary structure information of the gRNA; Subsequently, the gRNA sequence, target RNA sequence, and secondary structure of the gRNA were one-hot encoded and converted into binary information; Finally, the BERT pre-trained model was used to process the gRNA sequence, target RNA sequence, and secondary structure of gRNA to obtain sequence features.

7. The CRISPR gRNA design method based on the dynamic intracellular environment as described in claim 6, characterized in that, The specific steps for processing gRNA sequences, target RNA sequences, and gRNA secondary structures using the BERT pre-trained model are as follows: The target RNA sequence was 3-mer encoded using a BERT pre-trained model to obtain a deep feature representation, which was then dimensionality reduced. Subsequently, the one-hot encoded features of the gRNA sequence, target RNA sequence, and gRNA secondary structure were embedded and fused through a convolutional layer.

8. The CRISPR-Cas13d gRNA design method based on the dynamic intracellular environment as described in claim 5, characterized in that, The specific steps for evaluating gRNA efficiency prediction results by introducing an efficiency difference threshold are as follows: Proliferation screening experiments were conducted based on gRNA efficiency prediction results; An efficiency difference threshold was introduced to evaluate the results of the proliferation screening experiment.

9. A CRISPR-Cas13d gRNA design system based on the dynamic intracellular environment, characterized in that, include: The data acquisition module is configured to acquire secondary structure information of target RNA and distribution of RBP binding sites within the cell line; The model prediction module is configured to build a gRNA design model based on a deep learning algorithm. The gRNA design model is used to learn the gRNA recognition process, the RNA structure of the target site, and the relationship between the RBP binding process based on the secondary structure information of the target RNA in the cell line and the distribution of RBP binding sites, so as to obtain the gRNA efficiency prediction results. The evaluation adjustment module is configured to introduce an efficiency difference threshold to evaluate the gRNA efficiency prediction results.

10. The application of the CRISPR-Cas13d gRNA design system based on the intracellular dynamic environment as described in claim 9 in different cellular environments.