Method for identifying effectiveness of cas target points by combining gene large model and ml

By combining large-scale gene models with machine learning to identify the effectiveness of Cas targets, the shortcomings in assessing the sequence recognition ability of Cas proteins in gene fragments were addressed, thereby improving the accuracy and efficiency of gene editing and nucleic acid detection.

CN120636549BActive Publication Date: 2025-11-11ZHEJIANG LAB
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511131013.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-13
Publication Date
2025-11-11
Estimated Expiration
2045-08-13

AI Technical Summary

Technical Problem

Current technologies lack effective methods for assessing and predicting the ability of Cas proteins to recognize different sequences in gene fragments, which affects the accuracy of gene editing and nucleic acid detection.

Method used

A method combining large gene model and machine learning (ML) was used to identify the effectiveness of Cas targets. By constructing a detection array and incubation reaction, the evo2 deep learning model and Lasso regression were used to extract features from fluorescent probes and predict the recognition ability of Cas proteins.

Benefits of technology

This improves the accuracy of Cas protein in recognizing target sequences, providing more accurate target selection for gene editing and nucleic acid detection, and enhancing the reliability and efficiency of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120636549B_ABST
    Figure CN120636549B_ABST
Patent Text Reader

Abstract

This invention discloses a method for identifying the effectiveness of Cas targets using a large gene model combined with machine learning (ML), belonging to the fields of biotechnology and artificial intelligence. The method includes: constructing a detection array of A types of joint sequences and target joint sequences, the detection array having A detection units; adding one joint sequence to each detection unit, performing an incubation reaction, acquiring the fluorescence value of each detection unit and normalizing it; extracting key features from each joint sequence using a feature extraction model; constructing a tag set and a training set, constructing an ensemble model and training it using the tag set and training set; combining the feature extraction model and the trained ensemble model to form a machine learning model; for new target sequences and PAM sequences, combining them into a joint sequence to be analyzed, inputting it into the machine learning model to obtain a predicted fluorescence value, and judging the Cas protein's ability to recognize the joint sequence based on the fluorescence value. This invention can assist in targeted therapy detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the fields of biotechnology and artificial intelligence, specifically relating to a method for jointly identifying the effectiveness of Cas targets using a large gene model and machine learning. Background Technology

[0002] CRISPR-Cas (Clustered Regularly Interspaced Short Palindromic Repeats and CRISPR-associated proteins) is a natural adaptive immune system in bacteria and archaea, used to defend against the invasion of viruses (bacteriophages) or foreign nucleic acids. In recent years, this system has been engineered into highly efficient gene-editing tools (such as CRISPR-Cas12), widely used in biological research, medicine, and agriculture. Cas proteins (CRISPR-associated proteins) are key functional proteins in the CRISPR system, responsible for recognizing and cleaving foreign nucleic acids. crRNA binds to Cas proteins (such as Cas12) to form a CRISPR-Cas complex, which specifically recognizes and binds to complementary foreign DNA. The target DNA is then cleaved by the nuclease activity of the Cas proteins, thus providing immune defense. The CRISPR-Cas system's recognition and cleavage of target DNA relies on a specific short sequence—the protospaceradjacent motif (PAM). PAM is a necessary condition for Cas protein cleavage. Cas proteins (such as Cas9 and Cas12) must first recognize the PAM sequence on the target DNA before initiating subsequent DNA unwinding and crRNA complementarity verification. Furthermore, PAM helps the CRISPR system distinguish between exogenous DNA (containing PAM) and host DNA (which does not contain PAM in the CRISPR array), preventing the host genome from being mistakenly cleaved.

[0003] The CRISPR / Cas system includes several gene editing systems such as CRISPR / Cas9, CRISPR / Cas12a, and CRISPR / Cas13a. Among them, CRISPR-Cas12a is a Class 2 RNA-guided endonuclease that can be used to edit the mammalian genome. In September 2015, Zhang Feng's team (Zetsche B, Gootenberg JS, Abudayyeh OO, et al. Cpf1 is a single RNA-guided endonuclease of a class 2 CRISPR-Cas system. Cell, 2015, 163(3): 759-771.) identified a new Cas protein—Cas12a, which belongs to Class 2 Type V of the CRISPR-Cas system. The CRISPR / Cas12a system belongs to Class 2 Type V, is small in size, and has the ability to be both an RNA endonuclease and a DNA endonuclease. Cas12a specifically recognizes and cleaves complementary double-stranded DNA (dsDNA) (B, Z., et al., Cpf1 is a single RNA-guided endonuclease of a class 2 CRISPR-Cas system. Cell, 2015, 163(3): p.759-71.), and when the CRISPR / Cas12a protein recognizes and cleaves target double-stranded DNA in a sequence-specific manner, it can induce trans-cleavage activity of non-specific single-stranded DNA (ssDNA) (JS, C., et al., CRISPR-Cpf1 target binding unleashes indiscriminate single-strandedDNase activity. Science (New York, NY), 2018, 360(6387): p.436-439.). Cas12a can recognize thymine (T)-rich PAM sequences, expanding the editable range of the CRISPR / Cas system. The PAM of Cas12a is located upstream of the target sequence (5' end) and is usually a short motif of 2-6 bp, such as the PAM of LbCas12a, which is TTTV.

[0004] The Cas12a protein initiates target recognition by recognizing a PAM sequence upstream of the target DNA sequence. Upon PAM recognition, Cas12a unwinds the target DNA double helix, forming a triple-stranded R-loop structure. In this structure, the target strand (TS) of the target DNA hybridizes with the crRNA guide, while the non-target strand (NTS) becomes single-stranded. R-loop formation is a crucial step in Cas12a's target recognition, allowing it to further confirm the target DNA's specificity. Once the R-loop is formed and Cas12a has confirmed target specificity, the RuvC domain of Cas12a is activated, cleaving both strands of the target DNA. Furthermore, the flexibility of the REC2 domain plays a key role in Cas12a's DNA targeting and nuclease activation. It allows for the dynamic formation and extension of the R-loop, while simultaneously enhancing Cas12a's DNA targeting specificity through delayed contact formation.

[0005] The length, GC content, and sequence specificity of the target sequence affect the binding stability of the Cas protein. Generally, longer target sequences may provide more complementary pairing opportunities, thereby enhancing the binding ability of the Cas protein. Besides target length, the target sequence itself also affects the binding ability of the Cas protein, such as the GC content within the sequence. GC content affects DNA stability and the binding ability of the Cas protein. Both excessively high and low GC content can lead to a decrease in the binding ability of the Cas protein to the target DNA. For example, some studies have shown that excessively high or low GC content in Cas9 leads to reduced sgRNA editing activity (Doench, J., Hartenian, E., Graham, D. et al. Rational design of highly active sgRNAs for CRISPR-Cas9–mediated gene inactivation. Nat Biotechnol 32,1262–1267 (2014).). Cas12a exhibits some sequence specificity in the binding and cleavage of target DNA. Although Cas12a is relatively tolerant of GC content, specific sequence patterns may affect its binding efficiency. Certain DNA sequences may affect the binding of Cas12a to target DNA due to their secondary structure. For the Cas12a protein, the overall impact of the PAM and crRNA complementary binding regions on protein recognition ability has not yet been evaluated. Therefore, establishing a method to assess and predict the recognition ability of Cas protein at different sequences within gene fragments is of great significance for the application of Cas in gene editing and nucleic acid detection. Summary of the Invention

[0006] To address the problems in the prior art, this invention provides a method for jointly identifying the effectiveness of Cas targets using a large gene model and machine learning.

[0007] The technical solution of the present invention is as follows:

[0008] In a first aspect, this invention discloses a method for jointly identifying the effectiveness of Cas targets using a large gene model and machine learning, the method comprising the following steps:

[0009] 1) Select N Cas enzyme target binding regions, and bind the nucleotide sequences complementary to crRNA at their binding sites to M different combinations of PAM sequences to form N×M combined sequences;

[0010] 2) Construct a detection array targeting the combined sequence, wherein the detection array has N×M detection units, and each detection unit includes Cas protein, crRNA and fluorescent probe;

[0011] 3) Add a combination sequence to each detection unit and perform an incubation reaction; within the detection unit, the Cas protein, crRNA, and combination sequence form a ternary complex, which cleaves the fluorescent probe to emit fluorescence. Obtain the fluorescence value of each detection unit and normalize it.

[0012] 4) Key features are extracted from each joint sequence using a feature extraction model; specifically: each joint sequence is input into the evo2 deep learning model to obtain multidimensional features characterizing the biological properties of the joint sequence; the multidimensional features are standardized and PCA is used for dimensionality reduction; then Lasso regression is used to filter the dimensionality-reduced features to obtain key features;

[0013] 5) Construct a label set based on the joint sequence and its corresponding normalized fluorescence value; construct a training set based on the joint sequence and its corresponding key features; construct an ensemble model and train it using the label set and training set; combine the feature extraction model with the trained ensemble model to form a machine learning model;

[0014] 6) For new target sequences and PAM sequences, combine them into a combined sequence to be analyzed, and then input it into the machine learning model obtained in step 5) to obtain the predicted fluorescence value. Based on the predicted fluorescence value, determine the ability of the Cas protein to recognize the combined sequence.

[0015] In an optional embodiment of the present invention, the Cas enzyme targeting region is selected from the Cas enzyme targeting regions of target genes such as EMX1, DNMT1, MerS, eGFP-1, eGFP-3, and FANCF. These target genes are all genes with sequences known in the art and are commonly used to study the effectiveness of Cas proteins against targets.

[0016] Furthermore, the PAM sequence has a number of bases of 'a', each base position in the PAM sequence is any one of A / T / C / G, and the number of PAM sequence combinations M=4. a The selection of a is usually related to the Cas protein, and an adaptive selection of a can be made based on existing knowledge in the field.

[0017] Furthermore, the Cas protein is either Cas12a or Cas9. When the Cas protein is Cas12a, the number of bases in the PAM sequence corresponding to the Cas12a protein can be 4, and the number of combinations of PAM sequences is 256. When the Cas protein is Cas9, the number of bases in the PAM sequence corresponding to the Cas9 protein can be 3, and the number of combinations of PAM sequences is 96.

[0018] More preferably, the Cas protein in the incubation system in step 3) can be Cas12a protein. In the incubation system, the concentration of Cas12a is 10 ng / μL, the concentration of the fluorescent probe is 25 pM, the concentration of crRNA is 1 μM, and the volume ratio of the combined sequence to the system is 1:20. The incubation temperature is 37°C. The fluorescent probe is a single-stranded DNA reporter gene labeled with FAM and BHQ1.

[0019] Secondly, the present invention also discloses an apparatus for jointly identifying the effectiveness of Cas targets using a large gene model and ML, comprising a memory and one or more processors, wherein the memory stores executable code, and the one or more processors execute the executable code to implement the method for jointly identifying the effectiveness of Cas targets using a large gene model and ML.

[0020] Thirdly, the present invention also discloses a computer-readable storage medium having a program stored thereon, which, when executed by a processor, implements the method for jointly identifying the effectiveness of Cas targets using a large gene model and ML.

[0021] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0022] This invention, based on the in vitro paracleaning ability of the Cas protein, utilizes N complementary crRNA sequences and N×M combined sequences, each containing M PAM sequences, to generate N×M fluorescence values ​​of the Cas protein under different conditions. This invention employs an advanced evo2 deep learning model to characterize the combined sequences and combines machine learning to evaluate the Cas protein's ability to recognize different combined sequences. A key feature of this invention is that it groups the target regions bound to PAM and crRNA together, providing a holistic analysis of the Cas protein's effectiveness against the entire sequence. This invention provides a valuable reference for selecting suitable target sequences for the Cas protein, which is significant for gene editing and Cas enzyme-based nucleic acid detection. Attached Figure Description

[0023] Figure 1 This is a fluorescence thermogram of the Cas12a protease PAM sequence of MerS in an embodiment of the present invention.

[0024] Figure 2 This is a thermogram of the fluorescence value of the Cas12a protease PAM sequence of eGFP-1 in an embodiment of the present invention.

[0025] Figure 3 This is a thermogram of the fluorescence value of the Cas12a protease PAM sequence of eGFP-3 in an embodiment of the present invention.

[0026] Figure 4 This is a fluorescence thermogram of the Cas12a protease PAM sequence of FANCF in an embodiment of the present invention.

[0027] Figure 5 This is a fluorescence thermogram of the PAM sequence of the Cas12a protease of DNMT1 in an embodiment of the present invention.

[0028] Figure 6 This is a fluorescence thermogram of the Cas12a protease PAM sequence of EMX in an embodiment of the present invention;

[0029] Figure 7 This is a PCA dimensionality reduction result diagram of 1536 vectors with a length of 4096 after Evo2 characterization in this embodiment of the invention;

[0030] Figure 8 This is a graph showing the variance explanation ratio of the first S principal components in this embodiment of the invention.

[0031] Figure 9 This is a comparison chart of MSE and R2 indices for different regression methods;

[0032] Figure 10 It compares the predicted values ​​with the actual values ​​from different regression methods;

[0033] Figure 11 This is a comparison chart of the predicted and actual values ​​of the integrated model of this invention;

[0034] Figure 12 This is a graph showing the detection results of Cas12a protein at the G12C site;

[0035] Figure 13 This is an overall flowchart of the solution of the present invention. Detailed Implementation

[0036] The present invention will be further described and illustrated below with reference to specific embodiments. The embodiments described are merely examples of the content of this disclosure and do not limit the scope of the invention. The technical features of each embodiment in the present invention can be combined accordingly, provided that there is no mutual conflict.

[0037] To establish a method for assessing and predicting the recognition ability of Cas proteins at different sequences in gene fragments, this invention proposes a method for jointly identifying the effectiveness of Cas targets using a large gene model and machine learning.

[0038] like Figure 13 As shown, the method for jointly identifying the effectiveness of Cas targets using a large gene model and ML according to the present invention includes the following steps:

[0039] 1) Select N Cas enzyme target binding regions, and bind the nucleotide sequences complementary to crRNA on the binding regions to M different combinations of PAM sequences to form N×M combined sequences;

[0040] 2) Construct a detection array targeting the combined sequence, wherein the detection array has N×M detection units, and each detection unit includes Cas protein, crRNA and fluorescent probe;

[0041] 3) Add a combination sequence to each detection unit and perform an incubation reaction; within the detection unit, the Cas protein, crRNA, and combination sequence form a ternary complex, which cleaves the fluorescent probe to emit fluorescence. Obtain the fluorescence value of each detection unit and normalize it.

[0042] 4) Key features are extracted from each joint sequence using a feature extraction model; specifically: each joint sequence is input into the evo2 deep learning model to obtain multidimensional features characterizing the biological properties of the joint sequence; the multidimensional features are standardized and reduced in dimensionality using PCA; then Lasso regression is used to filter the reduced-dimensional features to obtain key features; wherein, the feature extraction model includes the evo2 deep learning model, a standardization module, a PCA module, and a Lasso regression model;

[0043] 5) Construct a label set based on the joint sequence and its corresponding normalized fluorescence value; construct a training set based on the joint sequence and its corresponding key features; construct an ensemble model and train it using the label set and training set; combine the feature extraction model with the trained ensemble model to form a machine learning model;

[0044] 6) For new target sequences and PAM sequences, combine them into a combined sequence to be analyzed, and then input it into the machine learning model obtained in step 5) to obtain predicted fluorescence values. Based on the predicted fluorescence values, determine the Cas protein's ability to recognize the combined sequence. The larger the predicted fluorescence value, the stronger the Cas protein's ability to recognize the combined sequence; the smaller the predicted fluorescence value, the weaker the Cas protein's ability to recognize the combined sequence.

[0045] The following examples, in conjunction with the accompanying drawings, further illustrate the method of using a large gene model and ML to jointly identify the effectiveness of Cas targets according to the present invention.

[0046] 1. Construction of experimental dataset

[0047] Utilizing the trans-cleavage activity of the Cas12a protein, this invention uses the fluorescence values ​​obtained from in vitro detection of different combined sequences of the Cas12a protein as data on its ability to recognize target sequences and PAM sequences.

[0048] First, target genes are selected. In this embodiment, the selected target genes are EMX1, DNMT1, MerS, eGFP-1, eGFP-3, and FANCF. Then, the target sequences at the target sites of the target genes are obtained. In this embodiment, the target sequence length is 23 nt. Next, a PAM sequence is constructed. In this embodiment, the PAM sequence contains 4 bases, each at any one of the bases A, T, C, or G. Therefore, the total number of PAM sequence combinations is 4. 4 This results in 256 different combinations of PAM sequences. Finally, the six target sequences are combined with the 256 different combinations of PAM sequences to form 6 × 256 (1536) combined sequences.

[0049] Each target sequence incorporates 256 different combinations of PAM sequences through gene mutation, resulting in combined sequences containing different PAM combinations (each target sequence has 256 different PAM detection samples). Since the PAM sequence contains 4 bases, the 256 PAM combinations cover all possible PAM combinations.

[0050] A detection array targeting the combined sequence was constructed, comprising 1536 detection units. In this embodiment, the detection units include: Cas12a protein, crRNA, and a fluorescent detection probe. The crRNA comprises a 21-nt repeat region interacting with Cas12a (a repeat sequence module, sequence (SEQ ID No. 19) UaaUUUcUacUaagUgUagaU) and a 23-nt spacer sequence complementary to the target DNA. The fluorescent detection probe is a single-stranded DNA reporter gene (5'-FAM-TTATT-BHQ1-3') labeled with FAM and BHQ1.

[0051] This application employs PCR annealing to generate Cas enzyme-bound dsDNA fragments. Using the EMX1 fragment as the target gene, the initial concentrations of the forward and reverse primers were 100 μM. 10 μL of the forward primer and 10 μL of the reverse primer were used to form a 20 μL annealing system. The entire annealing system was placed in a PCR instrument and incubated at 95°C for 2 minutes, followed by natural cooling to room temperature. The primers for the six target genes were processed in the same manner. After completing 256 reactions, 256 short-stranded DNA fragments containing both PAM and target sequences were obtained for testing. Then, based on the molar amount and concentration of the double strands, the fragments were diluted with water to a concentration of 5 × 10⁻⁶ mcg / μL. 19 Each molecule was used for subsequent experiments. The dsDNA fragments obtained by PCR annealing were all 55 bases in length. Table 1 lists the forward and reverse primer information for each target sequence, where N represents any one of the bases A / T / C / G.

[0052] Table 1. Information on forward and reverse primers

[0053]

[0054] This invention designs crRNAs targeting six different sequences, the sequences of which are shown in Table 2. The Cas12a protein is obtained as follows: the AsCas12a gene is cloned into the pET28a plasmid, expressed in E. coli, and purified for identification experiments.

[0055] Table 2 crRNA sequences

[0056]

[0057] The detection principle is as follows: When the conjugate sequence is added to the detection unit, the Cas12a protein forms a ternary complex with crRNA and the conjugate sequence. This ternary complex then exercises its trans-cleavage activity, cutting the fluorescent probe (a fluorescently labeled single-stranded DNA with a luminescent group and a quencher group at each end; after cleavage, the luminescent group is excited and emits light), thus producing fluorescence. The specificity of the Cas12a protein in recognizing the target sequence + PAM sequence is determined by the fluorescence intensity of each sample in the detection array.

[0058] The specific detection procedure was as follows: The experimental sample system consisted of 200 ng purified AsCas12a, 25 pM fluorescent probe, 1 μM crRNA, and 1 μL of binding sequence, with a final volume of 20 μL. The reaction was incubated at 37°C. In the detection system, once the Cas12a protein was activated by the target PAM sequence, the fluorescent probe ssDNA-FQ was cleaved. The fluorescence value of the detection system was measured using a full-wavelength microplate reader. The excitation wavelength of the full-wavelength microplate reader was 485 nm, and the emission wavelength was 520 nm. The signal was recorded every 30 seconds for 15 minutes.

[0059] A total of 1536 (256×6) reactions were conducted in the experiment, resulting in 1536 fluorescence values, which formed the dataset required for the experiment. The fluorescence values ​​of 256 samples containing different PAM sequences for each target sequence were normalized to obtain the PAM heatmap of the AsCas12a protein under a single target sequence. The results are as follows: Figures 1 to 6 As shown in the results, the Cas12a protein exhibits significant differences in its preference for different PAM sequences within the combined sequence. The detection of six target sequences can avoid the influence of target specificity on the detection results. In each heatmap, the higher the fluorescence value, the stronger the Cas12a protein's ability to recognize the corresponding PAM sequence or the entire PAM and target region at that location. Conversely, the lower the fluorescence value in the heatmap, the weaker the Cas12a protein's ability to recognize the corresponding PAM sequence. The PAM sequence or PAM sequence plus target region corresponding to that location is not suitable for editing regions of the Cas12a protein.

[0060] 2. Construction of a joint sequence sample set based on the experimental dataset

[0061] In genetics, a 4nt sequence composed of four basic nucleotides (A adenine, T thymine, G guanine, and C cytosine; for simplicity, one nucleotide is defined as 1 nt) has 256 possible combinations. This is because each position can be one of the four nucleotides, so the total number of combinations is 4 × 4 × 4 × 4 = 256. These different sequences encode genetic information in DNA and RNA, determining the traits and functions of organisms. Each of these 256 combinations is considered a potential PAM sequence and is defined as... , .

[0062] The initial step in Cas12a protein recognition of target DNA is the recognition of the neighboring motif (PAM), a necessary condition for Cas12a protein to bind to and cleave target DNA. After recognizing the PAM, Cas12a protein unwinds the target DNA double helix, forming a triple-stranded R-loop structure. In this structure, the target strand (TS) of the target DNA hybridizes with the crRNA guide, while the non-target strand (NTS) becomes single-stranded. Once the R-loop is formed and Cas12a protein confirms target specificity, the RuvC domain of Cas12a protein is activated, leading to cleavage of both strands of the target DNA. Cleavage typically begins with the non-target strand, followed by the target strand. The binding affinity of Cas protein to its target is influenced by various factors, with the PAM sequence being a key factor in Cas protein recognition and binding to target DNA. However, the target sequence also affects Cas protein binding affinity; for example, the GC content in the target sequence affects DNA stability and Cas protein binding affinity. Both excessively high and low GC content can reduce the binding affinity of Cas protein to target DNA. The PAM sequence and target sequence jointly influence the binding affinity of Cas12a protein, but the influence of the target sequence has been largely ignored in previous target site selection. This invention considers the PAM sequence and target sequence together as non-single influencing factors to evaluate the affinity of Cas12a protein. This comprehensive analysis better reflects the true target selectivity and provides useful guidance for selecting appropriate target regions in gene editing and nucleic acid detection scenarios.

[0063] Six 23nt target sequences from human, fluorescent protein, and viral gene target sites (MERS, EGFP1, EGFP3, DNMT1, EMX1, and FANCF) were combined with 256 different combinations of PAM sequences, resulting in a total of 256 × 6 = 1536 combinations. Specifically, when the target gene is MERS, the target sequence for EMX1 (SEQ ID No. 20) is GACATATGGAAAACGAACTATGT; when the target gene is EGFP1, the target sequence for EGFP1 (SEQ ID No. 21) is CGTCGCCGTCCAGCTCGACCAGG; when the target gene is EGFP3, the target sequence for EGFP3 (SEQ ID No. 22) is CTCAGGGCGGACTGGGTGCTCAG; and when the target gene is DNMT1, the target sequence for DNMT1 (SEQ ID No. 22) is... No. 23) is CTGAGTGGTCCATGTCTGTTACTC; when the target gene is EMX1, the target sequence of EMX1 (SEQ ID No. 24) is TCATCTGTGCCCCCCTCCCTCCCTG; when the target gene is FANCF, the target sequence of FANCF (SEQ ID No. 25) is ACCTTGGAGACGGCGACTCTCTG.

[0064] Because the actual fluorescence values ​​of the detection system are relatively large (e.g., the fluorescence value of the MERS combined sequence with PAM sequence TTTG (target gene MERS) is 681.04, the fluorescence value of the MERS combined sequence with PAM sequence TTTA (target gene MERS) is 1189.31, the fluorescence value of the FANCF combined sequence with PAM sequence GGCG (target gene FANCF) is 1223.33, and the fluorescence value of the FANCF combined sequence with PAM sequence GGCA (target gene FANCF) is 3860.79), it is not conducive to subsequent machine learning processing. Therefore, the maximum-minimum normalization method (a linear transformation) is used to map the actual fluorescence values ​​to the [0, 1] interval to achieve scaling, which is more conducive to model convergence and improves performance and generalization ability. After completing the fluorescence value normalization, a set of 1536 tags, with normalized fluorescence values ​​as labels, is constructed for potential PAM sequences combined with 6 target sequences.

[0065] In this embodiment, the normalized fluorescence value of the MERS combined sequence with PAM sequence TTTG (target gene MERS) is 0.0564, the normalized fluorescence value of the MERS combined sequence with PAM sequence TTTA (target gene MERS) is 0.0987, the normalized fluorescence value of the FANCF combined sequence with PAM sequence GGCG (target gene FANCF) is 0.1015, and the normalized fluorescence value of the FANCF combined sequence with PAM sequence GGCA (target gene FANCF) is 0.3207.

[0066] The formula for the maximum-minimum normalization method is shown below:

[0067] ;

[0068] in, The normalized fluorescence value; This is the original fluorescence value; This is the original minimum fluorescence value; This represents the original maximum fluorescence value.

[0069] 3. Nucleic acid sequence characterization based on the Evo2 large-scale model

[0070] The previous invention constructed a set of 1536 tags, each consisting of a potential PAM sequence and six target sequences, with normalized fluorescence values ​​as labels. Each sample can be viewed as a joint sequence composed of 27 nucleotides. However, machine learning cannot directly process nucleotide sequences and needs to convert them into numerical features. Traditional methods typically use one-hot encoding or statistical encoding methods (such as TF-IDF) to directly process the original nucleic acid sequence. However, due to the high complexity and variability of nucleic acid sequences, traditional methods have significant disadvantages, such as potentially leading to low computational efficiency, insufficient feature extraction, and poor model generalization ability. Furthermore, the high dimensionality and complex structure of the original sequence may also lead to overfitting, affecting the stability and prediction accuracy of the model.

[0071] The evo2 deep learning model is specifically designed for processing biological sequence data in this field. It can transform complex sequences into fixed-length vectors by learning the evolutionary and structural information of large amounts of nucleic acid sequences. These vectors retain the key features of the original sequence and enhance the ability to distinguish subtle differences between sequences, achieving efficient and accurate representation. Compared to directly processing raw nucleic acid sequences, the 4096-dimensional vectors generated by the evo2 deep learning model offer several advantages. Fixed-length vectors simplify subsequent data processing and improve compatibility with other machine learning or deep learning models. Simultaneously, this representation method effectively reduces data complexity, improves computational efficiency, and enhances the model's ability to capture sequence features. Furthermore, the output vectors of the evo2 deep learning model can more accurately capture the evolutionary and structural information of the sequence, thereby enhancing the accuracy and reliability of subsequent analysis.

[0072] This invention employs the advanced Evo2 deep learning model (Genome modeling and design across all domains of life with Evo 2. bioRxiv. Preprint posted online February 21, 2025. doi: 10.1101 / 2025.02.18.638918.) to characterize 27nt target DNA, generating a fixed-length vector of 4096. The 4096 dimensions of this vector represent the nucleotide sequence, encompassing a wide range of biological significance. Different dimensions may correspond to k-mer frequencies (e.g., 3-mer and 4-mer combination preferences), conserved motifs (e.g., transcription factor binding sites), or secondary structure tendencies (e.g., stem-loop, hairpin structures). Some dimensions may encode physicochemical properties such as electrostatic potential, hydrogen bonding ability, and hydrophobicity, affecting protein-DNA interactions (e.g., the binding strength of Cas12a to target DNA). In this way, this invention efficiently processes and analyzes nucleic acid sequence data, laying a solid foundation for bioinformatics research and applications.

[0073] In actual processing, the potential "Pi (4nt)" and "target sequence (23nt)" are directly spliced ​​together to form 1536 joint sequences of 27 nucleotides each. These joint sequences are then fed into the evo2 deep learning model, which outputs vectors of length 4096. These vectors represent the nucleotide sequences and encompass a wide range of biological significance.

[0074] The 1536 vectors of length 4096 represented by the Evo2 deep learning model were reduced to two dimensions (principal component 1 and principal component 2) using principal component analysis (PCA), and then plotted on a 2D image. The result is as follows. Figure 7 As shown. From Figure 7 It can be seen that there are relatively clear boundaries between the 256 feature points corresponding to the six target genes, which proves the effectiveness of the representation using the Evo2 deep learning model.

[0075] 4. Construction of the training sample set based on the representation of the Evo2 deep learning model

[0076] 4.1 Division of Training Set and Label Set

[0077] After representing the Evo2 deep learning model, a training set of 1536 samples was constructed, using the representation results of the Evo2 deep learning model as features, combining latent PAM sequences with six target sequences. All samples were used... express.

[0078] In step 2, after normalizing the fluorescence values, a set of 1536 tags, each using normalized fluorescence values ​​as labels, was constructed to bind potential PAM sequences to six target sequences. The tag set was then used... express.

[0079] 4.2 Training Set Feature Standardization

[0080] Because the numerical values ​​of the 4096-dimensional features vary to some extent, without standardization, features with larger numerical ranges may dominate during model training, thus reducing the model's sensitivity to features with smaller numerical ranges and ultimately affecting the model's performance and generalization ability. This is particularly important because many machine learning algorithms (such as distance-based algorithms and gradient descent optimization algorithms) are sensitive to feature scale.

[0081] The purpose of standardization is to scale data with different features to a relatively uniform scale, thereby eliminating bias caused by differences in magnitude between features and ensuring that each feature contributes equally to the model. The specific formula for standardization is as follows:

[0082] ;

[0083] in, Represents the original features of the i-th sample; This represents the standardized features of the i-th sample; and Let represent the mean and standard deviation of the original features of the i-th sample, respectively.

[0084] Calculate the mean and standard deviation for each sample using this formula, then subtract the mean from each element of the original feature of that sample and divide by the standard deviation. This ensures that the feature values ​​of each sample have zero mean and unit variance, thus completing the standardization of each sample.

[0085] 4.3 Standardized Training Set Features PCA Dimensionality Reduction

[0086] Because the number of features in the standardized training set of this invention (4096) is much greater than the number of samples (1536), distance calculation and pattern recognition between data points become difficult in high-dimensional space, and the model is prone to overfitting. Furthermore, high-dimensional data may contain a large number of features that are irrelevant to or weakly correlated with the target variable, which can interfere with model learning. For many machine learning algorithms (such as distance-based algorithms, support vector machines, etc.), high-dimensional features significantly increase the computational burden. Dimensionality reduction using PCA projects data into a low-dimensional space, preserving key information while reducing data complexity. This effectively removes noise and redundant information, highlighting the main sources of variation in the data, thereby enabling the model to have better generalization performance and computational efficiency.

[0087] This invention uses a 95% variance proportion to determine the principal components in PCA. During PCA dimensionality reduction, the variance explanation proportion for each principal component represents the proportion of information contained in that principal component relative to the total information. By calculating the cumulative variance explanation proportion, the percentage of information contained in the sum of the first S principal components relative to the total information in the entire dataset can be determined. In actual calculations, the cumulative variance explanation proportion is made to reach approximately 95%. This preserves most of the data's information while significantly reducing its dimensionality.

[0088] Assume there are a total of K principal component eigenvalues ​​(already arranged in descending order): Then the cumulative variance explained by the first S principal components (using W) S The calculation formula (represented by:) is:

[0089] ;

[0090] when The value of S is the final number of principal components. Figure 8 The graph shows the variance explanation ratio of the cumulative top S principal components in this invention. As can be seen from the graph, the final determined... The significance of these 5 values ​​is to project the original 4096-dimensional data onto a few orthogonal directions, so that the variance is maximized and the information loss is minimized in these directions.

[0091] 4.4 Training Set Feature Selection Based on Lasso Regression

[0092] In step 4.3, this invention has already compressed and refined the sample feature space using PCA dimensionality reduction. However, there is still a possibility of redundant features. To make the sample features more strongly correlated with the labels, this invention further employs Lasso regression to perform feature selection in the dimensionality-reduced feature space. As a linear regression model, Lasso regression introduces an L1 regularization term, which can compress some feature coefficients to zero, directly eliminating features that contribute little or are irrelevant to the model's predictive ability. This is equivalent to further screening key features based on dimensionality reduction, ensuring that the features used for the final model training are both concise and representative. Specifically, the optimal regularization parameter of the Lasso regression model is determined through cross-validation, and the dimension of key features is determined through the optimal regularization parameter. This step-by-step strategy combines the advantages of global information extraction from PCA and precise feature selection from Lasso, ultimately making the model training and prediction process more efficient and accurate, with stronger generalization ability, fully mining the value of data, and providing strong support for building efficient machine learning models. In this invention, the final key feature has a dimension of 4. As can be seen from the results, one of the features after PCA dimensionality reduction was not selected, indicating that the feature has a high correlation with other features and has no significant impact on the prediction results, so it was not selected.

[0093] 5. Machine Learning-Based Prediction of Fluorescence Values ​​in PAM Combination Sequences

[0094] 5.1 Training and Testing of Single Regression Model

[0095] After the preceding steps, the training set and label set of this invention become as follows:

[0096] ;

[0097] Since the label set in this problem is a quantified numerical value, it is a regression problem. To obtain a more reliable final model, this invention first tested 10 commonly used regression models (linear regression, regularization and variable selection via the elastic net, ridge regression, Lasso regression, classification and regression trees, random forests, gradient boosting regression, support vector regression, Bayesian regression, and XGBoost regression). The main test metrics for different models were MSE and R².

[0098] MSE is the average of the squared differences between the model's predicted values ​​and the actual values. It quantifies the average error between the model's predicted values ​​and the actual values; a smaller value is better. Its definition is:

[0099] ;

[0100] In the formula, n represents the number of samples (n=1536 in this embodiment). and These are the actual label value and the predicted label value, respectively.

[0101] Because MSE is sensitive to large errors, it may overemphasize the impact of outliers. To ensure the effectiveness of the model, this invention further employs the R² metric for model evaluation. This metric measures the model's ability to explain data variation; it represents the proportion of variation explained by the model relative to the total variation. Its value ranges from 0 to 1, with values ​​closer to 1 being better. Its definition is as follows:

[0102] ;

[0103] In the formula, n represents the number of samples (n=1536 in this embodiment). and These are the true label value and the predicted label value. It is the arithmetic mean of all the true label values.

[0104] Figure 9 The results show a comparison of MSE and R2 indices for different regression methods on the training set of this invention. From Figure 9 The graph shows that the MSE (Mean Sequence Size) of all 10 methods is relatively low, but it's also clear that Decision Tree Regression, Random Forest Regression, Gradient Boosting Regression, Support Vector Regression, and XGBoost Regression have better MSE scores than the other five methods. Similarly, looking at the R² (Relative Square of Indicators), Decision Tree Regression, Random Forest Regression, Gradient Boosting Regression, Support Vector Regression, and XGBoost Regression all significantly outperform the other five methods, with values ​​exceeding 0.7. This graph clearly demonstrates that these five methods are superior in performance.

[0105] Figure 10 The results show in detail the performance of 10 methods on the new training set samples of this invention (where blue represents the true value and red represents the predicted value). The results also show that the five methods of decision tree regression, random forest regression, gradient boosting regression, support vector regression and XGBoost regression are better because the two curves overlap more, indicating that the predicted value is closer to the true value.

[0106] 5.2 Regression Model Integration

[0107] In step 5.1, this invention tested 10 regression models and obtained relatively clear conclusions. However, a single model may overfit or underfit, limiting its generalization ability. An ensemble model, by combining predictions from multiple models, reduces the risk of overfitting and improves generalization ability and accuracy. Different models have different sensitivities to data features; ensemble integration can reduce the impact of individual model instability on the overall prediction. Furthermore, each model can capture different features and patterns in the data; ensemble integration can combine this information to form a more comprehensive prediction. Therefore, to improve prediction performance, enhance model robustness, and integrate information from various models, this invention constructs an ensemble model. The ensemble model consists of the five models obtained from the previous tests: decision tree regression, random forest regression, gradient boosting regression, support vector regression, and XGBoost regression. This invention uses the arithmetic mean of the predictions from these five models as the final prediction result. ,Right now:

[0108] ;

[0109] in, The prediction results of the decision tree regression model; The prediction results are from the random forest regression model; The prediction results of the gradient boosting regression model; The prediction results of the support vector regression model; This represents the prediction results from the XGBoost regression model.

[0110] The ensemble model is used to predict all sample sets, and the results are as follows: Figure 11 As shown, RMS = 0.0023 and R² = 0.9435 at this point. Both the quantitative indicators and the graphs show that the overall model performance is quite satisfactory. It should be noted that the yellow curve deviates somewhat from the blue curve; this is precisely the function of the ensemble model. Due to the incompleteness of the training data, a perfect overlap between the predicted and actual results indicates overfitting. A certain degree of deviation reflects the model's stronger generalization ability.

[0111] 6. Integrated models for predicting novel joint sequence fluorescence values

[0112] Cas12a protein, due to its trans-cleavage activity, is often developed for nucleic acid detection, such as virus detection and SNP detection. KRAS-regulated downstream pathways involve cell differentiation, proliferation, and metabolism, including the RAS-RAF-MEK-ERK (RAS-RAF pathway) and PI3K-AKT-mTOR. When KRAS is mutated, it locks into an active GTP-binding state, loses its hydrolytic enzyme activity, leading to dysregulation of the switching molecule and constitutive activation of KRAS and its downstream pathways. This can result in uncontrolled cell proliferation, impaired differentiation, and ultimately tumor formation. Therefore, Kras mutations are common in cancers, especially in non-small cell lung cancer. Kras has various mutation types, with KRas-G12C being the dominant one. However, the TTTV PAM site, used for Cas12a detection, is lacking near the KRas-G12C site. However, the 5'-gagctTgtggcgtaggcaagagt-3' sequence (SEQ ID No. 26) near the G12C site is a potential recognition target. In this invention, this fragment and the preceding PAM sequence are combined to use an ensemble model to predict the identifiability of this region. As can be seen from the results in Table 3, the Cas12a protein has a good ability to recognize this fragment.

[0113] Table 3. Prediction results of each regression model for the new sequence.

[0114]

[0115] KRas G12C mutation is a key driver of various cancers, and detecting this site can guide targeted therapy and determine whether a patient is suitable for KRas G12C inhibitors. Therefore, this invention, based on the prediction that the Cas protein has good recognition ability at the 5'-gttggagctTgtggcgtaggcaagagt-3' position, designed a nucleic acid detection reaction targeting the G12C mutation. The principle and method of the detection reaction are described in the experimental data acquisition instructions. A crRNA targeting the G12C mutation was designed. This crRNA can recognize the G12C sequence, but its recognition ability for the WT sequence is poor due to the presence of the mutated base, thus distinguishing between WT and G12C mutation information. The crRNA sequence (SEQ ID No. 28) is: 5'-UaaUUUcUacUaagUgUagaUgagcUUgUggcgUaggcaagagU-3', the reaction time is 15 min, and the fluorescence values ​​obtained are read. The results are as follows: Figure 12 As shown, the experimental results indicate that the AsCas12a protein can well recognize the 5'-gttggagctTgtggcgtaggcaagagt-3' sequence and can clearly distinguish between WT and G12C mutations.

[0116] In an embodiment of the present invention, an apparatus for jointly identifying the effectiveness of Cas targets using a large gene model and machine learning is also provided, comprising a memory and one or more processors, wherein the memory stores executable code, and the one or more processors execute the executable code to implement the method for jointly identifying the effectiveness of Cas targets using a large gene model and machine learning.

[0117] For the system embodiments, since they basically correspond to the method embodiments, relevant details can be found in the descriptions of the method embodiments; the implementation methods of the remaining modules will not be repeated here. The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of the present invention according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0118] The system embodiments of the present invention can be applied to any device with data processing capabilities, such as a computer or other similar device. The system embodiments can be implemented in software, hardware, or a combination of both. Taking software implementation as an example, as a logical device, it is formed by the processor of any data processing device loading the corresponding computer program instructions from non-volatile memory into memory for execution.

[0119] This invention also provides a computer-readable storage medium storing a program thereon, which, when executed by a processor, implements the method for jointly identifying the effectiveness of Cas targets using a large gene model and ML.

[0120] The computer-readable storage medium can be an internal storage unit of any data processing device described in any of the foregoing embodiments, such as a hard disk or memory. The computer-readable storage medium can also be an external storage device of any data processing device, such as a plug-in hard disk, smart media card (SMC), SD card, flash card, etc., equipped on the device. Furthermore, the computer-readable storage medium can include both internal storage units and external storage devices of any data processing device. The computer-readable storage medium is used to store the computer program and other programs and data required by the data processing device, and can also be used to temporarily store data that has been output or will be output.

[0121] Obviously, the embodiments and accompanying drawings described above are merely some examples of this application. Those skilled in the art can apply this application to other similar situations based on these drawings without any creative effort. Furthermore, it is understood that although the work done in this development process may be complex and lengthy, for those skilled in the art, certain design, manufacturing, or production modifications made based on the technical content disclosed in this application are merely conventional technical means and should not be considered as insufficient disclosure of this application. Several modifications and improvements can be made without departing from the concept of this application, and these all fall within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the appended claims.

Claims

1. A method for jointly identifying the effectiveness of Cas targets using a large gene model and machine learning, characterized in that, The method includes the following steps: 1) Select N Cas enzyme target binding regions, and bind the nucleotide sequences complementary to crRNA at their binding sites to M different combinations of PAM sequences to form N×M combined sequences; 2) Construct a detection array targeting the combined sequence, wherein the detection array has N×M detection units, and each detection unit includes Cas protein, crRNA and fluorescent probe; 3) Add a combination sequence to each detection unit and perform an incubation reaction; within the detection unit, the Cas protein, crRNA, and combination sequence form a ternary complex, which cleaves the fluorescent probe to emit fluorescence. Obtain the fluorescence value of each detection unit and normalize it. 4) Key features are extracted from each joint sequence using a feature extraction model; specifically: each joint sequence is input into the evo2 deep learning model to obtain multidimensional features representing the biological characteristics of the joint sequence; the multidimensional features are standardized and PCA is used for dimensionality reduction; then Lasso regression is used to filter the dimensionality-reduced features to obtain key features; 5) Construct a label set based on the joint sequence and its corresponding normalized fluorescence value; construct a training set based on the joint sequence and its corresponding key features; construct an ensemble model and train it using the label set and training set; combine the feature extraction model with the trained ensemble model to form a machine learning model; 6) For new target sequences and PAM sequences, combine them into a combined sequence to be analyzed, and then input it into the machine learning model obtained in step 5) to obtain the predicted fluorescence value. Based on the predicted fluorescence value, determine the ability of the Cas protein to recognize the combined sequence.

2. The method according to claim 1, characterized in that, The PAM sequence has a number of bases *a*, each base position in the PAM sequence is any one of A / T / C / G, and the number of PAM sequence combinations is M=4. a .

3. The method according to claim 1, characterized in that, The Cas protein mentioned is either Cas12a or Cas9.

4. The method according to claim 1, characterized in that, The fluorescent probe in step 2) is a single-stranded DNA reporter gene labeled with FAM and BHQ1, with the sequence TTATT.

5. The method according to claim 1, characterized in that, The crRNA in step 2) includes a 21nt repeat sequence that interacts with the Cas protein and a 23nt spacer sequence that is complementary to the Cas protease target binding sequence.

6. The method according to claim 1, characterized in that, The feature extraction model includes the evo2 deep learning model, a normalization module, a PCA module, and a Lasso regression model; The evo2 deep learning model yields a multidimensional feature dimension of 4096. A 95% variance ratio is used for PCA dimensionality reduction to determine the features after PCA dimensionality reduction. Cross-validation is used to determine the optimal regularization parameter for the Lasso regression model, and the optimal regularization parameter is used to determine the dimension of the final key features.

7. The method according to claim 1, characterized in that, The integrated model is used to predict fluorescence values. It consists of five regression models, and the predicted fluorescence value is the mean of the fluorescence prediction results of each of the five regression models. The five regression models are decision tree regression model, random forest regression model, gradient boosting regression model, support vector regression model and XGBoost regression model.

8. The method according to claim 1, characterized in that, The ability of Cas protein to recognize the combined sequence based on the predicted fluorescence value is as follows: the larger the predicted fluorescence value, the stronger the ability of Cas protein to recognize the combined sequence; the smaller the predicted fluorescence value, the weaker the ability of Cas protein to recognize the combined sequence.

9. A device for jointly identifying the effectiveness of Cas targets using a large gene model and ML, characterized in that, The device includes a memory and one or more processors, wherein the memory stores executable code, and the one or more processors execute the executable code to implement the method for jointly identifying the effectiveness of Cas targets using a large gene model and ML as described in any one of claims 1-8.

10. A computer-readable storage medium, characterized in that, It stores a program that, when executed by a processor, implements the method for jointly identifying the effectiveness of Cas targets using a large gene model and ML as described in any one of claims 1-8.

Citation Information

Patent Citations

  • Method for predicting compound target based on compound perturbation transcriptome and protein sequence and application

    CN119889424A

  • Cas protein recognition method and system based on protein language model and multi-scale convolutional neural network

    CN120072053A