A method and system for optimizing sgRNAs based on pathogenic single nucleotide variants in a multi-editing platform
Patent Information
- Application Number
- CN202610655304.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-13
- Publication Date
- 2026-08-18
AI Technical Summary
[0006]针对现有技术的以上缺陷或改进需求,本发明提供了一种基于致病单核苷酸变异的多编辑平台sgRNA优选方法,其目的在于,解决现有基于规则特征与浅层建模的方法仅能捕捉有限的人工特征,难以建模序列间复杂的非线性关系,导致对未实验验证序列的预测能力不足的技术问题,以及现有基于序列输入的深度学习方法无法充分捕捉序列的多尺度模式与全局依赖,导致对关键编辑区域及远程依赖关系的表征能力不足的技术问题,以及现有基于编辑上下文建模的碱基编辑预测方法缺乏对编辑前后序列变化关系的显示建模能力,难以刻画编辑前与编辑后的序列差异,从而限制了其预测能力的技术问题,以及现有三种方法均聚焦于单一指标或单一编辑平台,未提供综合排序依据,难以支持实际疾病场景下的编辑策略选择的技术问题
1、本发明由于采用了步骤(2-1)到步骤(2-5),其在Prism-on模型中引入位置增强模块、多尺度上下文特征提取模块、特征融合模块以及残差连接单元,能够从sgRNA序列中提取局部序列模式、多尺度上下文信息及全局依赖特征,因此能够解决现有基于规则特征与浅层建模的方法仅能捕捉有限人工特征、难以建模复杂非线性关系,导致对未实验验证序列预测能力不足的技术问题;
Smart Images

Figure CN122598780A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the interdisciplinary field of gene editing, bioinformatics and deep learning, and more specifically, relates to a method and system for optimizing sgRNA based on pathogenic single nucleotide variants in a multi-editing platform. Background Technology
[0002] With the deepening of genomics research, single-nucleotide variants (SNVs) are considered one of the most important pathogenic factors in human genetic diseases. Precise repair of pathogenic SNVs is a crucial foundation for gene therapy. In recent years, various gene editing technologies, such as the CRISPR / Cas9 system, adenine-based editors (ABEs), and cytosine-based editors (CBEs), have been continuously developed, providing diverse means for repairing different types of mutations. However, in practical applications, different gene editing technologies exhibit significant differences in editing mechanisms and applicable mutation types, and also face issues such as off-target effects and bystander editing. Against this backdrop, how to uniformly model, evaluate, and optimize candidate sgRNAs across multiple editing platforms has become a key issue in improving gene editing efficiency and safety.
[0003] Currently, several methods have been proposed to address the sgRNA optimization problem. The first category comprises rule-based feature and shallow modeling architectures. These methods typically rely on manually designed sequence features (such as GC content, base position preferences, and PAM features) and combine them with statistical or machine learning models to evaluate sgRNA activity. The second category consists of deep learning architectures based on sequence input. These methods encode and model sequences using convolutional neural networks, long short-term memory networks, or Transformer architectures, automatically extracting local patterns or long-range dependencies from the sequence. The third category is base editing prediction architectures based on editing context modeling. These methods introduce editing windows and base context information on top of sequence modeling. By learning features from local regions of the target sequence, they establish a mapping relationship between the sequence context and the distribution of editing results, thereby predicting the proportion of editing results.
[0004] However, all of the above-mentioned existing methods have some significant drawbacks: First, methods based on rule features and shallow modeling can only capture limited artificial features and are difficult to model complex nonlinear relationships between sequences, resulting in insufficient predictive ability for sequences that have not been experimentally verified. Second, deep learning methods based on sequence input cannot fully capture the multi-scale patterns and global dependencies of sequences, resulting in insufficient representation capabilities for key editing regions and long-range dependencies. Third, base editing prediction methods based on editing context modeling lack the ability to explicitly model the relationship between sequence changes before and after editing, making it difficult to characterize the sequence differences before and after editing, thus limiting their predictive ability.
[0005] Fourth, the three existing methods all focus on a single indicator or a single editing platform, without providing a comprehensive ranking basis, making it difficult to support the selection of editing strategies in actual disease scenarios. Summary of the Invention
[0006] To address the aforementioned deficiencies or improvement needs of existing technologies, this invention provides a method for optimizing sgRNA across multiple editing platforms based on pathogenic single nucleotide variants. The aim is to resolve the following technical issues: existing methods based on rule features and shallow modeling can only capture limited artificial features, making it difficult to model complex nonlinear relationships between sequences, resulting in insufficient predictive ability for sequences not experimentally validated; existing deep learning methods based on sequence input cannot fully capture multi-scale patterns and global dependencies of sequences, resulting in insufficient characterization of key editing regions and long-range dependencies; existing base editing prediction methods based on editing context modeling lack the ability to explicitly model the relationship between sequence changes before and after editing, making it difficult to characterize the differences between pre- and post-edited sequences, thus limiting their predictive ability; and existing methods all focus on a single indicator or a single editing platform, failing to provide comprehensive ranking criteria, making it difficult to support the selection of editing strategies in real-world disease scenarios.
[0007] To achieve the above objectives, according to one aspect of the present invention, a method for optimizing sgRNA based on pathogenic single nucleotide variants in a multi-editing platform is provided, comprising the following steps: (1) Obtain multiple single nucleotide variant (SNV) data, and perform preprocessing and candidate sgRNA generation processing on all SNV data in sequence to obtain a candidate sgRNA set; (2) Input the CRISPR / Cas9 candidate sgRNA set in the candidate sgRNA set obtained in step (1) into the pre-trained Prism-on model to obtain the target cleavage efficiency prediction value corresponding to the candidate sgRNA set; (3) Input the cytosine base editor candidate sgRNA set and the cytosine base editor candidate sgRNA set obtained in step (1) into the pre-trained Prism-be model to obtain the target editing efficiency prediction value and the bystander editing efficiency prediction value of the cytosine base editor candidate sgRNA set in the candidate sgRNA set, and the target editing efficiency prediction value and the bystander editing efficiency prediction value of the adenine base editor candidate sgRNA set. (4) Based on the predicted values of the target cleavage efficiency of the CRISPR / Cas9 candidate sgRNA set obtained in step (2), the predicted values of the target editing efficiency and the bystander editing efficiency of the cytosine base editor candidate sgRNA set obtained in step (3), and the predicted values of the target editing efficiency and the bystander editing efficiency of the adenine base editor candidate sgRNA set, the candidate sgRNA set is comprehensively sorted to obtain the optimal results of sgRNA for multiple editing platforms.
[0008] Preferably, step (1) includes the following sub-steps: (1-1) Obtain multiple SNV data from the public database ClinVar, select all mutation sites marked as pathogenic or potentially pathogenic from all SNV data, and obtain all variant data mapped to the GRCh38 reference genome. All mutation sites and all variant data constitute the initial pathogenic SNV dataset. (1-2) Perform variant function annotation processing on the initial pathogenic SNV dataset obtained in step (1-1) to obtain the annotation results of the mutation function impact type corresponding to each pathogenic SNV data; based on the annotation results corresponding to all pathogenic SNV data, select the set of function-related mutation sites consisting of all mutation sites located in the coding region or splice region from the initial pathogenic SNV dataset (the purpose is to improve the functional relevance and interpretability of subsequent editing design); (1-3) Based on the editable mutation types of the CRISPR / Cas9 system, the adenine base editor, and the cytosine base editor, classify the set of function-related mutation sites obtained in step (1-2) to obtain the set of mutation sites corresponding to the CRISPR / Cas9 system, the set of mutation sites corresponding to the adenine base editor, and the set of mutation sites corresponding to the cytosine base editor, respectively. (1-4) Obtain the reference genome sequence corresponding to the GRCh38 reference genome obtained in step (1-1) from the UCSC genome database. Based on the chromosomal position of each mutation site in the mutation site set corresponding to the CRISPR / Cas9 system, the mutation site set corresponding to the adenine base editor, and the mutation site set corresponding to the cytosine base editor obtained in step (1-3), perform local sequence extraction processing on the obtained reference genome sequence to obtain the local sequence set corresponding to the CRISPR / Cas9 system, the local sequence set corresponding to the adenine base editor, and the local sequence set corresponding to the cytosine base editor, respectively. (1-5) Based on the preset PAM sequence constraints of the protospacer sequence, perform PAM sequence scanning on the local sequence sets corresponding to the CRISPR / Cas9 system, the adenine base editor, and the cytosine base editor obtained in step (1-4) to obtain multiple PAM sequences corresponding to the CRISPR / Cas9 system, the adenine base editor, and the cytosine base editor. Extract a 20bp protospacer sequence from the adjacent region of each PAM sequence corresponding to the CRISPR / Cas9 system within its local sequence set as the corresponding sgRNA sequence. Concatenate all sgRNA sequences with their corresponding PAM sequences to form an initial candidate sgRNA set of 23nt length corresponding to the CRISPR / Cas9 system. Extract a 20bp protospacer sequence from the adjacent region of each PAM sequence corresponding to the adenine base editor within its local sequence set. A 20bp original spacer sequence is extracted from the adjacent region of the PAM sequence as the corresponding sgRNA sequence. All sgRNA sequences are spliced with their corresponding PAM sequences to form an initial candidate sgRNA set of 23nt length corresponding to the adenine base editor. A 20bp original spacer sequence is extracted from the adjacent region of each PAM sequence corresponding to the cytosine base editor in its local sequence set as the corresponding sgRNA sequence. All sgRNA sequences are spliced with their corresponding PAM sequences to form an initial candidate sgRNA set of 23nt length corresponding to the cytosine base editor. The specific sequence of the motif adjacent to the original spacer sequence is NGG, where N represents any one of the nucleotides A, T, C, and G, A is adenine, T is thymine, C is cytosine, and G is guanine. (1-6) The initial candidate sgRNA set corresponding to the CRISPR / Cas9 system obtained in step (1-5) is screened based on mutation site coverage to retain all candidate sgRNAs whose corresponding mutation sites are located within the sequence range corresponding to the candidate sgRNAs, forming the CRISPR / Cas9 candidate sgRNA set; the initial candidate sgRNA set corresponding to the adenine base editor and the initial candidate sgRNA set corresponding to the cytosine base editor obtained in step (1-5) are screened based on the editing window position to retain all candidate sgRNAs whose corresponding target bases are located within the editing window, thereby obtaining the adenine base editor candidate sgRNA set and the cytosine base editor candidate sgRNA set, respectively. The CRISPR / Cas9 candidate sgRNA set, the adenine base editor candidate sgRNA set, and the cytosine base editor candidate sgRNA set together constitute the candidate sgRNA set.
[0009] Preferably, the Prism-on model includes a sequence encoding module, a location enhancement module, a multi-scale contextual feature extraction module, a feature fusion module, and a regression prediction module; Step (2) includes the following sub-steps: (2-1) Input the CRISPR / Cas9 candidate sgRNA set in the candidate sgRNA set obtained in step (1) into the sequence coding module in the Prism-on model for encoding processing to obtain the sequence coding matrix; (2-2) Input the sequence encoding matrix obtained in step (2-1) into the position augmentation module in the Prism-on model for position augmentation processing to obtain the enhanced sequence feature representation. ; (2-3) Input the enhanced sequence feature representation obtained in step (2-2) into the multi-scale context feature extraction module in the Prism-on model for multi-scale context feature extraction processing to obtain multi-scale sequence features. ; (2-4) The multi-scale sequence features obtained in step (2-3) The input is processed by the feature fusion module in the Prism-on model to obtain fused features. And the fusion feature is obtained through residual connection units. and the multi-scale sequence features Element-by-element addition is performed to obtain the fused and enhanced sequence feature representation; (2-5) Input the fusion-enhanced sequence feature representation obtained in step (2-4) into the regression prediction module in the Prism-on model for regression prediction processing to obtain the target cleavage efficiency prediction value corresponding to the CRISPR / Cas9 candidate sgRNA set.
[0010] Preferably, the encoding process in step (2-1) uses one-hot encoding. Step (2-1) specifically involves first encoding the CRISPR / Cas9 candidate sgRNA set bit-by-bit using a one-hot encoding method. This means encoding the nucleotides A, T, C, and G into encoding vectors (1,0,0,0), (0,1,0,0), (0,0,1,0), and (0,0,0,1), respectively. The resulting encoding vectors for all nucleotides form a set with a size of [missing value]. The sequence encoding matrix; Step (2-2) involves introducing learnable positional encoding vectors into the sequence encoding matrix and then adding the sequence encoding matrix and the positional encoding vectors element-wise to obtain the enhanced sequence feature representation. ; in, For sequence encoding matrix, A learnable positional encoding vector; Steps (2-3) are as follows: First, local sequence patterns are extracted from the enhanced sequence feature representation through two one-dimensional convolutional layers; then, multi-scale context features are extracted from the extracted local sequence patterns through a multi-scale context aggregation structure to obtain multi-scale sequence features. Steps (2-4) are as follows: First, the multi-scale sequence features are concatenated along the channel dimension, and the concatenation result is input into the feature fusion module to obtain the fused feature. Then, the multi-scale sequence features and the fused feature are input together into the residual connection unit for element-wise addition to obtain the fused and enhanced sequence feature representation. : ; Steps (2-5) are as follows: First, the fused and enhanced sequence feature representation is flattened to obtain a one-dimensional feature vector. This one-dimensional feature vector is then sequentially input into two fully connected layers with output dimensions of 512 and 256, respectively, to obtain the processed features. Finally, the processed features are input into a single neuron linear output layer to obtain a continuous prediction result, which serves as the predicted value for the target cleavage efficiency corresponding to the CRISPR / Cas9 candidate sgRNA set.
[0011] Preferably, the Prism-be model includes a dual-sequence input encoding module, a differential feature construction module, a multi-scale convolutional feature extraction module, a channel attention enhancement module, and a regression prediction module; This step (3) includes the following sub-steps: (3-1) Construct the corresponding target sequence coding matrix and editing result sequence coding matrix based on the cytosine base editor candidate sgRNA set in the candidate sgRNA set obtained in step (1), and construct the corresponding target sequence coding matrix and editing result sequence coding matrix based on the adenine base editor candidate sgRNA set in the candidate sgRNA set obtained in step (1). The target sequence is the target sequence corresponding to the candidate sgRNA set, and the editing result sequence is the sequence obtained after performing all possible base substitutions on the editable bases in the editing window of the target sequence according to the editor type and editable base type corresponding to the candidate sgRNA set. (3-2) Input the first target sequence encoding matrix and the first editing result sequence encoding matrix corresponding to the adenine base editor candidate sgRNA set obtained in step (3-1) into the differential feature construction module in the Prism-be model for differential sequence feature construction processing to obtain the first differential sequence feature. The first target sequence encoding matrix, the first edited result sequence encoding matrix, and the first differential sequence feature are concatenated along the channel dimension into a first 12-channel input tensor. The second target sequence encoding matrix and the second edited result sequence encoding matrix corresponding to the cytosine base editor candidate sgRNA set obtained in step (3-1) are then input into the differential feature construction module in the Prism-be model for differential sequence feature construction to obtain the second differential sequence feature. The second target sequence encoding matrix, the second edit result sequence encoding matrix, and the second difference sequence features are concatenated in the channel dimension to form a second 12-channel input tensor. (3-3) Input the first 12-channel input tensor corresponding to the adenine base editor candidate sgRNA set obtained in step (3-2) into the multi-scale convolution feature extraction module of the Prism-be model and perform multi-scale convolution feature extraction processing to obtain the first multi-scale convolution feature representation; input the second 12-channel input tensor corresponding to the cytosine base editor candidate sgRNA set obtained in step (3-2) into the multi-scale convolution feature extraction module of the Prism-be model and perform multi-scale convolution feature extraction processing to obtain the second multi-scale convolution feature representation; (3-4) Input the first multi-scale convolutional feature representation corresponding to the adenine base editor candidate sgRNA set obtained in step (3-3) into the channel attention enhancement module in the Prism-be model for channel attention enhancement processing to obtain the feature representation after the first channel enhancement; input the second multi-scale convolutional feature representation corresponding to the cytosine base editor candidate sgRNA set obtained in step (3-3) into the channel attention enhancement module in the Prism-be model for channel attention enhancement processing to obtain the feature representation after the second channel enhancement; (3-5) Input the enhanced feature representation of the first channel corresponding to the adenine base editor candidate sgRNA set obtained in step (3-4) into the regression prediction module in the Prism-be model for regression prediction processing to obtain the predicted value of the editing result ratio corresponding to the adenine base editor candidate sgRNA set; input the enhanced feature representation of the second channel corresponding to the cytosine base editor candidate sgRNA set obtained in step (3-4) into the regression prediction module in the Prism-be model for regression prediction processing to obtain the predicted value of the editing result ratio corresponding to the cytosine base editor candidate sgRNA set; (3-6) Calculate the predicted editing result ratios of the candidate sgRNA set for adenine base editors obtained in step (3-5) to obtain the predicted target editing efficiency and the predicted bystander editing efficiency of the candidate sgRNA set for adenine base editors; calculate the predicted editing result ratios of the candidate sgRNA set for cytosine base editors obtained in step (3-5) to obtain the predicted target editing efficiency and the predicted bystander editing efficiency of the candidate sgRNA set for cytosine base editors.
[0012] Preferably, step (3-1) specifically involves: firstly, for the adenine base editor candidate sgRNA set, replacing the adenine bases located within the editing window in the corresponding target sequence with A to G to obtain the editing result sequence corresponding to each candidate sgRNA sequence in the adenine base editor candidate sgRNA set; secondly, for the cytosine base editor candidate sgRNA set, replacing the adenine bases located within the editing window in the corresponding target sequence with C to T to obtain the editing result sequence corresponding to each candidate sgRNA sequence in the cytosine base editor candidate sgRNA set. Then, the target sequence and the edited result sequence corresponding to the adenine base editor candidate sgRNA set are respectively input into the double sequence input encoding module in the Prism-be model for encoding processing to obtain the first target sequence encoding matrix and the first edited result sequence encoding matrix, respectively. The target sequence and the edited result sequence corresponding to the cytosine base editor candidate sgRNA set are respectively input into the double sequence input encoding module in the Prism-be model for encoding processing to obtain the second target sequence encoding matrix and the second edited result sequence encoding matrix, respectively. Specifically, the encoding process in this step uses one-hot encoding: The encoding process specifically involves first using one-hot encoding to encode each nucleotide of both the target sequence and the edited sequence bit by bit. Specifically, the nucleotides A, T, C, and G are encoded into encoding vectors (1,0,0,0), (0,1,0,0), (0,0,1,0), and (0,0,0,1), respectively. The encoding vectors corresponding to all nucleotides form a sequence of size [missing information]. The target sequence encoding matrix and the edited result sequence encoding matrix; Step (3-3) is as follows: First, the 12-channel input tensor is initially mapped using a one-dimensional convolutional layer with a kernel size of 3 to convert it into 32-channel initial convolutional features. Then, the 32-channel initial convolutional features are input into a batch normalization layer and a ReLU activation function to obtain 32-channel activation features. Subsequently, the 32-channel activation features are input into the first multi-scale convolutional unit to obtain a 48-channel multi-scale feature representation. Finally, the 48-channel multi-scale feature representation is input into the second multi-scale convolutional unit to obtain a 64-channel multi-scale feature representation. Each multi-scale convolutional unit employs the Inception architecture, comprising four parallel branches, namely... Convolutional branches, Convolutional branches, The convolutional branch and the max pooling branch are then combined. Finally, the outputs of each branch are concatenated along the channel dimension to obtain a multi-scale convolutional feature representation. Steps (3-4) are as follows: First, the multi-scale convolutional feature representation is input into a global average pooling unit for global average pooling to obtain a global statistical vector. Then, the global statistical vector is input into two fully connected layers and a sigmoid function to obtain channel weights between 0 and 1. Finally, the channel weights are multiplied channel-by-channel by the multi-scale convolutional feature representation to obtain the channel-enhanced feature representation. Specifically, after the first multi-scale convolutional unit, the channel attention enhancement module operates on 48 channels; after the second multi-scale convolutional unit, the channel attention enhancement module operates on 64 channels.
[0013] Preferably, step (3-5) specifically involves: first, flattening the enhanced channel features to obtain a one-dimensional feature vector; then, sequentially inputting the one-dimensional feature vector into two fully connected layers with output dimensions of 256 and 64 to obtain the processed features; finally, inputting the processed features into a single-neuron linear output layer to obtain a continuous prediction result, which serves as the predicted editing result ratio for the candidate sgRNA set. Step (3-6) specifically involves the following steps: First, based on the edited result sequences corresponding to the adenine base editor candidate sgRNA set and the cytosine base editor candidate sgRNA set obtained in step (3-1), the edited result sequences that have not undergone base substitution are selected as unedited result sequences. Then, based on the predicted edit result ratio obtained in step (3-5), the predicted edit result ratio corresponding to the unedited result sequences is obtained. And predict the proportion based on the edited result. Obtain the editor-in-chief's efficiency prediction: ; Then, based on the edited result sequences corresponding to the adenine base editor candidate sgRNA set and the cytosine base editor candidate sgRNA set obtained in step (3-1), the edited result sequences that have undergone target base substitution are selected as target edited result sequences. Simultaneously, the remaining edited result sequences in the edited result sequence set, excluding the unedited result sequences, are considered as edited result sequences. Based on the predicted edit result ratios corresponding to the candidate sgRNA sets obtained in step (3-5), the predicted edit result ratios of the target edited result sequences and the edited result sequences are obtained. Finally, the predicted target editing efficiency value is obtained based on the predicted edit result ratios of the target edited result sequences and the edited result sequences. : ; in, This represents the predicted proportion of edit results in the target edit result sequence. This represents the predicted proportion of edited results in the edited result sequence; Finally, based on the target editing efficiency prediction value Calculate the predicted value of bystander editing efficiency : ;
[0014] Preferably, step (4) specifically involves: for the CRISPR / Cas9 candidate sgRNA set in the candidate sgRNA set, sorting the CRISPR / Cas9 candidate sgRNA set according to the target cleavage efficiency prediction value obtained in step (2), and preferentially selecting candidate sgRNAs with higher target cleavage efficiency prediction values as the preferred sgRNAs corresponding to the CRISPR / Cas9 candidate sgRNA set; for the cytosine base editor candidate sgRNA set and the adenine base editor candidate sgRNA set in the candidate sgRNA set, comprehensively sorting the cytosine base editor candidate sgRNA set according to the target editing efficiency prediction value and the bystander editing efficiency prediction value corresponding to the cytosine base editor candidate sgRNA set obtained in step (3). Candidate sgRNAs with higher predicted target editing efficiency and lower predicted bystander editing efficiency are selected as preferred sgRNAs corresponding to the cytosine base editor candidate sgRNA set. Based on the adenine base editor candidate sgRNA set obtained in step (3), the candidate sgRNA set is comprehensively sorted, and candidate sgRNAs with higher predicted target editing efficiency and lower predicted bystander editing efficiency are selected as preferred sgRNAs corresponding to the adenine base editor candidate sgRNA set. The preferred sgRNAs corresponding to the CRISPR / Cas9 candidate sgRNA set, the preferred sgRNAs corresponding to the cytosine base editor candidate sgRNA set, and the preferred sgRNAs corresponding to the adenine base editor candidate sgRNA set together constitute the preferred sgRNA results for the multi-editing platform. The Prism-on model is obtained through the following steps: (5-1) Experimental data related to WT-SpCas9, eSpCas9 (1.1) and SpCas9-HF1 were obtained from publicly published high-fidelity Cas9 nucleotide sgRNA activity research data. The experimental data included different sgRNA sequences and their editing efficiency values measured under the corresponding Cas9 nuclease conditions. All experimental data were summarized into a CRISPR / Cas9 editing efficiency training dataset, and the CRISPR / Cas9 editing efficiency training dataset was divided into a training set, a validation set and a test set in a ratio of 7:1:2. (5-2) For each sample in the training set obtained in step (5-1), the sgRNA sequence in the sample is input into the sequence encoding module of the Prism-on model for one-hot encoding to obtain the sequence encoding matrix corresponding to the sample; the sequence encoding matrix is input into the position enhancement module of the Prism-on model to obtain the position encoding vector, and the sequence encoding matrix and the position encoding vector are added element by element to obtain the enhanced sequence feature representation corresponding to the sample; (5-3) For each sample in the training set obtained in step (5-1), the enhanced sequence feature representation corresponding to the sample obtained in step (5-2) is input into the multi-scale context feature extraction module of the Prism-on model for processing to obtain the multi-scale sequence feature corresponding to the sample. The multi-scale sequence feature is then input into the feature fusion module of the Prism-on model for channel concatenation and convolutional projection processing to obtain the fused feature. The multi-scale sequence feature and the fused feature are then input together into the residual connection unit for element-wise addition to obtain the fused enhanced sequence feature representation corresponding to the sample. (5-4) For each sample in the training set obtained in step (5-1), the fused and enhanced sequence feature representation of the sample obtained in step (5-3) is input into the regression prediction module of the Prism-on model for processing to obtain the predicted value of the sgRNA targeting cleavage efficiency of the sample. (5-5) For each sample in the training set obtained in step (5-1), the predicted value of the sgRNA targeting cleavage efficiency corresponding to the sample and the experimentally determined editing efficiency corresponding to the sample are input into the mean squared error loss function to obtain the training loss corresponding to the sample. (5-6) For each sample in the training set obtained in step (5-1), the Prism-on model is iteratively trained using gradient descent based on the training loss corresponding to the sample obtained in step (5-5). During the training process, the model training state is monitored and optimized using the validation set obtained in step (5-1) until the validation set loss tends to stabilize or the Prism-on model reaches the preset number of iterations. The optimal parameters of the Prism-on model during the training process are obtained, thus obtaining the finally trained Prism-on model.
[0015] Preferably, the Prism-be model is trained through the following steps: (6-1) Obtain experimental data related to ABE7.10, ABEmax, ABE8e, BE4, CBE4max and Target-AID from publicly published base editor sgRNA activity research data. The experimental data includes the target sequence, the edited result sequence and the true value of the edited result ratio determined by the corresponding base editor. All experimental data are summarized into a base editor training dataset, and the base editor training dataset is divided into training set, validation set and test set in a ratio of 7:1:2. (6-2) For each sample in the training set obtained in step (6-1), the target sequence and the edit result sequence in the sample are input into the double sequence input encoding module of the Prism-be model for one-hot encoding to obtain the target sequence encoding matrix and the edit result sequence encoding matrix corresponding to the sample. The target sequence encoding matrix and the edit result sequence encoding matrix are input into the differential feature construction module of the Prism-be model for processing to obtain the differential sequence features corresponding to the sample. The target sequence encoding matrix, the edit result sequence encoding matrix and the differential sequence features are concatenated in the channel dimension to obtain the input tensor corresponding to the sample. (6-3) For each sample in the training set obtained in step (6-1), the input tensor corresponding to the sample obtained in step (6-2) is input into the multi-scale convolutional feature extraction module of the Prism-be model for processing to obtain the multi-scale convolutional feature representation corresponding to the sample. (6-4) For each sample in the training set obtained in step (6-1), the multi-scale convolutional feature representation corresponding to the sample obtained in step (6-3) is input into the channel attention enhancement module in the Prism-be model for channel attention enhancement processing to obtain the final channel-enhanced feature representation corresponding to the sample. (6-5) For each sample in the training set obtained in step (6-1), the final channel-enhanced feature representation of the sample obtained in step (6-4) is input into the regression prediction module of the Prism-be model for processing to obtain the predicted value of the editing result ratio of the sample. (6-6) For each sample in the training set obtained in step (6-1), input the true value of the editing result ratio corresponding to the sample and the predicted value of the editing result ratio corresponding to the sample obtained in step (6-5) into the mean square error loss function to obtain the training loss corresponding to the sample. (6-7) For each sample in the training set obtained in step (6-1), the Prism-be model is iteratively trained using gradient descent based on the training loss corresponding to the sample obtained in step (6-5). During the training process, the model training state is monitored and optimized using the validation set obtained in step (6-1) until the validation set loss tends to stabilize or the Prism-be model reaches the preset number of iterations. The optimal parameters of the Prism-be model during the training process are obtained, thus obtaining the finally trained Prism-be model.
[0016] According to another aspect of the present invention, a multi-editing platform sgRNA optimization system based on pathogenic single nucleotide variants is provided, comprising the following modules: The first module is used to acquire multiple single nucleotide variant (SNV) data, and to preprocess all SNV data and generate candidate sgRNAs sequentially to obtain a candidate sgRNA set. The second module is used to input the CRISPR / Cas9 candidate sgRNA set from the candidate sgRNA set obtained by the first module into the pre-trained Prism-on model to obtain the target cleavage efficiency prediction value corresponding to the candidate sgRNA set. The third module is used to input the cytosine base editor candidate sgRNA set and the cytosine base editor candidate sgRNA set obtained from the first module into the pre-trained Prism-be model to obtain the target editing efficiency prediction value and the bystander editing efficiency prediction value corresponding to the cytosine base editor candidate sgRNA set in the candidate sgRNA set, as well as the target editing efficiency prediction value and the bystander editing efficiency prediction value corresponding to the adenine base editor candidate sgRNA set. The fourth module is used to perform comprehensive sorting of the candidate sgRNA sets based on the predicted target cleavage efficiency values of the CRISPR / Cas9 candidate sgRNA sets obtained in the second module, the predicted target editing efficiency values and bystander editing efficiency values of the cytosine base editor candidate sgRNA sets obtained in the third module, and the predicted target editing efficiency values and bystander editing efficiency values of the adenine base editor candidate sgRNA sets, in order to obtain the optimal sgRNA results for multiple editing platforms.
[0017] In summary, compared with the prior art, the above-described technical solutions conceived by this invention can achieve the following beneficial effects: 1. This invention, by employing steps (2-1) to (2-5), introduces a position enhancement module, a multi-scale contextual feature extraction module, a feature fusion module, and a residual connection unit into the Prism-on model. This enables the extraction of local sequence patterns, multi-scale contextual information, and global dependency features from sgRNA sequences. Therefore, it can solve the technical problem that existing methods based on regular features and shallow modeling can only capture limited artificial features and are difficult to model complex nonlinear relationships, resulting in insufficient prediction capabilities for sequences that have not been experimentally verified. 2. Because the present invention employs steps (2-1) to (2-5), it uses learnable position encoding, dilated convolution branches with different dilation rates in a multi-scale context aggregation structure, and global average pooling branches to jointly characterize key editing regions and long-range dependencies in sgRNA sequences. Therefore, it can solve the technical problem that existing deep learning methods based on sequence input cannot fully capture multi-scale patterns and global dependencies of sequences, resulting in insufficient characterization capabilities of key editing regions and long-range dependencies. 3. Because the present invention employs steps (3-1) to (3-6), its Prism-be model simultaneously constructs the target sequence encoding matrix and the edited result sequence encoding matrix, and explicitly represents the relationship of base changes before and after editing through the differential feature construction module. Combined with multi-scale convolution feature extraction and channel attention enhancement to extract local sequence patterns and editing context features at different scales, it can solve the technical problem that existing base editing prediction methods based on editing context modeling lack the ability to explicitly model the relationship of sequence changes before and after editing, making it difficult to characterize the differences between the sequences before and after editing, thus limiting the prediction ability. 4. Since the present invention adopts steps (1) to (4), it takes the pathogenic SNV as the center and performs unified generation, performance prediction and comprehensive ranking of candidate sgRNAs corresponding to the CRISPR / Cas9 system, adenine base editor and cytosine base editor. Therefore, it can solve the technical problem that the existing methods focus on a single indicator or a single editing platform, do not provide comprehensive ranking basis, and are difficult to support the selection of editing strategies in actual disease scenarios. 5. This invention proposes an sgRNA design framework centered on pathogenic SNVs, achieving unified modeling from mutation to editing strategies; 6. The Prism-on and Prism-be models of this invention achieve state-of-the-art prediction accuracy on various independent datasets, with their correlation metrics outperforming leading methods by up to 2.92%. 7. By enabling automation and visualization through a web platform, researchers can screen out a large number of low-activity sgRNAs before experiments, thereby shortening the cycle and saving costs to support the clinical translation of gene editing technology. Attached Figure Description
[0018] Figure 1 This is a schematic diagram of the overall process of the present invention for the optimal sgRNA selection method based on pathogenic single nucleotide variants in a multi-editing platform; Figure 2 This is a diagram of the model architecture used in this invention to predict the targeting cleavage efficiency of the CRISPR / Cas9 candidate sgRNA set. Figure 3 This is a model architecture diagram of the present invention used to predict the proportion of editing results corresponding to the candidate sgRN sets of adenine base editor and cytosine base editor. Detailed Implementation
[0019] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. Furthermore, the technical features involved in the various embodiments of this invention described below can be combined with each other as long as they do not conflict with each other.
[0020] The basic idea of this invention is to construct an sgRNA optimization system for multiple editing platforms, using pathogenic single nucleotide variants (SNVs) as the core entry point. First, starting with pathogenic SNVs, a unified generation of candidate sgRNA sets for different gene editing systems (CRISPR / Cas9 system, adenine base editor, and cytosine base editor) is achieved through systematic data preprocessing and constraints. Second, deep learning models are constructed to predict key performance indicators of sgRNAs for different editing mechanisms, including the targeting efficiency of the CRISPR / Cas9 system and the editing result ratio of the adenine base editor and the cytosine base editor. Finally, automated optimization of sgRNAs for multiple editing platforms is achieved through multi-index evaluation and ranking strategies. Through these methods, this invention realizes end-to-end modeling from pathogenic SNV identification, candidate sgRNA set generation, performance prediction to final optimization decision, thereby effectively improving the accuracy, systematicity, and practicality of sgRNA design.
[0021] like Figure 1 As shown, this invention provides a method for optimizing sgRNA on a multi-editing platform based on pathogenic single nucleotide variants, comprising the following steps: (1) Obtain multiple single-nucleotide variant (SNV) data, and perform preprocessing and candidate sgRNA generation processing on all SNV data in sequence to obtain a candidate sgRNA set; This step includes the following sub-steps: (1-1) Obtain multiple SNV data from the public database ClinVar, select all mutation sites marked as pathogenic or potentially pathogenic from all SNV data, and obtain all variant data mapped to the GRCh38 reference genome. All mutation sites and all variant data constitute the initial pathogenic SNV dataset. (1-2) Perform variant function annotation processing on the initial pathogenic SNV dataset obtained in step (1-1) to obtain the annotation results of the mutation function impact type corresponding to each pathogenic SNV data; based on the annotation results corresponding to all pathogenic SNV data, select the set of function-related mutation sites consisting of all mutation sites located in the coding region or splice region from the initial pathogenic SNV dataset (the purpose is to improve the functional relevance and interpretability of subsequent editing design); (1-3) Based on the editable mutation types of the CRISPR / Cas9 system, the adenine base editor, and the cytosine base editor, classify the set of function-related mutation sites obtained in step (1-2) to obtain the set of mutation sites corresponding to the CRISPR / Cas9 system, the set of mutation sites corresponding to the adenine base editor, and the set of mutation sites corresponding to the cytosine base editor, respectively. (1-4) Obtain the reference genome sequence corresponding to the GRCh38 reference genome obtained in step (1-1) from the UCSC genome database. Based on the chromosomal position of each mutation site in the mutation site set corresponding to the CRISPR / Cas9 system, the mutation site set corresponding to the adenine base editor, and the mutation site set corresponding to the cytosine base editor obtained in step (1-3) on the GRCh38 reference genome, perform local sequence extraction processing on the obtained reference genome sequence to obtain the local sequence set corresponding to the CRISPR / Cas9 system, the local sequence set corresponding to the adenine base editor, and the local sequence set corresponding to the cytosine base editor, respectively.
[0022] (1-5) Based on the preset protospacer adjacent motif (PAM) sequence constraints (the specific sequence of the protospacer adjacent motif is NGG, where N represents any one of A, T, C, and G, A is adenine, T is thymine, C is cytosine, and G is guanine), perform PAM sequence scanning on the local sequence sets corresponding to the CRISPR / Cas9 system, the adenine base editor, and the cytosine base editor obtained in step (1-4) to obtain multiple PAM sequences corresponding to the CRISPR / Cas9 system, the adenine base editor, and the cytosine base editor. Then, for each PAM sequence corresponding to the CRISPR / Cas9 system, perform PAM sequence scanning on the adjacent motifs in its local sequence set. A 20bp original spacer sequence was extracted from the region and used as the corresponding sgRNA sequence for the PAM sequence. All sgRNA sequences were spliced with their corresponding PAM sequences to form an initial candidate sgRNA set of 23nt length corresponding to the CRISPR / Cas9 system. A 20bp original spacer sequence was extracted from the adjacent region of each PAM sequence corresponding to the adenine base editor in its local sequence set and used as the corresponding sgRNA sequence for the PAM sequence. All sgRNA sequences were spliced with their corresponding PAM sequences to form an initial candidate sgRNA set of 23nt length corresponding to the adenine base editor. A 20bp original spacer sequence was extracted from the adjacent region of each PAM sequence corresponding to the cytosine base editor in its local sequence set and used as the corresponding sgRNA sequence for the PAM sequence. All sgRNA sequences were spliced with their corresponding PAM sequences to form an initial candidate sgRNA set of 23nt length corresponding to the cytosine base editor.
[0023] (1-6) The initial candidate sgRNA set corresponding to the CRISPR / Cas9 system obtained in step (1-5) is screened based on mutation site coverage to retain all candidate sgRNAs whose corresponding mutation sites are located within the sequence range corresponding to the candidate sgRNAs, forming the CRISPR / Cas9 candidate sgRNA set; the initial candidate sgRNA set corresponding to the adenine base editor and the initial candidate sgRNA set corresponding to the cytosine base editor obtained in step (1-5) are screened based on the editing window position to retain all candidate sgRNAs whose corresponding target bases are located within the editing window (preferably retain all candidate sgRNAs whose corresponding target bases are located between the 4th and 8th positions), thereby obtaining the adenine base editor candidate sgRNA set and the cytosine base editor candidate sgRNA set, respectively. The CRISPR / Cas9 candidate sgRNA set, the adenine base editor candidate sgRNA set, and the cytosine base editor candidate sgRNA set together constitute the candidate sgRNA set.
[0024] To illustrate the feasibility of the candidate sgRNA design steps for the multi-editing system, the mutation coverage and candidate sgRNA generation results under different editing systems were statistically analyzed, and the results are shown in Table 1.
[0025] Table 1: Statistical results of candidate sgRNA design under different editing systems The advantage of the above sub-steps (1-1) to (1-6) is that they enable the systematic generation and screening of candidate sgRNAs from multiple editing platforms, starting with pathogenic SNVs.
[0026] (2) Input the CRISPR / Cas9 candidate sgRNA set in the candidate sgRNA set obtained in step (1) into the pre-trained Prism-on model to obtain the target cleavage efficiency prediction value corresponding to the candidate sgRNA set; like Figure 2 As shown, the Prism-on model of this invention includes a sequence encoding module, a position enhancement module, a multi-scale contextual feature extraction module, a feature fusion module, and a regression prediction module. Step (2) includes the following sub-steps: (2-1) Input the CRISPR / Cas9 candidate sgRNA set in the candidate sgRNA set obtained in step (1) into the sequence coding module in the Prism-on model for encoding processing to obtain the sequence coding matrix (which serves as the basic input feature of the Prism-on model). Specifically, the encoding process in this step uses one-hot encoding. Furthermore, this step specifically involves first encoding the CRISPR / Cas9 candidate sgRNA set bit-by-bit using a one-hot encoding method, that is, encoding the nucleotides A, T, C, and G into encoding vectors (1,0,0,0), (0,1,0,0), (0,0,1,0), and (0,0,0,1), respectively. The size of the encoding vectors corresponding to all nucleotides is [size missing]. The sequence encoding matrix.
[0027] (2-2) Input the sequence encoding matrix obtained in step (2-1) into the position augmentation module in the Prism-on model for position augmentation processing to obtain the enhanced sequence feature representation. ; Specifically, this step involves introducing learnable positional encoding vectors into the sequence encoding matrix and then adding the sequence encoding matrix and the positional encoding vectors element-wise to obtain an enhanced sequence feature representation (used to enhance the Prism-on model's ability to characterize the importance of different base positions). ; in, For sequence encoding matrix, It is a learnable location encoding vector.
[0028] The advantage of step (2-2) is that by introducing learnable positional encoding, the model's ability to perceive the importance of different base positions can be enhanced, thereby better representing key editing regions.
[0029] (2-3) Input the enhanced sequence feature representation obtained in step (2-2) into the multi-scale context feature extraction module in the Prism-on model for multi-scale context feature extraction processing to obtain multi-scale sequence features. ; Specifically, this step involves first extracting local sequence patterns from the enhanced sequence feature representation using two one-dimensional convolutional layers; then, using a multi-scale context aggregation structure to extract multi-scale context features from the extracted local sequence patterns to obtain multi-scale sequence features.
[0030] This multi-scale context aggregation structure includes a Convolutional branches, multiple different dilatation rates (1, 2, 4, 6) The dilated convolution branch and a global average pooling branch.
[0031] The advantage of step (2-3) is that it can simultaneously extract local sequence patterns, multi-scale context information and global dependency features through the multi-scale context aggregation structure.
[0032] (2-4) The multi-scale sequence features obtained in step (2-3) The input is processed by the feature fusion module in the Prism-on model to obtain fused features. And the fusion feature is obtained through residual connection units. and the multi-scale sequence features Element-by-element addition is performed to obtain the fused and enhanced sequence feature representation; Specifically, this step involves first concatenating the multi-scale sequence features along the channel dimension, and then inputting the concatenation result into the feature fusion module (which uses...). The convolution projectes the concatenated result onto 32 channels to obtain the fused features; subsequently, the multi-scale sequence features and the fused features are input together into the residual connection unit for element-wise addition, thereby obtaining the fused and enhanced sequence feature representation. (Used to improve model stability and feature preservation ability): ; (2-5) Input the fusion-enhanced sequence feature representation obtained in step (2-4) into the regression prediction module in the Prism-on model for regression prediction processing to obtain the target cleavage efficiency prediction value corresponding to the CRISPR / Cas9 candidate sgRNA set; Specifically, this step involves first flattening the fused and enhanced sequence feature representation to obtain a one-dimensional feature vector. This one-dimensional feature vector is then sequentially input into two fully connected layers with output dimensions of 512 and 256, respectively (with dropout set to 0.3 after the first fully connected layer and 0.2 after the second fully connected layer to reduce overfitting risk and improve model generalization ability). This yields processed features. Finally, these processed features are input into a single-neuron linear output layer to obtain continuous prediction results, which serve as the predicted values for the target cleavage efficiency corresponding to the CRISPR / Cas9 candidate sgRNA set.
[0033] (3) Input the cytosine base editor candidate sgRNA set and the cytosine base editor candidate sgRNA set obtained in step (1) into the pre-trained Prism-be model to obtain the target editing efficiency prediction value and the bystander editing efficiency prediction value of the cytosine base editor candidate sgRNA set in the candidate sgRNA set, and the target editing efficiency prediction value and the bystander editing efficiency prediction value of the adenine base editor candidate sgRNA set. like Figure 3As shown, the Prism-be model of the present invention includes a dual-sequence input encoding module, a differential feature construction module, a multi-scale convolutional feature extraction module, a channel attention enhancement module, and a regression prediction module.
[0034] This step (3) includes the following sub-steps: (3-1) Construct the corresponding target sequence coding matrix and editing result sequence coding matrix based on the cytosine base editor candidate sgRNA set in the candidate sgRNA set obtained in step (1), and construct the corresponding target sequence coding matrix and editing result sequence coding matrix based on the adenine base editor candidate sgRNA set in the candidate sgRNA set obtained in step (1) (wherein the target sequence is the target sequence corresponding to the candidate sgRNA set, and the editing result sequence is the sequence obtained after performing all possible base substitutions on the editable bases in the target sequence located in the editing window according to the editor type and editable base type corresponding to the candidate sgRNA set).
[0035] Specifically, this step involves first, for the adenine base editor candidate sgRNA set, replacing the adenine bases located within the editing window in the corresponding target sequence with A to G to obtain the edited result sequence corresponding to each candidate sgRNA sequence in the adenine base editor candidate sgRNA set; and second, for the cytosine base editor candidate sgRNA set, replacing the adenine bases located within the editing window in the corresponding target sequence with C to T to obtain the edited result sequence corresponding to each candidate sgRNA sequence in the cytosine base editor candidate sgRNA set. Then, the target sequence and the edited result sequence corresponding to the adenine base editor candidate sgRNA set are respectively input into the double sequence input encoding module in the Prism-be model for encoding processing to obtain the first target sequence encoding matrix and the first edited result sequence encoding matrix, respectively. The target sequence and the edited result sequence corresponding to the cytosine base editor candidate sgRNA set are respectively input into the double sequence input encoding module in the Prism-be model for encoding processing to obtain the second target sequence encoding matrix and the second edited result sequence encoding matrix, respectively. Specifically, the encoding process in this step uses one-hot encoding: Furthermore, the encoding process specifically involves first using one-hot encoding to encode each bit of the target sequence and the edited sequence, that is, encoding the nucleotides A, T, C, and G into encoding vectors (1,0,0,0), (0,1,0,0), (0,0,1,0), and (0,0,0,1), respectively. The encoding vectors corresponding to all nucleotides form a sequence of size [missing information]. The target sequence encoding matrix and the edit result sequence encoding matrix.
[0036] By encoding the target sequence and the edited result sequence separately, the Prism-be model can simultaneously obtain the background information of the sequence before editing and the state information of the sequence after editing.
[0037] (3-2) Input the first target sequence encoding matrix and the first editing result sequence encoding matrix corresponding to the adenine base editor candidate sgRNA set obtained in step (3-1) into the differential feature construction module in the Prism-be model for differential sequence feature construction processing to obtain the first differential sequence feature. The first target sequence encoding matrix, the first edited result sequence encoding matrix, and the first differential sequence feature are concatenated along the channel dimension into a first 12-channel input tensor. The second target sequence encoding matrix and the second edited result sequence encoding matrix corresponding to the cytosine base editor candidate sgRNA set obtained in step (3-1) are then input into the differential feature construction module in the Prism-be model for differential sequence feature construction to obtain the second differential sequence feature. The second target sequence encoding matrix, the second edit result sequence encoding matrix, and the second difference sequence features are concatenated in the channel dimension to form a second 12-channel input tensor. Among them, differential sequence features Used to explicitly characterize base changes before and after editing, and equal to: ; in, and These represent the target sequence encoding matrix and the edited result sequence encoding matrix, respectively.
[0038] The advantage of step (3-2) is that by explicitly representing the relationship between base changes before and after editing through the differential feature construction module, the model's ability to distinguish between different editing results can be enhanced.
[0039] (3-3) Input the first 12-channel input tensor corresponding to the adenine base editor candidate sgRNA set obtained in step (3-2) into the multi-scale convolution feature extraction module of the Prism-be model and perform multi-scale convolution feature extraction processing to obtain the first multi-scale convolution feature representation; input the second 12-channel input tensor corresponding to the cytosine base editor candidate sgRNA set obtained in step (3-2) into the multi-scale convolution feature extraction module of the Prism-be model and perform multi-scale convolution feature extraction processing to obtain the second multi-scale convolution feature representation; Specifically, this step involves the following steps: First, a one-dimensional convolutional layer with a kernel size of 3 is used to initially map the 12-channel input tensor, converting it into initial 32-channel convolutional features. Then, these initial 32-channel convolutional features are input into a batch normalization layer and a ReLU activation function to obtain 32-channel activated features (improving feature representation and training stability). Subsequently, these 32-channel activated features are input into the first multi-scale convolutional unit to obtain a 48-channel multi-scale feature representation. Finally, this 48-channel multi-scale feature representation is input into the second multi-scale convolutional unit to obtain a 64-channel multi-scale feature representation.
[0040] Each multi-scale convolutional unit employs the Inception architecture, comprising four parallel branches, namely... Convolutional branches, Convolutional branches, The convolutional branch and the max-pooling branch are then used. Finally, the outputs of each branch are concatenated along the channel dimension to obtain a multi-scale convolutional feature representation.
[0041] The advantage of step (3-3) is that by extracting local sequence patterns and editing context features at different scales through the multi-scale convolution feature extraction module, the model's ability to represent sequence context information inside and outside the editing window can be improved.
[0042] (3-4) Input the first multi-scale convolutional feature representation corresponding to the adenine base editor candidate sgRNA set obtained in step (3-3) into the channel attention enhancement module in the Prism-be model for channel attention enhancement processing to obtain the feature representation after the first channel enhancement; input the second multi-scale convolutional feature representation corresponding to the cytosine base editor candidate sgRNA set obtained in step (3-3) into the channel attention enhancement module in the Prism-be model for channel attention enhancement processing to obtain the feature representation after the second channel enhancement; Specifically, this step involves first inputting the multi-scale convolutional feature representation into a global average pooling unit for global average pooling to obtain a global statistical vector; then, inputting this global statistical vector into two fully connected layers (the first fully connected layer is used for dimensionality reduction and combined with the ReLU activation function to extract inter-channel dependencies, and the second fully connected layer restores the features to the original number of channels) and a Sigmoid function to obtain channel weights between 0 and 1; finally, multiplying these channel weights with the multi-scale convolutional feature representation channel by channel (to achieve feature recalibration) to obtain the channel-enhanced feature representation; specifically, after the first multi-scale convolutional unit, the channel attention enhancement module operates on 48 channels; after the second multi-scale convolutional unit, the channel attention enhancement module operates on 64 channels.
[0043] (3-5) Input the enhanced feature representation of the first channel corresponding to the adenine base editor candidate sgRNA set obtained in step (3-4) into the regression prediction module in the Prism-be model for regression prediction processing to obtain the predicted value of the editing result ratio corresponding to the adenine base editor candidate sgRNA set; input the enhanced feature representation of the second channel corresponding to the cytosine base editor candidate sgRNA set obtained in step (3-4) into the regression prediction module in the Prism-be model for regression prediction processing to obtain the predicted value of the editing result ratio corresponding to the cytosine base editor candidate sgRNA set; Specifically, this step involves first flattening the enhanced channel features to obtain a one-dimensional feature vector; then, sequentially inputting this one-dimensional feature vector into two fully connected layers with output dimensions of 256 and 64 respectively (where dropout is set to 0.4 after the first fully connected layer and 0.3 after the second fully connected layer to reduce the risk of overfitting and improve the model's generalization ability) to obtain processed features; finally, inputting the processed features into a single-neuron linear output layer to obtain a continuous prediction result, which serves as the predicted editing result ratio for the candidate sgRNA set.
[0044] (3-6) Calculate the predicted editing result ratios of the candidate sgRNA set for adenine base editors obtained in step (3-5) to obtain the predicted target editing efficiency and the predicted bystander editing efficiency of the candidate sgRNA set for adenine base editors; calculate the predicted editing result ratios of the candidate sgRNA set for cytosine base editors obtained in step (3-5) to obtain the predicted target editing efficiency and the predicted bystander editing efficiency of the candidate sgRNA set for cytosine base editors. Specifically, this step involves first selecting the edited sequences corresponding to the adenine base editor candidate sgRNA set and the cytosine base editor candidate sgRNA set obtained in step (3-1) as unedited sequences, based on the edited sequences that did not undergo base substitution. Then, based on the predicted edited sequence proportion obtained in step (3-5), the predicted edited sequence proportion corresponding to the unedited sequences is obtained. And predict the proportion based on the edited result. Obtain the editor-in-chief's efficiency prediction: ; Then, based on the edited result sequences corresponding to the adenine base editor candidate sgRNA set and the cytosine base editor candidate sgRNA set obtained in step (3-1), the edited result sequences that have undergone target base substitution are selected as target edited result sequences. Simultaneously, the remaining edited result sequences in the edited result sequence set, excluding the unedited result sequences, are considered as edited result sequences. Based on the predicted edit result ratios corresponding to the candidate sgRNA sets obtained in step (3-5), the predicted edit result ratios of the target edited result sequences and the edited result sequences are obtained. Finally, the predicted target editing efficiency value is obtained based on the predicted edit result ratios of the target edited result sequences and the edited result sequences. : ; in, This represents the predicted proportion of edit results in the target edit result sequence. This represents the predicted proportion of edited results in the edited result sequence.
[0045] Finally, based on the target editing efficiency prediction value Calculate the predicted value of bystander editing efficiency (It is used to reflect the possibility of unintended editing occurring at non-target sites within the editing window): ; The above-mentioned indicators were used to calculate and process the target editing efficiency prediction value and the bystander editing efficiency prediction value of the candidate sgRNA set under different base editing platforms.
[0046] The advantage of step (3-6) is that by further calculating the target editing efficiency and bystander editing efficiency based on the predicted proportion of the editing results, the accuracy and security of the editing can be comprehensively considered.
[0047] (4) Based on the predicted values of the target cleavage efficiency of the CRISPR / Cas9 candidate sgRNA set obtained in step (2), the predicted values of the target editing efficiency and the bystander editing efficiency of the cytosine base editor candidate sgRNA set obtained in step (3), and the predicted values of the target editing efficiency and the bystander editing efficiency of the adenine base editor candidate sgRNA set, the candidate sgRNA set is comprehensively sorted to obtain the optimal results of sgRNA for multiple editing platforms.
[0048] Specifically, for the CRISPR / Cas9 candidate sgRNA set in the candidate sgRNA set, the CRISPR / Cas9 candidate sgRNA set is sorted according to the target cleavage efficiency prediction value obtained in step (2), and the candidate sgRNA with the higher target cleavage efficiency prediction value is selected as the preferred sgRNA corresponding to the CRISPR / Cas9 candidate sgRNA set; for the cytosine base editor candidate sgRNA set and the adenine base editor candidate sgRNA set in the candidate sgRNA set, the cytosine base editor candidate sgRNA is sorted according to the target cleavage efficiency prediction value obtained in step (3). The predicted values of target editing efficiency and bystander editing efficiency corresponding to the NA set are used to comprehensively rank the cytosine base editor candidate sgRNA set, and candidate sgRNAs with higher predicted target editing efficiency and lower predicted bystander editing efficiency are preferentially selected as the preferred sgRNAs corresponding to the cytosine base editor candidate sgRNA set. Based on the adenine base editor candidate sgRNA set obtained in step (3), the sgRNAs with higher predicted target editing efficiency and lower predicted bystander editing efficiency are preferentially selected as the preferred sgRNAs corresponding to the adenine base editor candidate sgRNA set. The preferred sgRNAs corresponding to the CRISPR / Cas9 candidate sgRNA set, the preferred sgRNAs corresponding to the cytosine base editor candidate sgRNA set, and the preferred sgRNAs corresponding to the adenine base editor candidate sgRNA set together constitute the sgRNA optimization result for the multi-editing platform.
[0049] The advantage of this step (4) is that by uniformly sorting the results of multiple indicators under different editing platforms, the comprehensive optimization of sgRNAs of multiple editing platforms can be achieved, thereby supporting the selection of the optimal editing strategy for specific pathogenic mutations.
[0050] The Prism-on model of this invention is obtained through the following steps: (5-1) Obtain experimental data related to WT-SpCas9, eSpCas9(1.1) and SpCas9-HF1 from publicly published high-fidelity Cas9 nucleotide sgRNA activity research data (the experimental data includes different sgRNA sequences and their editing efficiency values measured under the corresponding Cas9 nuclease conditions), summarize all experimental data into a CRISPR / Cas9 editing efficiency training dataset, and divide the CRISPR / Cas9 editing efficiency training dataset into a training set, a validation set and a test set in a 7:1:2 ratio; (5-2) For each sample in the training set obtained in step (5-1), the sgRNA sequence in the sample is input into the sequence encoding module of the Prism-on model for one-hot encoding to obtain the sequence encoding matrix corresponding to the sample; the sequence encoding matrix is input into the position enhancement module of the Prism-on model to obtain the position encoding vector, and the sequence encoding matrix and the position encoding vector are added element by element to obtain the enhanced sequence feature representation corresponding to the sample.
[0051] (5-3) For each sample in the training set obtained in step (5-1), the enhanced sequence feature representation corresponding to the sample obtained in step (5-2) is input into the multi-scale context feature extraction module of the Prism-on model for processing to obtain the multi-scale sequence feature corresponding to the sample. The multi-scale sequence feature is then input into the feature fusion module of the Prism-on model for channel concatenation and convolutional projection processing to obtain the fused feature. The multi-scale sequence feature and the fused feature are then input together into the residual connection unit for element-wise addition to obtain the fused enhanced sequence feature representation corresponding to the sample.
[0052] (5-4) For each sample in the training set obtained in step (5-1), the fused and enhanced sequence feature representation of the sample obtained in step (5-3) is input into the regression prediction module of the Prism-on model for processing to obtain the predicted value of the sgRNA targeting cleavage efficiency of the sample.
[0053] (5-5) For each sample in the training set obtained in step (5-1), the predicted value of the sgRNA targeting cleavage efficiency corresponding to the sample and the experimentally determined editing efficiency corresponding to the sample are input into the mean squared error loss function to obtain the training loss corresponding to the sample. (5-6) For each sample in the training set obtained in step (5-1), the Prism-on model is iteratively trained using gradient descent based on the training loss corresponding to the sample obtained in step (5-5). During the training process, the model training state is monitored and optimized using the validation set obtained in step (5-1) until the validation set loss tends to stabilize or the Prism-on model reaches the preset number of iterations (200 times in this invention). The optimal parameters of the Prism-on model during the training process are obtained, thereby obtaining the finally trained Prism-on model.
[0054] To verify the predictive performance of the Prism-on model of this invention, the test set of this invention was input into the trained Prism-on model for testing, and the Spearman correlation coefficient was used to measure the correlation between the predicted sgRNA targeting cleavage efficiency output by the model and the experimentally determined editing efficiency. Subsequently, the Prism-on model of this invention was compared with existing methods such as TransCrispr, PLM-CRISPR, and DeepMEns on the WT-SpCas9, eSpCas9(1.1), and SpCas9-HF1 datasets. The results are detailed in Table 2.
[0055] According to the experimental results in Table 2, the Prism-on model of this invention has high correlation on multiple datasets. The Spearman correlation coefficients on the WT-SpCas9, eSpCas9(1.1) and SpCas9-HF1 datasets are 0.8640, 0.8446 and 0.8538, respectively, which are better than other methods, indicating that the model has good prediction accuracy.
[0056] Table 2: Performance Comparison of CRISPR / Cas9 Editing Efficiency Prediction Models (Spearman) The Prism-be model of this invention is obtained through the following steps: (6-1) Obtain experimental data related to ABE7.10, ABEmax, ABE8e, BE4, CBE4max and Target-AID from publicly published base editor sgRNA activity research data (the experimental data includes the target sequence, the edited result sequence and the actual value of the edited result ratio determined by the corresponding base editor), summarize all experimental data into a base editor training dataset, and divide the base editor training dataset into training set, validation set and test set in a ratio of 7:1:2; (6-2) For each sample in the training set obtained in step (6-1), the target sequence and the edit result sequence in the sample are input into the dual sequence input encoding module of the Prism-be model for one-hot encoding to obtain the target sequence encoding matrix and the edit result sequence encoding matrix corresponding to the sample. The target sequence encoding matrix and the edit result sequence encoding matrix are input into the differential feature construction module of the Prism-be model for processing to obtain the differential sequence features corresponding to the sample. The target sequence encoding matrix, the edit result sequence encoding matrix and the differential sequence features are concatenated in the channel dimension to obtain the input tensor corresponding to the sample.
[0057] (6-3) For each sample in the training set obtained in step (6-1), the input tensor corresponding to the sample obtained in step (6-2) is input into the multi-scale convolutional feature extraction module of the Prism-be model for processing to obtain the multi-scale convolutional feature representation corresponding to the sample. (6-4) For each sample in the training set obtained in step (6-1), the multi-scale convolutional feature representation corresponding to the sample obtained in step (6-3) is input into the channel attention enhancement module in the Prism-be model for channel attention enhancement processing to obtain the final channel-enhanced feature representation corresponding to the sample.
[0058] (6-5) For each sample in the training set obtained in step (6-1), the final channel-enhanced feature representation of the sample obtained in step (6-4) is input into the regression prediction module of the Prism-be model for processing to obtain the predicted value of the editing result ratio of the sample.
[0059] (6-6) For each sample in the training set obtained in step (6-1), input the true value of the editing result ratio corresponding to the sample and the predicted value of the editing result ratio corresponding to the sample obtained in step (6-5) into the mean square error loss function to obtain the training loss corresponding to the sample. (6-7) For each sample in the training set obtained in step (6-1), the Prism-be model is iteratively trained using gradient descent based on the training loss corresponding to the sample obtained in step (6-5). During the training process, the model training state is monitored and optimized using the validation set obtained in step (6-1) until the validation set loss tends to stabilize or the Prism-be model reaches the preset number of iterations (200 times in this invention). The optimal parameters of the Prism-be model during the training process are obtained, thereby obtaining the finally trained Prism-be model.
[0060] To verify the predictive performance of the Prism-be model of this invention, the test set of this invention was input into the trained Prism-be model for testing, and the Spearman correlation coefficient was used to measure the correlation between the predicted value of the edit result ratio and the true value of the edit result ratio output by the model. Subsequently, the Prism-be model of this invention was compared with existing methods such as BEDeepon, DeepBE, and CRISPRon on the ABEmax, ABE8e, BE4, CBE4max, and Target-AID datasets. The results are detailed in Table 3.
[0061] According to the experimental results in Table 3, the Prism-be model of this invention has high correlation on multiple datasets. The Spearman correlation coefficients on the ABEmax, ABE8e, BE4, CBE4max and Target-AID datasets are 0.8584, 0.8790, 0.6597, 0.8466, 0.8783 and 0.8631, respectively, which are better than other methods, indicating that the model has good prediction accuracy.
[0062] Table 3: Performance Comparison of Base Editor Prediction Models (Spearman) To verify the application effect of the model described in this invention in actual disease mutation repair, sgRNA sequences reported in the literature and with experimental measurement results were selected as external validation samples. After obtaining the model's prediction results, the predicted values of the sgRNA sequences were compared with the experimental measurement values reported in the literature, and the Spearman correlation coefficient was used to evaluate the correlation between the prediction results and the experimental results, so as to quantify the model's predictive ability for actual therapeutic sgRNA editing performance.
[0063] To further illustrate the predictive performance of the model at the specific sequence level, representative sgRNA samples from three editing systems—CRISPR / Cas9, ABE, and CBE—were selected for analysis, and the results are shown in Table 4. The Spearman correlation coefficients for the three editing systems were 0.8059, 0.6545, and 0.3333, respectively.
[0064] Table 4: Information on typical sgRNA samples Under different editing systems, the predicted results of the model of this invention for sgRNA editing efficiency show high consistency with experimental measurements. This indicates that the model described in this invention can maintain good ordering consistency, that is, it can correctly distinguish between highly active and inactive sgRNAs, thereby meeting the requirements of sgRNA optimization tasks.
[0065] Therefore, by introducing an external validation method based on publicly available literature data, this invention not only verifies the accuracy of the model's prediction results, but also further illustrates the application value of the method in actual disease-related gene editing design.
[0066] Those skilled in the art will readily understand that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for optimizing sgRNA on a multi-editing platform based on pathogenic single nucleotide variants, characterized in that, Includes the following steps: (1) Obtain multiple single nucleotide variant (SNV) data, and perform preprocessing and candidate sgRNA generation processing on all SNV data in sequence to obtain a candidate sgRNA set; (2) Input the CRISPR / Cas9 candidate sgRNA set in the candidate sgRNA set obtained in step (1) into the pre-trained Prism-on model to obtain the target cleavage efficiency prediction value corresponding to the candidate sgRNA set; (3) Input the cytosine base editor candidate sgRNA set and the cytosine base editor candidate sgRNA set obtained in step (1) into the pre-trained Prism-be model to obtain the target editing efficiency prediction value and the bystander editing efficiency prediction value of the cytosine base editor candidate sgRNA set in the candidate sgRNA set, and the target editing efficiency prediction value and the bystander editing efficiency prediction value of the adenine base editor candidate sgRNA set. (4) Based on the predicted values of the target cleavage efficiency of the CRISPR / Cas9 candidate sgRNA set obtained in step (2), the predicted values of the target editing efficiency and the bystander editing efficiency of the cytosine base editor candidate sgRNA set obtained in step (3), and the predicted values of the target editing efficiency and the bystander editing efficiency of the adenine base editor candidate sgRNA set, the candidate sgRNA set is comprehensively sorted to obtain the optimal results of sgRNA for multiple editing platforms.
2. The method for optimizing sgRNA based on pathogenic single nucleotide variants on a multi-editing platform according to claim 1, characterized in that, Step (1) includes the following sub-steps: (1-1) Obtain multiple SNV data from the public database ClinVar, select all mutation sites marked as pathogenic or potentially pathogenic from all SNV data, and obtain all variant data mapped to the GRCh38 reference genome. All mutation sites and all variant data constitute the initial pathogenic SNV dataset. (1-2) Perform variant function annotation processing on the initial pathogenic SNV dataset obtained in step (1-1) to obtain the annotation results of the mutation function impact type corresponding to each pathogenic SNV data; based on the annotation results corresponding to all pathogenic SNV data, select the set of function-related mutation sites consisting of all mutation sites located in the coding region or splice region from the initial pathogenic SNV dataset (the purpose is to improve the functional relevance and interpretability of subsequent editing design); (1-3) Based on the editable mutation types of the CRISPR / Cas9 system, the adenine base editor, and the cytosine base editor, classify the set of function-related mutation sites obtained in step (1-2) to obtain the set of mutation sites corresponding to the CRISPR / Cas9 system, the set of mutation sites corresponding to the adenine base editor, and the set of mutation sites corresponding to the cytosine base editor, respectively. (1-4) Obtain the reference genome sequence corresponding to the GRCh38 reference genome obtained in step (1-1) from the UCSC genome database. Based on the chromosomal position of each mutation site in the mutation site set corresponding to the CRISPR / Cas9 system, the mutation site set corresponding to the adenine base editor, and the mutation site set corresponding to the cytosine base editor obtained in step (1-3), perform local sequence extraction processing on the obtained reference genome sequence to obtain the local sequence set corresponding to the CRISPR / Cas9 system, the local sequence set corresponding to the adenine base editor, and the local sequence set corresponding to the cytosine base editor, respectively. (1-5) Based on the preset PAM sequence constraints of the protospacer sequence, perform PAM sequence scanning on the local sequence sets corresponding to the CRISPR / Cas9 system, the adenine base editor, and the cytosine base editor obtained in step (1-4) to obtain multiple PAM sequences corresponding to the CRISPR / Cas9 system, the adenine base editor, and the cytosine base editor. Extract a 20bp protospacer sequence from the adjacent region of each PAM sequence corresponding to the CRISPR / Cas9 system within its local sequence set as the corresponding sgRNA sequence. Concatenate all sgRNA sequences with their corresponding PAM sequences to form an initial candidate sgRNA set of 23nt length corresponding to the CRISPR / Cas9 system. Extract a 20bp protospacer sequence from the adjacent region of each PAM sequence corresponding to the adenine base editor within its local sequence set. A 20bp original spacer sequence is extracted from the adjacent region of the PAM sequence as the corresponding sgRNA sequence. All sgRNA sequences are spliced with their corresponding PAM sequences to form an initial candidate sgRNA set of 23nt length corresponding to the adenine base editor. A 20bp original spacer sequence is extracted from the adjacent region of each PAM sequence corresponding to the cytosine base editor in its local sequence set as the corresponding sgRNA sequence. All sgRNA sequences are spliced with their corresponding PAM sequences to form an initial candidate sgRNA set of 23nt length corresponding to the cytosine base editor. The specific sequence of the motif adjacent to the original spacer sequence is NGG, where N represents any one of the nucleotides A, T, C, and G, A is adenine, T is thymine, C is cytosine, and G is guanine. (1-6) The initial candidate sgRNA set corresponding to the CRISPR / Cas9 system obtained in step (1-5) is screened based on mutation site coverage to retain all candidate sgRNAs whose corresponding mutation sites are located within the sequence range corresponding to the candidate sgRNAs, forming the CRISPR / Cas9 candidate sgRNA set; the initial candidate sgRNA set corresponding to the adenine base editor and the initial candidate sgRNA set corresponding to the cytosine base editor obtained in step (1-5) are screened based on the editing window position to retain all candidate sgRNAs whose corresponding target bases are located within the editing window, thereby obtaining the adenine base editor candidate sgRNA set and the cytosine base editor candidate sgRNA set, respectively. The CRISPR / Cas9 candidate sgRNA set, the adenine base editor candidate sgRNA set, and the cytosine base editor candidate sgRNA set together constitute the candidate sgRNA set.
3. The method for selecting the optimal sgRNA for a multi-editing platform based on pathogenic single nucleotide variants according to claim 1 or 2, characterized in that, The Prism-on model includes a sequence encoding module, a location augmentation module, a multi-scale contextual feature extraction module, a feature fusion module, and a regression prediction module. Step (2) includes the following sub-steps: (2-1) Input the CRISPR / Cas9 candidate sgRNA set in the candidate sgRNA set obtained in step (1) into the sequence coding module in the Prism-on model for encoding processing to obtain the sequence coding matrix; (2-2) Input the sequence encoding matrix obtained in step (2-1) into the position augmentation module in the Prism-on model for position augmentation processing to obtain the enhanced sequence feature representation. ; (2-3) Input the enhanced sequence feature representation obtained in step (2-2) into the multi-scale context feature extraction module in the Prism-on model for multi-scale context feature extraction processing to obtain multi-scale sequence features. ; (2-4) The multi-scale sequence features obtained in step (2-3) The input is processed by the feature fusion module in the Prism-on model to obtain fused features. And the fusion feature is obtained through residual connection units. and the multi-scale sequence features Element-by-element addition is performed to obtain the fused and enhanced sequence feature representation; (2-5) Input the fusion-enhanced sequence feature representation obtained in step (2-4) into the regression prediction module in the Prism-on model for regression prediction processing to obtain the target cleavage efficiency prediction value corresponding to the CRISPR / Cas9 candidate sgRNA set.
4. The preferred method for sgRNA based on pathogenic single nucleotide variants on a multi-editing platform according to any one of claims 1 to 3, characterized in that, The encoding process in step (2-1) uses one-hot encoding. Step (2-1) specifically involves first encoding the CRISPR / Cas9 candidate sgRNA set bit-by-bit using a one-hot encoding method. This means encoding the nucleotides A, T, C, and G into encoding vectors (1,0,0,0), (0,1,0,0), (0,0,1,0), and (0,0,0,1), respectively. The resulting encoding vectors for all nucleotides form a set with a size of [missing value]. The sequence encoding matrix; Step (2-2) involves introducing learnable positional encoding vectors into the sequence encoding matrix and then adding the sequence encoding matrix and the positional encoding vectors element-wise to obtain the enhanced sequence feature representation. ; in, For sequence encoding matrix, A learnable positional encoding vector; Steps (2-3) are as follows: First, local sequence patterns are extracted from the enhanced sequence feature representation through two one-dimensional convolutional layers; then, multi-scale context features are extracted from the extracted local sequence patterns through a multi-scale context aggregation structure to obtain multi-scale sequence features. Steps (2-4) are as follows: First, the multi-scale sequence features are concatenated along the channel dimension, and the concatenation result is input into the feature fusion module to obtain the fused feature. Then, the multi-scale sequence features and the fused feature are input together into the residual connection unit for element-wise addition to obtain the fused and enhanced sequence feature representation. : ; Steps (2-5) are as follows: First, the fused and enhanced sequence feature representation is flattened to obtain a one-dimensional feature vector. This one-dimensional feature vector is then sequentially input into two fully connected layers with output dimensions of 512 and 256, respectively, to obtain the processed features. Finally, the processed features are input into a single neuron linear output layer to obtain a continuous prediction result, which serves as the predicted value for the target cleavage efficiency corresponding to the CRISPR / Cas9 candidate sgRNA set.
5. The method for optimizing sgRNA based on pathogenic single nucleotide variants in a multi-editing platform according to claim 4, characterized in that, The Prism-be model includes a dual-sequence input encoding module, a differential feature construction module, a multi-scale convolutional feature extraction module, a channel attention enhancement module, and a regression prediction module; This step (3) includes the following sub-steps: (3-1) Construct the corresponding target sequence coding matrix and editing result sequence coding matrix based on the cytosine base editor candidate sgRNA set in the candidate sgRNA set obtained in step (1), and construct the corresponding target sequence coding matrix and editing result sequence coding matrix based on the adenine base editor candidate sgRNA set in the candidate sgRNA set obtained in step (1). The target sequence is the target sequence corresponding to the candidate sgRNA set, and the editing result sequence is the sequence obtained after performing all possible base substitutions on the editable bases in the editing window of the target sequence according to the editor type and editable base type corresponding to the candidate sgRNA set. (3-2) Input the first target sequence encoding matrix and the first editing result sequence encoding matrix corresponding to the adenine base editor candidate sgRNA set obtained in step (3-1) into the differential feature construction module in the Prism-be model for differential sequence feature construction processing to obtain the first differential sequence feature. The first target sequence encoding matrix, the first edited result sequence encoding matrix, and the first differential sequence feature are concatenated along the channel dimension into a first 12-channel input tensor. The second target sequence encoding matrix and the second edited result sequence encoding matrix corresponding to the cytosine base editor candidate sgRNA set obtained in step (3-1) are then input into the differential feature construction module in the Prism-be model for differential sequence feature construction to obtain the second differential sequence feature. The second target sequence encoding matrix, the second edit result sequence encoding matrix, and the second difference sequence features are concatenated in the channel dimension to form a second 12-channel input tensor. (3-3) Input the first 12-channel input tensor corresponding to the adenine base editor candidate sgRNA set obtained in step (3-2) into the multi-scale convolution feature extraction module of the Prism-be model and perform multi-scale convolution feature extraction processing to obtain the first multi-scale convolution feature representation; input the second 12-channel input tensor corresponding to the cytosine base editor candidate sgRNA set obtained in step (3-2) into the multi-scale convolution feature extraction module of the Prism-be model and perform multi-scale convolution feature extraction processing to obtain the second multi-scale convolution feature representation; (3-4) Input the first multi-scale convolutional feature representation corresponding to the adenine base editor candidate sgRNA set obtained in step (3-3) into the channel attention enhancement module in the Prism-be model for channel attention enhancement processing to obtain the feature representation after the first channel enhancement; The second multi-scale convolutional feature representation corresponding to the cytosine base editor candidate sgRNA set obtained in step (3-3) is input into the channel attention enhancement module in the Prism-be model for channel attention enhancement processing to obtain the second channel enhanced feature representation; (3-5) Input the enhanced feature representation of the first channel corresponding to the adenine base editor candidate sgRNA set obtained in step (3-4) into the regression prediction module in the Prism-be model for regression prediction processing to obtain the predicted value of the editing result ratio corresponding to the adenine base editor candidate sgRNA set; input the enhanced feature representation of the second channel corresponding to the cytosine base editor candidate sgRNA set obtained in step (3-4) into the regression prediction module in the Prism-be model for regression prediction processing to obtain the predicted value of the editing result ratio corresponding to the cytosine base editor candidate sgRNA set; (3-6) Calculate the predicted editing result ratio of the candidate sgRNA set for adenine base editor obtained in step (3-5) to obtain the predicted target editing efficiency and the predicted bystander editing efficiency of the candidate sgRNA set for adenine base editor. The predicted editing results ratios of the candidate sgRNA sets for cytosine base editors obtained in steps (3-5) are calculated to obtain the predicted target editing efficiency and the predicted bystander editing efficiency of the candidate sgRNA sets for cytosine base editors.
6. The method for optimizing sgRNA based on pathogenic single nucleotide variants on a multi-editing platform according to claim 5, characterized in that, Step (3-1) specifically involves the following steps: First, for the adenine base editor candidate sgRNA set, the adenine bases located within the editing window in the corresponding target sequence are replaced with A to G to obtain the edited result sequence corresponding to each candidate sgRNA sequence in the adenine base editor candidate sgRNA set; for the cytosine base editor candidate sgRNA set, the adenine bases located within the editing window in the corresponding target sequence are replaced with C to T to obtain the edited result sequence corresponding to each candidate sgRNA sequence in the cytosine base editor candidate sgRNA set. Then, the target sequence and the edited result sequence corresponding to the adenine base editor candidate sgRNA set are respectively input into the double sequence input encoding module in the Prism-be model for encoding processing to obtain the first target sequence encoding matrix and the first edited result sequence encoding matrix, respectively. The target sequence and the edited result sequence corresponding to the cytosine base editor candidate sgRNA set are respectively input into the double sequence input encoding module in the Prism-be model for encoding processing to obtain the second target sequence encoding matrix and the second edited result sequence encoding matrix, respectively. Specifically, the encoding process in this step uses one-hot encoding: The encoding process specifically involves first using one-hot encoding to encode each nucleotide of both the target sequence and the edited sequence bit by bit. Specifically, the nucleotides A, T, C, and G are encoded into encoding vectors (1,0,0,0), (0,1,0,0), (0,0,1,0), and (0,0,0,1), respectively. The encoding vectors corresponding to all nucleotides form a sequence of size [missing information]. The target sequence encoding matrix and the edited result sequence encoding matrix; Step (3-3) is as follows: First, the 12-channel input tensor is initially mapped through a one-dimensional convolutional layer with a kernel size of 3 to convert the 12-channel input tensor into 32-channel initial convolutional features; then, the 32-channel initial convolutional features are input into a batch normalization layer and a ReLU activation function to obtain 32-channel activation features; subsequently, the 32-channel activation features are input into the first multi-scale convolutional unit to obtain a 48-channel multi-scale feature representation. Then, the 48-channel multi-scale feature representation is input into the second multi-scale convolutional unit to obtain a 64-channel multi-scale feature representation; Each multi-scale convolutional unit employs the Inception architecture, comprising four parallel branches, namely... Convolutional branches, Convolutional branches, The convolutional branch and the max pooling branch are then combined. Finally, the outputs of each branch are concatenated along the channel dimension to obtain a multi-scale convolutional feature representation. Steps (3-4) are as follows: First, the multi-scale convolutional feature representation is input into a global average pooling unit for global average pooling to obtain a global statistical vector. Then, the global statistical vector is input into two fully connected layers and a sigmoid function to obtain channel weights between 0 and 1. Finally, the channel weights are multiplied channel-by-channel by the multi-scale convolutional feature representation to obtain the channel-enhanced feature representation. Specifically, after the first multi-scale convolutional unit, the channel attention enhancement module operates on 48 channels; after the second multi-scale convolutional unit, the channel attention enhancement module operates on 64 channels.
7. The method for selecting the optimal sgRNA for a multi-editing platform based on pathogenic single nucleotide variants according to claim 6, characterized in that, Steps (3-5) are as follows: First, the enhanced features of the channel are flattened to obtain a one-dimensional feature vector; then, the one-dimensional feature vector is sequentially input into two fully connected layers with output dimensions of 256 and 64 respectively to obtain the processed features; finally, the processed features are input into a single neuron linear output layer to obtain a continuous prediction result, which is used as the predicted value of the editing result ratio corresponding to the candidate sgRNA set. Step (3-6) specifically involves the following steps: First, based on the edited result sequences corresponding to the adenine base editor candidate sgRNA set and the cytosine base editor candidate sgRNA set obtained in step (3-1), the edited result sequences that have not undergone base substitution are selected as unedited result sequences. Then, based on the predicted edit result ratio obtained in step (3-5), the predicted edit result ratio corresponding to the unedited result sequences is obtained. And predict the proportion based on the edited result. Obtain the editor-in-chief's efficiency prediction: ; Then, based on the edited result sequences corresponding to the adenine base editor candidate sgRNA set and the cytosine base editor candidate sgRNA set obtained in step (3-1), the edited result sequences that have undergone target base substitution are selected as target edited result sequences. Simultaneously, the remaining edited result sequences in the edited result sequence set, excluding the unedited result sequences, are considered as edited result sequences. Based on the predicted edit result ratios corresponding to the candidate sgRNA sets obtained in step (3-5), the predicted edit result ratios of the target edited result sequences and the edited result sequences are obtained. Finally, the predicted target editing efficiency value is obtained based on the predicted edit result ratios of the target edited result sequences and the edited result sequences. : ; in, This represents the predicted proportion of edit results in the target edit result sequence. This represents the predicted proportion of edited results in the edited result sequence; Finally, based on the target editing efficiency prediction value Calculate the predicted value of bystander editing efficiency : ;。 8. The method for selecting the optimal sgRNA for a multi-editing platform based on pathogenic single nucleotide variants according to claim 7, characterized in that, Step (4) specifically involves: for the CRISPR / Cas9 candidate sgRNA set within the candidate sgRNA set, ranking the CRISPR / Cas9 candidate sgRNA set according to the predicted target cleavage efficiency obtained in step (2), and prioritizing candidate sgRNAs with higher predicted target cleavage efficiency as the preferred sgRNAs corresponding to the CRISPR / Cas9 candidate sgRNA set; for the cytosine base editor candidate sgRNA set and the adenine base editor candidate sgRNA set within the candidate sgRNA set, comprehensively ranking the cytosine base editor candidate sgRNA set according to the predicted target editing efficiency and the predicted bystander editing efficiency corresponding to the cytosine base editor candidate sgRNA set obtained in step (3), and prioritizing... First, candidate sgRNAs with high predicted target editing efficiency and low predicted bystander editing efficiency are selected as the preferred sgRNAs corresponding to the cytosine base editor candidate sgRNA set. Based on the adenine base editor candidate sgRNA set obtained in step (3), the candidate sgRNAs with high predicted target editing efficiency and low predicted bystander editing efficiency are selected as the preferred sgRNAs corresponding to the adenine base editor candidate sgRNA set. The preferred sgRNAs corresponding to the CRISPR / Cas9 candidate sgRNA set, the preferred sgRNAs corresponding to the cytosine base editor candidate sgRNA set, and the preferred sgRNAs corresponding to the adenine base editor candidate sgRNA set together constitute the preferred sgRNA results for the multi-editing platform. The Prism-on model is obtained through the following steps: (5-1) Experimental data related to WT-SpCas9, eSpCas9 (1.1) and SpCas9-HF1 were obtained from publicly published high-fidelity Cas9 nucleotide sgRNA activity research data. The experimental data included different sgRNA sequences and their editing efficiency values measured under the corresponding Cas9 nuclease conditions. All experimental data were summarized into a CRISPR / Cas9 editing efficiency training dataset, and the CRISPR / Cas9 editing efficiency training dataset was divided into a training set, a validation set and a test set in a ratio of 7:1:
2. (5-2) For each sample in the training set obtained in step (5-1), the sgRNA sequence in the sample is input into the sequence encoding module of the Prism-on model for one-hot encoding to obtain the sequence encoding matrix corresponding to the sample; the sequence encoding matrix is input into the position enhancement module of the Prism-on model to obtain the position encoding vector, and the sequence encoding matrix and the position encoding vector are added element by element to obtain the enhanced sequence feature representation corresponding to the sample; (5-3) For each sample in the training set obtained in step (5-1), the enhanced sequence feature representation corresponding to the sample obtained in step (5-2) is input into the multi-scale context feature extraction module of the Prism-on model for processing to obtain the multi-scale sequence feature corresponding to the sample. The multi-scale sequence feature is then input into the feature fusion module of the Prism-on model for channel concatenation and convolutional projection processing to obtain the fused feature. The multi-scale sequence feature and the fused feature are then input together into the residual connection unit for element-wise addition to obtain the fused enhanced sequence feature representation corresponding to the sample. (5-4) For each sample in the training set obtained in step (5-1), the fused and enhanced sequence feature representation of the sample obtained in step (5-3) is input into the regression prediction module of the Prism-on model for processing to obtain the predicted value of the sgRNA targeting cleavage efficiency of the sample. (5-5) For each sample in the training set obtained in step (5-1), the predicted value of the sgRNA targeting cleavage efficiency corresponding to the sample and the experimentally determined editing efficiency corresponding to the sample are input into the mean squared error loss function to obtain the training loss corresponding to the sample. (5-6) For each sample in the training set obtained in step (5-1), the Prism-on model is iteratively trained using gradient descent based on the training loss corresponding to the sample obtained in step (5-5). During the training process, the model training state is monitored and optimized using the validation set obtained in step (5-1) until the validation set loss tends to stabilize or the Prism-on model reaches the preset number of iterations. The optimal parameters of the Prism-on model during the training process are obtained, thus obtaining the finally trained Prism-on model.
9. The method for selecting the optimal sgRNA for a multi-editing platform based on pathogenic single nucleotide variants according to claim 8, characterized in that, The Prism-be model is trained using the following steps: (6-1) Obtain experimental data related to ABE7.10, ABEmax, ABE8e, BE4, CBE4max and Target-AID from publicly published base editor sgRNA activity research data. The experimental data includes the target sequence, the edited result sequence and the true value of the edited result ratio determined by the corresponding base editor. All experimental data are summarized into a base editor training dataset, and the base editor training dataset is divided into training set, validation set and test set in a ratio of 7:1:
2. (6-2) For each sample in the training set obtained in step (6-1), the target sequence and the edit result sequence in the sample are input into the double sequence input encoding module of the Prism-be model for one-hot encoding to obtain the target sequence encoding matrix and the edit result sequence encoding matrix corresponding to the sample. The target sequence encoding matrix and the edit result sequence encoding matrix are input into the differential feature construction module of the Prism-be model for processing to obtain the differential sequence features corresponding to the sample. The target sequence encoding matrix, the edit result sequence encoding matrix and the differential sequence features are concatenated in the channel dimension to obtain the input tensor corresponding to the sample. (6-3) For each sample in the training set obtained in step (6-1), the input tensor corresponding to the sample obtained in step (6-2) is input into the multi-scale convolutional feature extraction module of the Prism-be model for processing to obtain the multi-scale convolutional feature representation corresponding to the sample. (6-4) For each sample in the training set obtained in step (6-1), the multi-scale convolutional feature representation corresponding to the sample obtained in step (6-3) is input into the channel attention enhancement module in the Prism-be model for channel attention enhancement processing to obtain the final channel-enhanced feature representation corresponding to the sample. (6-5) For each sample in the training set obtained in step (6-1), the final channel-enhanced feature representation of the sample obtained in step (6-4) is input into the regression prediction module of the Prism-be model for processing to obtain the predicted value of the editing result ratio of the sample. (6-6) For each sample in the training set obtained in step (6-1), input the true value of the editing result ratio corresponding to the sample and the predicted value of the editing result ratio corresponding to the sample obtained in step (6-5) into the mean square error loss function to obtain the training loss corresponding to the sample. (6-7) For each sample in the training set obtained in step (6-1), the Prism-be model is iteratively trained using gradient descent based on the training loss corresponding to the sample obtained in step (6-5). During the training process, the model training state is monitored and optimized using the validation set obtained in step (6-1) until the validation set loss tends to stabilize or the Prism-be model reaches the preset number of iterations. The optimal parameters of the Prism-be model during the training process are obtained, thus obtaining the finally trained Prism-be model.
10. A multi-editing platform sgRNA optimization system based on pathogenic single nucleotide variants, characterized in that, Includes the following modules: The first module is used to acquire multiple single nucleotide variant (SNV) data, and to preprocess all SNV data and generate candidate sgRNAs sequentially to obtain a candidate sgRNA set. The second module is used to input the CRISPR / Cas9 candidate sgRNA set from the candidate sgRNA set obtained by the first module into the pre-trained Prism-on model to obtain the target cleavage efficiency prediction value corresponding to the candidate sgRNA set. The third module is used to input the cytosine base editor candidate sgRNA set and the cytosine base editor candidate sgRNA set obtained from the first module into the pre-trained Prism-be model to obtain the target editing efficiency prediction value and the bystander editing efficiency prediction value corresponding to the cytosine base editor candidate sgRNA set in the candidate sgRNA set, as well as the target editing efficiency prediction value and the bystander editing efficiency prediction value corresponding to the adenine base editor candidate sgRNA set. The fourth module is used to perform comprehensive sorting of the candidate sgRNA sets based on the predicted target cleavage efficiency values of the CRISPR / Cas9 candidate sgRNA sets obtained in the second module, the predicted target editing efficiency values and bystander editing efficiency values of the cytosine base editor candidate sgRNA sets obtained in the third module, and the predicted target editing efficiency values and bystander editing efficiency values of the adenine base editor candidate sgRNA sets, in order to obtain the optimal sgRNA results for multiple editing platforms.