A single-celled 6 A prediction method for methylation profile

By constructing a machine learning model based on single-cell m6A regulatory factors and cis-m6A sequence features, the complexity and accuracy problems of single-cell m6A methylation calculation in existing technologies are solved, and fast and accurate m6A methylation spectrum prediction is achieved, which is suitable for single-cell research on large-scale samples.

CN116364191BActive Publication Date: 2025-09-05GUANGXI MEDICAL UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310126069.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-16
Publication Date
2025-09-05
Estimated Expiration
2043-02-16

AI Technical Summary

Technical Problem

Existing single-cell m6A sequencing technologies such as scDART-seq are complex to operate and cannot be used for in vivo research. They are also not widely applicable to m6A methylation calculations at the single-cell level and cannot accurately identify the distribution of m6A in different RNAs or individual cells, resulting in an inability to fully understand its role in normal cell function and disease pathogenesis.

Method used

Using machine learning technology, a prediction model based on the expression profile of single-cell m6A regulatory factors and the combined characteristics of cis-m6A sequences was established. The GEO database was used to collect data, normalized and grouped into co-expression networks. An RFR model was constructed to predict m6A methylation profiles. HOMER was used to calculate the m6A motif values ​​and correlations for model training and validation.

Benefits of technology

It achieves fast and accurate single-cell m6A methylation profile prediction, with good model fitting effect, applicable to large-scale samples, fast prediction speed, high accuracy, simplified operation process, independent of experimental operation, and applicable to multiple data sources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116364191B_ABST
    Figure CN116364191B_ABST
Patent Text Reader

Abstract

The present invention discloses a method based on trans m 6 A regulatory factor expression profile and cis-m 6 A sequence combination feature predicts m in single cells 6 A methylation profiling method, including the establishment of single-cell m 6 A methylation profile prediction model and single-cell m 6 A methylation profile prediction. The batch prediction model constructed by this invention has good fitting effect. The difference between the true value and the predicted value is defined as true when it is within the range of -0.5 to 0.5, and false when it is outside the range. When the receiver operating characteristic (ROC) curve is plotted, the average area under the curve (AUC) of 4162 models reaches 0.91. The prediction speed is also fast, and the prediction time for samples containing tens of thousands of single cells only takes a few minutes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of genomic bioinformatics, and in particular relates to a method for detecting the presence of trans mRNAs in single cells. 6 A regulatory factor expression levels and cis-m 6 A sequence combination feature predicts m at the single-cell level 6 A. Abundance method. Background Art

[0002] Single-cell sequencing technology has garnered widespread attention since its initial introduction in 2009. Simply put, single-cell sequencing is a new technology that performs high-throughput sequencing analysis of the genome, transcriptome, and epigenome at the individual cell level. Its workflow primarily involves four steps: single-cell isolation, whole-genome amplification, high-throughput sequencing, and data analysis. Compared to traditional sequencing based on multi-cell signal averages, single-cell sequencing can capture individual cells from mixed tissues and obtain their genetic structure and gene expression status, revealing intercellular heterogeneity.

[0003] N6-methyladenosine (m 6 A) is the most abundant and most extensive RNA epigenetic modification within mRNA, and plays an important role in the fields of tumors, developmental biology, microbiology, neuroscience, etc., and is becoming a hot topic in life science research. 6 A modification involved in m 6 There are more than ten A regulatory factors, including: "encoder" (writer) METTL3, METTL14, WTAP, etc., "eraser" (eraser) FTO, ALKBH5 and "reader" (reader) YTH protein, IGF2BP1 / 2 / 3, HNRNPA2B1, HNRNPC, HNRNPG, etc., which can recognize chemical modification sites. Among them, the function of the encoder complex is to add methyl groups to nucleotides, and m 6 A is also regulated by DRACH sequence characteristics. Currently, the technology of calculating gene expression at the single-cell level using scRNA-seq is very mature. 6 As the most abundant methylation modification on RNA, A participates in many important cell biological processes and shows great value in the treatment of clinical diseases. 6 A mainstream high-throughput sequencing method is antibody-based MeRIP-seq, but this sequencing method has obvious disadvantages. It requires a large amount of RNA to build the library and is particularly not applicable to single-cell level research.

[0004] The scDART-seq technology developed in 2022 uses the m6A reader YTH protein to connect to the RNA C to U editing enzyme APOBEC1 to detect the m6A mutation sites caused by the YTH-APOBEC1 complex at the single cell level (https: / / doi.org / 10.1016 / j.molcel.2021.12.038). However, this technology is very complicated to operate and requires overexpression of the APOBEC1 gene. Therefore, it can only be used for cell line research and cannot be used for in vivo research. 6 A content is low, so sequencing methods are used to solve single-cell m 6 A's computational problem is impossible to solve. It can only be solved through machine learning prediction.

[0005] In addition, scDART-seq can only partially detect single-cell m 6 A, and the operation requires overexpression of APOBEC1, so it can only be used for cell line research, not for living research. 6 A content is low, the operation is complex and the reproducibility is poor. Therefore, scDART-seq only partially solves the single-cell m 6 The calculation problem of A, there is currently no widely applicable single-cell m 6 A sequencing technology meets the current research needs. The current main theory is that m 6 A methylation is affected by these cis-m 6 A regulatory factor and trans sequence characteristics of the regulation, so the use of but the cellular level is sufficient cis m 6 A regulatory factor and trans sequence characteristics can theoretically predict m 6 A level, but no one has achieved it yet. Identify m 6 The distribution of A in different RNAs or in single cells is crucial for understanding m 6 How A contributes to normal cell function and disease pathogenesis is crucial, and m 6 A has high intercellular heterogeneity, so the m 6 The development of A calculation method is very necessary. Summary of the Invention

[0006] In view of this, the present invention provides a method based on trans m 6 A regulatory factor expression profile and cis-m 6 A sequence combination feature predicts all m 6 A method to rapidly and accurately measure single-cell m 6 A. Methylation prediction.

[0007] The specific technical solutions include the following:

[0008] A single-celled 6 A method for predicting methylation profiles, including establishing single-cell m 6 A methylation profile prediction model and single-cell m 6 A methylation profile prediction, in which single cell m 6 A methylation profile prediction model includes the following steps:

[0009] Step 1: Collect all m from GEO (Gene Expression Omnibus) database 6 A regulatory factor transcriptome data and m 6 A-seq data, and then extract all m 6 A trans-regulatory factor and cis-motif sequence information, and the corresponding m 6 A regulatory network relationship of methylation levels;

[0010] Step 2: Get all m 6 Gene expression profile values ​​and m of A regulatory factors 6 A methylation spectrum value, and perform standardization conversion;

[0011] Step 3: The processed single cell m 6 A matrix was used to group the co-expression network, and HOMER was used to calculate the m value of genes in each group. 6 A motif value, the proportion of each A / C / G / T in the sequence that meets the DRACH pattern and has the smallest p-value is organized into a matrix, and compared with the single cell m 6 A regulatory factor matrix merge;

[0012] Step 4: Single cell m 6 A performs dimensionality reduction and calculates single cell m 6 A methylation profile and m 6 A regulatory factor correlation, and obtain a one-to-one correspondence between single-cell m6A trans-regulatory factor-cis motif matrix and single-cell m 6 A methylation profile matrix;

[0013] The obtained pairwise matching matrix is ​​modeled and the single cell m 6 A trans-regulatory factor and cis-motif matrix is ​​used as model input, and single cell m 6 A methylation spectrum matrix is ​​used as the target variable for supervised learning; all samples are divided into training set and test set, and the training set is subjected to 5-fold cross-validation; given the parameters, the training set is used to build the model, and the test set is used to verify the model regression prediction performance.

[0014] Preferably, after step 2, the obtained m is determined 6A. Whether there are a large number of missing values ​​in the methylation profile data, samples and genes with a missing rate > 10% are directly deleted, and positions with a missing rate < 10% are filled with the mean of the row mean and the column mean. The calculation formula is:

[0015]

[0016] Among them, Z ij represents the missing value coordinates, is the mean of the row where the missing value is located, is the mean of the column where the missing value is located.

[0017] Preferably, the step 4 further includes the following steps:

[0018] Step 4.1: Optimal parameter screening. Use grid search to calculate the optimal hyperparameters of the RFR model. The parameter n_estimators is the number of RFR model base estimators, and max_depth is the maximum depth of the RFR model base estimators. Branches exceeding the maximum depth will be pruned. The parameter grid is set to:

[0019] ′max′ depth :range(2, 15, 1) (2)

[0020] 'n' estimators :range(50,300,5) (3)

[0021] In cross-validation, the model is constructed using the combination of each two parameters n_estimators and max_depth, and then the coefficient of determination (R-Square, R 2 ):

[0022]

[0023] The numerator represents the sum of the squared differences between the true value and the predicted value, and the denominator represents the sum of the squared differences between the true value and the mean. 2 The value range is [0,1], R 2 The larger it is, the better the model fitting effect is;

[0024] Step 4.2: Use the optimal parameters and the training set data to build the model, and finally use the test set to evaluate the model. The evaluation indicators include the expected variance score, the coefficient of determination (R 2), mean absolute error (MAE), mean squared error (MSE) and median absolute error (MedSE); in the test set, the true target value is defined as y, and the average value of the true target value is The predicted target value is The calculation formula for the above evaluation indicators is:

[0025]

[0026]

[0027]

[0028]

[0029]

[0030] The expected variance, R 2 Returns a value between [0,1]. The closer to 1, the more the independent variable can explain the variance of the dependent variable. The smaller the value, the worse the effect.

[0031] MAE, MSE, and MedSE are used to evaluate the closeness between the predicted results and the actual data set. The smaller the value, the better the fitting effect.

[0032] Preferably, in step 4.1, for the screening of the optimal parameters, for each parameter combination, each verification of the 5-fold cross validation generates 1 R 2 , a total of 5 verifications are performed to generate 5 R 2 ; After that, the average R of 5 verifications is selected 2 The highest parameter combination is taken as the optimal parameter.

[0033] Preferably, all final prediction models constructed with the optimal parameter combination are saved and used for prediction.

[0034] In one embodiment, single cell m 6 A methylation profile prediction includes the following steps:

[0035] Step 5: Obtain the single-cell sequencing results of the sample and extract m 6 A. Regulatory factor gene expression matrix;

[0036] Step 6: Get the m corresponding to the sample 6 A methylation profile and preprocessing;

[0037] Step 7: Calculate the m of genes in each group6 A motif value, and m 6 A. Regulatory factor data merging;

[0038] Step 8: Based on the single cell m 6 A. Regulatory factor and motif expression levels on single cell m 6 A methylation profile is predicted.

[0039] The technical solution of the present invention has the following advantages:

[0040] 1) The single cell m disclosed in the present invention 6 A methylation profile prediction method is a genome-wide m 6 The calculation technology of A methylation spectrum is based on machine learning technology and has a fast prediction speed. The prediction model constructed using the present invention can quickly predict large-scale samples. The prediction time for samples containing tens of thousands of single cells only takes a few minutes.

[0041] 2) The batch prediction model constructed by the present invention has good fitting effect and high accuracy. The difference between the true value and the predicted value is defined as true if it is within the range of -0.5 to 0.5, and false if it is outside the range. The ROC curve is drawn, and the average AUC of 4162 models can reach 0.91.

[0042] 3) The present invention is an integrated model constructed based on a tree model, which does not require molecular manipulation of YTH similar to scDART-seq. It is easy to use, convenient for data processing, and easy to promote.

[0043] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application so that it can be implemented in accordance with the contents of the specification, and to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the following is a detailed description of the preferred embodiment of the present application in conjunction with the accompanying drawings.

[0044] Based on the detailed description of the specific embodiments of the present application in conjunction with the accompanying drawings below, those skilled in the art will become more aware of the above and other objects, advantages and features of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following will briefly introduce the drawings required for the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings without inventive work. In all drawings, similar elements or parts are generally identified by similar reference numerals. In the drawings, the elements or parts are not necessarily drawn according to the actual scale.

[0046] Figure 1 The invention discloses the establishment of single cell m 6 A. Schematic diagram of the methylation profile prediction model process;

[0047] Figure 2 The single cell m disclosed in the present invention 6 A. Schematic diagram of the process of model construction and prediction method of methylation profile;

[0048] Figure 3 The single cell m disclosed in the present invention 6 A Box plot of model performance indicators of methylation profile prediction methods. DETAILED DESCRIPTION

[0049] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of the embodiments. In the following description, specific details such as specific configurations and components are provided only to help fully understand the embodiments of the present application. Therefore, it should be clear to those skilled in the art that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present application. In addition, for clarity and brevity, the description of known functions and structures has been omitted in the embodiments.

[0050] The term "and / or" in this article is only a description of the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can mean: A exists alone, B exists alone, and A and B exist at the same time. The term " / and" in this article describes another type of association object relationship, indicating that two relationships can exist. For example, A / and B can mean: A exists alone, and A and B exist alone. In addition, the character " / " in this article generally indicates that the previous and subsequent associated objects are in an "or" relationship.

[0051] It should also be noted that, in this document, relational terms such as first and second are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "include," "comprises," or any other variations thereof are intended to cover non-exclusive inclusion.

[0052] Example 1

[0053] This example introduces a single cell m 6 A. Prediction method of methylation profile.

[0054] Reference Attachment Figure 1 、 2 As shown, Figure 1 The invention discloses the establishment of single cell m 6 A single cell m6A methylation profile prediction model flow chart. A single cell m6A methylation profile prediction method includes establishing a single cell m6A methylation profile prediction model. 6 A methylation profile prediction model and single-cell m 6 A methylation profile prediction, in which single cell m 6 A methylation profile prediction model includes the following steps:

[0055] Step 1: Collect all m from GEO (Gene Expression Omnibus) database 6 A regulatory factor transcriptome data and m 6 A-seq data, and then extract all m 6 A trans-regulatory factor and cis-motif sequence information, and the corresponding m 6 A Regulatory network relationship of methylation levels;

[0056] Step 2: Get all m 6 Gene expression profile values ​​and m of A regulatory factors 6 A methylation spectrum value, and perform standardization conversion;

[0057] Among them, after step 2, it is necessary to judge the m 6 A. Whether there are a large number of missing values ​​in the methylation profile data, samples and genes with a missing rate > 10% are directly deleted, and positions with a missing rate < 10% are filled with the mean of the row mean and the column mean. The calculation formula is:

[0058]

[0059] Among them, Z ij represents the missing value coordinates, is the mean of the row where the missing value is located, is the mean of the column where the missing value is located.

[0060] Step 3: The processed single cell m 6 A matrix was used to group the co-expression network, and HOMER was used to calculate the m value of genes in each group. 6 A motif value, the proportion of each A / C / G / T in the sequence that meets the DRACH pattern and has the smallest p-value is organized into a matrix, and compared with the single cell m 6 A regulatory factor matrix merge;

[0061] Step 4: Single cell m 6 A performs dimensionality reduction and calculates single cell m 6 A methylation profile and m 6A regulatory factor correlation, to obtain a one-to-one correspondence of single cell m 6 A trans-regulatory factor-cis-motif matrix and single-cell m 6 A methylation profile matrix;

[0062] The obtained pairwise matching matrix is ​​modeled and the single cell m 6 A trans-regulatory factor and cis-motif matrix is ​​used as model input, and single cell m 6 A methylation spectrum matrix is ​​used as the target variable for supervised learning; all samples are divided into training set and test set, and the training set is subjected to 5-fold cross-validation; given the parameters, the training set is used to build the model, and the test set is used to verify the model regression prediction performance.

[0063] The single cell m constructed by the present invention 6 A methylation profile prediction model has a good fitting effect and can effectively perform single-cell m 6 A. Methylation profile prediction.

[0064] Example 2

[0065] On the basis of Example 1, combined Figure 2 , Figure 2 The single cell m disclosed in this application 6 A Schematic diagram of the process of model construction and prediction method of methylation profile.

[0066] For step 4, we put 29173 m 6 A performs dimensionality reduction based on co-expression levels, and then calculates single-cell m 6 A methylation profile and m 6 A regulatory factor correlation (https: / / doi.org / 10.1093 / nar / gkz1206), after data processing, 38 groups of one-to-one corresponding single cell m 6 A trans-regulatory factor-cis-motif matrix and single-cell m 6 A methylation profile matrix, the obtained pairwise matching matrix is ​​modeled, single cell m 6 A regulatory factor and motif matrix are used as model input, single cell m 6 A methylation spectrum matrix was used as the target variable for supervised learning. 70% of all samples were divided into training sets and 30% into test sets. The training sets were subjected to 5-fold cross-validation, that is, the training sets were randomly divided into 5 groups, one of which was used as the validation set in sequence, and the remaining 4 groups were used as training sets.

[0067] Step 4 also includes the following steps: Step 4.1, Optimal Parameter Screening: Use grid search to calculate the optimal hyperparameters of the RFR model. The parameter n_estimators is the number of RFR model base estimators, and max_depth is the maximum depth of the RFR model base estimators. Branches exceeding the maximum depth will be pruned. The parameter grid is set to:

[0068] ′max′ depth :range(2, 15, 1) (2)

[0069] 'n' estimators :range(50,300,5) (3)

[0070] In cross-validation, the model is constructed using the combination of each two parameters n_estimators and max_depth, and then the coefficient of determination (R-Square, R 2 ):

[0071]

[0072] The numerator represents the sum of the squared differences between the true value and the predicted value, and the denominator represents the sum of the squared differences between the true value and the mean. 2 The value range is [0,1], R 2 The larger it is, the better the model fitting effect is.

[0073] Step 4.2: Use the optimal parameters and the training set data to build the model, and finally use the test set to evaluate the model. The evaluation indicators include the expected variance score, the coefficient of determination (R 2 ), mean absolute error (MAE), mean squared error (MSE), and median absolute error (MedSE). In the test set, the true target value is defined as y, and the mean of the true target values ​​is The predicted target value is The calculation formula for the above evaluation indicators is:

[0074]

[0075]

[0076]

[0077]

[0078]

[0079] The expected variance, R 2 Returns a value between [0,1]. The closer it is to 1, the more the independent variable can explain the variance of the dependent variable. The smaller the value, the worse the effect. MAE, MSE, and MedSE are used to evaluate the closeness between the predicted results and the actual data set. The smaller the value, the better the fitting effect.

[0080] In 4162 m 6 Among the 4162 regression models constructed by site A, 3063 models, accounting for 73.6% of all models, indicate that the RFR model is effective for batch prediction of m 6 A has a better prediction effect.

[0081] Furthermore, for the screening of the optimal parameters in step 4: for each parameter combination, each verification of the 5-fold cross validation generates 1 R 2 , a total of 5 verifications are performed to generate 5 R 2 ; After that, the average R of 5 verifications is selected 2 The highest parameter combination is taken as the optimal parameter.

[0082] Furthermore, all the final prediction models constructed with the optimal parameter combination are saved and predicted, and the final prediction model can be saved as a .pkl file. 6 A trans-regulatory factor and cis-motif matrix, input saved as .pkl file, can predict the m of single cell 6 A.

[0083] Through this embodiment and the combination Figure 3 Public single cell m 6 A box plot of the model performance indicators of the prediction method of methylation spectrum. The present invention has a good fitting effect and can effectively perform single-cell m 6 A methylation spectrum prediction: The prediction model constructed using the present invention can quickly predict large-scale samples. The prediction time for 100 samples only takes a few seconds, and the prediction speed is fast.

[0084] Example 3

[0085] Based on the prediction model constructed in Example 2 above, single cell m 6 A methylation profile prediction includes the following steps:

[0086] Step 5: Obtain the single-cell sequencing results of the sample and extract m 6 A. Regulatory factor gene expression matrix;

[0087] Step 6: Get the m corresponding to the sample 6A methylation profile and preprocessing;

[0088] Step 7: Calculate the m of genes in each group 6 A motif value, and m 6 A. Regulatory factor data merging;

[0089] Step 8: Based on the single cell m 6 A. Regulatory factor and motif expression levels on single cell m 6 A methylation profile is predicted.

[0090] Furthermore, HOMER is a set of tools based on C++ and Perl languages ​​for motif search and second-generation data analysis. Here, we use findMotifGenome.pl in HOMER to calculate the m of single-cell sequencing data. 6 A motif, set the parameters as follows according to the instructions on the official website:

[0091] findMotifsGenome.pl sample.bedhg38"analysis_output / "-size 200-mask-rna

[0092] The input file is the bed file of the sample, the reference gene set version used is hg38, "analysis_output / " is the custom name of the output folder, the parameter -size sets the motif length to 200bp, and the parameter -mask selects RNA sequences.

[0093] Furthermore, from the output homerResults.html file, find the low complexity and p-value-significant motif that conforms to the DRACH (D = A / G / U; R = A / G; H = A / C / U) pattern, extract the specific probability of each nucleotide (A / C / G / T), and merge it into the m obtained in the previous step. 6 A regulatory factor expression matrix.

[0094] Furthermore, the RFR prediction model .pkl file saved was used to predict the single cell m 6 A regulatory factor + motif data use the corresponding single cell m 6 A prediction model.

[0095] Furthermore, according to the single cell m 6 A regulatory factor + motif expression level affects the single cell m 6 A methylation profile for predictive applications:

[0096] Input file organization single cell m 6A regulatory factor and motif data are in standard input format, with one row representing one gene and one column representing one cell. Missing values ​​are filled with the mean of the row and column means.

[0097] Import the input data into the model in sequence, run the model and output the single cell m 6 A Methylation profile.

[0098] The batch prediction model constructed by the present invention has good fitting effect and high accuracy. Moreover, the present invention is an integrated model constructed based on a tree model, does not require experimental operation or other processing on YTH, is easy to use, and facilitates data processing.

[0099] Compared with the scDART-seq technology developed in 2022, it can only partially detect single-cell m 6 A, and scDART-seq is a method to partially solve the single-cell m 6 The problem with A detection is that it is not m 6 A's computational technology; in addition, APOBEC1 must be overexpressed in the operation, and sites where APOBEC1-YTH protein has low abundance or cannot bind will not be detected. The operation is complicated and has poor reproducibility and a high false positive rate.

[0100] In addition, the present invention is not platform-dependent and can perform predictions on data from multiple sources.

[0101] The above description is merely a preferred embodiment of the present application and does not limit the scope of protection of the present application. For those skilled in the art, various modifications and variations of the present application are possible. Any changes, modifications, substitutions, integrations, and parameter changes to these embodiments, which are made within the spirit and principles of the present application through conventional substitutions or that achieve the same functions without departing from the principles and spirit of the present application, fall within the scope of protection of the present application.

Claims

1. A single-celled m 6 A method for predicting a methylation profile, characterized in that: Including the establishment of single cell m 6 A methylation profile prediction RFR model and single cell m 6 A methylation profile prediction, in which single cell m 6 A methylation profile prediction RFR model includes the following steps: Step 1: Collect all m from GEO (Gene Expression Omnibus) database 6 A regulatory factor transcriptome data and m 6 A-seq data, and then extract all m 6 A trans-regulatory factor and cis-motif sequence information, and the corresponding m 6 A regulatory network relationship of methylation levels; Step 2: Get all m 6 Gene expression profile values ​​and m of A regulatory factors 6 A methylation spectrum value, and perform standardization conversion; Step 3: The processed single cell m 6 A matrix was used to group the co-expression network, and HOMER was used to calculate the m value of genes in each group. 6 A motif value, the proportion of each A / C / G / T in the sequence that meets the DRACH pattern and has the smallest p-value is organized into a matrix, and compared with the single cell m 6 A regulatory factor matrix merge; Step 4: Single cell m 6 A performs dimensionality reduction and calculates single cell m 6 A methylation profile and m 6 A regulatory factor correlation, to obtain a one-to-one correspondence of single cell m 6 A trans-regulatory factor-cis-motif matrix and single-cell m 6 A methylation profile matrix; The RFR model is constructed for the obtained pairwise matching matrix, and the single cell m 6 A trans-regulatory factor and cis-motif matrix is ​​used as input to the RFR model, and single-cell m 6 A methylation spectrum matrix is ​​used as the target variable of supervised learning; all samples are divided into training set and test set, and 5-fold cross-validation is performed on the training set; given parameters, the training set is used to build the RFR model, and the test set is used to verify the regression prediction performance of the RFR model.

2. A single cell m according to claim 1 6 A method for predicting a methylation profile, characterized in that: After step 2, determine the obtained m 6 A. Check whether there are a large number of missing values ​​in the methylation profile data. Samples and genes with a missing rate > 10% are directly deleted, and positions with a missing rate < 10% are filled using the formula. The calculation formula is: Among them, Z ij represents the missing value coordinates, is the mean of the row where the missing value is located, is the mean of the column where the missing value is located.

3. A single cell m according to claim 1 6 A method for predicting a methylation profile, characterized in that: The step 4 also includes the following steps: Step 4.1: Optimal parameter screening. Use grid search to calculate the optimal hyperparameters of the RFR model. The parameter n_estimators is the number of RFR model base estimators, and max_depth is the maximum depth of the RFR model base estimators. Branches exceeding the maximum depth will be pruned. The parameter grid is set to: ′max′ depth :range(2,15,1) (2) ′n′ estimators :range(50,300,5) (3) In cross-validation, the RFR model is constructed using the combination of each two parameters n_estimators and max_depth, and then the coefficient of determination (R-Square, R 2 ): The numerator represents the sum of the squared differences between the true value and the predicted value, and the denominator represents the sum of the squared differences between the true value and the mean. 2 The value range is [0,1], R 2 The larger it is, the better the RFR model fitting effect is; Step 4.2: Use the optimal parameters and the training set data to build the RFR model, and finally use the test set to evaluate the RFR model. The evaluation indicators include the expected variance score, the coefficient of determination (R 2 ), mean absolute error (MAE), mean squared error (MSE) and median absolute error (MedSE); in the test set, the true target value is defined as y, and the average value of the true target value is The predicted target value is The calculation formula for the above evaluation indicators is: The expected variance, R 2 Returns a value between [0,1]. The closer to 1, the more the independent variable can explain the variance of the dependent variable. The smaller the value, the worse the effect. MAE, MSE, and MedSE are used to evaluate the closeness between the predicted results and the actual data set. The smaller the value, the better the fitting effect.

4. A single cell m according to claim 3 6 A method for predicting a methylation profile, characterized in that: In step 4.1, for the screening of optimal parameters, for each parameter combination, each validation of the 5-fold cross validation generates 1 R 2 , a total of 5 verifications are performed to generate 5 R 2 ; After that, the average R of 5 verifications is selected 2 The highest parameter combination is taken as the optimal parameter.

5. A single cell m according to any one of claims 3 or 4 6 A method for predicting a methylation profile, characterized in that: All final prediction RFR models constructed with the optimal parameter combinations are saved and used for prediction.

6. A single cell m according to claim 1 6 A method for predicting a methylation profile, characterized in that: Based on this RFR model, single-cell m 6 A methylation profile prediction includes the following steps: Step 5: Obtain the single-cell sequencing results of the sample and extract m 6 A. Regulatory factor gene expression matrix; Step 6: Get the m corresponding to the sample 6 A methylation profile and preprocessing; Step 7: Calculate the m of genes in each group 6 A motif value, and m 6 A. Regulatory factor data merging; Step 8: Based on the single cell m 6 A. Regulatory factor and motif expression levels on single cell m 6 A methylation profile is predicted.

Citation Information

Patent Citations

  • Single-cell ATAC-seq data analysis method

    CN110544509A

  • Higher-Order Network Embedding

    US20200177466A1