A method for evaluating the quality of pluripotent stem cells

Through transcriptome and whole genome sequencing combined with bioinformatics methods, a variety of data are integrated to evaluate the pluripotency, singularity and genetic stability of pluripotent stem cells, solving the subjectivity and blindness of evaluating the quality of pluripotent stem cells in the prior art, and achieving rapid and accurate evaluation results.

CN113658636BActive Publication Date: 2025-08-05FUTURE HOMO SAPIENS INST OF REGENERATIVE MEDICINE CO LTD (FHSR) +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110830156.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-07-22
Publication Date
2025-08-05
Estimated Expiration
2041-07-22

AI Technical Summary

Technical Problem

The existing technology cannot quickly and accurately evaluate the quality of pluripotent stem cells, resulting in hindering research and clinical applications. The evaluation methods are subjective and blind, and cannot form a scientific system.

Method used

By conducting transcriptome and whole genome sequencing of the cells to be tested, combining gene expression statistics and DNA variation analysis, a variety of data were integrated using bioinformatics methods to construct a pluripotent stem cell database to evaluate the pluripotency, singularity and genetic stability of cells to form a comprehensive evaluation system.

Benefits of technology

It achieves rapid and accurate evaluation of pluripotent stem cells, reduces experimental procedures, improves the scientificity and reliability of the evaluation, and can screen out high-quality pluripotent stem cells.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113658636B_ABST
    Figure CN113658636B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for evaluating the quality of pluripotent stem cells, the method comprising the following steps: sequencing the cells to be tested to obtain transcriptome expression pedigree data and whole genome sequence data; based on the transcriptome expression pedigree data, performing statistics on the expression of genes to obtain a cell quality score of 1; based on the whole genome sequencing data, performing whole genome DNA variation analysis; according to the obtained cell quality score and the whole genome DNA variation analysis results, obtaining a cell quality score of 2 to evaluate the quality of pluripotent stem cells. The method of the present invention evaluates cell quality by comprehensively analyzing many aspects, and the present invention reduces a large number of experimental processes by using sequencing methods, and only evaluates the quality of pluripotent stem cells from the level of transcriptome sequencing and whole genome sequencing, which quickly shortens the time, and the data results obtained are reliable and accurate through statistical processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of bioinformatics, and in particular relates to a method for evaluating the quality of pluripotent stem cells. Background Art

[0002] With the rapid growth of human pluripotent stem cells, especially induced pluripotent stem cells (iPSC), there is an urgent need for a method that can well evaluate the quality of existing pluripotent stem cells. At present, the mainstream evaluation method is to evaluate the quality of cells separately from aspects such as cell morphology, pluripotency verification, karyotype detection and genetic stability marker detection. This method requires a combination of multiple technical means, such as qualified cell culture technicians, animal experiments (mouse teratoma experiments), tissue section staining, qPCR technology (fluorescence quantitative PCR), etc. In addition to the long wait time for a set of process evaluations, it must also consume a lot of manpower and material resources. Moreover, the method of verifying cell quality based on mainstream experiments is highly subjective and blind, and cannot form an accurate scientific system. It has now seriously hindered the research and clinical development of pluripotent stem cells.

[0003] Related technical research has demonstrated that qPCR can be used to verify the genetic stability of pluripotent stem cells. Molecular markers on the surface of hESC cells can clarify cell function. Using fluorescent responses of genes effectively demonstrates the genetic stability of pluripotent stem cells and can be used as a criterion for evaluating cell quality. Furthermore, transcriptome sequencing results for gene expression are highly correlated with qPCR results, and the principles behind these results are largely the same. Summary of the Invention

[0004] The present invention aims to solve at least one of the technical problems existing in the above-mentioned prior art. To this end, the present invention provides a method for evaluating the quality of pluripotent stem cells, which can quickly evaluate the quality of clinical-grade cells.

[0005] According to one aspect of the present invention, a method for evaluating the quality of pluripotent stem cells is provided, the method comprising the following steps:

[0006] S1. Sequence the cells to be tested to obtain transcriptome expression profile data and whole genome sequence data;

[0007] S2. Based on the transcriptome expression profile data of the cells to be tested, gene expression statistics are performed, and the gene expression statistical results are analyzed for pluripotency, singularity, and cell characterization values; the results of the cell pluripotency analysis, singularity analysis, and cell characterization value analysis are then integrated to obtain a cell quality score of 1; and whole-genome DNA variation analysis is performed based on the whole-genome sequence data of the cells to be tested;

[0008] S3. Based on the cell quality score obtained in step S2 and the results of the whole-genome DNA variation analysis, a cell quality score 2 is obtained to evaluate the cell quality of the pluripotent stem cells.

[0009] In some embodiments of the present invention, the pluripotent stem cells include embryonic stem cells and induced pluripotent stem cells.

[0010] In some embodiments of the present invention, in step S1, the sequence determination includes one of transcriptome sequencing, whole genome sequencing, targeted sequencing and multiplex PCR sequencing, gene chip and qPCR.

[0011] In some embodiments of the present invention, the transcriptome expression profile data includes transcriptome sequencing data.

[0012] In some embodiments of the present invention, in step S2, the statistics of gene expression levels further include performing data analysis and quality control on transcriptome expression profile data, and comparing the obtained data with a reference genome.

[0013] In some embodiments of the present invention, the data analysis and quality control include filtering low-quality data using fastp software.

[0014] In some embodiments of the present invention, the statistics of gene expression levels further include using fastp software to filter the original low-quality RNA-Seq reads of the pair-end, assembling and aligning the original RNA-Seq reads of the pair-end, aligning the reference genome, and counting the gene expression levels. Preferably, fastp software is used to filter low-quality data, HISAT2 software is used to assemble and align, and the reference genome is aligned using default parameters.

[0015] In some embodiments of the present invention, a method for evaluating the quality of multiple groups of pluripotent stem cells comprises the following steps:

[0016] (1) Sequence multiple groups of cells to be tested to obtain transcriptome expression profile data;

[0017] (2) Based on the transcriptome expression spectrum data of the cells to be tested, the gene expression statistics are performed, and the gene expression statistical results are analyzed for pluripotency, singularity, and cell characterization value respectively; then the results of the cell pluripotency analysis, singularity analysis, and cell characterization value analysis are integrated; the cell quality score 1 is obtained, and the cells are ranked;

[0018] (3) Select the top 10% of cells to be tested for sequencing to obtain whole-genome sequence data, and analyze the DNA variation of the whole genome based on the whole-genome sequence data;

[0019] (4) Based on the cell mass fractions obtained in steps (2) and (3) and the results of the whole-genome DNA variation analysis, a cell mass fraction of 2 is obtained to evaluate the cell quality of the pluripotent stem cells; the plurality of groups is greater than 10 pluripotent stem cell groups.

[0020] In some embodiments of the present invention, the reference genome is a human reference genome.

[0021] In some embodiments of the present invention, the reference genome is the human reference genome GRCh37d5.

[0022] In some embodiments of the present invention, the statistical gene expression amount is calculated using one of htseq-count software and featurecount software.

[0023] In some embodiments of the present invention, the statistical gene expression level is performed using the default parameters of htseq-count (v0.13.5) software (raw count).

[0024] In some embodiments of the present invention, the pluripotency analysis includes the following steps: using the R statistical language to convert the raw count data in the transcriptome of pluripotent stem cells and non-pluripotent stem cells, the conversion method using TDM software, the purpose of this method is to integrate sample data from different platforms, and convert the converted raw count data into data consistent with the microarray chip, using lumi (v2.42.0) software to batch correct all data, using NMF software to perform NMF (non-negative matrix decomposition) on the obtained data matrix, and then extracting the characteristic values of the pluripotent cell and non-pluripotent cell data through a machine learning model of logistic regression, adjusting the parameters according to the accuracy of the training set and the validation set, determining the formula of the classification model and calculating the pluripotency score of the unknown cells.

[0025] In some embodiments of the present invention, the singularity analysis includes the following steps: using the R statistical language to convert the raw count data obtained from the ESC (embryonic stem cell) transcriptome, the conversion method using TDM software to convert the converted raw count data into data consistent with the microarray chip. The purpose of this method is to integrate sample data from different platforms, using lumi (v2.42.0) software to batch correct all data, using NMF software to perform NMF (non-negative matrix decomposition) on the obtained data matrix, and then calculating a single score as the singularity score by comparing the residuals and root mean square error (RMSE) of the decomposed V matrix of the unknown sample and the embryonic stem cell sample.

[0026] In some embodiments of the present invention, obtaining the cell characterization value includes obtaining the expression levels of relevant cell genetic stability genes and cell function genes from transcriptome expression profile data, and calculating the baseline deviation value from the embryonic stem cell sample using R statistical language.

[0027] In some embodiments of the present invention, the cell function genes include embryonic stem cell characteristic genes and iPS cell characteristic genes; the embryonic stem cell characteristic genes include SOX2, OCT4, NANOG, SSEA-4, TRA-1-60, TRA-1-81 and SSEA-1; the iPS cell characteristic genes include SSEA3, SSEA4, TRA-1-60, TRA-1-81, OCT4 and NANOG.

[0028] In some embodiments of the present invention, the cell genetically stable genes include tert, TET1, TET3, Sirt1, CHK1, Oct4-endo, OCT4, Nanog and P53.

[0029] In some embodiments of the present invention, the results of the integrated cell pluripotency analysis, the singularity analysis, and the cell characterization value analysis are analyzed using the following formula:

[0030]

[0031]

[0032] C=∑F i -----------Formula 3;

[0033] Among them, F in formula 1 i is the score for evaluating cell quality, is the score for evaluating pluripotency, and X i Refers to the value calculated under a certain cell quality evaluation, X max is the maximum threshold of the cell quality evaluation design, X min is the minimum threshold of the cell evaluation design, q represents the number of cell quality evaluation aspects; F in formula 2 i It is an evaluation score of cell quality evaluation, which is one of the singularity and cell characterization values. i Refers to the value calculated under a certain cell quality evaluation, X max is the maximum threshold of the cell quality evaluation design, X min is the minimum threshold of the cell evaluation design, q represents the number of cell quality evaluation aspects; C in formula 3 represents the cell quality score 1.

[0034] In some embodiments of the present invention, the analysis of the cell characterization value includes the following steps: obtaining the expression levels of relevant cell genetic stability genes and cell function genes from the transcriptome expression spectrum data, and statistically calculating the baseline deviation value of the gene expression level and the expression level of the gene in the embryonic stem cell sample.

[0035] In some embodiments of the present invention, the statistical calculation of the deviation between the gene expression level and the baseline of the embryonic stem cell sample is performed using the R statistical language.

[0036] In some embodiments of the present invention, in step S3, the processing of the whole genome data further includes the steps of performing data analysis and quality control on the whole genome data, and comparing the obtained data with the reference genome.

[0037] In some embodiments of the present invention, the data analysis and quality control include filtering low-quality data using fastp software.

[0038] In some embodiments of the present invention, whole genome data processing also includes using fastp software to process and clean pair-end original raw reads (FASTQ) data, assembling and aligning the pair-end original raw reads, aligning the reference genome, and analyzing the variation of the whole genome. Preferably, fastp software is used for processing and analysis, BWA software is used for assembly and alignment, the reference genome is aligned using default parameters, and the DNA variation of the whole genome is analyzed.

[0039] In some embodiments of the present invention, the reference genome is a human reference genome.

[0040] In some embodiments of the present invention, the reference genome is the human reference genome GRCh37d5.

[0041] In some embodiments of the present invention, in step S3, the whole genome DNA variation analysis includes point mutation analysis, insertion / deletion analysis, copy number variation analysis, and large-segment chromosome variation analysis.

[0042] In some embodiments of the present invention, in step S3, the whole genome DNA variation analysis uses GATK4 software and cnvnator software.

[0043] In some embodiments of the present invention, the whole genome DNA variation analysis uses the HaplotypeCaller component of the GATK4 software to perform SNPs / Indel calling on the bam file after a series of processing, and the obtained vcf file is annotated at the disease level using a professional database.

[0044] In some embodiments of the present invention, the whole genome DNA variation analysis uses cnvnator software to perform sliding window analysis on the sorted bam files, and the obtained vcf files are also annotated at the disease level using a professional database.

[0045] In some embodiments of the present invention, whole-genome DNA variation analysis includes analyzing the total number of cellular DNA mutations, the number of genetic diseases, the number of potentially harmful mutations, and the number of benign mutations in the whole-genome sequence data, and then integrating the results of the analysis of the total number of cellular DNA mutations, the number of genetic diseases, the number of potentially harmful mutations, and the number of benign mutations to obtain a genetic variation score for the cell.

[0046] In some embodiments of the present invention, the results of genome-wide DNA variation analysis are calculated using the following formula:

[0047]

[0048] V=∑F i ------------Formula five.

[0049] Among them, F in formula 4 i is the score of cell variation evaluation, X i The number of variations under a certain variation factor, X max The maximum value of each factor of variation, X min is the minimum value of each variant factor, and v in Formula 5 represents the genetic variation score.

[0050] In some embodiments of the present invention, the variation factors include the total number of mutations, the number of genetic diseases, the number of potentially harmful variations, and the number of benign variations; X max Each group of pluripotent stem cells X under each variable factor i The minimum value in X min The number of pluripotent stem cell mutations in each group under each mutation factor X i The minimum value in .

[0051] In some embodiments of the present invention, the method further comprises a step of culturing the cells before sequencing, wherein the step comprises culturing the cells to be tested using Essential 8™ culture medium.

[0052] In some embodiments of the present invention, the chromosomal variation includes chromosomal deletion, duplication, translocation, and inversion.

[0053] A method for constructing a pluripotent stem cell database comprises the following steps: constructing a pluripotent stem cell database by integrating transcriptome data of different pluripotent stem cells and non-pluripotent stem cells.

[0054] A method for constructing a pluripotent stem cell database comprises the following steps: using the R statistical language to convert raw count data in the transcriptomes of pluripotent stem cells and non-pluripotent stem cells, integrating sample data from different sequencing platforms, and converting the converted raw count data into data consistent with a microarray chip; and using lumi software to perform batch correction on all data to construct a pluripotent stem cell database.

[0055] According to the embodiments of the present invention, there are at least the following beneficial effects: the present invention scheme performs bioinformatics analysis on transcriptome gene data and whole genome data, integrates the results of the bioinformatics analysis of the obtained transcriptome gene data and whole genome data, and quickly calculates the cell quality score; the method of the present invention is simple, and the present invention scheme is the first to comprehensively evaluate the quality of cells through information on cell function and cell genome variation, and can effectively screen the quality of induced pluripotent stem cells (iPSC) or embryonic stem cells (ESC). In addition, the present invention uses sequencing to reduce a large number of experimental processes, and only evaluates the quality of pluripotent stem cells from the level of transcriptome sequencing and whole genome sequencing, which not only quickly shortens the time, but also the data results obtained are reliable and accurate through statistical processing. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] The present invention will be further described below with reference to the accompanying drawings and embodiments, in which:

[0057] Figure 1 Flowchart of bioinformatics analysis of transcriptome expression profile data in Example 1 of the present invention;

[0058] Figure 2 This is a flow chart of the bioinformatics analysis of whole genome data in Example 1 of the present invention;

[0059] Figure 3 This is a diagram of pluripotency analysis of cells in Example 1 of the present invention;

[0060] Figure 4 This is a flow chart for evaluating the quality of pluripotent stem cells in Example 1 of the present invention. DETAILED DESCRIPTION

[0061] The following will clearly and completely describe the concept and technical effects of the present invention in conjunction with the embodiments to fully understand the purpose, features and effects of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, other embodiments obtained by those skilled in the art without creative work are all within the scope of protection of the present invention.

[0062] Cell culture medium: Essential 8TM medium (Thermo Fisher SCIENTIFIC); sequencing platform selected MGISEQ-2000 (purchased from BGI) and its supporting whole genome and transcriptome library reagents; total RNA extraction and DNA extraction were selected Kit and PureLink™ Kit (Thermo Fisher SCIENTIFIC).

[0063] Example 1

[0064] This embodiment provides a method for evaluating the quality of pluripotent stem cells, comprising the following steps:

[0065] 1. After the cells to be tested were cultured with Essential 8TM medium, total RNA was extracted using a total RNA extraction kit ( Kit) and a DNA extraction kit (PureLink™ kit) were used to extract DNA and RNA, and the extracted DNA and RNA were sequenced. The sequencing platform selected was MGISEQ-2000, and BGI and its supporting whole genome and transcriptome library reagents were used; the transcriptome expression spectrum data and whole genome sequence data were obtained (if multiple groups of pluripotent stem cells are evaluated, RNA is first extracted for sequencing, and the obtained transcriptome expression spectrum data are analyzed, and the multiple groups of pluripotent stem cells are ranked, and the top 10% of the cells are selected for whole genome sequence determination. Multiple groups, i.e., pluripotent stem cell groups, are greater than 10 groups).

[0066] 2. Analyze the transcriptome expression profile data obtained in step 1: Use fastp software to process and analyze paired-end raw RNA-Seq reads (FASTQ). Paired-end reads are assembled and aligned using HISAT2 software and aligned to the human reference genome (GRCh37d5) using default parameters. Gene expression statistics (raw counts) are performed using htseq-count (v0.13.5) software with default parameters. The bioinformatics analysis process for transcriptome expression profile data is as follows: Figure 1 shown.

[0067] 3. Analyze the whole genome data obtained in step 1: Use fastp software to process and clean the pair-end original raw reads (FASTQ) data, use the pair-end reads to assemble and align with the reference genome using BWA software (v0.7.1), and use the default parameters to align the human reference genome (GRCh37d5) to form a sam file. Use samtools (v1.12) software to convert the sam file to a bam file to facilitate file storage and sorting of the internal genome sequence. The whole genome DNA variation analysis calls the HaplotypeCaller component of the GATK4 software to perform SNPs / Indel calling on the bam file after a series of processing, and the obtained vcf file is annotated at the disease level using a professional database. The CNV (copy number variation) process uses cnvnator software to perform sliding window analysis on the sorted bam file, and the obtained vcf file is also annotated at the disease level using a professional database; the process of bioinformatics analysis of whole genome data is as follows Figure 2 shown.

[0068] 4. Cell quality analysis

[0069] (1) Analysis of pluripotency

[0070] The R statistical language was used to transform the raw count data from the transcriptomes of pluripotent stem cells and non-pluripotent stem cells. The transformation method used TDM software (Thompson, 2017). The purpose of this method is to integrate sample data from different sequencing platforms and convert the converted raw count data into data consistent with the microarray chip. All data were batch corrected using lumi (v2.42.0) software. The obtained data matrix was subjected to NMF (non-negative matrix factorization) using NMF software. The characteristic values of the pluripotent cell and non-pluripotent cell data were extracted using a logistic regression machine learning model. The parameters were adjusted based on the accuracy of the training set and the validation set. The highest threshold of pluripotency for the logistic regression model constructed based on the current data set was 94.21, and the lowest threshold was 23.344 (as shown in Figure 2). Figure 3 As shown in the figure, cells with a score higher than 23.344 are all pluripotent cells, and cells with a score lower than this threshold or even negative are non-pluripotent cells. The classification model is determined and the pluripotency score of unknown cells is calculated. It can be seen from the figure that the classification model of the present invention can accurately distinguish pluripotent stem cells from non-pluripotent stem cells.

[0071] (2) Analysis of singularity

[0072] The raw count data obtained from the known ESC (embryonic stem cell) transcriptome were converted using the R statistical language. The conversion method used TDM software (Thompson, 2017) to convert the converted raw count data into data consistent with the microarray chip. The purpose of this method is to integrate sample data from different platforms. All data were batch corrected using lumi (v2.42.0) software. After performing NMF (non-negative matrix decomposition) on the obtained data matrix using NMF software, a single score was calculated by comparing the residuals and root mean square error (RMSE) of the V matrix of the decomposed unknown sample and the embryonic stem cell sample (the high-quality embryonic stem cell sample from the data set constructed in step (1)) as the singularity score.

[0073] (3) Analysis of cell characterization values

[0074] The expression levels of relevant cell genetic stability genes and cell function genes are obtained from the transcriptome expression profile data. The relevant genes include embryonic stem cell characteristic genes (SOX2, OCT4, NANOG, SSEA-4, TRA-1-60, TRA-1-81 and non-expressed SSEA-1 gene), iPS cell characteristic genes (SSEA3, SSEA4, TRA-1-60, TRA-1-81, OCT4 and NANOG), cell genetic stability genes (ert, TET1, TET3, Sirt1, CHK1, Oct4-endo, OCT4, Nanog and P53), and the baseline deviation value of the gene expression level from the embryonic stem cell sample (from the high-quality embryonic stem cell sample in the data set constructed in step (1)) is calculated using R statistical language.

[0075] (4) Comprehensive evaluation and quantifiable analysis

[0076] The deviation values of the above-mentioned pluripotency, singularity, and cell characterization molecules are statistically converted into a cell quality score, and a comprehensive ranking is performed. The ranking formula is as follows:

[0077]

[0078]

[0079] C=∑F i -----------Formula 3;

[0080] Among them, F in formula 1 i is the score for evaluating cell quality, is the score for evaluating pluripotency, and X i Refers to the value calculated under a certain cell quality evaluation, X max is the maximum threshold of the cell quality evaluation design, Xmin is the minimum threshold of the cell evaluation design, q represents the number of cell quality evaluation aspects; F in formula 2 i It is an evaluation score of cell quality, which is one of the singularity and cell characterization aspects. i Refers to the value calculated under a certain cell quality evaluation, X max is the maximum threshold of the cell quality evaluation design, X min is the minimum threshold of the cell evaluation design, q represents the number of cell quality evaluation aspects (q is 3); C in formula 3 represents the cell quality score 1.

[0081] The highest threshold in pluripotency analysis was 94.21, and the lowest threshold was 23.344; the highest threshold in singularity analysis was 1.17, and the lowest threshold was 0.46; the highest threshold in cell characterization value analysis was 1, and the lowest threshold was 0.

[0082] 5. Whole-genome DNA variation analysis (disease analysis at the genetic level)

[0083] After ranking by cell quality score 1, the top 10% of cells that require large-scale screening are selected for whole-genome sequencing to save sequencing costs. The sequencing depth is guaranteed to be greater than or equal to 30X, in line with the generally accepted sequencing standards. The disease annotation results of CNVs and SNPs / Indels are then classified according to the "ACMG Standards and Guidelines for the Classification of Genetic Variations", cells with genetic risks are marked, and relevant literature is consulted to confirm authenticity. Cells at the genetic level are graded and scored to confirm that the new mutations are not from the donor, and the number of such mutations is counted. In combination with the "ACMG Standards and Guidelines for the Classification of Genetic Variations", harmful mutations and potentially benign mutations are counted, with the more such mutations present, the more dangerous the cells are.

[0084]

[0085] V=∑F i ------------Formula five.

[0086] Among them, F in formula 4 i is the score of cell variation evaluation, X i The number of variations under a certain variation factor, X max The maximum value of each factor of variation, X min is the minimum value of each variant factor, and v in Formula 5 represents the genetic variation score.

[0087] The variation factors include the total number of mutations, the number of genetic diseases, the number of potentially harmful mutations, and the number of benign mutations; X max The minimum value of each group of pluripotent stem cells Xi under each variable factor; Xmin The minimum value of the number of pluripotent stem cell mutations Xi in each group under each mutation factor.

[0088] 6. Comprehensive evaluation and analysis

[0089] The average of the values (C) obtained in step 4 and the values (V) obtained in step 5 is taken as the cell mass score 2, and the ranking is performed again.

[0090] The process of the method for evaluating the quality of pluripotent cells in this embodiment is as follows: Figure 4 shown.

[0091] Test Example 1

[0092] The above method was used to evaluate the quality of 20 stem cell samples of unknown pluripotency, as shown in Table 1. The above cell abbreviations were subjected to transcriptome analysis. The culture medium of the 20 groups of pluripotent stem cell samples in Table 1 was selected from Essential8 TM Culture medium was used for culture, and the sequencing platform selected was MGISEQ-2000 (BGI) and its supporting whole genome and transcriptome library reagents; total RNA extraction and DNA extraction were selected Kits and PureLink TM Kit (ThermoFisher SCIENTIFIC). Transcriptome expression profile data and whole genome sequence data were obtained.

[0093] The transcriptome expression profile data obtained from C1-C20 cells were used to assess stem cell quality using the method for assessing pluripotent stem cell quality in Example 1. Table 2 shows the data obtained from bioinformatics analysis of the transcriptome expression profile data. As can be seen from the table, different genes are expressed at varying levels in different stem cells. R statistical software was used to analyze the pluripotency, singularity, and deviation of cell-characterizing molecules. The scores were then calculated to produce a cell quality score of 1, and the top 10% of cells were ranked. The results are shown in Table 3.

[0094] The whole genome sequence data obtained from C1-C20 cells were used to assess stem cell quality according to the method for assessing pluripotent stem cell quality in Example 1. The whole genome sequencing data was subjected to bioinformatics analysis and disease-level annotation, and statistical variation data and genetic variation scores were calculated as shown in Table 4 below.

[0095] The quality of 20 stem cell samples of unknown pluripotency was evaluated as shown in Table 5. After taking the average of the cell quality score 1 and the genetic variation score to form the cell quality score 2, they were sorted. It can be seen that the best cell in this batch is C6.

[0096] Table 1

[0097] Sample name Cell type C1-C60 iPS cells

[0098] Table 2

[0099]

[0100]

[0101]

[0102]

[0103]

[0104] Table 3

[0105]

[0106]

[0107] Table 4

[0108]

[0109]

[0110] Table 5

[0111]

[0112]

[0113] Cell quality assessment is a mixed bag. Is there a scientifically sound system for its implementation? Furthermore, methods based on mainstream experiments to verify cell quality are highly subjective and uninformed, resulting in significant time and expense and potentially inaccurate results. Related technologies propose methods for verifying pluripotency and singularity, chromosomal aberrations, cellular genetic stability, and cell function, all based on single-model assessments, lacking a cohesive process framework.

[0114] In the evaluation of the pluripotency and singularity of cells, related technologies use machine learning to predict the pluripotency of unknown cell samples based on microarray gene expression datasets. Unfortunately, the drawbacks of this method have become increasingly prominent. Existing sequencing systems have begun to phase out the original microarray sequencing technology and have been replaced by a new generation of sequencing technology, making it impossible for the original dataset to correctly evaluate the existing data. In addition, this technical approach was limited by the understanding of embryonic stem cells and iPS cells at the time, the defects of cell culture technology, and differences between races and genders. The paper did not make certain distinctions. Mixed experimental data is prone to erroneous evaluations and is not allowed for clinical-grade cell quality assessments.

[0115] In the assessment of genetic variation, a method has been developed that can evaluate chromosomal aberrations based on transcriptome expression data. However, transcriptome expression profiles represent only a portion of the whole genome data, and the genetic variation cannot be explained at the disease level. Technically, expression profile data cannot accurately represent microdeletions and multiple deletions in some genomes. This can lead to the appearance that some chromosomal deletions do not cause disease, and the lack of precision can also result in diseases not being detected, leading to incorrect assessments of cell quality and safety.

[0116] In the evaluation of cell-characteristic molecules (cell genetic stability and cell function), experiments based on qPCR technology need to be carried out separately, but the evaluation of transcriptome data can also obtain valid data.

[0117] A single evaluation method can easily overlook other equally important aspects of cell quality, and an innovative evaluation method that can integrate the above factors is urgently needed.

[0118] Currently, the mouse teratoma assay remains the gold standard for evaluating pluripotency and the core defining characteristic of all pluripotent stem cells (PSCs). However, the lack of standardized experimental methods for assessing this characteristic has drawn scrutiny from researchers. The mouse teratoma assay primarily involves the generation of well-differentiated teratomas following injection of pluripotent stem cells into immunodeficient mice. Although this method is nonquantitative and subjective, skilled pathologists mastering teratoma histology can distinguish between tumors composed primarily of poorly differentiated neuroectoderm and cystic masses composed of highly differentiated tissue from all three embryonic germ layers. The former presents as a malignant tumor, resembling a teratoma, while the latter presents as an encapsulated, benign mass, representing a true teratoma derived from pluripotent stem cells. However, the tedious and time-consuming nature of the assay, coupled with the use of animals and the requirement for expert pathological evaluation, limits the use of the teratoma assay as a routine screening tool. With the development of microarray (chip) sequencing technology, some studies have adopted a cost-effective, animal-free alternative to teratoma detection to assess the pluripotency of human cells. This method can predict the pluripotency of unknown cell samples through machine learning based on the gene expression dataset of the microarray. Unfortunately, the disadvantages of this method have become increasingly prominent. The existing sequencing system has begun to phase out the original microarray sequencing technology and replace it with the second-generation sequencing technology or the third-generation sequencing technology, making it impossible to correctly evaluate the existing data with the original dataset. In addition, due to the limitations of the understanding of embryonic stem cells and iPS cells at the time, the defects of cell culture technology, and the differences between races and genders, if this method is still used to evaluate Chinese embryonic stem cells and iPS cells, there will be a large bias, resulting in frequent inaccuracies.

[0119] For the detection of chromosome karyotype, most experiments still use tissue section staining technology, which can visually observe the structure and number of chromosomes under a microscope. However, when a large number of pluripotent stem cells need to be screened for clinical and scientific research purposes, this method is no longer the best choice. Related technologies have developed a method that can evaluate chromosome aberrations based on transcriptional profile data, using the average gene expression value of the total sample or the degree of variation of SNPs (single nucleotide polymorphisms) as a standard for measuring whether chromosome variation has occurred. Compared with the copy number variation and mutation analysis results under whole genome sequencing, whole genome sequencing covers a wider range and has higher resolution accuracy. The results of a single sample can be annotated to the disease level, which is incomparable to transcriptome data.

[0120] While the embodiments of the present invention have been described in detail above with reference to the accompanying drawings, the present invention is not limited to the embodiments described above. Various modifications may be made within the scope of knowledge possessed by a person skilled in the art without departing from the spirit of the present invention. Furthermore, the embodiments of the present invention and the features thereof may be combined with one another unless there is a conflict.

Claims

1. A method for evaluating the quality of pluripotent stem cells, characterized in that: The method comprises the following steps: S1. Sequence the cells to be tested to obtain transcriptome expression profile data and whole genome sequence data; S2. Based on the transcriptome expression profile data of the cells to be tested, the gene expression statistics are performed, and the gene expression statistical results are respectively analyzed for pluripotency analysis, singularity analysis, and cell characterization value analysis. Then, the results of the cell pluripotency analysis, singularity analysis, and cell characterization value analysis are integrated to obtain the cell quality score 1; The following formula is used to integrate the results of the analysis of cell pluripotency, analysis of singularity, and analysis of cell characterization value: --Formula 1; -Formula 2; -----------Formula 3; Among them, in formula 1 F i is the score for evaluating cell quality, is the score for evaluating pluripotency, X i It refers to the value calculated under the quality evaluation of a certain cell, that is, the result of the pluripotency analysis. X max This is the maximum threshold of the cell quality evaluation design, specifically 94.

21. X min is the minimum threshold of the cell evaluation design, specifically 23.344, q represents the number of cell quality evaluation aspects, specifically 3; in formula 2 F i It is an evaluation score of cell quality evaluation, which is one of the singularity and cell characterization values. X i refers to the value calculated under the quality evaluation of a certain cell, that is, the result of the singularity analysis, X max is the maximum threshold of the cell quality evaluation design, specifically 1.17, X min is the minimum threshold of the cell evaluation design, specifically 0.46; when X i refers to the value calculated under the quality evaluation of a cell, that is, the result of the analysis of the cell characterization value, X max is the maximum threshold of the cell quality evaluation design, specifically 1, X min It is the minimum threshold of the cell evaluation design, specifically 0, q represents the number of cell quality evaluation aspects 3; C in formula 3 represents the cell quality score 1; The analysis of the results of the pluripotency analysis includes the following steps: using R statistical language to convert the raw count data in the transcriptome of pluripotent stem cells and non-pluripotent stem cells, the conversion method using TDM software, the purpose of this method is to integrate sample data from different platforms, and convert the converted raw count data into data consistent with the microarray chip, using lumiv2.42.0 software to batch correct all data, using NMF software to perform NMF (non-negative matrix decomposition) on the obtained data matrix, and then extracting the characteristic values of the pluripotent cell and non-pluripotent cell data through a logistic regression machine learning model, adjusting parameters based on the accuracy of the training set and the validation set, determining the formula of the classification model and calculating the pluripotency score of the unknown cell. X i ; The singularity analysis includes the following steps: using the R statistical language to perform data conversion on the raw count data obtained from the embryonic stem cell transcriptome, using TDM software to convert the converted raw count data into data consistent with the microarray chip. The purpose of this method is to integrate sample data from different platforms, using lumi v2.42.0 software to perform batch correction on all data, using NMF software to perform NMF (non-negative matrix factorization) on the obtained data matrix, and then calculating a single score by comparing the residual and root mean square error (RMSE) of the decomposed V matrix of the unknown sample and the embryonic stem cell sample to use as the singularity score. X i ; The acquisition of the cell characterization value includes obtaining the expression levels of relevant cell genetic stability genes and cell function genes from the transcriptome expression spectrum data, and calculating the baseline deviation value from the embryonic stem cell sample using the R statistical language; At the same time, a whole-genome DNA variation analysis is performed based on the whole-genome sequence data of the cells to be tested to obtain a genetic variation score; the whole-genome DNA variation analysis includes point mutation analysis, insertion and deletion analysis, copy number variation analysis, and large-segment chromosome variation analysis; The results of the genome-wide DNA variation analysis were calculated using the following formula: --------Formula 4; ------------ Formula 5; Among them, Formula 4 F i is the score for cell variation evaluation, X i The number of variations under a certain variation factor, X max is the maximum value of each factor of variation, X min is the minimum value of each variation factor, where v in Formula 5 represents the genetic variation score; the variation factors include the total number of mutations, the number of genetic diseases, the number of potentially harmful variations, and the number of benign variations; S3. The cell quality of the pluripotent stem cells is evaluated by averaging the cell quality score 1 obtained in step S2 and the genetic variation score obtained by genome-wide DNA variation analysis to obtain a cell quality score 2.

2. The method according to claim 1, characterized in that In step S1, the statistics of gene expression also include the steps of performing data analysis and quality control on the transcriptome expression spectrum data, and comparing the obtained data with the reference genome.

3. The method according to claim 2, characterized in that The reference genome is the human reference genome.

4. The method according to claim 1, wherein In step S1, the statistical gene expression level is calculated using one of htseq-count software and featurecount software.

5. The method according to claim 1, characterized in that The cell function genes include embryonic stem cell characterization genes and iPS cell characterization genes; the embryonic stem cell characterization genes include SOX2 、 OCT4 、 NANOG 、 SSEA-4 、 TRA-1- 60 、 TRA-1-81 and SSEA-1 ; The iPS cell characterization genes include SSEA3 、 SSEA4 、 TRA-1-60 、 TRA-1-81 、 OCT4 and NANOG .

6. The method according to claim 1, characterized in that The cell genetically stable genes include tert 、 TET1 、 TET3 、 Sirt1 、 CHK1 、 Oct4-endo 、 OCT4 、 Nanog and P53 .

7. The method according to claim 1, characterized in that In step S3, the processing of the whole genome data also includes the steps of performing data analysis and quality control on the whole genome data, and comparing the obtained data with the reference genome.

Citation Information

Patent Citations

  • Cell characterisation

    CN103403180A

  • Method for Assessing the Quality of Various Cells Including Induced Pluripotent Stem Cells

    US20180163269A1