A method, apparatus and system for predicting cell cycle time

By extracting the characteristic gene expression matrix from cell sequencing data and calculating the likelihood function, combined with Markov chain Monte Carlo sampling, the problems of time-consuming, labor-intensive, and low-resolution methods in existing technologies are solved, achieving efficient and accurate cell cycle time prediction.

CN116153398BActive Publication Date: 2026-03-31WEST CHINA HOSPITAL SICHUAN UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-22
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing technologies are time-consuming and labor-intensive in acquiring cell cycle stage information, have a limited number of cells to process, are invasive, and provide low-resolution information that makes it difficult to accurately characterize cell cycle stages and affects downstream analysis.

Method used

By extracting the characteristic gene expression matrix from cell sequencing expression data, fitting the cosine function and calculating the likelihood function, and combining Markov chain Monte Carlo sampling, a stationary Markov chain is constructed to determine the time period of the cell.

Benefits of technology

It enables efficient and accurate prediction of cell cycle time, applicable to single-cell, cell line, or tissue samples, with accuracy superior to existing methods, and is suitable for unlabeled cell sequencing data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116153398B_ABST
    Figure CN116153398B_ABST
Patent Text Reader

Abstract

The application discloses a method for predicting cell cycle time, which comprises the following steps: extracting a characteristic gene expression matrix from obtained cell sequencing expression data, fitting the characteristic gene expression matrix by using a cosine function, calculating a likelihood function of the characteristic gene expression, and obtaining a continuous time sequence of a cell cycle stage or a cycle stage according to maximum likelihood function. The method is applicable to any unlabeled cell sequencing data, and the accuracy is obviously superior to that of existing prediction methods.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a method, device, and system for predicting cell cycle time. Background Technology

[0002] The cell cycle is a series of events that occur during cell growth and division, and a complete cell cycle includes four distinct phases. The cell cycle is known to participate in fundamental biological processes, and its important relationship with development and differentiation has been demonstrated in various systems (e.g., humans, plants, and mice). Furthermore, dysregulation of the cell cycle is a major characteristic of cancer and an important target for cancer therapy. For single-cell transcriptome data, besides technical noise, cell type and cell cycle are undoubtedly the biggest causes of gene expression differences. Especially in the aforementioned developmental or cancer treatment studies, the cell cycle is often considered a confounding factor. Therefore, accurately characterizing cell cycle phases, quantifying cell cycle effects, and correcting for downstream analyses are crucial for a thorough understanding of biological issues.

[0003] Currently, obtaining information about cell cycle stages through experimental methods is a common approach. Firstly, flow cytometry staining is used, typically employing flow cytometry to sort and enrich cells at different cycle stages, followed by library construction and sequencing. This method works by using DNA dyes to label cells and then sorting them based on the varying DNA content at different cell cycle stages. Secondly, imaging-based fluorescence detection methods are used. These methods rely on genetic manipulation strategies, inserting fluorescent probes (FUCCI technology) into differentially expressed genes at different cell cycle stages, combined with fluorescence imaging to determine the cell cycle. However, these methods are time-consuming and labor-intensive, process a limited number of cells, and are invasive, potentially interfering with subsequent cell studies. Furthermore, these methods yield discrete classifications of the four cell cycle stages, representing a low-resolution characterization. Summary of the Invention

[0004] To address the aforementioned problems, this invention provides a method for predicting cell cycle time. This method involves extracting a feature gene expression matrix from acquired cell sequencing expression data, fitting the feature gene expression matrix using a cosine function, calculating the likelihood function of the feature gene expression, and determining the cell's time period based on maximizing the likelihood function.

[0005] Furthermore, the formula for calculating the likelihood function is as follows:

[0006]

[0007] Wherein, P(x ij |t)=NormPDF(x ij g i (t); σi )

[0008] g i (t)=S i (k)

[0009] k = t*(N+1)

[0010]

[0011] is the phase of the curve corresponding to the cosine function, with a value range of [0, 2π]; N is the number of equally spaced segments in [0, 2π]; I is the number of periodic features of the matrix; J is the number of samples; X is the feature gene expression matrix.

[0012] Furthermore, the likelihood function is obtained by constructing and stationarily distributing a Markov chain through Markov chain Monte Carlo sampling, thereby maximizing the likelihood function.

[0013] Furthermore, the cells are derived from single cells, cell lines, or tissue samples.

[0014] Furthermore, it is achieved through a cell cycle time prediction model, which includes the following steps:

[0015] S1: Collect cell sequencing expression data and preprocess it to obtain cycle gene expression matrix information;

[0016] S2: Use singular value decomposition to reduce the dimensionality of the gene expression matrix information obtained in step S1, and then use the Shannon entropy of the dataset to extract the matrix of feature gene cells.

[0017] S3: Fit the feature gene expression matrix using a cosine function and calculate the likelihood function of feature gene expression. Determine the time period of the cell based on maximizing the likelihood function.

[0018] Furthermore, the preprocessing described in S1 involves first standardizing the cell sequencing expression data, then transforming the standardized data using log2, and finally extracting the cell cycle gene expression matrix information using the prior cell cycle gene database Cyclebase 3.0.

[0019] The "standardization" described in this invention is based on the fundamental principle that subtracting the mean from a numerical value and then dividing by its standard deviation yields data with a mean of 0 and a standard deviation of 1, following a standard normal distribution. After data standardization, the original data is all within the same order of magnitude, making it suitable for comprehensive comparative evaluation.

[0020] The range of the "time period in which the cell is located" in this invention is [0, 2π], where [0, 2π] corresponds to one cell cycle.

[0021] The present invention also provides an apparatus for predicting cell cycle time, the apparatus comprising an acquisition unit, a processing unit, and an output unit;

[0022] The acquisition unit is used to acquire cell sequencing expression data;

[0023] The processing unit obtains cell cycle characteristics by performing the method described above;

[0024] The output unit is used to output the time of the cell cycle phase, with a numerical range of 0-2π.

[0025] Furthermore, the cells are derived from single cells, cell lines, or tissue samples.

[0026] The present invention also provides a computer-readable storage medium having a computer program stored thereon, the computer program being used to cause a computer to perform the aforementioned method for predicting cell cycle time and / or the aforementioned steps.

[0027] Finally, this invention provides a system for predicting cell cycle time, comprising the following devices connected via a data cable and / or data interface:

[0028] Cell sequencing expression data acquisition and / or input device;

[0029] The aforementioned equipment.

[0030] This invention provides a method for predicting cell cycle time. This method is based on cell cycle information mining from cell sequencing data. It leverages the heterogeneity of cell populations to model and predict cell cycle time, saving time and effort. Experiments have shown that this invention is applicable to any unlabeled cell sequencing data, and its accuracy is significantly better than existing prediction methods.

[0031] Obviously, based on the above description of the present invention, and according to common technical knowledge and conventional methods in the field, various other modifications, substitutions or alterations can be made without departing from the basic technical concept of the present invention.

[0032] The following detailed embodiments further illustrate the above-described content of the present invention. However, this should not be construed as limiting the scope of the present invention to the following examples. All technologies implemented based on the above-described content of the present invention fall within the scope of the present invention. Attached Figure Description

[0033] Figure 1 Computer devices for predicting cell cycle time

[0034] Figure 2Accuracy comparison bar chart (scP1: prediction model of this invention; Cyclum: comparison prediction model 1; reCAT comparison prediction model 2; Cyclops: comparison prediction model 3; mESC: mouse embryonic stem cells; hESC: human embryonic stem cells) Detailed Implementation

[0035] It should be noted that the algorithms for data acquisition, transmission, storage and processing steps not specifically described in the embodiments, as well as the hardware structures and circuit connections not specifically described, can all be implemented using content already disclosed in the prior art.

[0036] Example 1: Constructing a Cell Cycle Time Prediction Model

[0037] S1: Collect cell sequencing expression data from single cells, cell lines, or tissue samples, standardize the data first, then transform the standardized data using log2, and finally extract the cycle gene expression matrix information using the cycle gene information in the prior cell cycle gene database Cyclebase 3.0;

[0038] S2: Use singular value decomposition to reduce the dimensionality of the gene expression matrix information obtained in step S1, and then use the Shannon entropy of the dataset to extract the matrix of characteristic gene expression.

[0039] S3: Fit the matrix of characteristic gene expression obtained in step S2 with a cosine function, calculate the likelihood function of characteristic gene expression, determine the phase of the cosine function by maximizing the likelihood function, and then determine the cell cycle time.

[0040] That is, the phase of the cosine curve corresponding to the cosine function of each characteristic gene (Eigengene). The value range of is [0, 2π]. Divide the 2π range into N equal parts (where N is a hyperparameter, such as 10000). The period time range of each part is from... arrive K takes values ​​in the range [0, N]; assuming the number of samples is J, the number of periodic features in the matrix obtained in step S2 is set to I. Eigengene expression level is represented by X, for any group... Combinations (each combination has I elements) The likelihood function of the model is as follows:

[0041]

[0042] in,

[0043] P(x ij |t)=NormPDF(x ij g i (t); σi )

[0044] g i (t)=S i (k)

[0045] k = t*(N+1)

[0046] For the cosine function curve of each characteristic gene (Eigengene), we have the following:

[0047]

[0048] Markov Chain Monte Carlo (MCMC) sampling is employed. For the target distribution to be sampled, a Markov chain is constructed such that the stationary distribution of this Markov chain is the target distribution. That is, starting from any initial state, the state transition sequence obtained by following the Markov chain will converge to the target distribution, thereby maximizing the likelihood. The combination of these factors determines the time phase within the 2π range in which each sample falls.

[0049] Example 2: Device for predicting cell cycle time

[0050] The device includes the following functional modules: acquisition unit, processing unit, and analysis unit.

[0051] The acquisition unit is used to acquire cell sequencing expression data.

[0052] The processing unit is used to input cell sequencing expression data into the cell cycle time prediction model of Example 1 and perform calculations;

[0053] The output unit is used to output the specific stage within the 2π range where the cell is located, as determined by the cell cycle time prediction model, in order to determine the time of the cell cycle stage.

[0054] Example 3: Computer device for predicting cell cycle time

[0055] A computer device for predicting cell cycle time, with the following structure: Figure 1 As shown, it includes at least one processor 501 and a storage device 502 connected to the at least one processor. In this embodiment, the storage device 502 has a computer program that can be executed by the at least one processor 501.

[0056] The processor 501 is the computing and control center of the cell cycle time determination device. It is connected to various parts of the lung lesion detection device through various interfaces and lines. It can execute instructions stored in the memory 502 or call data stored in the memory 502 to realize the function of cell cycle time prediction.

[0057] Optionally, processor 501 includes one or more processing units, and processor 501 may integrate an application processor and a modem. The application processor primarily handles various multimedia information such as the user interface and result detection interface, facilitating doctors to directly view the results of intelligent testing from the computer terminal. The modem primarily handles the translation and transmission of computer digital signals, facilitating communication between multiple computers. It is worth noting that the aforementioned modem may also not be integrated into processor 501.

[0058] Processor 501 can be any type of processor capable of implementing or executing the methods, steps, and logic block diagrams disclosed in the embodiments of this application, including central processing units (CPUs), digital signal processors, application-specific integrated circuits (ASICs), field-programmable gate arrays, or other programmable logic devices. The steps of the methods disclosed in the examples of this application can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules within the processor.

[0059] Optionally, the memory 502 includes multiple storage units arranged in order of unit number. In this embodiment, the memory can be divided into two main categories: main memory and auxiliary memory. The main memory can directly exchange information with the processor 501, storing or retrieving various types of information according to the address of the storage unit. The auxiliary memory cannot be directly accessed by the processor 501 and is usually used as non-volatile memory.

[0060] The memory 502 can be any medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, including flash memory, hard disk, multimedia card, card-type memory, random access memory (RAM), static random access memory (SRAM), programmable read-only memory (PROM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), magnetic memory, magnetic disk, optical disk, etc. It can also be a circuit or any other device capable of performing storage functions for storing program instructions and / or data.

[0061] Note that in some embodiments, the processor 501 and the memory 502 can be implemented on the same chip, while in other embodiments, they can be implemented on separate chips. This embodiment does not limit the specific connection medium between the processor 501 and the memory 502. Figure 1 Taking the bus connection between the processor 501 and the memory 502 as an example, it can be divided into data bus, address bus and control bus according to the type of information transmitted.

[0062] The following experimental examples further illustrate the beneficial effects of the present invention.

[0063] Experimental Example 1: Comparison of the Prediction Model of the Present Invention with Existing Prediction Models

[0064] I. Prediction Device

[0065] The predictive device of the present invention: the device described in Example 2;

[0066] Contrast prediction device 1: The prediction model in the processing unit of embodiment 2 is replaced with contrast prediction model 1, which is a cell cycle time prediction model based on an autoencoder.

[0067] The Cyclum algorithm is derived from the following literature: Liang S, Wang F, Han J, et al. Latent periodicprocess inference from single-cell RNA-seq data[J]. Nature communications, 2020, 11(1): 1-8.

[0068] Comparison prediction device 2: The prediction model in the processing unit of embodiment 2 is replaced with comparison prediction model 2. This comparison prediction model (reCAT algorithm) is a model that uses the principle of solving the traveling salesman problem (TSP) to predict cell cycle time.

[0069] The reCAT algorithm comes from the following literature: Liu Z, Lou H, Xie K, et al. Reconstructing cell cyclepseudo time-series via single-cell transcriptome data[J]. Nature Communications, 2017, 8(1): 1-9.

[0070] Contrast prediction device 3: The prediction model in the processing unit of embodiment 2 is replaced with contrast prediction model 2, which is a cell cycle time prediction model based on an autoencoder.

[0071] Anafi RC literature comes from: Francey LJ, Hogenesch JB, et al. CYCLOPS reveals human transcriptional rhythms in health and disease [J]. Proceedings of the National Academy of Sciences, 2017, 114(20): 5312-5317.

[0072] II. Cell cycle time prediction

[0073] Validation was performed using multiple datasets, including single-cell transcriptome data from human and mouse cell samples. These datasets were input into the aforementioned prediction device, and the accuracy of cell cycle times with output values ​​ranging from 0 to 2π was calculated using labeled data. The results are as follows: Figure 2 As shown, from Figure 2 It is evident that the prediction model of this invention has a significant advantage in accuracy compared to existing prediction models.

[0074] In summary, the cell cycle information mining method based on cell sequencing data in this invention utilizes the heterogeneity of cell populations to model and predict the cell cycle, saving time and effort. Experiments have demonstrated that this invention is applicable to any unlabeled cell sequencing data, and its accuracy is significantly better than existing prediction methods.

Claims

1. A method of predicting cell cycle time, the method comprising: It is from the obtained cell sequencing expression data to extract the feature gene expression matrix, adopts the cosine function to fit the feature gene expression matrix, and calculates the likelihood function of the feature gene expression, and determines the time period where the cell is located according to the maximum likelihood function; ​ The calculation formula of the likelihood function is: wherein ; φ is the phase of the curve corresponding to the cosine function, and the value range is [0, 2π]; N is the equal interval fraction of [0, 2π]; I is the periodic characteristic number of the matrix; J is the sample number; X is the feature gene expression matrix; The likelihood function is obtained by Markov chain Monte Carlo sampling, constructing and stationary distribution Markov chain, and obtaining the maximum likelihood function.

2. The method of claim 1, wherein: The cell is a single cell, a cell line or a cell derived from a tissue sample.

3. The method according to any one of claims 1 to 2, characterized in that: It is realized by a cell cycle time prediction model, and the cell cycle time prediction model comprises the following steps: S1: collecting cell sequencing expression data, and preprocessing to obtain periodic gene expression matrix information; S2: dimensionality reduction of the gene expression matrix information obtained in step S1 by singular value decomposition, and obtaining the feature gene expression matrix by calculating the Shannon entropy of the data set; S3: adopting the cosine function to fit the feature gene expression matrix, and calculating the likelihood function of the feature gene expression, and determining the time period where the cell is located according to the maximum likelihood function.

4. The method of claim 3, wherein: The preprocessing in S1 is to standardize the cell sequencing expression data first, then to adopt log2 conversion of the standardized data, and finally to extract the periodic gene expression matrix information by using the prior cell cycle gene database Cyclebase 3.

0.

5. An apparatus for predicting cell cycle time, characterized by: The device comprises an acquisition unit, a processing unit and an output unit; The acquisition unit is used to acquire cell sequencing expression data; The processing unit is used to execute the method of any one of claims 1-4 to obtain the cell cycle time; The output unit is used to output the time of the cell cycle stage, and the numerical range is 0-2π.

6. The apparatus of claim 5, wherein: The cell is a single cell, a cell line or a cell derived from a tissue sample.

7. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and the computer program is used to make a computer execute the method for predicting the cell cycle time as claimed in any one of claims 1-4.

8. A system for predicting cell cycle time, characterized in that, The device comprises a data line and / or a data interface, and a cell sequencing expression data acquisition and / or input device.

Citation Information

Patent Citations

  • Methods for determining spatial and temporal gene expression dynamics in single cells

    US20190218276A1