A cell type identification method based on cell migration trajectory

CN122598749APending Publication Date: 2026-08-18PEKING UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202610727064.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-25
Publication Date
2026-08-18

AI Technical Summary

Technical Problem

[0006]针对现有技术中细胞类型鉴定方法存在的需要破坏活体生物样本、依赖繁琐的生化分析、存在信号延迟、缺乏长期可观察性等技术缺陷,本发明旨在提供一种基于细胞迁移轨迹的细胞类型鉴定方法

Benefits of technology

[0029] (1) Extremely low sample size dependence and advantages of minimally invasive/non-destructive testing: Traditional biochemical assays (such as single-cell transcriptome sequencing, single-cell proteome sequencing, and immunoblotting) usually require lysis and destruction of millions (10) of samples. 5 ~10 6 The sheer number of cells presents a significant limitation to its application in scarce samples. This invention, relying on hypothesis testing statistical methods and employing non-invasive techniques, allows for high-confidence cell type identification with only a few dozen effective cells in a minute sample. In clinical applications, this technological breakthrough means that pathological monitoring can be completed with only extremely small puncture biopsies or minute amounts of body fluid, significantly reducing trauma to patients (minimally invasive) and conserving extremely valuable live biological samples in clinical medicine and basic science.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122598749A_ABST
    Figure CN122598749A_ABST
Patent Text Reader

Abstract

The application discloses a cell type identification method based on cell migration trajectory and belongs to the technical field of cell type identification. In view of the defects that traditional biochemical detection is strong in destructiveness and large in sample quantity requirement at a molecular scale, the application extracts and quantifies cell migration characteristics through microscopic imaging of a biological living body sample, single cell migration trajectory tracking and statistical testing, and realizes nondestructive identification of different cell types which are difficult to distinguish in appearance and morphology under naked eyes accurately and quickly. The method takes the behavior of cells at a cell scale as an anchor point, realizes high-confidence distinction under the conditions of a lower sample quantity, minimal detection damage and high automation, is particularly suitable for low-light-toxicity two-dimensional time-lapse images, can also be compatible with three-dimensional conditions, has excellent universality and expansibility, and can be widely applied to basic scientific researches which need to observe living biological samples for a long time, and can also be applied to clinical pathological practices, for example, is used for dynamic monitoring of cells in living pathological samples.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the technical field of cell type identification, specifically relating to a low-cost, accurate, rapid and non-destructive method for identifying cell types in living biological tissues, organs, individuals and cell line samples based on quantitative characteristics of cell migration trajectories. Background Technology

[0002] Life is often composed of many different types of cells. Generally speaking, multicellular animals typically contain cell types such as nerve, throat, muscle, intestine, and skin, thus constructing a complex and orderly life system [Börner et al. Nat. Cell Biol., 2021, 23:1117–1128]. These cell types play an irreplaceable role in maintaining tissue homeostasis and responding to changes in the external environment in a healthy organism through division of labor and cooperation; however, when cells undergo malignant transformation (such as in a state of gene mutation, drug stimulation, etc.), they are often accompanied by abnormalities in morphology and physiological function, becoming diseased cell types. In the fields of tissue and organ regeneration and transplantation, controlling the development and differentiation of stem cells in vitro using biochemical and physical stimulation, and directing them to be induced into specific cell types, has become an important clinical pathway for treating tissue and organ diseases. This includes, but is not limited to, inducing hematopoietic stem cells to treat leukemia [Copelan. N. Engl. J. Med., 2006, 354:1813–1826], pancreatic β-cell therapy for type 1 diabetes [Pagliuca et al. Cell, 2014, 159:428–439], cardiomyocyte repair of damaged cardiac regions after myocardial infarction [Chong et al. Nature, 2014, 510:273–277], and hair follicle stem cell intervention for hair loss [Hsu et al. Nat. Med., 2014, 20:847–856]. Although different cell types in the later stages of development and differentiation form distinctly different morphological structures and physiological functions (e.g., networked nerves efficiently transmit information, and spread-out skin efficiently dissipates heat), early cell development and differentiation usually do not exhibit significantly different morphologies and functions, making them difficult to distinguish with the naked eye. Identification often requires biochemical methods such as fluorescent labeling and transcriptome / proteome sequencing. Due to the high delay in biochemical reactions, the high operability of personnel / instruments, and the high invasiveness of biological samples involved in biochemical methods, scenarios such as disease diagnosis and drug screening that rely on cell type identification often face problems such as high cost and lack of timeliness [Clevers. Cell, 2016, 165:1586–1597]. Therefore, the inexpensive, accurate, rapid, and non-destructive identification of specific cell types or the health and abnormal states of cells has significant scientific and applied value for revealing the fundamental laws of life activities, realizing the diagnosis of major diseases, and promoting the directed induction of stem cells and drug screening.

[0003] Currently, various molecular-scale biochemical techniques have been developed in this field for cell type identification. Biochemical methods such as gene activity fluorescent labeling [Barker et al. Nature, 2007, 449:1003–1007], immunofluorescence labeling [Coons et al. Experimental Biology and Medicine, 1941, 47:200–202], and single-cell transcriptome / proteome sequencing [Macosko et al. Cell, 2015, 161:1202–1214] can provide refined characterization of cell types. However, these methods are not only expensive in terms of experimental reagents and cumbersome in terms of data analysis procedures, but also have high molecular signal delays (for example, due to the kinetic limitations of chromophore folding and oxidation, the maturation to luminescence of conventional fluorescent proteins may require a lag of up to half an hour [Shaner et al. Nature Methods, 2005, 2:905–909]), but also usually require cell fixation, irreversible exogenous invasive labeling, or even lysis [Li et al. Brief. Bioinform., 2025,26:bbaf207]. This destructive identification is often accompanied by the complete loss of cell viability, making it impossible to retain the original live sample for subsequent observation. On the other hand, fluorescent labeling and other technologies can be combined with existing high-resolution live-cell imaging technologies to achieve the extraction of three-dimensional spatial features of cellular substructures (such as cell membrane, mitochondria, Golgi apparatus, etc.). For example, the invention patent application with publication number CN121215019A discloses a method for cell fate differentiation and differentiation prediction based on three-dimensional morphology, which discloses a three-dimensional cell membrane boundary extraction scheme based on image segmentation algorithm. However, this type of approach requires high-resolution microscopes to obtain accurate spatial features, relies on expensive experimental equipment, relatively cumbersome imaging parameter tuning, and specialized image segmentation algorithms, thus limiting its applicability. In addition, three-dimensional imaging methods using fluorescent labeling must perform Z-axis tomography with high-power excitation light, which significantly accumulates phototoxicity and photobleaching, irreversibly damaging live biological samples and their subsequent observability; weak lasers can cause a significant decrease in image signal-to-noise ratio and blurring of the image, leading to increased errors in downstream feature extraction [Stephens & Allan. Science, 2003, 300:82–86; Cao et al. Nat. Commun., 2020, 11:6254; Guan et al. Membranes, 2024, 14:137; Guan et al. Nat. Commun., 2025, 16:3700].Therefore, there is an urgent need in this field to develop a high-throughput cell type identification method that is low-cost, easy to operate, highly automated, and causes minimal sample damage. This method allows practitioners to perform time-lapse imaging of live samples in a near-non-destructive manner under two-dimensional imaging conditions, thereby effectively meeting the pressing needs of current clinical medicine and regenerative cell engineering scenarios.

[0004] Cell migration is one of the most fundamental cellular behaviors in living organisms, occurring throughout the entire life cycle and playing a pivotal role in many key physiological and pathological phenomena [Ridley et al. Science, 2003, 302:1704–1709]. Physiologically, in healthy individuals, the rapid migration of fibroblasts, keratinocytes, and immune cells to the damaged site during wound healing is a crucial step in rebuilding the barrier and clearing infection; a decrease in migration rate can lead to delayed healing or scar formation [Martin. Science, 1997, 276:75–81]. Pathologically, the uncontrolled migration of cancer cells allows them to breach the stromal barrier and spread to distant tissues and organs; their extraordinary invasiveness is a key driver of rapid tumor deterioration [Hanahan & Weinberg. Cell, 2011, 144:646–674]. Numerous basic and clinical studies have shown that, due to fundamental differences in cytoskeleton remodeling dynamics, transmembrane adhesion molecule expression abundance, and cellular energy metabolism patterns, the cellular migration characteristics (such as distance and rate) of cells of different developmental and differentiation types, or even the same cell type, between healthy and disease states, may be used to identify cell types. This has significant scientific research value and broad clinical application prospects for early disease diagnosis and anti-tumor drug screening. Unlike three-dimensional spatial feature extraction, cell migration can be feature extracted using only two-dimensional monolayer or projection imaging, significantly reducing photostimulation and phototoxicity (by about two orders of magnitude) [Icha et al. BioEssays, 2017, 39:10.1002 / bies.201700003; Guan et al. Nat. Commun.,2025, 16:3700; Guan et al., Commun. Biol., 2025, 9:8].

[0005] In recent years, with the rapid evolution of microscopic optical hardware and computer vision algorithms, existing technologies have the foundation for high-precision capture of live cell migration trajectories over large areas and for extended periods. Simply put, through live-cell imaging technology and automated cell tracking algorithms, the spatiotemporal positional changes of single cells can be quantitatively captured, thereby obtaining migration characteristics. However, there is still no mature technical solution for deeply mining this trajectory data and constructing standardized and practical evaluation models. Against this backdrop, this invention proposes a cell type identification method that integrates high efficiency, small sample size, and automated computation by systematically measuring and analyzing core kinematic parameters such as cell migration distance and rate. This provides a novel methodological support for specific applications such as non-destructive cell type identification, rapid screening in clinical pathology, and regenerative medicine engineering. Summary of the Invention

[0006] To address the shortcomings of existing cell type identification methods, such as the need to destroy live biological samples, reliance on cumbersome biochemical analyses, signal delays, and lack of long-term observability, this invention aims to provide a cell type identification method based on cell migration trajectories. This invention utilizes microscopic imaging of live biological samples, single-cell migration trajectory tracking, and statistical testing to accurately and quickly achieve non-destructive identification of different cell types with relatively low sample size requirements by extracting and quantifying cell migration characteristics. This method is applicable to two-dimensional image data with low light stimulation, low phototoxicity, and low photobleaching, and can also be extended to three-dimensional image data, demonstrating excellent versatility and scalability.

[0007] To achieve the above objectives, the present invention adopts the following technical solution:

[0008] A cell type identification method based on cell migration trajectory, such as Figure 1 As shown, it includes the following steps:

[0009] S1: Obtain two sets of target live biological samples, one set as a pre-identification training set with known prior type information, and the other set as a test set to be identified with unknown type, and place them in a culture environment that maintains their physiological activity.

[0010] S2: Label the cells or subcellular structures (such as the nucleus and cell membrane) of the target cell population in the live biological sample (including all training and test sets) using label-free or specific labeling methods to establish a spatial reference target for computer vision to identify single cells.

[0011] S3: Use microscopic imaging technology to continuously capture two-dimensional or three-dimensional images of the spatial reference target to obtain time-lapse images of cells or their subcellular structures (such as cell nuclei and cell membranes) containing the target cell population.

[0012] S4: Using a cell tracking algorithm, extract the spatial dynamic behavior (i.e., single-cell migration trajectory) of the spatial reference target in the time-lapse images obtained in step S3 for all cell populations in the training and test sets, and quantify specific migration features.

[0013] S5: For the samples in the training set, firstly, obtain prior information about cell types through other physical or biochemical methods to provide a reference standard for identification; then, select two cell types in the training set to construct two cell populations whose types are known based on prior information and their migration characteristics. , ;

[0014] S6: Migration characteristics of the two cell populations obtained in step S5 , Based on the sample size, select an appropriate hypothesis testing statistical method (e.g., a two-sided Wilcoxon rank-sum test for small sample sizes, and a Z-statistic hypothesis test for large sample sizes). Conduct multiple independent hypothesis testing experiments with different sample sizes to obtain the mean migration characteristics of the two cell populations and determine the minimum sample size A required to achieve high-confidence identification of the two cell types. min ;

[0015] S7: Based on the mean migration characteristics of the two cell populations of known types obtained in step S6 and the minimum sample size A required for identification. min In the test set of unknown types to be identified, the minimum sample size of two cell populations is randomly selected respectively. The appropriate statistical method is selected according to the sample size (for example, when the sample size A < 10, the two-sided Wilcoxon rank-sum test is selected; when A ≥ 10, the Z statistic hypothesis test is selected). Hypothesis testing analysis is performed based on the migration features extracted in step S4, and the cells are classified according to the mean value of the migration features, and the cell type identification results are output.

[0016] Furthermore, the live biological samples mentioned in step S1 include, but are not limited to: all or part of the cell populations of model animals (such as Caenorhabditis elegans, salamanders, Xenopus laevis, hydra, mice, etc.) embryos or even individuals, all or part of the cell populations of in vitro constructed embryo-like and organoids, all or part of the cell populations cultured on microfluidic chips, and human body fluids or solid tissues and organs obtained clinically.

[0017] Furthermore, the label-free or specific labeling methods described in step S2 are non-destructive and low-invasive live cell labeling methods, including but not limited to: label-free methods that do not rely on external dyes, probes, or molecular markers; labeling methods using live cell fluorescent dyes, gene knock-in expressed fluorescent fusion proteins (such as GFP, mCherry), or quantum dots or nanoprobes with extremely low cytotoxicity, and continuous imaging. Simultaneously, the target cell or its subcellular structures include entities that can be spatially located and tracked, such as the nucleus, cell membrane, cytoskeleton, and endoplasmic reticulum.

[0018] Furthermore, the microscopic imaging technology described in step S3 covers two-dimensional or three-dimensional time-lapse imaging systems, specifically including: for unlabeled cases, the imaging system includes, but is not limited to, ordinary optical microscopes (bright-field and dark-field microscopes), phase-contrast or differential interference difference microscopes, digital holographic quantitative phase imaging systems, coherent Raman scattering imaging systems, photothermal microscopic imaging systems, etc.; for cases with specific labeling using fluorescent dyes, fluorescent fusion proteins, quantum dots, or nanoprobes, the imaging system includes, but is not limited to, ordinary fluorescence microscopes, wide-field fluorescence microscopes, laser scanning confocal microscopes, two-photon excitation microscopes, light-sheet illumination microscopes, etc.

[0019] Furthermore, the cell tracking algorithm described in step S4 primarily establishes a correspondence between different time points based on the spatial coordinates and temporal evolution of the target cell or its subcellular structural signals (such as reference targets like the cell nucleus and cell membrane), thereby achieving single-cell migration trajectory tracking. Specifically, an automated live cell nucleus tracking system can be employed, such as threshold segmentation and connected component analysis, active contour models, Kalman filter tracking, and deep learning target detection algorithms. The migration features are extracted and quantified from the migration trajectory, including but not limited to: total migration distance, effective migration distance, instantaneous migration rate, average migration rate, instantaneous migration deflection angle, orientation angle dispersion, and orientation index.

[0020] Furthermore, the other physical or biochemical methods described in step S5 include cell population separation methods such as lineage-based single-cell amplification, gene expression labeling, single-cell sequencing, immunofluorescence labeling, flow cytometry, multi-omics analysis, and cell identity tracking algorithms, to obtain prior information about cell populations of known types.

[0021] Further, step S6 selects an appropriate hypothesis testing statistical method based on the sample size, such as a two-tailed Wilcoxon rank-sum test or a Z-statistic hypothesis test, and conducts multiple independent hypothesis testing experiments with different sample sizes, specifically including:

[0022] For the two cell populations in the training set, random sampling is performed to automatically select a sample sub-cell population of size A. The migration features are used as input parameters for hypothesis testing (a two-sided Wilcoxon rank-sum test is selected when the sample size A < 10; a Z-statistic hypothesis test is selected when A ≥ 10), and the significance level P of a single comparison is calculated. i and the mean of the migration features of the two types of samples. , , where i represents the i-th independent test.

[0023] The system performs B independent random sampling tests consecutively and calculates the mean significance level of all independent tests after log transformation. Its calculation formula is (where lg is the logarithm with base 10, P) i (where the significance level is the value obtained from the i-th independent test); calculate the average of the mean transfer features from all independent tests. and

[0024] By gradually adjusting the value of A through a traversal algorithm, the average significance level is reached after B independent random sampling tests. Stable less than the preset threshold P threshold (When the discrimination index is as strict as possible, for example, 0.01, meaning there is only a 1% probability of misidentification), output A as the minimum sample size A required to successfully identify these two specific cell types. min Compare the average of the mean values ​​of the migration features under this sample size. and The size is used to classify it into the corresponding cell type.

[0025] Furthermore, the cell type identification process in step S7 specifically includes:

[0026] The selection and sampling of two cell populations from the test set included two cell populations of unknown type to be identified, specifically including two cell populations of unknown type from the same biological live sample or two biological live samples.

[0027] After obtaining single-cell migration trajectory data from these two cell populations using microscopic imaging techniques, the minimum sample size A for the two cell populations obtained in step S6 is then determined. min In the two cell populations to be identified, the number of samples drawn is A≥A. min For the test sample, select an appropriate hypothesis testing method based on the sample size (two-sided Wilcoxon rank-sum test when sample size A < 10; Z-statistic hypothesis test when A ≥ 10) to compare the significance of its migration characteristics; if the calculated significance level satisfies Require( If the discrimination is sufficient (e.g., 0.05), then the cell type can be effectively identified in the application scenario. Based on the mean value of migration characteristics, the two unknown cell population samples can be classified into the corresponding types, or it can be determined whether the two unknown cell populations belong to different types, thereby achieving the goals of early pathological monitoring or cell drug efficacy screening in small sample scenarios.

[0028] Compared with the prior art, the present invention has the following outstanding substantive features and significant progress:

[0029] (1) Extremely low sample size dependence and advantages of minimally invasive / non-destructive testing: Traditional biochemical assays (such as single-cell transcriptome sequencing, single-cell proteome sequencing, and immunoblotting) usually require lysis and destruction of millions (10) of samples. 5 ~10 6 The sheer number of cells presents a significant limitation to its application in scarce samples. This invention, relying on hypothesis testing statistical methods and employing non-invasive techniques, allows for high-confidence cell type identification with only a few dozen effective cells in a minute sample. In clinical applications, this technological breakthrough means that pathological monitoring can be completed with only extremely small puncture biopsies or minute amounts of body fluid, significantly reducing trauma to patients (minimally invasive) and conserving extremely valuable live biological samples in clinical medicine and basic science.

[0030] (2) Enhancing universality and scalability by using cell-specific behavior as an anchor: This invention overcomes the bottlenecks of traditional methods, which heavily rely on cell type-specific molecular markers for delayed activation and leakage of upstream gene expression. It avoids destructive experimental operations such as physical lysis and chemical fixation, and uses the most basic and intuitive cell-specific behavior, such as single-cell migration trajectory, as the core anchor. In the original physiological microenvironment, this method preserves the most natural spatiotemporal movement state of living cells. It can be widely applied not only to basic disciplines such as developmental biology and stem cell biology that require identification of whether cell type changes have occurred, but also to clinical pathological practice, such as monitoring abnormal invasive behavior of malignant tumors and the failure of targeted chemotactic function of immune cells.

[0031] (3) Spatiotemporal data fidelity: Unlike traditional technologies that can only acquire a single "time section" of a cell and lose temporal information, this invention can flexibly target and image multiple subcellular structures such as the cell nucleus, cell membrane, cytoskeleton, and endoplasmic reticulum. Therefore, it can make full use of existing fluorescent markers of cells or their subcellular structures for microscopic imaging, or use the intrinsic physical or chemical properties of the sample for label-free imaging, without being limited to specific imaging methods, and is especially suitable for long-term in vivo imaging of living cells, tissues, organs, and individuals in biomedicine. Ultimately, the information fully utilized for cell type identification is the single-cell migration trajectory across time in the time-lapse image data, rather than an independent spatial point on the traditional "time section".

[0032] (4) Low photostimulation, low phototoxicity, and low photobleaching of two-dimensional data: Significantly different from the method disclosed in CN121215019A, this invention overcomes the limitation of relying solely on three-dimensional morphology for feature reconstruction and provides the possibility of quantitative data extraction of migration features using two-dimensional imaging methods. Since intensive Z-axis optical slicing scanning is not required, the two-dimensional imaging method significantly reduces the total exposure time and radiation dose compared to the three-dimensional case, reducing the photostimulation, phototoxicity, and photobleaching effects on the sample by approximately two orders of magnitude. This not only maximizes the preservation of the original physiological and developmental activity of the live sample but also better meets the needs of long-term continuous observation scenarios in most current clinical medicine and basic biology.

[0033] (5) Efficient utilization of cross-dimensional data and scenario generalization: CN121215019A is highly dependent on the three-dimensional spatial features of cell morphology, and the acquisition of its dimensional features is inevitably constrained by stringent microscopic optical resolution, extremely high signal-to-noise ratio fluorescent labeling, and high development and computational costs of three-dimensional image segmentation algorithms. This invention, by fully utilizing time-lapse imaging, maximizes the use of temporal dimension information and reduces the requirement for the three-dimensional information of the entire cell (a volume in space) to the zero-dimensional information of the cell's spatial reference target (a point in space). This strategy significantly lowers the resolution threshold of microscopic hardware and greatly reduces the system's absolute dependence on the accuracy of image segmentation algorithms, improving the robustness of the results. Therefore, this method has good compatibility in both two-dimensional and three-dimensional test data and can flexibly and adaptively switch quantitative evaluation methods according to different application scenarios and equipment conditions.

[0034] (6) High degree of automation and elimination of subjective bias: This invention has constructed a mature technical process that covers the integrated operation of in vivo time-lapse microscopy, computer-automated trajectory tracking, and algorithmic statistics and analysis. This highly automated data processing process can eliminate the problem of subjective observation bias caused by the heavy reliance on the experience of experimenters in traditional biological experiments. At the same time, it lowers the threshold for the popularization of this technology in grassroots testing institutions or third-party medical laboratories, ensuring the extremely high reproducibility and objectivity of the analysis results.

[0035] (7) Economic benefits and time efficiency: By avoiding the use of expensive experimental reagents and complex biochemical treatments, the present invention has the advantage of low cost per detection; in addition, the efficient synchronization of cell migration trajectory image acquisition and statistical algorithm can shorten the experimental cycle, and gain an extremely valuable time window for rapid screening in clinical emergency and high-throughput screening of massive amounts of drugs in the laboratory. Attached Figure Description

[0036] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below.

[0037] Figure 1 This is a schematic diagram illustrating the steps of the cell type identification method based on cell migration trajectory of the present invention.

[0038] Figure 2 This is a flowchart illustrating the cell type identification method based on cell migration trajectory of the present invention.

[0039] Figure 3 This document provides a schematic diagram of cell migration trajectories and an explanation of the calculation methods for different quantitative characteristics.

[0040] Figure 4 This is a schematic diagram of the process in Example 1, which uses the cell nucleus as a spatial reference target, establishes the minimum sample size for identifying cell types based on the instantaneous migration rate on a two-dimensional focal plane, and identifies unknown muscle and skin cell populations that are respectively derived from the division and expansion of a single progenitor cell.

[0041] Figure 5 The results of the minimum sample size benchmark for cell type identification in Example 1 are as follows: A represents the pairwise cross-validation results of all data points for the five cell types (using the Z-statistic hypothesis test to determine the instantaneous migration rate with the cell nucleus as the spatial reference target on the two-dimensional focal plane); B represents the sample size test results of pairwise independent sampling of the five cell types, i.e., the significance level P of random independent sampling test of sample number A repeated B (B=1000) times. i C represents unknown cell lineage trees of muscle (left, Cap lineage) and skin (right, Cpa lineage), which are derived from the division and expansion of a single progenitor cell, respectively.

[0042] Figure 6 This is a schematic diagram of the process in Example 2, which uses the centroid of the region enclosed by the cell membrane as a spatial reference target point, establishes the minimum sample size for identifying cell types based on the three-dimensional overall migration distance as a migration feature, and identifies unknown muscle and intestinal cell populations that are respectively derived from the division and expansion of a single progenitor cell.

[0043] Figure 7 The results of the minimum sample size benchmark for cell type identification in Example 2 are as follows: A represents the pairwise cross-identification results of all data points for the five cell types (using the Z-statistic hypothesis test with the centroid of the three-dimensional cell membrane enclosed region as the spatial reference target for the overall migration distance); B represents the sample size test results of pairwise independent sampling of the five cell types, i.e., the significance level P of random independent sampling test of sample number A repeated B (B=1000) times. iC represents unknown cell lineage trees of muscle (left, lineage D) and intestine (right, lineage E), which are derived from the division and expansion of a single progenitor cell, respectively. Detailed Implementation

[0044] The technical process of the present invention will now be described in detail and completely with reference to the accompanying drawings and embodiments, so that those skilled in the art can have a deeper understanding of the feasibility and operability of the technical solution of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0045] This invention provides a cell type identification method based on cell migration trajectory, the overall workflow of which is as follows: Figure 2 As shown, the specific steps include:

[0046] S1: Collect live biological samples according to specific application scenarios. One part serves as a training set for known cell types, and the other part as a test set for unknown cell types. The live biological samples encompass biological materials across species and scales, specifically including: embryos or specific tissues of model organisms such as *C. elegans*, salamanders, Xenopus laevis, hydra, and mice; embryos or meristems of plant seeds; artificially constructed embryo-like, organoid, or microfluidic chip cell culture systems; and clinically obtained human body fluids (such as blood and pleural effusion) or puncture tissue samples. After collection, the samples must be placed in a suitable culture environment (such as an incubator with constant temperature, humidity, and specific gas concentrations) that maintains their physiological activity and natural cell migration state.

[0047] S2: Based on the physical characteristics of the sample and observation requirements, select appropriate non-destructive, low-invasive live-cell labeling methods to label the cells or subcellular structures of the target cell population to establish spatial reference targets. The labeling methods include, but are not limited to, label-free methods that do not rely on external dyes, probes, or molecular markers; labeling methods using live-cell fluorescent dyes (such as DAPI / Hoechst, DiO / DiI / DiD, Calcein-AM, etc.), gene knock-in expressed fluorescent fusion proteins (such as GFP, mCherry), and quantum dot or nanoprobe labeling methods with extremely low cytotoxicity; and continuous imaging within a fixed field of view. The target cells or subcellular structures include entities with high contrast and capable of stable spatial positioning and tracking, such as the nucleus, cell membrane, cytoskeleton, or endoplasmic reticulum.

[0048] S3: Use microscopic imaging technology to perform non-destructive imaging of the space reference target (including all training and test sets). To accommodate the physical differences in thickness, transmittance, and scattering characteristics of various living samples, the microscopic imaging encompasses two-dimensional or three-dimensional imaging systems. Specifically, for observation scenarios that do not require specific staining (label-free), classic optical microscopes (such as bright-field and dark-field microscopes) [Zernike. Science, 1955, 121:345–349], differential interferometry phase contrast microscopy [Wang et al. Nat. Commun., 2023,14:2063], or high-resolution digital holographic quantitative phase imaging systems [Popescu et al. Opt. Lett., 2006, 31:775–777], coherent Raman scattering imaging [Freudiger et al. Science, 2008, 322:1857–1861], and photothermal microscopy [Zhang et al. Sci. Adv., 2016, ] can be used. [2:e1600521] etc.; For high-contrast observation scenarios that rely on the imaging of specific targets, i.e., imaging systems based on fluorescent dyes, fluorescent fusion proteins, quantum dots, and nanoprobes, ordinary fluorescence microscopes, wide-field fluorescence microscopes [Lichtman et al. Nat. Methods, 2005, 2:910–919], laser scanning confocal microscopes [White et al. J. Cell Biol., 1987, 105:41–48], two-photon excitation microscopes [Denk et al. Science, 1990, 248:73–76], or light-sheet illumination microscopes [Huisken et al. Science, 2004, 30:1007–1009] can be selected. In this process, it is necessary to set an appropriate spatial resolution (i.e., pixel size) according to the physical size of the sample, and set a matching temporal resolution (i.e., frame rate) according to the migration speed of the target cells, so as to obtain time-lapse image data containing the continuous dynamic behavior of the target.

[0049] S4: Using a cell tracking algorithm, extract the spatial dynamic behavior of the spatial reference target from the time-lapse images obtained in step S3 for all cell populations in the training and test sets. Specific tracking algorithms can refer to threshold segmentation and connected component analysis, active contour models, Kalman filter tracking, and deep learning target detection algorithms. Extract the spatial coordinates of the target frame by frame to obtain the time-lapse image data of the target object, and extract the single-cell migration trajectories of all cell populations in the training and test sets. Define a set of spatial trajectories formed by the target object within the observation period, with a total of n frames. At frame τ, the spatiotemporal coordinate vector of the target object is... ,in Let t be the spatial coordinates of the image data of the target object. τ Let be the time corresponding to the target object in the τth frame. For a two-dimensional imaging system, ,in For the extension of the set of real numbers R in two-dimensional space, that is, all two-dimensional coordinate points (x, y, z) τ , y τ A set of ); for a three-dimensional imaging system, ,in For the extension of the set of real numbers R in three-dimensional space, that is, all three-dimensional coordinate points (x, y, z) τ , y τ , z τ The system automatically calculates and quantifies various migration features at the single-cell level based on the spatial vector sequence given by the aforementioned reference target. Figure 3 Specifically, including but not limited to:

[0050] (1) Total Distance (D): The total actual path length of the target object's migration trajectory during the entire tracking period, which is equal to the sum of the displacement vector magnitudes between adjacent frames.

[0051]

[0052] in, The spatiotemporal coordinate vector representing the target object. Let t be the spatial coordinates of the image data of the target object. τ Let τ be the time corresponding to the target object in the τth frame, and n be the total number of frames in the spatial trajectory set.

[0053] (2) Effective Displacement (D) eff ): The net linear displacement of the target object from the starting position to the ending position is equal to the magnitude of the displacement vectors at the initial and ending positions.

[0054]

[0055] in, and The spatiotemporal coordinate vectors representing the initial and final positions of the target object, respectively. and The image data spatial coordinates of the initial and final positions of the target object, t1 and t2, are respectively. n These represent the times corresponding to the first and last frames tracked for the target object.

[0056] (3) Instantaneous migration rate (v(t)) τThe velocity of the target object within a certain time interval is equal to the magnitude of the displacement vectors of two adjacent frames divided by the time resolution.

[0057]

[0058] in, The spatiotemporal coordinate vector representing the target object. Let t be the spatial coordinates of the image data of the target object. τ Let τ be the time corresponding to the target object in the τth frame.

[0059] (4) Average migration rate : The mathematical expectation of the rate over the entire tracking period, i.e., the mean of the instantaneous migration rates over all time intervals.

[0060]

[0061] in, The spatiotemporal coordinate vector representing the target object. Let t be the spatial coordinates of the image data of the target object. τ Let τ be the time corresponding to the target object in the τth frame, and n be the total number of frames in the spatial trajectory set.

[0062] (5) Instantaneous migration deflection angle (θ(t)) τ The angle between the displacement vectors of two consecutive time intervals reflects the degree of change in the direction of motion at that moment.

[0063]

[0064] in, The spatiotemporal coordinate vector representing the target object. Let t be the spatial coordinates of the image data of the target object. τ Let τ be the time corresponding to the target object in the τth frame.

[0065] (6) Var(θ) of Angle: The variance of instantaneous migration deflection angle at all times.

[0066]

[0067] in, Represents θ(t) τ The average of θ(t) over the total number of frames in the spatial trajectory set. τ ) for t τ With t τ+1 The angle between the displacement vectors at time intervals, where n is the total number of frames in the spatial trajectory set.

[0068] (7) Directionality Ratio (DR): A dimensionless index characterizing the directionality of cell migration, defined as the effective migration distance D. eff The ratio to the total migration distance D. The range is from 0 to 1, with a tendency towards 1 indicating highly linear directional movement and a tendency towards 0 indicating disordered random Brownian motion.

[0069]

[0070] in, The spatiotemporal coordinate vector representing the target object. Let t be the spatial coordinates of the image data of the target object. τ Let τ be the time corresponding to the target object in the τth frame, and n be the total number of frames in the spatial trajectory set.

[0071] S5: To establish the recognition benchmark for this algorithm, this invention needs to integrate standard data with known prior information as training or comparison objects. For training set samples in specific application scenarios, the system needs to introduce existing high-precision biochemical analysis or cell lineage construction methods in parallel to identify real cell types and obtain known cell types as reference standards. The verification methods include: computational methods such as cell lineage construction and identification based on cell division behavior [Guignard et al., Science, 2020, 369:eaar5663]; or experimental methods such as gene activity fluorescent labeling [Imai et al., Science, 2006, 312:1183–1187]. Based on the above prior information, the obtained migration trajectory dataset is labeled in the system for cells of different types or healthy and disease states, thereby constructing a cell set containing multiple "known types" as a standard dataset.

[0072] S6: To minimize sample dependence in unknown detections (test set), this system optimizes statistical parameters based on the aforementioned training set. Its core mechanism is as follows:

[0073] In step S5, migration features of two cell populations with different known types are extracted. Random Monte Carlo sampling is then performed on the two datasets within each cell population, automatically selecting a subset of samples A at a time. For the i-th test, a hypothesis test is performed on the selected subset, and the significance level P for a single comparison is calculated. i and the mean migration features of the two groups of samples. and When A < 10, since the small sample size may not follow a normal distribution, a two-tailed Wilcoxon rank-sum test is used to estimate the significance level P. iWhen A ≥ 10, the sample tends to follow a normal distribution. The result of the two-tailed Wilcoxon rank-sum test infinitely approximates the Z-statistic hypothesis test. Therefore, the Z-statistic hypothesis test is used to estimate the significance level P. i .

[0074] The system performs B independent random sampling tests consecutively and calculates the mean significance level of all independent tests after log transformation. It is used to determine the discrimination of a population, and its calculation formula is:

[0075]

[0076] Where lg is the logarithm with base 10, P i Let be the significance level obtained from the i-th independent random sampling. Calculate the average of the mean values ​​of the migration features from all independent experiments. and ;

[0077] By iterating and gradually adjusting the value of A, the average significance level is obtained after B random samplings. Stable less than the preset statistical threshold P threshold When the discrimination index is as strict as possible, for example, 0.01, output A at this time as the minimum independent sample size benchmark A required to successfully identify the two cell types in this specific application scenario. min Compare the average of the mean values ​​of the migration features under this sample size. and The size is used to classify it into the corresponding cell type.

[0078] S7: In actual testing or clinical applications, repeat step S1 to obtain a population of live cells of unknown type or suspected physiological abnormalities (test set). Based on the mean migration characteristics of the two cell types with known cell types obtained in step S6 and the minimum independent sample size benchmark A... min Automatically extract sample numbers A ≥ A from the migration trajectory dataset of the cell populations to be identified (two groups of unknown cell populations of unknown type or with abnormal physiological state). min The test set was used to perform multiple rounds of significance comparison of the transfer features of the two methods using a two-sided Wilcoxon rank-sum test; if the calculated significance level consistently satisfies... Require( A sufficient discrimination index (e.g., 0.05) is required to effectively identify cell types in the application scenario. Unknown samples can be categorized into corresponding types based on the mean of the migration features. This method can be used to label cell types or determine whether cells in unknown test groups have undergone significant abnormal transformation, outputting health or disease status monitoring and assessment results, thereby achieving goals such as early pathological monitoring or precise screening of cell-drug efficacy in small sample scenarios.

[0079] In summary, the technical solution of this invention effectively replaces some cumbersome and destructive biochemical tests by integrating intuitive and highly digitized cell behavior tags into statistical testing. This avoids high demands on experimental equipment, significantly reducing the cost of experimental consumables for basic scientists, while preserving the spatiotemporal continuity of cell information and physiological activity in its native microenvironment. This system innovatively solves the challenge of cell type identification under small sample size limitations and significantly reduces the light stimulation, phototoxicity, and photobleaching associated with time-lapse imaging by utilizing two-dimensional imaging and zero-dimensional spatial reference targets, enabling long-term, non-destructive cell type identification. Cell position tracking, rather than morphological segmentation, reduces the high demands on computational algorithms and provides generalizability and platform scalability across multiple scenarios and fields. In basic life science research, it is suitable for monitoring embryonic morphogenesis in developmental biology; in clinical pathology practice, it can be directly used as a low-cost, high-throughput screening tool to assist clinicians in achieving highly objective and subjective minimally invasive rapid pathological screening (e.g., rapid determination of the abnormal invasive potential of early malignant tumors, and dynamic pathological assessment of the failure of immune cell targeted chemotaxis function); in the field of regenerative cell engineering, it has important application value for the identification of stem cell directed induction.

[0080] To verify that cell migration characteristics such as "overall migration distance" and "instantaneous migration rate" can effectively identify different types of cells, the present invention provides the following basic verification examples to capture the differences in migration characteristics exhibited by different types of cells, thereby verifying the feasibility and operability of the method of the present invention.

[0081] The following examples use a publicly available dataset of live continuous microscopic images of Caenorhabditis elegans embryos [Guan et al. Nat. Commun., 2025, 16:3700] as the basic test data to examine the heterogeneity of migration characteristics when embryos develop different cell types. This dataset includes embryos from ≤4 cell stages to ≥550 cell stages, sufficient to cover >95% of cells during embryonic development. Caenorhabditis elegans is a classic model organism in modern cell and developmental biology research. Its cell types (including but not limited to nerve, pharynx, muscle, skin, and intestine) are highly conserved with those of humans, and its genome is highly homologous to humans, providing a highly reliable and reliable data standard for the feasibility and operability verification of the methods of this invention.

[0082] S1: Eight hermaphroditic Caenorhabditis elegans embryos were collected as initial live biological samples and transferred to nematode growth medium agar plates inoculated with Escherichia coli. The embryos were incubated in a standard constant temperature and humidity incubator at 20°C to maintain normal physiological development.

[0083] S2: To achieve precise spatial tracking and morphological segmentation, this embodiment introduces a dual-color fluorescent label by constructing a specific transgenic strain (such as strain ZZY0861). Specifically: Green fluorescent protein (GFP) is fused with a histone (HIS-72) fragment to target and label the cell nucleus, providing a usable spatial reference anchor for subsequent cell migration trajectories; simultaneously, the membrane-specific domain PH (PLC1delta1) is fused with red fluorescent protein (mCherry) for expression, providing a usable spatial reference anchor for subsequent cell migration trajectories.

[0084] S3: At a constant ambient temperature of 20°C, the sample was placed under a laser scanning confocal microscope (Leica SP8 system) equipped with a high numerical aperture objective lens for time-lapse imaging. For example... Figure 4 and Figure 6 As shown, the fertilized eggs of Caenorhabditis elegans are oval in shape, with an average long axis of about 40-50 μm and a short axis of about 25-30 μm. The total time from fertilization to hatching is about 15 hours [Riddle et al., C. elegans II. 2nd edition., 1997].

[0085] To match the spatial and temporal scales of cell migration, the following imaging parameters were set: spatial resolution on the XY two-dimensional focal plane was set to 0.09 μm / pixel (712 pixels in the X direction and 512 pixels in the Y direction); spatial resolution on the Z-axis depth direction was set to 0.42 μm / pixel (92 pixels in the Z direction); temporal resolution was set to Δt ≤ 1.5 minutes. A complete stereoscopic image was captured. Acquisition was continuously performed at 240 time points (covering the entire process from ≤4 cell stage to ≥550 cell stage), simultaneously collecting images of GFP (nucleus) and mCherry (cell membrane) channels.

[0086] The following implementation examples will further illustrate two application scenarios of the present invention for cell type identification based on Caenorhabditis elegans embryo samples processed in steps S1-S3 above, in conjunction with the accompanying drawings.

[0087] Example 1

[0088] This embodiment aims to verify that the physical characteristic of "two-dimensional cell instantaneous migration rate" can be used to significantly distinguish the different types of cell populations mentioned above. See below. Figure 4 The steps of this embodiment are described in detail below:

[0089] S4: Using the GFP channel image acquired in step S3, the projection position of the cell nucleus on the XY focal plane is taken as the image input. Cell nucleus tracking is performed according to the StarryNite and AceTree methods. The tracking results are used to determine the cell's identity [Murray et al., Nat. Protoc., 2006, 1:1468–1476; Boyle et al., BMC Bioinformatics, 2006, 7:275; Bao et al., Proc. Natl. Acad. Sci. USA,2006, 103:2707–2712]. The spatiotemporal coordinate vector of the cell nucleus migration trajectory is output. Calculate the two-dimensional instantaneous migration rate of a single cell:

[0090]

[0091] Where x and y represent the spatial position of the target object in a two-dimensional Cartesian coordinate system, and t τ Let τ be the time corresponding to the target object in the τth frame.

[0092] S5: Because *C. elegans* exhibits an invariant developmental lineage among individuals, cell identity based on cell division behavior can be used for direct cell type labeling [Murray et al., Nat. Protoc., 2006, 1:1468–1476; Boyle et al., BMC Bioinformatics, 2006, 7:275; Bao et al., Proc. Natl. Acad.Sci. USA, 2006, 103:2707–2712]. Using "prior labels," a total of 89,744 cell data points belonging to the neural type (two-dimensional instantaneous cell migration rate data), 34,549 cell data points belonging to the pharyngeal type, 43,586 cell data points belonging to the muscle type, 63,873 cell data points belonging to the skin type, and 14,786 cell data points belonging to the intestinal type were obtained, constructing a "known type" cell population as a standard dataset.

[0093] S6: To minimize sample dependence in unknown detections (test set), this system, based on the cell populations given in step S5, performs pairwise cross-identification (e.g., muscle cells vs. intestinal cells) to find the minimum independent sample size benchmark A required to distinguish these two cell types. minA samples (A = 10, 20, 40, 60, 80, 100, 200, 400, 600, 800, 1000…) are randomly selected from each of the two categories to conduct a hypothesis test, and the significance level P for a single comparison is calculated. i and the mean migration features of the two groups of samples. and When A < 10, the sample is not normally distributed, and a two-tailed Wilcoxon rank-sum test is used; when A ≥ 10, the sample is nearly normally distributed, and the Z-statistic hypothesis test is used to estimate the significance level P. i To eliminate the randomness of extreme values ​​in a single sampling, the random independent sampling test is repeated B times (B=1000) for each sample size. The system calculates the log-geometric mean of the significance levels of all independent tests as the average significance level:

[0094]

[0095] Simultaneously calculate the average of the mean transfer characteristics of all independent experiments. and As the system iterates and the deduction increases, the value of A increases. The first stable value is less than a rigorous statistical confidence threshold. When the discrimination is as strict as possible, output and record the current A as the minimum independent sample size benchmark A for distinguishing these two cell types. min Compare the average of the mean values ​​of the migration features under this sample size. and The size of the sample determines its classification into the corresponding sample group.

[0096] like Figure 4 and Figure 5 As shown, step S6 is specifically implemented in identifying the five main cell types as follows.

[0097] First, pairwise Z-statistic hypothesis tests were performed using samples of all known cell types labeled in the prior tags. All cell type combinations were found to be effectively identified with the existing sample size (Table 1 and...). Figure 5 (A)

[0098]

[0099] Furthermore, by changing the sample size A and increasing it to 10,000, the minimum effective sample sizes in each pairing were as follows: nerve / pharynx (2,000, high values ​​point to nerve, low values ​​point to pharynx), nerve / muscle (2,000, high values ​​point to nerve, low values ​​point to muscle), nerve / gut (2,000, high values ​​point to gut, low values ​​point to nerve), pharynx / gut (400, high values ​​point to gut, low values ​​point to pharynx), pharynx / skin (2,000, high values ​​point to skin, low values ​​point to pharynx), muscle / gut (400, high values ​​point to gut, low values ​​point to muscle), muscle / skin (2,000, high values ​​point to skin, low values ​​point to muscle), and gut / skin (2,000, high values ​​point to gut, low values ​​point to skin). Among these, the two-dimensional instantaneous migration rate with the cell nucleus as the reference point showed good distinguishability in the pharynx / gut and muscle / gut pairs (Tables 2 and 3). Figure 5 (B)

[0100]

[0101]

[0102] S7: For two cell populations of unknown types, the mean migration characteristics of the two cell types with known types obtained in step S6 and the minimum independent sample size benchmark A can be used as a basis. min To test whether the two belong to different cell types, samples A ≥ A were drawn from the migration trajectory datasets of the two cell populations. min On the test set, multiple rounds of significance comparison of the migration characteristics of the two were conducted using appropriate hypothesis testing methods (two-sided Wilcoxon rank-sum test when sample size A < 10; Z-statistic hypothesis test when A ≥ 10); if the calculated significance level consistently satisfies Require( If the discrimination is sufficient, it proves that the two cell populations belong to different types. By measuring the magnitude of the feature mean, unknown samples can be classified into the corresponding types, thus enabling specific cell type identification for any cell population data in the same application scenario.

[0103] Step S7, which distinguishes between the two groups of cell types, muscle and skin, is implemented as follows.

[0104] A muscle cell population derived from the continuous division and expansion of a single Cap progenitor cell was selected as the test sample set, containing 31 cell types, all belonging to the muscle type, with a total of 7,712 data points. A skin cell population derived from the continuous proliferation of a single Cpa progenitor cell was selected as the test sample set, containing 15 cell types, all belonging to the skin type, with a total of 7,119 data points. The cell lineages of the two types are as follows: Figure 5 As shown in C.

[0105] In the two groups of cells to be tested, samples A≥A were drawn respectively. min =2,000 samples, randomly sampled 10 times, and the significance of the two two-dimensional instantaneous cell migration rates with the cell nucleus as the reference point was compared using the Z-statistic hypothesis test (Table 4). The test results showed that the significance level of all 10 tests met the requirements. The requirement is that the average significance level is 4.474 × 10⁻⁶. -18 The high values ​​indicate that the skin is the primary target and the low values ​​indicate that the two single progenitor cells and their progeny cells belong to the muscle and skin types, respectively.

[0106]

[0107] Example 2

[0108] This embodiment aims to verify that the physical characteristic of "three-dimensional total cell migration distance" can be used to significantly distinguish the different types of cell populations mentioned above. See below. Figure 6 The steps of this embodiment are described in detail below:

[0109] S4: For the mCherry channel image acquired in step S3, the distribution of the cell membrane in XYZ space is used as the image input, and cell membrane segmentation is performed according to the CMap method [Guan et al. Nat. Commun., 2025, 16:3700]; for the GFP channel image acquired in step S3, the position of the cell nucleus in XYZ space is used as the image input, and cell nucleus tracking is performed according to the StarryNite and AceTree methods. The tracking results are used to determine the cell identity [Murray et al., Nat.Protoc., 2006, 1:1468–1476; Boyle et al., BMC Bioinformatics, 2006, 7:275; Bao et al., Proc. Natl. Acad. Sci. USA, 2006, 103:2707–2712], and the spatiotemporal coordinate vector of the cell membrane migration trajectory is output. Calculate the total three-dimensional migration distance of a single cell:

[0110] Define Ω as the set of three-dimensional voxels occupied by a successfully segmented single-celled individual at time τ. τ It contains N voxels, where the spatial physical coordinates of the k-th voxel are: Therefore, the three-dimensional centroid coordinate vector of this single cell at this moment is strictly defined as the arithmetic mean of the coordinates of all voxels within that region:

[0111]

[0112] After extracting the centroid sequence across the entire time axis, the three-dimensional total migration distance of a single cell, using the centroid of the region enclosed by the cell membrane as the tracking reference point, is calculated using the following formula:

[0113]

[0114] Where x, y, and z represent the spatial position of the target object in a three-dimensional Cartesian coordinate system, and t τ Let τ be the time corresponding to the target object in the τth frame, n be the total number of frames tracked during the cell's life cycle, and the absolute value sign indicates the Euclidean modulus of the three-dimensional spatial displacement vector between two adjacent frames.

[0115] S5: Because *C. elegans* exhibits an invariant developmental lineage among individuals, cell identity based on cell division behavior can be used for direct cell type labeling [Murray et al., Nat. Protoc., 2006, 1:1468–1476; Boyle et al., BMC Bioinformatics, 2006, 7:275; Bao et al., Proc. Natl. Acad.Sci. USA, 2006, 103:2707–2712]. Using "prior labels," a total of 770 cell data points belonging to the neural type (3D overall migration distance data), 289 cell data points belonging to the pharyngeal type, 513 cell data points belonging to the muscle type, 602 cell data points belonging to the skin type, and 148 cell data points belonging to the intestinal type were obtained, constructing a "known type" cell population as a standard dataset.

[0116] S6: To minimize sample dependence in unknown detections (test set), this system, based on the cell populations given in step S5, performs pairwise cross-identification (e.g., nerve cells vs. pharyngeal cells) to find the minimum independent sample size benchmark A required to distinguish these two cell types. min A samples (A=2, 4, 6, 8, 10, 20, 40, 60, 80, 100…) were randomly selected from each of the two groups for hypothesis testing. The significance level P for a single comparison was calculated. i and the mean migration features of the two groups of samples. and When A < 10, the sample is not normally distributed, and a two-tailed Wilcoxon rank-sum test is used; when A ≥ 10, the sample is nearly normally distributed, and the Z-statistic hypothesis test is used to estimate the significance level P. iTo eliminate the randomness of extreme values ​​in a single sampling, the random independent sampling experiment for each sample size is repeated B times (B=1000). The system calculates the log-geometric mean of the significance levels of all independent experiments as the average significance level:

[0117]

[0118] Simultaneously calculate the average of the mean transfer characteristics of all independent experiments. and As the system iterates and the deduction increases, the value of A increases. The first stable value is less than a rigorous statistical confidence threshold. When the discrimination is as strict as possible, output and record the current A as the minimum independent sample size benchmark A for distinguishing these two cell types. min Compare the average of the mean values ​​of the migration features under this sample size. and The size of the sample determines its classification into the corresponding sample group.

[0119] like Figure 6 and Figure 7 As shown, step S6 is implemented in the following specific way to distinguish the five main fates.

[0120] First, pairwise Z-statistic hypothesis tests were performed using all samples with known fates marked in the prior labels. All cell type combinations were found to be effectively distinguishable with the existing sample size (Table 5 and...). Figure 7 (A)

[0121]

[0122] Furthermore, by changing the sample size A and increasing it to 200, the minimum effective sample sizes in each pairing were as follows: nerve / pharynx (60, high values ​​point to nerve, low values ​​point to pharynx), pharynx / muscle (80, high values ​​point to muscle, low values ​​point to pharynx), pharynx / gut (40, high values ​​point to gut, low values ​​point to pharynx), pharynx / skin (40, high values ​​point to skin, low values ​​point to pharynx), and muscle / gut (100, high values ​​point to gut, low values ​​point to muscle). Among these, the three-dimensional overall migration distance, with the centroid of the region enclosed by the cell membrane as the reference point, showed good distinguishability in pharynx / gut, pharynx / skin, nerve / pharynx, pharynx / muscle, and muscle / gut (Tables 6 and 7). Figure 7 (B)

[0123]

[0124] S7: For two cell populations of unknown types, the migration characteristics of the two cell types with known types obtained in step S6 and the minimum independent sample size benchmark A can be used as a basis for analysis. min To test whether the two belong to different cell types, samples A ≥ A were drawn from the migration trajectory datasets of the two cell populations. min On the test set, multiple rounds of significance comparison of the migration characteristics of the two were conducted using appropriate hypothesis testing methods (two-sided Wilcoxon rank-sum test when sample size A < 10; Z-statistic hypothesis test when A ≥ 10); if the calculated significance level consistently satisfies Require( If the mean of the migration feature is sufficient (as long as the discrimination is sufficient), then it proves that the two cell populations belong to different types. By measuring the magnitude of the mean of the migration feature, unknown samples can be classified into the corresponding types, thus enabling specific cell type identification for any cell population data in the same application scenario.

[0125] Step S7, which distinguishes between muscle and intestinal cell types, is implemented as follows.

[0126] A muscle cell population derived from the continuous proliferation of a single progenitor cell (D) was selected as the test sample set, containing 39 cells, all of which belong to the muscle type, totaling 143 samples from 8 independent embryos. Similarly, an intestinal cell population derived from the continuous proliferation of a single progenitor cell (E) was selected as the test sample set, containing 39 cells, all of which belong to the intestinal type, totaling 148 samples from 8 independent embryos. The cell lineages of both types are shown below. Figure 7 As shown in C.

[0127] In the two groups of cells to be tested, samples A≥A were drawn respectively. min =100 samples, randomly sampled 10 times, and the significance of the three-dimensional overall migration distance of the two with the centroid of the region enclosed by the cell membrane as the reference point was compared using the Z-statistic hypothesis test (Table 8). The test results showed that the significance level of all 10 tests met the requirements. The requirements were met, with an average significance level of 0.001. High values ​​indicated the gut and low values ​​indicated the muscle, suggesting that the cell types of the two single progenitor cells and their progeny belonged to the muscle and gut, respectively.

[0128]

[0129] The above embodiments illustrate and describe the feasibility and operability of the present invention, helping readers understand the implementation principle of the present invention. However, as mentioned above, it should be understood that the scope of protection of the present invention is not limited to the embodiments disclosed herein, and can also be applied to other biological systems and environments. Through the above description or improvements in related fields, those skilled in the art can make other combinations and modifications based on the technical teachings disclosed in the present invention. Any modifications and alterations made by those skilled in the art that do not depart from the spirit of the present invention should be within the scope of protection of the appended claims.

Claims

1. A method for cell type identification based on cell migration trajectory, comprising the following steps: S1: Obtain two sets of target live biological samples, one set as a pre-identification training set with known prior type information, and the other set as a test set to be identified with unknown type, and place them in a culture environment that maintains their physiological activity. S2: Label the cells or subcellular structures of the target cell population in the live biological sample using label-free or specific labeling methods to establish a spatial reference target for computer vision to identify single cells. S3: Use microscopic imaging technology to continuously capture two-dimensional or three-dimensional images of the spatial reference target to obtain time-lapse images of cells or subcellular structures containing the target cell population; S4: Using a cell tracking algorithm, extract the spatial dynamic behavior of the spatial reference target in the time-lapse images obtained in step S3, i.e., the single-cell migration trajectory, for all cell populations in the training and test sets, and quantify specific migration features. S5: For the samples in the training set, firstly, obtain prior information about cell types through other physical or biochemical methods to provide a reference standard for identification; then, select two cell types in the training set to construct two cell populations whose types are known based on prior information and their migration characteristics. , ; S6: Migration characteristics of the two cell populations obtained in step S5 , Based on the sample size, an appropriate hypothesis testing statistical method was selected, and multiple independent hypothesis testing experiments were conducted with different sample sizes to obtain the mean migration characteristics of the two cell populations. The minimum sample size A required to achieve high-confidence identification of the two cell types was then determined. min ; S7: Based on the mean migration characteristics of the two cell populations of known types obtained in step S6 and the minimum sample size A required for identification. min In the test set of unknown types, the minimum sample size of two cell populations is randomly selected. According to the sample size, an appropriate hypothesis testing statistical method is selected. Hypothesis testing analysis is performed based on the migration features extracted in step S4. The cells are classified according to the mean value of the migration features, and the cell type identification results are output.

2. The cell type identification method as described in claim 1, characterized in that, The live biological samples mentioned in step S1 are selected from: all or part of the cell populations of model animal embryos or even individuals, all or part of the cell populations of in vitro constructed embryo-like and organoids, all or part of the cell populations cultured on microfluidic chips, and human body fluids or solid tissues and organs.

3. The cell type identification method as described in claim 1, characterized in that, In step S2, labeling of live cells is achieved using label-free or specific labeling methods, including: label-free methods that do not rely on external dyes, probes, or molecular markers; and labeling methods using live cell fluorescent dyes, gene knock-in expressed fluorescent fusion proteins, or quantum dots or nanoprobes with extremely low cytotoxicity. The cells or their subcellular structures are selected from entities that can be spatially located and tracked, including the cell nucleus, cell membrane, cytoskeleton, and endoplasmic reticulum.

4. The cell type identification method as described in claim 1, characterized in that, The microscopic imaging techniques described in step S3 cover two-dimensional or three-dimensional time-lapse imaging systems, wherein: for the unlabeled case, the imaging system includes a conventional optical microscope, a phase difference or differential interference difference microscope, a digital holographic quantitative phase imaging system, a coherent Raman scattering imaging system, and a photothermal microscopic imaging system; for the case with specific labels, the imaging system includes a conventional fluorescence microscope, a wide-field fluorescence microscope, a laser scanning confocal microscope, a two-photon excitation microscope, and a light-sheet illumination microscope.

5. The cell type identification method as described in claim 1, characterized in that, The cell tracking algorithm described in step S4 is based on the spatial coordinates and temporal evolution of the target cell or its subcellular structure signal. It employs one or more of the following algorithms for live cell tracking: threshold segmentation and connected component analysis, active contour model, Kalman filter tracking, and deep learning target detection. The migration features are extracted and quantified by the migration trajectory, including at least one of the following: total migration distance, effective migration distance, instantaneous migration rate, average migration rate, instantaneous migration deflection angle, orientation angle dispersion, and orientation index.

6. The cell type identification method as described in claim 5, characterized in that, The specific migration feature mentioned in step S4 is selected from: (1) Overall migration distance D: The total actual path length of the target object's migration trajectory during the entire tracking period, which is equal to the sum of the magnitudes of the displacement vectors between adjacent frames; in, The spatiotemporal coordinate vector representing the target object. t represents the two-dimensional or three-dimensional spatial coordinates of the target object image data. τ Let τ be the time corresponding to the target object in the τth frame, and n be the total number of frames in the spatial trajectory set; (2) Effective migration distance D eff The net linear displacement of the target object from the starting position to the ending position is equal to the magnitude of the displacement vectors at the initial and ending positions. in, and The spatiotemporal coordinate vectors representing the initial and final positions of the target object, respectively. and The image data spatial coordinates of the initial and final positions of the target object, t1 and t2, are respectively. n These are the times corresponding to the first and last frames tracked for the target object, respectively. (3) Instantaneous migration rate v(t) τ The velocity of a target object within a certain time interval is equal to the magnitude of the displacement vectors of two adjacent frames divided by the time resolution. in, The spatiotemporal coordinate vector representing the target object. Let t be the spatial coordinates of the image data of the target object. τ The time corresponding to the target object in the τth frame; (4) Average migration rate : The mathematical expectation of the rate over the entire tracking period, i.e., the average instantaneous migration rate over all time intervals; in, The spatiotemporal coordinate vector representing the target object. Let t be the spatial coordinates of the image data of the target object. τ Let τ be the time corresponding to the target object in the τth frame, and n be the total number of frames in the spatial trajectory set; (5) Instantaneous migration deflection angle θ(t) τ ): The angle between the displacement vectors of two consecutive time intervals, reflecting the degree of drastic change in the direction of motion at that moment; in, The spatiotemporal coordinate vector representing the target object. Let t be the spatial coordinates of the image data of the target object. τ The time corresponding to the target object in the τth frame; (6) Direction angle dispersion Var(θ): The variance of the instantaneous migration deflection angle at all times; in, Represents θ(t) τ The average of θ(t) over the total number of frames in the spatial trajectory set. τ ) for t τ With t τ+1 The angle between the displacement vectors of the time interval, where n is the total number of frames in the spatial trajectory set; (7) Orientation index DR: A dimensionless index characterizing the orientation of cell migration, defined as the effective migration distance D eff The ratio to the total migration distance D; the range is from 0 to 1, with a tendency towards 1 indicating highly linear directional movement and a tendency towards 0 indicating disordered random Brownian motion; in, The spatiotemporal coordinate vector representing the target object. Let t be the spatial coordinates of the image data of the target object. τ Let τ be the time corresponding to the target object in the τth frame, and n be the total number of frames in the spatial trajectory set.

7. The cell type identification method as described in claim 1, characterized in that, Other physical or biochemical methods mentioned in step S5 include lineage-based single-cell amplification, gene expression labeling, single-cell sequencing, immunofluorescence labeling, flow cytometry, multi-omics analysis, and cell identity tracking algorithms, based on one or more of these methods to obtain prior information on cell populations of known types.

8. The cell type identification method as described in claim 1, characterized in that, In step S6, an appropriate statistical method for hypothesis testing is selected based on the sample size, and multiple independent hypothesis testing experiments are conducted under different sample sizes, including: For the two cell populations in the training set, random sampling is performed to automatically select a sample sub-cell population of size A; the migration feature is used as an input parameter to perform hypothesis testing, and the significance level P of a single comparison is calculated. i and the mean of the migration features of the two types of samples. , , where i represents the i-th independent test; The system performs B independent random sampling tests consecutively and calculates the mean significance level of all independent tests after log transformation. Where lg is the logarithm with base 10, and P i Let be the significance level obtained from the i-th independent test; calculate the average of the mean transfer features from all independent tests. and ; By gradually adjusting the value of A through a traversal algorithm, the average significance level is reached after B independent random sampling tests. Stable less than the preset threshold P threshold At this point, output A as the minimum sample size A required to successfully identify these two specific cell types. min Compare the average of the mean values ​​of the migration features under this sample size. and The size is used to classify it into the corresponding cell type.

9. The cell identification method as described in claim 8, characterized in that, The cell type identification process in step S7 includes: Two cell populations of unknown type were selected from the test set and sampled. The sample included two cell populations of unknown type to be identified, specifically including two cell populations of unknown type in the same biological live sample or two biological live samples. After obtaining single-cell migration trajectory data from two unknown cell populations using microscopic imaging techniques, the minimum sample size A for the two cell populations obtained in step S6 is then used. min In these two cell populations to be identified, the number of samples drawn is A≥A. min For the test sample, select an appropriate hypothesis testing statistical method based on the sample size to compare the significance of its migration characteristics; if the calculated significance level satisfies... Requirements, among which Then, based on the mean value of migration characteristics, the two unknown cell population samples are classified into the corresponding types, or it is determined whether the two unknown cell populations belong to different types.

10. The cell type identification method as described in claim 8 or 9, characterized in that, In steps S6 and S7, an appropriate statistical method for hypothesis testing is selected based on the sample size A. Specifically, when A < 10, a non-parametric two-sided Wilcoxon rank-sum test is used; when A ≥ 10, the Z-statistic hypothesis test is used.

Citation Information

Patent Citations

  • Cell fate distinguishing and differentiation prediction method based on three-dimensional morphology

    CN121215019A