AI-based human germ cell development and culture system evaluation method

By constructing standardized reference maps and AI models, the problems of data fragmentation and inconsistent evaluation standards in germ cell research have been solved, high-precision prediction of germ cell development stages and culture system evaluation have been achieved, and data integration and prediction accuracy have been improved.

CN120613015AActive Publication Date: 2025-09-09NANJING MEDICAL UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511110678.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-08
Publication Date
2025-09-09
Estimated Expiration
2045-08-08

AI Technical Summary

Technical Problem

In existing technologies, single-cell reference maps for human germ cell research are fragmented, data processing standards are inconsistent, the accuracy of developmental stage identification is insufficient, and there is a lack of quantitative standards for in vitro culture system evaluation, resulting in the inability to establish a unified and comparable reference map and a highly accurate developmental prediction model.

Method used

A standardized reference map was constructed, and a three-level quality screening system and AI model were adopted, including single-cell annotation variational inference, k-nearest neighbor joint prediction and feature selection optimization, and the in vitro culture system was evaluated by combining k-nearest neighbor distance and meta-cell-Pearson correlation coefficient.

Benefits of technology

It has realized a full-process solution from single-cell reference map construction to culture system evaluation, established a high-precision prediction model for germ cell development stages, provided comprehensive data integration and optimization clues, and improved data comparability and prediction accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120613015A_ABST
    Figure CN120613015A_ABST
Patent Text Reader

Abstract

The invention discloses an AI-based human germ cell development and culture system evaluation method, which comprises the following steps: step 1, constructing a standardized reference map, step 2, establishing a germ cell development stage prediction model which comprises a transfer learning-based single cell annotation variation inference prediction module, a united prediction module combining single cell annotation variation inference with k nearest neighbor and a cell type classification module for feature selection optimization realize collaborative decision through a three-level confidence system; and step 3, in-vitro culture system evaluation, including construction of double indexes for quantitative evaluation and analysis of evaluation results. According to the method, a whole-process solution from single cell reference map construction, development stage prediction to culture system evaluation is realized, and the problems of reference data fragmentation, inaccurate development stage annotation and lack of quantitative standards of culture system data in the current germ cell research field are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an AI-based method for evaluating human germ cell development and culture systems, and belongs to the technical field of single-cell transcriptome data analysis. Background Art

[0002] Currently, three core technical challenges urgently need to be addressed in the evaluation of germ cell developmental stages and in vitro culture systems based on the human germ cell single-cell reference atlas: First, the fragmentation of existing research data, the varying data processing standards, and the inconsistent germ cell classification criteria make it impossible to establish a unified and comparable reference atlas; second, existing annotation tools lack sufficient accuracy in discerning germ cell developmental stages, making it difficult to construct highly accurate developmental prediction models; and finally, existing in vitro culture system evaluation methods have significant flaws, lacking the single-cell resolution and multi-dimensional evaluation metrics comparable to in vivo developmental data. These technical bottlenecks have severely hampered the development of human germ cell research and clinical applications. Summary of the Invention

[0003] The purpose of the present invention is to provide an AI-based method for evaluating human germ cell development and culture systems. This method realizes a full-process solution from single-cell reference map construction, developmental stage prediction to culture system evaluation, and solves the three key technical bottlenecks in the current field of germ cell research: fragmentation of reference data, inaccurate developmental stage annotation, and lack of quantitative standards for culture system data.

[0004] To achieve the above objectives, the technical solution adopted by the present invention is: an AI-based method for evaluating human germ cell development and culture system, which comprises the following steps:

[0005] Step 1: Build a standardized reference map, including:

[0006] (1.1) Obtain multiple sets of single-cell transcriptome raw data;

[0007] (1.2) Differentiated processing based on sequencing technology type and raw data format;

[0008] (1.3) A three-level quality screening system was used to filter the data;

[0009] (1.4) Based on developmental stage and sex characteristics, the germ cell data were divided into three groups: prenatal male germ cell group, prenatal female germ cell group, and postnatal male germ cell group;

[0010] (1.5) Analyze each group independently, remove abnormal cell clusters, and annotate the cell clusters;

[0011] (1.6) Integrate the annotation results with the known germ cell development process to build a unified developmental hierarchical annotation system;

[0012] Step 2: Build a germ cell development stage prediction model, including: a single-cell annotation variational inference prediction module based on transfer learning, a joint prediction module based on single-cell annotation variational inference combined with k-nearest neighbor prediction, and a cell type classification module optimized by feature selection. These three modules achieve collaborative decision-making through a three-level confidence system.

[0013] Step 3: In vitro culture system evaluation, including:

[0014] (3.1) Construct a dual index for quantitative evaluation, where the dual index is the k-nearest neighbor distance and the metacell-Pearson correlation coefficient;

[0015] (3.2) Interpretation of evaluation results: Based on a dual-index evaluation, the degree of similarity between cultured cells and normally developing cells in vivo and the quality of the culture system are determined.

[0016] Furthermore, in step 1, the three-level quality screening system strictly filters the data including:

[0017] Individual-level screening: retain individual sample data of individuals under 45 years old, with a normal BMI and no abnormalities in gametogenesis;

[0018] Library-level processing: Double-cell recognition algorithm is used to automatically detect and remove double-cell interference;

[0019] Cell-level filtering: Differentiated filtering criteria are set based on the sequencing technology type. The criteria are as follows:

[0020] For 10x Genomics Chromium sequencing technology data, cells that simultaneously meet the following conditions are retained: (1) the number of genes detected is >1000 and <8000; (2) the proportion of mitochondrial gene expression is <20%;

[0021] For STRT-seq sequencing technology data, cells that met the following conditions were retained: the number of gene detections was >2000 and <12000.

[0022] Furthermore, in step 1, the process of independently analyzing each group is as follows:

[0023] (1.51) Scanpy software was used to calculate gene variation in each dataset, sorting the genes from high to low, and selecting the top 2000 genes as input features for the single-cell variational inference model;

[0024] (1.52) Data representation learning, the steps are:

[0025] (1.521) Using a single-cell variational inference model, the count matrix of 2000 highly variable genes was input;

[0026] (1.522) Set dataset and library to be considered for batch correction;

[0027] (1.523) Training a single-cell variational inference model and extracting 32-dimensional latent space features;

[0028] (1.524) Clustering and uniform manifold approximation and projection dimensionality reduction analysis in latent space;

[0029] (1.53) In the projection dimensionality reduction analysis, the following abnormal cell clusters are identified and eliminated according to the set criteria:

[0030] The median number of gene detections in the cell cluster is less than the 5th percentile of the entire data set (calculated in ascending order, the same below);

[0031] The median transcript count of the cell cluster is less than the 5th percentile of the entire dataset;

[0032] The average mitochondrial gene expression of the cell cluster is greater than the 95th percentile of the entire dataset;

[0033] The cell clusters lacked expression of known germ cell marker genes;

[0034] (1.54) After removing abnormal cell clusters, repeat step 1.52 to optimize data quality;

[0035] (1.55) Cell type annotation:

[0036] Cell cluster annotation based on marker gene expression patterns and original annotation information;

[0037] Annotation of prenatal male and female germ cell groups to fine developmental stages;

[0038] The data of the postnatal male group were first roughly divided into spermatogonia / spermatids / spermatids, and then fine annotation was achieved through iterative analysis by repeating steps 1.52 to 1.55.

[0039] Furthermore, in step 2, the single-cell annotation variational inference prediction module based on transfer learning includes a single-cell variational inference model and a single-cell annotation variational inference model, and its training includes:

[0040] Pre-training optimization phase: Use the optimized hyperparameter configuration to pre-train the single-cell variational inference model. The specific parameter configuration is 32 latent space dimensions, 0.2 neuron dropout rate, and 1024 training batch size.

[0041] Hierarchical transfer learning stage: Single-cell annotation variational inference model training is performed based on the multi-level annotated germ cell development map, and the maximum training rounds are set to 20 rounds.

[0042] Furthermore, in step 2, the steps of executing the joint prediction module of single-cell annotation variational inference combined with k-nearest neighbor include:

[0043] Joint embedding space construction: Unknown sample data and reference datasets are jointly mapped into a 32-dimensional latent space to eliminate batch effects;

[0044] k-nearest neighbor classification prediction: In the latent space, the Euclidean distance between the unknown cell and the reference cell is calculated, the 10 nearest reference cells are selected, the cell subtype frequencies of the reference cells are counted, and the cell subtype with the highest frequency is selected as the predicted identity of the unknown cell.

[0045] Furthermore, in step 2, the steps of executing the cell type classification model optimized by feature selection include:

[0046] Double gene screening: In the first round, feature genes are screened through stochastic gradient descent. In the second round, the top 300 genes by weight are selected for final classification training by cell type. This screening allows the model to focus on the most discriminative features.

[0047] Class-specific modeling: A one-to-many multi-classification strategy is used to establish a dedicated classifier for each cell subtype, ultimately predicting the cell subtype identity of unknown cells by probability.

[0048] Furthermore, in step 2, based on the prediction results of the three modules, a multi-module consensus decision-making mechanism is constructed and a three-level confidence evaluation system is established:

[0049] High confidence result: If the three modules predict the cell subtype of the unknown cell in exactly the same way, the predicted identity is output and considered to be a high confidence result;

[0050] Medium confidence result: If any two of the three modules predict the same subtype of the unknown cell, the predicted identity is output and considered to be a medium confidence result;

[0051] Low confidence result: If the three modules predict different subtypes of unknown cells, it is considered a low confidence result and the cell type cannot be output.

[0052] Furthermore, in step 3, the cell stability evaluation based on k-nearest neighbor distance is as follows:

[0053] Joint embedding space construction: The single-cell annotation variational inference prediction module in the germ cell development stage prediction model is used to map in vitro cultured cells and in vivo reference dataset cells into a 32-dimensional latent feature space to ensure data comparability;

[0054] k-nearest neighbor distance metric calculation: For each cultured cell, calculate its Euclidean distance to all cells in the reference dataset, select the 10 closest reference cells, and then calculate the average of the Euclidean distances between the cultured cell and the 10 closest reference cells.

[0055] Furthermore, in step 3, the similarity evaluation based on the metacell-Pearson correlation coefficient is as follows:

[0056] Gene screening: A germ cell development stage prediction model was used to predefine 2,000 highly discriminative characteristic genes to ensure analysis specificity;

[0057] Construction of reference expression profile: For each cultured cell, determine its 10 nearest neighbor reference cells in the latent space, and calculate the average expression value of these reference cells on the characteristic genes as the reference expression profile;

[0058] Similarity quantification: The Pearson correlation coefficient was used to assess the overall similarity between the cultured cells and the reference expression profile.

[0059] Furthermore, in step 3, the degree of similarity between the cultured cells and the normally developed cells in vivo is determined based on the following threshold standards:

[0060] High-quality cell population: k-nearest neighbor distance < 1.5 and correlation coefficient > 0.5, indicating that the cultured cells are highly similar to normally developed cells in vivo;

[0061] Deviant cell populations:

[0062] (1) k nearest neighbor distance > 1.5 or correlation coefficient < 0.5, indicating that the cultured cells have low similarity to the normally developed cells in vivo;

[0063] (2) k nearest neighbor distance > 2 or correlation coefficient < 0.3, indicating that the similarity between cultured cells and normally developed cells in vivo is very low;

[0064] Then, the proportion of high-quality cell populations cultured in vitro was calculated, and the quality of the culture system was evaluated: the higher the proportion of high-quality cell populations and the later the developmental stage, the more effective the culture system was considered to be; otherwise, the culture system was considered to need further optimization.

[0065] The beneficial effects of the present invention are as follows:

[0066] (1) Data integration: We have constructed the most comprehensive human germ cell development map to date, covering the entire developmental process from primordial germ cells to mature gametes, and have made data from different studies comparable through standardized processing procedures.

[0067] (2) Model prediction: A high-precision germ cell development stage prediction model was established, with a prediction accuracy of 99.5% at the first annotation level and 93% at the most detailed annotation level. In addition, the model significantly outperformed other existing tools in three other typical model evaluation indicators: Cohen's kappa coefficient, macro-average F1 score, and micro-average F1 score.

[0068] (3) Application value: It can complete the whole process analysis from gene expression matrix to culture evaluation results, and can be successfully applied to real in vitro culture data to provide optimization clues. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] Figure 1 is a flow chart of the evaluation method of the present invention;

[0070] Figure 2 is a schematic diagram of the standardized reference atlas of human germ cells constructed by the present invention;

[0071] Figure 3 is a flowchart of constructing a germ cell development stage prediction model of the present invention;

[0072] Figure 4 2 is a graph comparing the prediction performance of the germ cell development stage prediction model of the present invention with that of existing tools;

[0073] Figure 5 This is an example of the results of the germ cell development stage prediction model of the present invention for evaluating the quality of in vitro culture data. DETAILED DESCRIPTION

[0074] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0075] like Figures 1 to 5 , an AI-based method for evaluating human germ cell development and culture systems, mainly includes the following three steps:

[0076] Step 1: Build a standardized reference map: Establish a unified data processing process and a strict three-level quality control system;

[0077] Step 2: Build a high-precision germ cell development stage prediction model, including: a single-cell annotation variational inference prediction module based on transfer learning, a joint prediction module combining single-cell annotation variational inference with k-nearest neighbor prediction, and a cell type classification model optimized by feature selection. These three modules achieve collaborative decision-making through a three-level confidence system.

[0078] Step 3: Evaluation of the in vitro culture system: Based on the dual-index evaluation of k-nearest neighbor distance and meta-cell-Pearson correlation coefficient, the degree of similarity between the cultured cells and the normally developed cells in vivo is determined.

[0079] The specific implementation process of step 1 is as follows:

[0080] 1.1. Obtain multiple sets of single-cell transcriptome raw data from international public databases (GEO, ENA).

[0081] 1.2. Differentiated processing procedures are adopted according to the sequencing technology type and data format:

[0082] 1.2.1, 10x Genomics Chromium sequencing technology fastq format data: Use the official analysis software CellRanger to perform alignment and quantification based on the official 2024-A version of the transcriptome reference set.

[0083] 1.2.2. 10x Genomics Chromium sequencing technology bam format data: Use the official tool bam2fastq to convert to fastq format, and then process according to the process in 1.2.1.

[0084] 1.2.3. Unique molecular identifier count format data of 10x Genomics Chromium sequencing technology. Based on the correspondence between gene database identifiers and standard gene identifiers, gene identifiers are uniformly converted to the 2024-A standard nomenclature.

[0085] 1.2.4, STRT-seq sequencing technology fastq format data: open source analysis process is used for alignment and quantification based on the same reference genome 2024-A.

[0086] 1.3, use a three-level quality screening system to strictly filter the data:

[0087] 1.3.1, Individual level screening: retain individual sample data of individuals under 45 years old, with BMI within the normal range and no abnormalities in gametogenesis.

[0088] 1.3.2, Library-level processing: A two-cell recognition algorithm is used to automatically detect and remove two-cell interference for 10x Genomics Chromium sequencing technology data.

[0089] 1.3.3, Cell-level filtering: Set differentiated filtering standards based on the sequencing technology type. The standards are as follows:

[0090] For 10x Genomics Chromium sequencing technology data, cells that simultaneously meet the following conditions are retained: (1) the number of genes detected is >1000 and <8000; (2) the proportion of mitochondrial gene expression is <20%;

[0091] For STRT-seq sequencing technology data, cells that met the following conditions were retained: the number of gene detections was >2000 and <12000.

[0092] 1.4. Based on developmental stage and sex characteristics, the germ cell data were divided into the following three groups: prenatal male germ cell group, prenatal female germ cell group, and postnatal male germ cell group.

[0093] 1.5, perform the following analysis independently for each group:

[0094] 1.5.1. Scanpy software was used to calculate gene variation per dataset, sorting the data from high to low. The top 2000 genes were selected as input features for the single-cell variational inference model.

[0095] 1.5.2, Data representation learning, the steps are:

[0096] 1.5.2.1, using the single-cell variational inference model, input the count matrix of 2000 highly variable genes;

[0097] 1.5.2.2, Set up datasets and libraries as batch calibration considerations;

[0098] 1.5.2.3, train a single-cell variational inference model and extract 32-dimensional latent space features;

[0099] 1.5.2.4, Clustering and Uniform Manifold Approximation and Projection Dimensionality Reduction Analysis in Latent Space;

[0100] 1.5.3. In the projection dimensionality reduction analysis, the following abnormal cell clusters were identified and eliminated according to the set criteria:

[0101] The median number of gene detections in the cell cluster is less than the 5th percentile of the entire data set (calculated in ascending order, the same below);

[0102] The median transcript count of the cell cluster is less than the 5th percentile of the entire dataset;

[0103] The average mitochondrial gene expression of the cell cluster is greater than the 95th percentile of the entire dataset;

[0104] The cell clusters lacked expression of known germ cell marker genes;

[0105] 1.5.4, after removing abnormal cell clusters, repeat step 1.52 to optimize data quality;

[0106] 1.5.5, Cell type annotation:

[0107] Cell cluster annotation based on marker gene expression patterns and original annotation information;

[0108] Annotation of prenatal male and female germ cell groups to fine developmental stages;

[0109] The data of the postnatal male group were first roughly divided into spermatogonia / spermatids / spermatids, and then fine annotation was achieved through iterative analysis by repeating steps 1.52 to 1.55.

[0110] 1.6. Integrate the refined annotation results with the known germ cell development process to construct a unified developmental hierarchical annotation system. For male germ cells, a four-level developmental annotation system was constructed. For female germ cells, a three-level developmental annotation system was constructed.

[0111] In step 2, the establishment of the germ cell development stage prediction model includes:

[0112] 2.1, Core module architecture of the prediction model:

[0113] 2.1.1, Single-cell annotation variational inference prediction module based on transfer learning:

[0114] 2.1.1.1. Pre-training Optimization Phase: A single-cell variational inference model was pre-trained using optimized hyperparameters, with a latent space dimension of 32, a neuron dropout rate of 0.2, and a training batch size of 1024. This parameter combination has been tested to achieve the best balance between classification accuracy and model generalization.

[0115] 2.1.1.2, Hierarchical Transfer Learning Phase: A single-cell annotation variational inference model is trained based on a multi-level annotated germ cell developmental atlas, with a maximum training epoch of 20. This design overcomes the limitations of traditional single-level annotation models and can simultaneously meet the needs of coarse-grained lineage analysis and fine-grained developmental stage identification.

[0116] 2.1.2, Single-cell annotation variational inference combined with k-nearest neighbor joint prediction module:

[0117] 2.1.2.1, Joint Embedding Space Construction: First, the unknown sample data and the reference dataset are jointly mapped into a 32-dimensional latent space to effectively eliminate batch effects;

[0118] 2.1.2.2, k-nearest neighbor classification prediction: Calculate the Euclidean distance between the unknown cell and the reference cells, select the 10 closest reference cells, count the subtype frequencies of the reference cells, and select the subtype with the highest frequency as the predicted identity of the unknown cell. This method effectively solves the problem of classifying cell boundaries during developmental transitions.

[0119] 2.1.3, Feature selection optimized cell type classification module:

[0120] 2.1.3.1, Double gene screening: In the first round, feature genes are screened by stochastic gradient descent. In the second round, the top 300 genes by weight are selected for final classification training by cell type. This screening allows the model to focus on the most discriminative features.

[0121] 2.1.3.2 Class-Specific Modeling: A one-to-many multi-classification strategy is used to build a dedicated classifier for each cell subtype, ultimately predicting the cell subtype identity of unknown cells through probability.

[0122] 2.2. Based on the prediction results of the three modules, a multi-module consensus decision-making mechanism is constructed and a three-level confidence assessment system is established:

[0123] 2.2.1, High confidence result: If the three modules predict the cell subtype of the unknown cell in exactly the same way, the predicted identity is output and considered to be a high confidence result;

[0124] 2.2.2, Medium confidence result: If any two of the three modules predict the same subtype of an unknown cell, the predicted identity is output and considered to be a medium confidence result;

[0125] 2.2.3, Low confidence result: If the three modules predict different subtypes of unknown cells, it is considered a low confidence result and the cell type cannot be output.

[0126] In step 3, the in vitro culture assessment specifically includes:

[0127] 3.1, Construct a dual-index quantitative evaluation method:

[0128] 3.1.1, Cell stability evaluation based on k-nearest neighbor distance:

[0129] 3.1.1.1. Joint embedding space construction: The single-cell annotation variational inference prediction module in the germ cell development stage prediction model is used to jointly map in vitro cultured cells and in vivo reference dataset cells into a 32-dimensional latent feature space to ensure data comparability.

[0130] 3.1.1.2. Calculation of k-nearest neighbor distance index: For each cultured cell, calculate its Euclidean distance to all cells in the reference dataset, select the 10 closest reference cells, and then calculate the average of the Euclidean distances between the cultured cell and the 10 closest reference cells.

[0131] 3.1.2, Similarity evaluation based on metacell-Pearson correlation coefficient:

[0132] 3.1.2.1. Gene screening: A germ cell development stage prediction model was used to predefine 2,000 highly discriminative characteristic genes to ensure analysis specificity.

[0133] 3.1.2.2, Reference expression profile construction: For each cultured cell, determine its 10 nearest neighbor reference cells in the latent space, and calculate the average expression value of these reference cells on the characteristic genes as the reference expression profile.

[0134] 3.1.2.3, Similarity quantification: Pearson correlation coefficient was used to evaluate the overall similarity between the expression profiles of cultured cells and the reference.

[0135] 3.2, Analysis and application of evaluation results:

[0136] 3.2.1. The similarity between cultured cells and normally developed cells in vivo and the quality of the culture system should be judged based on the following threshold standards:

[0137] 3.2.1.1, High-quality cell population: k-nearest neighbor distance < 1.5 and correlation coefficient > 0.5, indicating that the cultured cells are highly similar to normally developed cells in vivo.

[0138] 3.2.1.2, Deviation from the cell population:

[0139] (1) k nearest neighbor distance > 1.5 or correlation coefficient < 0.5, indicating that the cultured cells have low similarity to the normally developed cells in vivo;

[0140] (2) A k-nearest neighbor distance > 2 or a correlation coefficient < 0.3 indicates that the cultured cells have a low similarity to the normally developing cells in vivo.

[0141] Then, the proportion of high-quality cell populations cultured in vitro was calculated, and the quality of the culture system was evaluated: the higher the proportion of high-quality cell populations and the later the developmental stage, the more effective the culture system was considered to be; otherwise, the culture system was considered to need further optimization.

[0142] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the above embodiments do not limit the scope of protection of the present invention in any form. All technical solutions obtained by equivalent substitution, etc., fall within the scope of protection of the present invention. Parts not covered by the present invention are the same as the existing technology or can be implemented using existing technology.

Claims

1. An AI-based method for evaluating human germ cell development and culture system, characterized in that: The steps include: Step 1: Build a standardized reference map, including: (1.1) Obtain multiple sets of single-cell transcriptome raw data; (1.2) Differentiated processing based on sequencing technology type and raw data format; (1.3) A three-level quality screening system was used to filter the data; (1.4) Based on developmental stage and sex characteristics, the germ cell data were divided into three groups: prenatal male germ cell group, prenatal female germ cell group, and postnatal male germ cell group; (1.5) Analyze each group independently, remove abnormal cell clusters, and annotate the cell clusters; (1.6) Integrate the annotation results with the known germ cell development process to build a unified developmental hierarchical annotation system; Step 2: Build a germ cell development stage prediction model, including: a single-cell annotation variational inference prediction module based on transfer learning, a joint prediction module based on single-cell annotation variational inference combined with k-nearest neighbor prediction, and a cell type classification module optimized by feature selection. These three modules achieve collaborative decision-making through a three-level confidence system. Step 3: In vitro culture system evaluation, including: (3.1) Construct a dual index for quantitative evaluation, where the dual index is the k-nearest neighbor distance and the metacell-Pearson correlation coefficient; (3.2) Interpretation of evaluation results: Based on a dual-index evaluation, the degree of similarity between cultured cells and normally developing cells in vivo and the quality of the culture system are determined.

2. The AI-based method for evaluating human germ cell development and culture system according to claim 1, characterized in that: In step 1, the data is strictly filtered by a three-level quality screening system, including: Individual-level screening: retain individual sample data of individuals under 45 years old, with a normal BMI and no abnormalities in gametogenesis; Library-level processing: Double-cell recognition algorithm is used to automatically detect and remove double-cell interference; Cell-level filtering: Differentiated filtering criteria are set based on the sequencing technology type. The criteria are as follows: For 10x Genomics Chromium sequencing technology data, cells that simultaneously meet the following conditions are retained: (1) the number of genes detected is >1000 and <8000; (2) the proportion of mitochondrial gene expression is <20%; For STRT-seq sequencing technology data, cells that meet the following conditions are retained: the number of gene detection is >2000 and <12000.

3. The AI-based method for evaluating human germ cell development and culture system according to claim 1, wherein: In step 1, each group is analyzed independently, and the process is as follows: (1.51) Scanpy software was used to calculate gene variation in each dataset, sorting the genes from high to low, and selecting the top 2000 genes as input features for the single-cell variational inference model; (1.52) Data representation learning, the steps are: (1.521) Using a single-cell variational inference model, the count matrix of 2000 highly variable genes was input; (1.522) Set dataset and library to be considered for batch correction; (1.523) Training a single-cell variational inference model and extracting 32-dimensional latent space features; (1.524) Clustering and uniform manifold approximation and projection dimensionality reduction analysis in latent space; (1.53) In the projection dimensionality reduction analysis, the following abnormal cell clusters are identified and eliminated according to the set criteria: The median number of gene detections in the cell cluster is less than the 5th percentile of the entire data set; The median transcript count of the cell cluster is less than the 5th percentile of the entire dataset; The average mitochondrial gene expression of the cell cluster is greater than the 95th percentile of the entire dataset; The cell clusters lacked expression of known germ cell marker genes; (1.54) After removing abnormal cell clusters, repeat step 1.52 to optimize data quality; (1.55) Cell type annotation: Cell cluster annotation based on marker gene expression patterns and original annotation information; Annotation of prenatal male and female germ cell groups to fine developmental stages; The data of the postnatal male group were first roughly divided into spermatogonia / spermatids / spermatids, and then fine annotation was achieved through iterative analysis by repeating steps 1.52 to 1.

55.

4. The AI-based method for evaluating human germ cell development and culture system according to claim 1, wherein: In step 2, the single-cell annotation variational inference prediction module based on transfer learning includes a single-cell variational inference model and a single-cell annotation variational inference model, and its training includes: Pre-training optimization phase: Use the optimized hyperparameter configuration to pre-train the single-cell variational inference model. The specific parameter configuration is 32 latent space dimensions, 0.2 neuron dropout rate, and 1024 training batch size. Hierarchical transfer learning stage: Single-cell annotation variational inference model training is performed based on the multi-level annotated germ cell development map, and the maximum training rounds are set to 20 rounds.

5. The AI-based method for evaluating human germ cell development and culture system according to claim 1, wherein: In step 2, the steps of executing the joint prediction module of single-cell annotation variational inference combined with k-nearest neighbor include: Joint embedding space construction: first, the unknown sample data and the reference dataset are jointly mapped into a 32-dimensional latent space to eliminate batch effects; k-nearest neighbor classification prediction: Then, in the latent space, the Euclidean distance between the unknown cell and the reference cell is calculated, the 10 nearest reference cells are selected, the cell subtype frequencies of the reference cells are counted, and the cell subtype with the highest frequency is selected as the predicted identity of the unknown cell.

6. The AI-based method for evaluating human germ cell development and culture system according to claim 1, wherein: In step 2, the steps of executing the cell type classification model optimized by feature selection include: Double gene screening: In the first round, feature genes are screened through stochastic gradient descent. In the second round, the top 300 genes by weight are selected for final classification training by cell type. This screening allows the model to focus on the most discriminative features. Class-specific modeling: A one-to-many multi-classification strategy is used to establish a dedicated classifier for each cell subtype, ultimately predicting the cell subtype identity of unknown cells by probability.

7. The AI-based method for evaluating human germ cell development and culture system according to claim 1, wherein: In step 2, based on the prediction results of the three modules, a multi-module consensus decision-making mechanism is constructed and a three-level confidence evaluation system is established: High confidence result: If the three modules predict the cell subtype of the unknown cell in exactly the same way, the predicted identity is output and considered to be a high confidence result; Medium confidence result: If any two of the three modules predict the same subtype of the unknown cell, the predicted identity is output and considered to be a medium confidence result; Low confidence result: If the three modules predict different subtypes of unknown cells, it is considered a low confidence result and the cell type cannot be output.

8. The AI-based method for evaluating human germ cell development and culture system according to claim 1, wherein: In step 3, the cell stability evaluation based on k-nearest neighbor distance is as follows: Joint embedding space construction: The single-cell annotation variational inference prediction module in the germ cell development stage prediction model is used to map in vitro cultured cells and in vivo reference dataset cells into a 32-dimensional latent feature space to ensure data comparability; k-nearest neighbor distance metric calculation: For each cultured cell, calculate its Euclidean distance to all cells in the reference dataset, select the 10 closest reference cells, and then calculate the average of the Euclidean distances between the cultured cell and the 10 closest reference cells.

9. The AI-based method for evaluating human germ cell development and culture system according to claim 1, wherein: In step 3, the similarity evaluation based on the metacell-Pearson correlation coefficient is as follows: Gene screening: A germ cell development stage prediction model was used to predefine 2,000 highly discriminative characteristic genes to ensure analysis specificity; Construction of reference expression profile: For each cultured cell, determine its 10 nearest neighbor reference cells in the latent space, and calculate the average expression value of these reference cells on the characteristic genes as the reference expression profile; Similarity quantification: The Pearson correlation coefficient was used to assess the overall similarity between the cultured cells and the reference expression profile.

10. The AI-based method for evaluating human germ cell development and culture system according to claim 1, wherein: In step 3, the degree of similarity between the cultured cells and the normally developed cells in vivo is determined based on the following threshold standards: High-quality cell population: k-nearest neighbor distance < 1.5 and correlation coefficient > 0.5, indicating that the cultured cells are highly similar to normally developed cells in vivo; Deviant cell populations: (1) k nearest neighbor distance > 1.5 or correlation coefficient < 0.5, indicating that the cultured cells have low similarity to the normally developed cells in vivo; (2) k nearest neighbor distance > 2 or correlation coefficient < 0.3, indicating that the similarity between cultured cells and normally developed cells in vivo is very low; Then, the proportion of high-quality cell populations cultured in vitro was calculated, and the quality of the culture system was evaluated: the higher the proportion of high-quality cell populations and the later the developmental stage, the more effective the culture system was considered to be; otherwise, the culture system was considered to need further optimization.

Citation Information

Patent Citations

  • Digital cell development trajectory tracking method, system, equipment and medium

    CN118824370A

  • ScRNA-seq data clustering method, system and device based on ZINB distribution and graph attention

    CN120432017A

  • Systems, software, and methods for multiomic single cell classification and prediction and longitudinal trajectory analysis

    US20240249839A1