Ruminant parenchymal hepatic cell subtype analysis system

By integrating and analyzing multidimensional evidence from hepatocytes of ruminants, the problem of accuracy in identifying hepatocyte subtypes in ruminants has been solved, achieving highly reliable cell subtype identification and overcoming the lack of species specificity in existing technologies.

CN121306256APending Publication Date: 2026-01-09HENAN UNIV OF ANIMAL HUSBANDRY & ECONOMY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511609096.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-05
Publication Date
2026-01-09

AI Technical Summary

Technical Problem

The lack of effective species-specific markers and identification methods in current technologies has led to insufficient research on hepatocyte subtypes in ruminants, hindering a deeper understanding of the physiological and pathological state of dairy cow livers.

Method used

A ruminant hepatocyte subtype analysis system was used to identify hepatocyte subtypes by integrating multidimensional evidence through single-cell or single-cell nuclear transcriptome sequencing data, combined with cell type-specific markers, cell cycle analysis, differentially expressed gene analysis, and pseudo-time series analysis.

Benefits of technology

This study achieved highly reliable identification of hepatic parenchymal cell subtypes in ruminants, transforming the conclusion from probabilistic to certain. Over 90% of cell clusters were identified with clear and consistent subtypes, avoiding the uncertainty of relying on prior markers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121306256A_ABST
    Figure CN121306256A_ABST
Patent Text Reader

Abstract

The invention discloses a ruminant parenchymal hepatic cell subtype analysis system, which belongs to the technical field of bioinformatics and animal genetics, and comprises a data processing module, which is used for executing cell typing, cell cycle analysis, quasi-time sequence analysis and function enrichment analysis based on single cell sequencing data, and integrating the multi-dimensional analysis results based on a weighted decision rule to carry out cross validation and identification on the subtype of the parenchymal hepatic cells. The method provided by the invention overcomes the defect that in the prior art, human and mouse cells are directly used for marking ruminants to cause inaccurate subtypes of the parenchymal hepatic cells, and provides a reliable tool for accurate identification of subtypes of the parenchymal hepatic cells of the ruminants by establishing a set of special analysis system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the fields of bioinformatics and animal genetics, and particularly relates to a subtype analysis system for hepatocytes in ruminants. Background Technology

[0002] The liver is a core organ for metabolism, composed of a highly heterogeneous population of cells. Hepatocytes, as the main functional cells of the liver, also exhibit different functional subtypes. In studies on humans and mice, hepatocytes are classified into different subtypes based on their spatial location within the liver lobules (portal zone, intermediate zone, central venous zone). Each subtype plays a different role in physiological processes such as glucose metabolism, lipid metabolism, detoxification, and biosynthesis. The identification of these subtypes is highly dependent on their specific molecular markers. After years of research, the markers for hepatocyte subtypes in humans and mice have become relatively mature and well-defined, allowing researchers to directly utilize these known markers for accurate subtype identification of hepatocytes in both humans and mice.

[0003] However, the situation is quite different for ruminants such as dairy cows. Due to genetic and physiological differences between species, directly applying mature cell markers from humans or mice to identify cellular subtypes in ruminants carries the risk of insufficient accuracy and reliability. Currently, research on hepatic parenchymal cell subtypes in ruminants, particularly dairy cows, is lacking, with a shortage of fully validated, species-specific subtype markers and identification methods. This technological deficiency severely hinders a deeper understanding of the precise regulatory mechanisms of the dairy cow liver under physiological and pathological conditions at the cellular level.

[0004] Therefore, there is an urgent need in this field to develop a technique for identifying hepatocyte subtypes in ruminants that does not rely on directly applying markers from other species. Summary of the Invention

[0005] This invention proposes a subtype analysis system for hepatocytes in ruminants to address the problems existing in the prior art.

[0006] To achieve the above objectives, the present invention provides a ruminant hepatocyte subtype analysis system, comprising a data processing module, wherein the data processing module includes: The cell typing unit is used to classify ruminant liver cells based on single-cell or single-cell nuclear transcriptome sequencing data using cell type-specific markers to delineate hepatocyte clusters. The cell cycle analysis unit is used to perform cell cycle analysis on the liver parenchymal cell clusters to obtain the cell cycle distribution characteristics of each cell cluster. The trajectory analysis unit is used to perform differential gene expression analysis on the hepatic parenchymal cell clusters to screen key genes and perform pseudo-temporal analysis based on their expression matrix, thereby obtaining the position information of each cell cluster in the cell state development trajectory. The functional analysis unit is used to perform functional enrichment analysis on the hepatocyte clusters to determine their representative biological pathways. An integrated identification unit is used to integrate the cell cluster segmentation results, periodic distribution characteristics, trajectory location information, and functional enrichment results, and to perform cross-validation through weighted decision rules to identify hepatocyte subtypes.

[0007] Optionally, the classification of ruminant liver cells using cell type-specific markers includes: Based on single-cell or single-nucleus sequencing data, we screened the top genes that were highly expressed in each cell population through differential expression analysis, and combined them with known classic cell type markers to classify independent cell types. Unsupervised clustering of hepatocytes was performed to classify cell clusters; Based on the expression profiles of the Top genes in each cell cluster and their matching degree with known markers of hepatocyte subpopulations, ruminant hepatocyte subpopulations were identified.

[0008] Optionally, cell cycle analysis of hepatocyte clusters may include: Calculate the cycle score for each cell using the cyclone function in the scran package of the R language; Each cell is assigned to G1, S, or G2 / M phase based on the aforementioned cycle score; The proportion of cells in different stages of the cell cycle in each hepatic parenchymal cell cluster was statistically analyzed to determine its cell cycle distribution characteristics.

[0009] Optionally, differential gene expression analysis of hepatocyte clusters to screen for key genes includes: In the differential expression analysis of hepatocyte clusters, significantly differentially expressed genes that meet the conditions of logFC and adjusted p-value are selected as key genes.

[0010] Optionally, the pseudo-time series analysis based on the expression matrix of key genes includes: Based on the expression matrix of the key genes, the trajectory of cell state evolution was constructed using the Monocle software package; Locate the position of each cell cluster in the trajectory to obtain its trajectory position information.

[0011] Optional, functional enrichment analysis of hepatocyte clusters includes: For each cluster of hepatocytes, several genes with the most significant differential expression in the differential expression analysis were selected as Top genes. KEGG pathway enrichment analysis was performed on the Top genes to identify representative biological pathways for each cell cluster.

[0012] Optionally, the cross-validation using weighted decision rules includes: Each analytical dimension in the cell cluster segmentation results, periodic distribution characteristics, trajectory location information, and functional enrichment results is considered as an independent source of evidence. Assess the degree of consistency between conclusions from different sources of evidence to determine the final subtype.

[0013] Optionally, the weighted decision rule further includes: When a cell cluster is located at the terminal position in the pseudo-time trajectory and its functional enrichment results show that pathways related to terminal differentiation are significantly enriched, the weight of its terminal subtype determination is increased. When the proportion of G2 / M phase cells in a certain cell cluster is relatively high, the weight for determining its proliferation-related subtype is increased.

[0014] Compared with the prior art, the present invention has the following advantages and technical effects: Traditional methods rely solely on the similarity of marker expression for inference, resulting in singular conclusions and high uncertainty. This invention integrates multidimensional evidence, including cell cycle, pseudo-chronological trajectories, and functional pathways, for cross-validation, transforming identification conclusions from isolated "speculation points" into a mutually supportive "evidence network." Practical verification shows that for the same set of data, traditional single-marker methods may fail to provide definitive conclusions or yield contradictory results for approximately 60% of cell clusters. In contrast, this invention, through multidimensional integration, successfully identified clear and consistent subtypes for over 90% of cell clusters, achieving a qualitative leap in reliability from "probability" to "certainty." Furthermore, this invention does not rely on a priori list of markers that may exhibit species differences; instead, it autonomously analyzes conserved genes and the functional state of the cells themselves to determine subtypes. Attached Figure Description

[0015] The accompanying drawings, which form part of this application, are used to provide a further understanding of this application. The illustrative embodiments and descriptions of this application are used to explain this application and do not constitute an undue limitation of this application. In the drawings: Figure 1 This is a system operation flowchart of an embodiment of the present invention, where HC represents a healthy dairy cow and CK represents a dairy cow with ketosis; Figure 2 This is an identification and analysis diagram of 14 cell types in bovine liver tissue according to an embodiment of the present invention; Figure 3 The top 10 genes of 7 cell clusters of bovine hepatocytes in this embodiment of the invention; Figure 4The cell cluster markers for bovine hepatocytes in this embodiment of the invention are as follows: ALB: albumin; CRP: C-reactive protein; HAMP: hepcidin antimicrobial peptide; APOE: apolipoprotein E; CP: viral capsid protein; A2M: α-2-macroglobulin; CYP1A2: cytochrome P450 enzyme 1A2; CYP2E1: cytochrome P450 enzyme 2E1; SDS: serine dehydratase; PCK1: phosphoenolpyruvate carboxykinase 1; AR: androgen receptor; HSD11B1: hydroxysteroid 11-β dehydrogenase; DPP4: dipeptidyl peptidase-4; IGFBP1: insulin-like growth factor binding protein; ASL: argininosuccinate lyase; CYP3A4: cytochrome P450 3A4 enzyme; AQP9: aquaporin 9; CYP2C19: cytochrome P450 enzyme 2C19; CYP7A1: cholesterol 7α-hydroxylase; Zone 1 / Periportal: portal zone cells; Zone 2 / Middle zone: middle zone cells; Zone 3 / CentralVenous: central venous zone cells; Figure 5 This is an analytical diagram of seven cell clusters of bovine hepatocytes according to an embodiment of the present invention; Figure 6 This shows the proportion of the three cell cycles in the seven cell clusters of bovine hepatocytes in an embodiment of the present invention. Figure 7 This is a pseudo-time series analysis diagram of seven cell clusters of bovine hepatocytes in an embodiment of the present invention. Figure 8 This is a functional analysis diagram of three cell subpopulations of bovine hepatocytes according to an embodiment of the present invention. Detailed Implementation

[0016] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.

[0017] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.

[0018] Example 1 like Figure 1 As shown, this embodiment provides a ruminant hepatocyte subtype analysis system, including a data processing module, which includes: The cell typing unit is used to classify ruminant liver cells based on single-cell and / or single-cell nuclear transcriptome sequencing data using cell type-specific markers to delineate hepatocyte clusters. The cell cycle analysis unit is used to perform cell cycle analysis on the liver parenchymal cell clusters to obtain the cell cycle distribution characteristics of each cell cluster. The trajectory analysis unit is used to perform differential gene expression analysis on the hepatic parenchymal cell clusters to screen key genes and perform pseudo-temporal analysis based on their expression matrix, thereby obtaining the position information of each cell cluster in the cell state development trajectory. The functional analysis unit is used to perform functional enrichment analysis on the hepatocyte clusters to determine their representative biological pathways. An integrated identification unit is used to integrate the cell cluster segmentation results, periodic distribution characteristics, trajectory location information, and functional enrichment results, and to perform cross-validation through weighted decision rules to identify hepatocyte subtypes.

[0019] The corresponding operation process is as follows: Based on single-cell / single-cell nuclear transcriptome sequencing data, the sequencing reads were compared with the ruminant reference genome; by analyzing conserved genes highly expressed in specific cell populations, liver cell type-specific markers were identified, and these markers were used to classify ruminant liver cells. Subsequently, hepatocytes were extracted from all cells for unsupervised clustering and hepatocyte clusters were divided. Cell cycle analysis was performed on the segmented hepatocyte clusters. The cycle score of each cell was calculated using the cyclone function in the scran package of R language. Based on the score, each cell was assigned to the G1, S or G2 / M phase. The proportion of cells in different cycle phases in each hepatocyte cluster was counted to determine its cycle distribution characteristics. Differential gene expression analysis was performed on the segmented hepatocyte clusters to screen for significantly upregulated key genes as input features for pseudo-temporal analysis. Based on the expression matrix of these key genes, the Monocle software package was used to construct the trajectory of cell state development and locate the position of each cell cluster in the trajectory. Differential gene expression analysis was performed on the segmented hepatocyte clusters to select the top genes in each cell cluster that ranked first according to the |log2| fold change (logFC) and were statistically significant (p-value). KEGG pathway enrichment analysis was performed on these genes to determine their representative biological pathways and functions. Based on a set of weighted decision rules, the system integrates four aspects of information: cell cluster division, cell cycle distribution, pseudo-temporal trajectory position, and functional enrichment results, and performs cross-validation to make a highly reliable final identification of hepatocyte subtypes.

[0020] Furthermore, ruminant liver cells were classified using cell type-specific markers, and hepatic parenchymal cell clusters were divided into: The cellular landscape of ruminant liver tissue was obtained based on single-cell / single-nucleus sequencing data. Differential expression analysis was used to screen for highly expressed Top genes (defined as genes with high average expression levels and statistical significance) in each cell group. Combined with known classical cell type markers, independent cell types were identified. Furthermore, unsupervised clustering of hepatocytes was performed to divide cell clusters. The expression profiles of Top genes in each cell cluster and their matching degree with known hepatocyte subpopulation markers provided support for the subpopulation classification of ruminant hepatocytes.

[0021] Furthermore, pseudo-temporal analysis was performed on the segmented hepatic parenchymal cell clusters, including: The pseudo-temporal analysis used the Monocle software package to perform machine learning based on the expression patterns of key genes, simulating the dynamic changes in the developmental process over time. The selection criteria for key genes were: differentially expressed genes in hepatocyte clusters that met the criteria of |logFC|>0.5 and adjusted p-value (adj.p.val)<0.05.

[0022] Furthermore, functional enrichment analysis was performed on the segmented hepatocyte clusters to determine their representative biological functions, including: KEGG enrichment analysis of top genes in cell subpopulations was performed to determine the function of hepatocyte subtypes, thus providing sufficient evidence for hepatocyte subtype identification. The selection logic for top genes was as follows: for each hepatocyte cluster, the top 100 genes with logFC values ​​in differential expression analysis and adj.p.val < 0.05 were selected for subsequent functional enrichment analysis.

[0023] Furthermore, the core of the weighted decision rule lies in treating each analytical dimension as an independent source of evidence, and determining the final subtype by assessing the consistency of conclusions from different sources of evidence. Through this set of rules, the system transforms the originally independent multidimensional data analysis into a structured and interpretable decision-making process, ultimately outputting the identification results for each cluster of hepatocytes. Specific principles include: Strong Evidence Prioritization: Certain analytical dimensions are highly indicative of specific subtypes. Spatial location in pseudo-time-series trajectories and core metabolic pathways in functional enrichment analysis are considered strong evidence distinguishing between portal and central venous subtypes. Consistent Convergence: When two or more strong pieces of evidence point to a unified subtype conclusion, a high-confidence identification can be made. Weak Evidence Supporting the Conclusion: Other analytical dimensions serve as weak or repetitive evidence, used to explain anomalies or enhance the reliability of the conclusion. The portal region is typically rich in cells with proliferative potential; therefore, a relatively high proportion of G2 / M phase cells can serve as supporting evidence.

[0024] The specific steps are as follows: Data input: The system receives raw data (FASTQ file) from single-cell transcriptome sequencing of bovine liver cells from the 10x Genomics platform.

[0025] Data quality control and cell typing: The data processing module used the Seurat software package for quality control, filtering out low-quality cell nuclei. Samples were quality checked using official software, with reads aligned to the bovine reference genome (Bos taurus, AssemblyARS-UCD1.3) to obtain Cell Ranger quality control. Cells with a retained gene count greater than 200, UMI count greater than 1000, log10GenesPerUMI greater than 0.7, and mitochondrial genome proportion less than 0.2 were further quality controlled as high-quality cells. Double-cell removal was then performed using DoubletFinder software for downstream analysis. Batch effects were eliminated using the MNN (Shared Nearest Neighbor) dimensionality reduction algorithm. Based on the MNN dimensionality reduction results, the single-cell clusters were visualized using the UMAP (Unified Manifold Approximation and Projection) algorithm, with the SNN (Shared Nearest Neighbor) clustering algorithm used to obtain the optimal cell clusters. Subsequently, classical conserved markers were used for preliminary typing of all cells, and hepatocytes were extracted for re-dimensionality reduction and clustering, resulting in several hepatocyte clusters.

[0026] Cell cycle analysis: Cell cycle analysis was performed on the segmented hepatocyte clusters. The cycle score of each cell was calculated using the cyclone function in the scran package of R language. Based on the score, each cell was assigned to the G1, S or G2 / M phase. The proportion of cells in different cycle stages in each hepatocyte cluster was counted to determine its cycle distribution characteristics. Pseudo-temporal analysis: The module calls the Monocle algorithm, using the standardized and log-normalized gene expression matrix of all hepatocytes as input. First, the dispersionTable function is used to select genes with high dispersion as sorting features; then, the dimensionality is reduced using the reduceDimension function (using the DDRTree method); finally, the orderCells function is used to construct a pseudo-temporal trajectory map at the single-cell level.

[0027] Functional enrichment analysis: This module performs differentially expressed gene analysis on each hepatocyte cluster. The analysis method uses Seurat's FindAllMarkers function, with the default Wilcoxon rank-sum test. The criteria for screening significantly differentially expressed genes are: mean logarithmic change (avg_log2FC) > 0.5 and a Bonferroni-corrected p-value (adj.p.val) < 0.05. Significant genes meeting these criteria from each cell cluster are submitted to the KEGG database for hypergeometric distribution testing to reveal the significantly enriched biological pathways and functions of each cell cluster.

[0028] Integrated Assessment and Output: The system ultimately generates a structured integrated report. The report presents all analysis results and provides the final assessment conclusion. In this way, the system achieves fully automated, highly reliable analysis from data to conclusions.

[0029] The integrated report specifically includes the following core elements: (1) Cell landscape map: UMAP map showing the distribution of all hepatic parenchymal cell clusters.

[0030] (2) Cell cycle distribution diagram: a pie chart showing the proportion of each cell cluster in the cell cycle.

[0031] (3) Pseudo-time trajectory diagram: Displays the trajectory diagram of cell development path and the location of each cell cluster.

[0032] (4) Functional enrichment map: Visualization results of Top enriched pathways in each cell cluster.

[0033] (5) Integrated identification conclusion table: Each cell cluster is clearly listed in tabular form, and its evidence in all the above analyses is combined to give the final subtype automated identification conclusion according to the weighted decision rule.

[0034] In this way, the system completes a fully automated and highly reliable analysis process from raw data to multi-dimensional visualization results, and finally to the identification conclusion.

[0035] The following experiment was conducted using dairy cows as an example: (1) Quality inspection and control of sequencing samples from dairy cow liver tissue: In this embodiment, liver tissue from two dairy cows underwent quality control. To ensure high-quality sequencing data, RNA quality strictly adhered to industry-standard single-cell sequencing library construction criteria: RNA integrity index (RIN) ≥ 7 and 28S / 18S ribosomal RNA ratio ≥ 0.7. This standard effectively guarantees the integrity of RNA fragments and avoids negative impacts from degradation products on the efficiency of subsequent single-cell library construction. After passing quality control, single-cell nuclear suspensions were prepared for single-cell nuclear transcriptome sequencing (snRNA-seq) on the 10x Genomics platform.

[0036] The 10x Genomics official software, Cell Ranger, was used for data preprocessing and initial quality control of the samples. It integrates STAR software to align reads to the bovine reference genome (Bos taurus, Assembly ARS-UCD1.3). The quality control steps output key metrics, including: the number of high-quality cells identified by cell barcodes, median genes per cell, valid barcode percentage, Q30 base percentage, and friction reads. These outputs directly assess the success of the experiment and determine whether the data can proceed to downstream analysis. In this implementation case, our acceptance criteria were: Median genes per cell > 1,000, Valid barcode > 90%, Q30 > 85%, and Friction reads > 70%. snRNA-seq captured liver tissue cells from 11,294 healthy dairy cows and 28,332 dairy cows with ketosis. The average number of reads per cell was 34,608 and 19,550, respectively, and the total number of genes per sample was 22,994 and 24,514, respectively. The overall cell sample data can be used for subsequent analysis.

[0037] Building upon initial quality control using Cell Ranger, more refined cell-level quality control was performed in the R environment using the Seurat software package. This step aimed to filter background noise and low-quality nuclei, with standards based on widely accepted empirical values: nuclei with a unique gene count (nFeature_RNA) greater than 200 were retained to filter out cell debris with low RNA content; nuclei with a UMI count (nCount_RNA) greater than 1000 were retained to ensure sufficient transcript coverage; and the percentage of mitochondrial genes (percent.mt) was limited to below 20% to exclude background mitochondrial RNA infiltration due to nuclear membrane damage. In this implementation, after applying the above filtering standards, the number of high-quality nuclei obtained were 8,380 and 25,053, respectively (nucleus pass rates ranging from 74.21% to 88.43%). The average library size (number of UMIs detected per cell nucleus) was 6,100 (range: 5,052–7,147), and the average number of genes detected per cell nucleus was 2,155 (range: 1,958–2,352).

[0038] (2) Determination of cell type markers in dairy cow liver tissue: We performed an unbiased examination of the cellular landscape of bovine livers using snRNA-seq. To integrate data from different individuals, we employed the nearest neighbor (MNN) algorithm for batch effect correction, specifically using the FindIntegrationAnchors and IntegrateData functions (default parameters) in Seurat to eliminate non-biotechnical bias between individuals. Subsequently, we performed principal component analysis (PCA) on the integrated data, and based on the first 20 principal components, we used the UMAP (Uniform Manifold Approximation and Projection) algorithm for nonlinear dimensionality reduction visualization, with the core parameters n.neighbors set to 30 and min.dist set to 0.3, to clearly show the natural distribution of cell populations. In this implementation case, we obtained transcriptomes of 33,433 cell nuclei from the livers of two bovine donors, identifying 17 clusters in the bovine liver tissue. Presto performed differential tests on the cell populations (screening criteria: logfc.threshold > 0, min.pct > 0.25) to obtain all marker genes for each cell population. Markers were sorted according to gene_diff, and the top 10 markers were retrieved. Based on previous research, markers for various cell types were integrated and their expression was analyzed. Specific expression details are as follows: Figure 2 As shown.

[0039] Based on this, considering the top marker / specific expression genes of each cell group and the markers based on liver tissue cell types in previous studies (Table 1), the cell types of the 17 clusters in this embodiment are ( Figure 2 The identification marker provides data support.

[0040] Table 1

[0041] Taking into account the identification of marker genes for cell populations ( Figure 2 In this embodiment, the 17 cell groups of bovine liver tissue are divided into 14 cell types.

[0042] The markers for hepatocytes were HNF4A (Hepatocyte Nuclear Factor 4 Alpha) and APOB (Apolipoprotein B). APOC3 (Apolipoprotein C3), APOC2 (Apolipoprotein C2), and ALB (Albumin) were used as auxiliary references.

[0043] The markers for hepatic sinusoidal endothelial cells are PECAM1 (also known as CD31, Platelet Endothelial Cell Adhesion Molecule-1) and LYVE1 (Lymphatic Vessel Endothelial Receptor-1). OIT3 (Oncoprotein Induced Transcript 3) and DNASE1I3 (Deoxyribonuclease-like 3) can also serve as auxiliary markers.

[0044] The marker for vascular endothelial cells is VWF (von Willebrand factor), and PECAM1 should also be considered.

[0045] The markers for T_NK cells are CD3E (Cluster of Differentiation 3), KLRK1 (Killer Cell Lectin Like Receptor K1), and CCL5 (also known as RANTES, CC Motif Chemokine Ligand 5).

[0046] The markers for B cells are CD79B (Cluster of Differentiation 79B), CD19 (Cluster of Differentiation 19), and MS4A1 (MembraneSpanning 4-Domains A1).

[0047] The markers for plasma cells are: IGHG1 (Immunoglobulin Heavy Constant Gamma 1) and MZB1 (Marginal Zone B and B1 Cell Specific Protein).

[0048] The markers for neutrophils are TGM3 (transglutaminase 3), S100A8 (S100 calcium binding protein A8), and S100A9 (S100 calcium binding protein A9).

[0049] The markers for monocytes are VCAN (Versican) and FCN1 (Ficolin 1), and the expression level of FCGR3A (Fc Fragment Of Igg Receptor IIIa) can also be considered.

[0050] The markers for macrophages are CD163 (Cluster of Differentiation 163) and C1QC (Complement C1q C Chain).

[0051] The markers for Kupffer cells are: CD5L (Cluster of Differentiation 5L) and VSIG4 (V Set And Ig Domain-Containing 4).

[0052] The markers for traditional dendritic cells 1 are: CLEC9A (C-Type Lectin Domain Containing 9A) and ANPEP (Alanyl Aminopeptidase, Membrane), while also considering the IRF8 (Interferon regulatory factor-8) gene.

[0053] The markers for traditional dendritic cells 2 are: FCRL5 (Fc Receptor Like 5) and PLD4 (Phospholipase D Family Member 4).

[0054] The markers for hepatic stellate cells are DCN (decorin proteoglycan), ANGPTL6 (angiopoietin-like factor 6), and COLEC11 (Collectin Subfamily Member 11).

[0055] The markers for fibroblasts were COL1A1 (Collagen Type I Alpha 1 chain) and ACTA2 (Actin Alpha 2), while also considering the expression of HP (Haptoglobin), C3 (Complement C3), and JUND (JunDProto-Oncogene / AP-1 Transcription Factor Subunit).

[0056] (3) Analysis of hepatocyte subset markers in dairy cows: Similarly, after eliminating batch effects and visualizing cell clusters, in this embodiment, bovine hepatic parenchymal cells were found to be divided into 7 cell clusters ( Figure 3 And display the top 10 markers for each cell subpopulation ( Figure 3 Based on previous research, we integrated and analyzed the expression of markers / top genes in human and mouse hepatocyte subsets. The marker and top genes are shown in Table 2, and their specific expression is as follows: Figure 4 As shown. Considering the top / markers for each hepatocyte subset ( Figure 3 ) and based on hepatocyte subset markers from previous studies (Table 2, Figure 4 ), to identify markers for cell subpopulations of the 7 cell clusters in this study ( Figure 4 Provide data support.

[0057] Table 2

[0058] Comprehensive consideration of markers for identifying hepatic parenchymal cell subsets ( Figure 4 In this implementation case, the seven cell subpopulations of bovine liver parenchymal cells were divided into three subpopulation types.

[0059] The first category consists of zone 1 hepatocytes, also known as portal cells (Zone 1, Periportal). Previous studies have found that classic marker / top genes in humans and mice include ALB, ASS1 (Argininosuccinate Synthetase 1), CRP (C-reactive protein), ASL, CP (Coat Protein), A2M (Alpha-2-Macroglobulin), SDS (Serine Dehydratase), and PCK1 (Phosphoenolpyruvate Carboxykinase 1), serving as markers and top genes in the portal zone. Indeed, we also found that the human and mouse marker ALB, the human top gene CRP, and the human top gene HAMP exhibit similar high expression patterns in bovine hepatocyte clusters 1, 7, and 3. These genes can serve as markers for portal hepatocytes in bovine liver tissue. Figure 4 ).

[0060] The second category consists of hepatocytes in Zone 3, also known as central venous cells. Previous studies have indicated that mouse CYP7A1 (cholesterol 7-Alpha Hydroxylase), human CYP2C19 (Cytochrome P450 Family 2 Subfamily C Member 19), human CYP3A4 (Cytochrome P450 Family 3 Subfamily A Member 4), mouse IGFBP1 (Insulin Like Growth Factor Binding Protein 1), human DPP4 (Dipeptidyl peptidase-4), and human AR (Androgen Receptor) can localize central venous hepatocytes (cell cluster 5). Similarly, these six markers are all highly expressed in sub-culster5 of bovine hepatocytes. Figure 4 ).

[0061] The third type of cell consists of cells located in the portal vein and central vein regions, defined as intermediate zone cells, also known as Zone 2 cells. Previous studies have mentioned Zone 2 markers / top genes: human CYP1A2 (Cytochrome P450 Family 1 Subfamily A Member 2), human CYP2E1 (Cytochrome P450 Family 2 Subfamily E Member 1), human APOE (Apolipoprotein E), human HAMP (Hepcidin Antimicrobial Peptide), human HSD11B1 (Hydroxysteroid 11-Beta Dehydrogenase 1), and human AQP9 (Aquaporin 9). Human CYP1A2 and CYP2E1 exhibit similar high expression patterns in bovine hepatocyte clusters 4, 2, and 6. Figure 4 ).

[0062] In this embodiment, considering the differences between species, the three regions of bovine hepatocytes may contain differentially expressed markers compared to those in humans and mice. This is because we noted that the top genes CP, A2M, and SDS in the human portal region are highly expressed in the intermediate region of bovine hepatocytes. PCK1 in the human and mouse portal regions is expressed in both the intermediate and central venous regions of bovine hepatocytes. ASL in the mouse portal region is highly expressed in the central venous region of bovine hepatocytes. The human intermediate region marker APOE is highly expressed in the portal region of bovine hepatocytes. Human HSD11B1 and AQP9 may be markers for the central venous region of bovine hepatocytes. Figure 4 Therefore, further analysis is needed to provide evidence for the analysis of hepatocyte subsets.

[0063] (4) Cell cycle analysis of bovine liver parenchymal cell clusters: Our cycle analysis revealed that the G1 phase was primarily located in cell clusters 1, 7, and 3, which may be related to hepatic progenitor cells within the bile canals and Hering ducts of the portal region. This also indirectly reflects that cell clusters 1, 7, and 3 are composed of portal hepatocytes. Figure 6 Furthermore, in human adult liver homeostasis, hepatocytes are in a resting state and proliferate slowly. However, cell cycle database analysis shows that the number of G1 cells in bovine liver parenchyma is relatively low. This may be because the current cell cycle database only collects human and mouse datasets, and existing analyses are based on homologous transformation of human data; therefore, the current cell cycle analysis results are for reference only.

[0064] (5) Pseudo-temporal analysis of bovine liver parenchymal cell clusters: Pseudo-temporal analysis was performed on seven cell clusters of hepatocytes, dividing them into two branches starting from cell cluster 1: the HC group (HC_Branch) and the CK group (CK_Branch). Figure 7 The two branches, from right to left, are the portal zone, intermediate zone, and central venous zone. Previous studies have suggested that hepatocytes in different regions of mice and humans perform different functions. For example, portal zone hepatocytes perform functions such as bio-oxidation (human and mouse), glycogen synthesis (human and mouse), fatty acid degradation (mouse), oxidative phosphorylation (mouse), Ras signaling pathway (mouse), complement and coagulation cascade (mouse), cholesterol and sterol biosynthesis (human and mouse), and immune pathway activation (human and mouse). Central venous zone hepatocytes perform functions such as glycolysis (human and mouse), exogenous metabolism (human and mouse), detoxification (human and mouse), glutamine synthesis (human and mouse), lipogenesis (human), drug metabolism (human), P450 pathway (human), Wnt pathway activation (human), and bile production (mouse). Intermediate zone hepatocytes in humans perform similar functions to those in the central venous zone, such as P450 pathway activation, Wnt pathway activation, and drug metabolism.

[0065] (6) Identification of hepatocyte subtypes based on cell cluster analysis combined with cell function: Based on hepatocyte subset analysis, pseudo-time series analysis, and top gene function enrichment analysis, the cellular functions of three regions of bovine hepatocytes were determined. Figure 8 Although the functional pathways differ slightly due to species differences, the functions of the three regions of bovine hepatocytes are similar to those in humans and mice.

[0066] (7) Analysis of subpopulations of bovine hepatocytes: Hepatocytes are the main component of liver cells, playing a crucial role in processes such as gluconeogenesis / glycolysis, lipolysis / fat synthesis, and detoxification. In humans and mice, hepatocytes exhibit different functions based on their location within the hepatic acini, and are classified by region.

[0067] In this embodiment, bovine hepatocytes were found to be divided into 7 cell clusters. Through classical marker analysis, cell cycle analysis, pseudo-chronological analysis, and cell function analysis, these 7 cell clusters were initially identified as 3 independent cell subpopulations: portal zone hepatocytes (clusters 1, 7, and 3), intermediate zone hepatocytes (clusters 4, 2, and 6), and central venous zone hepatocytes (cluster 5).

[0068] During implementation, cell typing, cell cycle analysis, pseudo-time series analysis, and functional enrichment analysis based on single-cell sequencing data were performed. The results of these multidimensional analyses were then integrated to cross-validate and identify hepatocyte subtypes. This application overcomes the inaccuracies in hepatocyte typing caused by directly applying human and mouse cell markers to ruminants in existing technologies. By establishing a proprietary analysis system, it provides a reliable tool for the accurate identification of hepatocyte subtypes in ruminants.

[0069] The above are merely preferred embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A system for analyzing subtypes of hepatocytes in ruminants, characterized in that, The data processing module includes: The cell typing unit is used to classify ruminant liver cells based on single-cell or single-cell nuclear transcriptome sequencing data using cell type-specific markers to delineate hepatocyte clusters. The cell cycle analysis unit is used to perform cell cycle analysis on the liver parenchymal cell clusters to obtain the cell cycle distribution characteristics of each cell cluster. The trajectory analysis unit is used to perform differential gene expression analysis on the hepatic parenchymal cell clusters to screen key genes and perform pseudo-temporal analysis based on their expression matrix, thereby obtaining the position information of each cell cluster in the cell state development trajectory. The functional analysis unit is used to perform functional enrichment analysis on the hepatocyte clusters to determine their representative biological pathways. An integrated identification unit is used to integrate the cell cluster segmentation results, periodic distribution characteristics, trajectory location information, and functional enrichment results, and to perform cross-validation through weighted decision rules to identify hepatocyte subtypes.

2. The system according to claim 1, characterized in that, The method of typing ruminant liver cells using cell type-specific markers includes: Based on single-cell or single-nucleus sequencing data, we screened the top genes that were highly expressed in each cell population through differential expression analysis, and combined them with known classic cell type markers to classify independent cell types. Unsupervised clustering of hepatocytes was performed to classify cell clusters; Based on the expression profiles of the Top genes in each cell cluster and their matching degree with known markers of hepatocyte subpopulations, ruminant hepatocyte subpopulations were identified.

3. The system according to claim 1, characterized in that, Cell cycle analysis of hepatocyte clusters includes: Calculate the cycle score for each cell using the cyclone function in the scran package of the R language; Each cell is assigned to G1, S, or G2 / M phase based on the aforementioned cycle score; The proportion of cells in different stages of the cell cycle in each hepatic parenchymal cell cluster was statistically analyzed to determine its cell cycle distribution characteristics.

4. The system according to claim 1, characterized in that, Differential gene expression analysis of hepatocyte clusters to screen for key genes includes: In the differential expression analysis of hepatocyte clusters, significantly differentially expressed genes that meet the conditions of logFC and adjusted p-value are selected as key genes.

5. The system according to claim 1, characterized in that, The pseudo-time series analysis based on the expression matrix of key genes includes: Based on the expression matrix of the key genes, the trajectory of cell state evolution was constructed using the Monocle software package; Locate the position of each cell cluster in the trajectory to obtain its trajectory position information.

6. The system according to claim 1, characterized in that, Functional enrichment analysis of hepatocyte clusters included: For each cluster of hepatocytes, several genes with the most significant differential expression in the differential expression analysis were selected as Top genes. KEGG pathway enrichment analysis was performed on the Top genes to identify representative biological pathways for each cell cluster.

7. The system according to claim 1, characterized in that, The cross-validation using weighted decision rules includes: Each analytical dimension in the cell cluster segmentation results, periodic distribution characteristics, trajectory location information, and functional enrichment results is considered as an independent source of evidence. Assess the degree of consistency between conclusions from different sources of evidence to determine the final subtype.

8. The system according to claim 7, characterized in that, The weighted decision rule also includes: When a cell cluster is located at the terminal position in the pseudo-time trajectory and its functional enrichment results show that pathways related to terminal differentiation are significantly enriched, the weight of its terminal subtype determination is increased. When the proportion of G2 / M phase cells in a certain cell cluster is relatively high, the weight for determining its proliferation-related subtype is increased.