Cell annotation method based on pulmonary fibrosis scRNA-Seq data
By constructing a dictionary voting matrix and similarity score calculation method, combined with the Seurat tool and the λ value regulation mechanism, the problem of insufficient cell annotation accuracy in single-cell data analysis of pulmonary fibrosis is solved, efficient and automated cell annotation is achieved, and the accuracy and stability of annotation is improved, and important data support is provided for the research of pulmonary fibrosis diseases.
Patent Information
- Application Number
- CN202510274805.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-27
AI Technical Summary
The prior art has problems such as insufficient cell annotation accuracy, incomplete marker gene construction and standardized evaluation system in the analysis of single-cell transcriptome data related to pulmonary fibrosis, which makes it difficult to achieve high-precision cell annotation and deep characterization.
A cell annotation method based on pulmonary fibrosis scRNA-Seq data is proposed, and automated cell type recognition is achieved by constructing a dictionary voting matrix and similarity score calculation. This method combines the Seurat tool to extract cluster-related Marker genes and introduces a λ value regulation mechanism to ensure the accuracy and robustness of the annotation results.
Efficient and automated cellular annotation are achieved, which significantly reduces the time and cost of traditional manual annotation, improves the stability and consistency of annotations, and provides important data support for the diagnosis and research of pulmonary fibrosis diseases.
Smart Images

Figure CN120220828A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to healthcare informatics, and particularly to a cell annotation method based on scRNA-Seq data of pulmonary fibrosis Background Art
[0002] With the continuous development of the economy and society, the problem of urban air pollution has become increasingly serious. Especially for people living in industrial cities, long-term exposure to harmful substances is likely to form chronic lung diseases, especially pulmonary fibrosis. City A in northern China is famous for its rich polymetallic symbiotic ore resources, especially rare earth reserves, and is also the main ore supply place for several large steel groups. In the production processes such as mining, ore dressing, and smelting in the mining areas and factory areas, workers inhale a large amount of industrial dust (mainly SiO2, iron, and rare earth compounds) every day. In recent years, pneumoconiosis has been effectively controlled, but there are still many patients suffering from pulmonary fibrosis caused by pneumoconiosis. In addition, smoking and bad living habits have also exacerbated the occurrence of pulmonary fibrosis, resulting in an increasing incidence rate year by year and gradually tending to be younger
[0003] The typical symptoms of pulmonary fibrosis include dyspnea, which may lead to hypoxia and acidosis in severe cases, and even endanger life. The fibrosis problem after organ transplantation is also a common and serious medical challenge. Since the pathogenic mechanisms of industrial dust and haze are highly similar, the research results related to pulmonary fibrosis have important reference value for understanding the pathogenic mechanism of haze and effective prevention and treatment. Therefore, it is of great significance to deeply study the pathogenesis of pulmonary fibrosis, especially its applications in gene expression regulation, industrial production, and residents' health
[0004] Although there are currently some related studies on the mechanism of pulmonary fibrosis, its molecular mechanism is still not fully understood. Most of the existing studies focus on understanding from the perspective of molecular gene expression, and the research on the molecular mechanism based on single-cell transcriptome (scRNA-seq) technology is still relatively scarce. In recent years, although some progress has been made in cell annotation work based on scRNA-seq technology, many challenges still remain. The combination of manual annotation and automatic annotation technologies shows natural advantages in cell subtype recognition. Although there are some studies integrating manual and automatic annotation, such as SCSA, scCATCH, SingleR, and SciBet, etc., these tools still have problems with insufficient accuracy. Especially in cell subtype screening, marker gene construction, and improvement of the standardized evaluation system, there are still many problems to be solved. There is still a certain gap between the existing research and high-precision cell annotation and in-depth characterization. Cell annotation work should focus on accuracy rather than quantity. Accurate cell annotation is the cornerstone of subsequent single-cell transcriptome data analysis, which can provide valuable resources for the diagnosis and prevention research of specific diseases and provide important new perspectives for understanding biological mechanisms and genomic diversity Summary of the Invention
[0005] In view of this, the present invention proposes a cell annotation method based on scRNA-Seq data of pulmonary fibrosis, which realizes the automatic cell type recognition of pulmonary fibrosis single-cell data by constructing a dictionary voting matrix and calculating similarity scores. This method first extracts clustering-related Marker genes based on the Seurat tool, and combines with a dictionary of cell types specific to pulmonary fibrosis to construct a voting matrix to quantify the characteristic expressions of different cell types. Subsequently, the optimal cell type is calculated using a similarity score matrix and a weighted scoring matrix, and a λ-value regulation mechanism is introduced to ensure the accuracy and robustness of the annotation results. In addition, the present invention has developed an online analysis platform for pulmonary fibrosis cell annotation based on the RShiny framework, providing functions such as UMAP visualization, score heatmap, and Delta distribution analysis, so as to provide efficient and reliable data support for single-cell research and precision medicine applications related to pulmonary fibrosis.
[0006] The research and development of the cell annotation method based on scRNA-Seq data of pulmonary fibrosis includes the following steps:
[0007] S1. Data source: Construct a dynamic dust-exposed mouse model to obtain scRNA-Seq data related to pulmonary fibrosis;
[0008] S2. Data preprocessing: Perform data preprocessing, including removing low-quality data and eliminating batch effects;
[0009] S3. Construction of the pulmonary fibrosis annotation dictionary: Sort and correct the marker genes according to the importance of the literature, screen out the marker genes (including duplicates), classify them into four levels: cell function, cell line, cell type, and cell subtype, and organize the annotation dictionary file related to pulmonary fibrosis;
[0010] S4. Development of the algorithm: Construct a cell annotation process related to pulmonary fibrosis based on scRNA-seq data, through a similarity matrix, manually organize the cell-gene association dictionary related to pulmonary fibrosis, and optimize the evaluation algorithm;
[0011] S5. Construction of the FPCAM database: Build an online analysis platform - FPCAM based on R shiny. As an optimization method,
[0012] S1. Data source: All collected lung tissue samples were numbered and grouped, and each experimental group contained three independent biological replicates to ensure the reliability and statistical significance of the experimental results. To further study the cellular heterogeneity and gene expression profiles in these samples, all samples were subsequently sent to BGI for single-cell RNA sequencing (scRNA-seq), a technique that can provide high-resolution gene expression data for in-depth analysis of the potential effects of silica dust exposure on mouse lung cell types and functions.
[0013] S2. Data preprocessing: The first step in single-cell RNA sequencing (scRNA-seq) data analysis is quality control of the raw data. Low-quality reads and adapter sequences are removed. In addition, cells with low quality also need to be removed, usually by filtering out cells with too few genes or transcripts. After data quality control, the next step is to align the read sequences to the reference genome. The purpose of alignment is to determine the position of each cell's transcript in the genome and provide a basis for subsequent quantification of gene expression levels. After alignment, the expression level of each gene in each cell is extracted by counting the number of reads of each gene in each cell. To eliminate the effects of sequencing depth and other technical biases, normalization is required. The normalized data can more accurately reflect the true expression levels of genes. In addition to normalization, another common step is to remove batch effects, especially when the data come from multiple experimental batches. After completing quality control, alignment, counting, normalization, and batch effect correction, the resulting gene expression matrix is usually the basis for subsequent analysis. Through the above steps, the finally generated gene expression matrix not only provides quantitative data on cell gene expression but also provides a solid foundation for subsequent cell type annotation, differential expression analysis, pathway enrichment analysis, etc. The expression information of each cell is normalized and converted into a format suitable for higher-order analysis, ultimately enabling the revelation of the heterogeneity and functions of the single-cell population.
[0014] S3. Construction of the pulmonary fibrosis annotation dictionary: Review existing scientific literature, collect information on known marker genes, and extract relevant gene expression data from public databases (such as GEO, EGA, SRA, etc.). The custom database comprehensively collates the data from a total of 6 articles by Sikkema, Madissoon, Sun, etc., sorts the marker genes according to the importance of the literature, and corrects them. The 879 selected marker genes (including duplicates) are classified at four levels: cell function, cell line, cell type, and cell subtype, covering 79 cell subtypes. Finally, the top 10 marker genes are selected by a voting system to construct a dictionary database. Based on the frequency of use of the marker genes, the weighted score of each gene is calculated, and this database is used as the custom database for subsequent cell type annotation. This database finally contains 631 marker genes.
[0015] S4. Algorithm development: First, construct an expression matrix through scRNA-seq analysis. The abscissa is the cell cluster, the ordinate is the Marker gene, and each element in the matrix is the expression value of different Marker genes in different cell clusters. The number of Marker genes corresponding to each cell cluster is different. Secondly, construct a dictionary voting matrix. The abscissa is the Marker gene, the ordinate is the cell category, and each element in the matrix is the vote count of each Marker gene in different cell categories. The number of Marker genes in different cell categories is different. Establish a similarity score matrix calculated based on the marker gene and the gene expression ratio. Determine the optimal threshold of λ. Finally, the cell will be annotated according to the cell type corresponding to the λ1 value. If there is a second-highest value λ2, an additional cell type corresponding to the second-highest value λ2 will be added for annotation on the original basis. It should be noted that the λ value may be 0 or negative, in which case the λ value is meaningless. For the evaluation of the model annotation ability, we can use the accuracy to measure the matching degree between the cell type predicted by the algorithm and the true annotated cell type.
[0016] S5. FPCAM Database - FPCAM Setup: FPCAM provides an integrated user interface designed to simplify the processing and analysis workflow of single - cell RNA sequencing (scRNA - seq) data. The menu bar on the left side of the web page includes multiple functional modules such as "Cell Annotation", "UMAP Visualization", "Score Heatmap", and "Delta Distribution", and each module corresponds to different analysis tasks. Users can select appropriate analysis modules according to their research needs, and the page guides users to complete data upload and processing through an intuitive interaction method. In each analysis module, users first need to upload data files that meet the requirements, usually single - cell data in Excel format or H5 format. After the file is uploaded, users can click the "Run" button on the page to start the data processing and analysis workflow. This workflow covers steps such as data quality control, dimensionality reduction, normalization, clustering, and cell type annotation to ensure the accuracy and reproducibility of the data. Each analysis module displays the analysis results through visualization charts. For example, UMAP plots are used to show the spatial distribution of cell populations, score heatmaps are used to show the expression patterns of cells on various genes, and Delta distribution plots reflect the differences between cell types. The result display part provides real - time analysis feedback for users, and users can visually observe the analysis results through the charts. At the same time, the analysis results can be saved through the download button on the page for users to further in - depth analysis. Generally speaking, FPCAM provides convenient support for the efficient processing and analysis of single - cell RNA sequencing data by simplifying the data analysis workflow, providing visualization tools, and result download functions. Its design takes into account both the operational convenience of users and ensures the efficiency and scientific nature of data processing and analysis, and can provide a solid data analysis foundation for related research.
[0017] Advantages of the present invention
[0018] The present invention can achieve efficient and automated cell annotation through a cell annotation method based on scRNA-Seq data of pulmonary fibrosis, combined with the voting matrix of the Marker gene dictionary and the calculation of similarity scores, providing accurate cell classification results for the analysis of single-cell data of pulmonary fibrosis. This method significantly reduces the time and cost required for traditional manual cell annotation, while improving the stability and consistency of annotation. In addition, this method can further reveal the roles of different cell types in the process of pulmonary fibrosis, facilitating the study of disease mechanisms. The online analysis platform constructed based on the R Shiny framework not only provides dynamic and interactive cell annotation calculations, but also can visually present analysis results such as UMAP distribution, score heatmap, and Delta distribution, greatly enhancing the intuitiveness and operability of data analysis. The model and calculation framework of the present invention can be extended to other respiratory diseases and the integration and analysis of multi-omics data, providing important technical support and research tools for the precise diagnosis of pulmonary fibrosis and the development of new treatment strategies. The platform described in the present invention provides researchers with an efficient and interactive analysis tool, which can be used for in-depth analysis of single-cell data related to pulmonary fibrosis, providing data support for disease research and precision medicine.
[0019] The technical route of the analysis of the present invention mainly includes the following four parts:
[0020] (1) Construction of a dynamic dust-exposed mouse model and acquisition of scRNA-Seq data;
[0021] (2) Bioinformatics analysis such as preprocessing of scRNA-seq data;
[0022] (3) Development of a cell annotation algorithm for pulmonary fibrosis;
[0023] (4) Development of the FPCAM online web page and downstream bioinformatics analysis. Brief Description of the Drawings
[0024] Figure 1 . is the technical route map of the present invention.
[0025] Figure 2 . is the flowchart of scRNA-seq data analysis.
[0026] Figure 3 . is the schematic diagram of the development of the cell annotation algorithm.
[0027] Figure 4 . is the overall framework diagram of the FPCAM online data analysis platform. Detailed Description of the Invention
[0028] Next, the development process of this technology will be introduced in detail with reference to the accompanying drawings. It should be noted that this embodiment is based on the premise of this technical solution, and gives detailed implementation manners and specific operation processes, but the protection scope of the present invention is not limited to this embodiment.
[0029] Example 1 provides a cell annotation method based on single-cell RNA sequencing (scRNA-Seq) data of pulmonary fibrosis
[0030] First, comprehensively collect single-cell sequencing data related to pulmonary fibrosis, integrate known cell types and their marker genes (Marker genes) in different databases, and construct a cell type-specific dictionary based on these data. Subsequently, using computational biology methods, combined with the dictionary voting matrix and similarity score calculation model, an automated cell annotation method is established to achieve accurate identification and classification of different cell populations in the pulmonary fibrosis microenvironment. This method can effectively improve the accuracy and efficiency of single-cell data analysis, and provide important support for the research of the disease mechanism, accurate diagnosis and personalized treatment of pulmonary fibrosis.
[0031] As Figure 1 shown, the specific process of the method includes:
[0032] S1. Construction of a dynamic dust-exposed mouse model
[0033] Specific pathogen-free (SPF)-grade C57BL / 6 mice, male, 8-10 weeks old, weighing 20-25
[0034] g, were purchased from Beijing Speywood Biotechnology Co., Ltd. The animal breeding temperature was controlled at 23-26°C, the air humidity was controlled at 55±5%, the light source was turned on and off regularly to ensure a 12-hour day-night cycle. After one week of adaptive feeding, the mice were subjected to experiments. Based on the previous work, the dynamic dust-exposed method was used to make the mice inhale suspended dust particles naturally, and the experimental conditions such as the exposure concentration of micron-sized silica particles, the daily exposure time, and the total number of stress days were explored. The degree of pulmonary fibrosis lesions in the mice was determined by methods such as pathological morphological observation and detection of key index factors to provide a stable animal model for the experiment.
[0035] The mice in the experimental group were sacrificed by cervical dislocation at each corresponding time period. The abdomen was opened, lung tissues were taken, and rinsed with normal saline. The left lung lobe was separated and fixed in 4% (volume fraction, the same below) paraformaldehyde fixative for more than 24 h. The right lung lobe was taken and stored in a cryotube in liquid nitrogen. After taking out, it was laid flat in an embedding cassette, dehydrated through multiple steps of xylene and ethanol with different concentration gradients, embedded with liquid paraffin, cooled, and sectioned, and then placed on a drying table. After the sections were routinely dewaxed, staining was then started. The main operation steps of haematoxylin-eosin (HE) staining include: haematoxylin staining, differentiation with differentiating solution, eosin staining, and dehydration, clearing, and mounting. The main operation steps of Masson staining include:
[0036] Preparing iron haematoxylin staining solution, differentiating with acidic ethanol differentiating solution, blueing with Masson blueing solution, staining with ponceau fuchsin, staining with aniline blue, rapidly dehydrating with ethanol, putting into xylene, and sectioning. Enzyme-linked immunosorbent assay was used to detect the contents of inflammatory factors (TNF-α, TGF-β, IL-6, FGF) and hydroxyproline (HYP) in lung tissues.
[0037] S2. scRNA-seq detection of lung tissues of mice exposed to dust dynamically
[0038] The experimental design included two main groups: the control group and the dust-exposed group. The mice in the control group did not receive any form of treatment throughout the experiment, while the mice in the dust-exposed group were exposed to silica dust for 4 hours per day for a continuous exposure period of 128 days. All the collected lung tissue samples were numbered and grouped, and each experiment included three independent biological replicates to ensure the reliability and statistical significance of the experimental results. To further study the cell heterogeneity and gene expression profiles in these samples, all the samples were subsequently sent to BGI for second-generation single-cell sequencing, which can provide high-resolution gene expression data to deeply analyze the potential effects of silica dust exposure on the cell types and functions in the lungs of mice.
[0039] S3. Preprocessing of scRNA-seq data
[0040] The preprocessing analysis of scRNA-seq data included eight steps, as Figure 2 shown.
[0041] (1) Statistics of sequencing results
[0042] Perform basic statistical analysis on the raw sequence data obtained by high-throughput sequencing. According to the functions and sources of each part in the library construction process, it can be divided into three categories: Barcode sequences for identifying different cells, UMI sequences for counting gene expression levels, and insert fragment sequences (i.e., mRNA read sequences) for sequencing. The first step in single-cell RNA sequencing (scRNA-seq) data analysis is to perform quality control on the raw data. The raw data is usually stored in FASTQ format, which contains transcript information for each cell. By using tools such as FastQC, quality assessment can be performed on the FASTQ file to check the data quality, including read length, GC content, sequencing quality, etc. If problems are found in the data, such as low-quality reads or adapter contamination, tools such as Trimmomatic or Cutadapt can be used for trimming to remove low-quality reads and adapter sequences. In addition, cells with relatively low quality also need to be removed, usually by filtering out cells with too few genes or transcripts.
[0043] (2) Statistics of alignment results
[0044] After data quality control, the next step is to align the read sequences to the reference genome. The purpose of the alignment is to determine the positions of transcripts of each cell in the genome and provide a basis for subsequent quantification of gene expression levels. Commonly used alignment tools include STAR and HISAT2. The output after alignment is usually a BAM format file, which contains the positions and alignment information of each read. According to the QC standards, after basic filtering of the raw data, the STAR software is used to align the RNA reads to the reference genome and count the proportions of various regions on the genome that the reads are aligned to.
[0045] (3) Quality control and data filtering
[0046] Cell filtering parameters: (1) Filter out cells with less than 200 identified genes or more than 90% of the maximum number of genes; (2) Filter out the top 15% of cells in terms of the proportion of mitochondrial reads; (3) Correct the influence of the cell cycle.
[0047] (4) Library quality
[0048] Based on the alignment of reads to the reference genome and the statistics of specific UMIs, the expression level of each gene in each cell can be obtained; at the same time, based on the alignment position information, the proportion of reads in different regions of the genome can be obtained; during the bioinformatics process, cells are filtered, and the number of cells before and after filtering is also counted, and these indicators can all be used as auxiliary judgment bases for library quality.
[0049] (5) Gene expression matrix
[0050] It is necessary to extract the expression level of each gene in each cell from the BAM file, usually by calculating the number of reads of each gene in each cell. Tools such as featureCounts or HTSeq can automate this process and generate a gene expression matrix. Commonly used normalization methods include TPM (transcripts per million) and CPM (counts per million). In addition, methods such as ComBat and Harmony can effectively remove batch effects, thereby improving the comparability between different experimental data.
[0051] (6) Single-sample cell clustering
[0052] Based on the results of dimensionality reduction, cluster analysis is performed on the cells. Commonly used clustering methods include PCA, tSNE, UMAP, and UAMP is a currently popular method. The clustering results are as follows. Each point represents a cell, and points with close spatial distances indicate that the gene expression patterns of these cells are relatively similar; the color is determined by the classification obtained from the graph-based clustering algorithm. The algorithm believes that the cells within the same cluster have the closest expression patterns, and the same color represents the same type of cells. Different algorithms will present different clustering effects, and UMAP is used for clustering by default.
[0053] (7) Single-sample Marker gene identification
[0054] By comparing the gene expression of the cells in each cluster with all the cells in the remaining other clusters, we can identify the Marker genes unique to the cluster. Usually, we use t-test or limma.
[0055] (8) Single-sample cell type annotation
[0056] In the actual process of scRNA-seq data processing, cell type annotation is a key step in subsequent analysis and research. Currently, commonly used cell annotation methods include SingleR, scMatch, CellAssign, and Garnett, etc. We are committed to developing a more accurate cell annotation method, FPCAM, and comparing it with existing annotation methods (such as SCSA, scCATCH, SingleR, and SciBet, etc.).
[0057] S4. Construction of the pulmonary fibrosis annotation dictionary
[0058] Review the existing scientific literature, collect the information of known marker genes, and extract relevant gene expression data from public databases (such as GEO, EGA, SRA, etc.). Our custom database comprehensively collates the data from six literatures by Sikkema, Madissoon, Sun, etc., and sorts and corrects the marker genes according to the primary and secondary order of the literature materials (see Table 1). The selected marker genes are classified according to four levels: the function of the cell, the cell line to which it belongs, the cell type, and the cell subtype. A total of 879 marker genes (without removing duplicates) are obtained and divided into 79 cell subtypes. Finally, a voting mechanism is used to select the top 10 marker genes as the first dataset; then, the marker genes are sorted according to the frequency of occurrence in all databases, and the marker genes that appear twice or more are selected to form the second dataset. The two databases will be used as custom dictionaries for annotating cell types in subsequent analyses. The first dataset contains 628 marker genes, and the second dataset contains 118 marker genes. The score of each marker gene is weighted according to its frequency of occurrence in the database. In the subsequent work of cell annotation and data analysis, these two databases are used to annotate and label cell types. The genes in the first database are mainly used for annotating key cell types, while the genes in the second database are used as auxiliary information to further improve the accuracy of annotation.
[0059] Table 1. Construction of the dictionary based on the primary and secondary order of the literature materials
[0060]
[0061]
[0062] S5. Establishment of the similarity score matrix and determination of the optimal λ threshold
[0063] This part mainly includes constructing the dictionary of lung fibrosis cell-gene associations, constructing the weighted scoring matrix, and introducing the optimal λ threshold (see Figure 3 ).
[0064] First, an expression matrix is constructed through scRNA-seq analysis. The abscissa is the cell cluster, and the ordinate is the Marker gene. Each element in the matrix is the expression value of different Marker genes in different cell clusters. The number of Marker genes i corresponding to each cell cluster is different (see Table 2), and the calculation formula is shown in Equation (1).
[0065] Table 2. Expression matrix
[0066] Gene\Cluster Cluster 1 Cluster 2 Cluster 3 Cluster j Gene 1 Gene 2 …… Gene i
[0067]
[0068] Secondly, construct a dictionary voting matrix. The abscissa is the Marker gene, and the ordinate is the cell type. Each element in the matrix is the voting number of each Marker gene in different cell types. Different cell types have different numbers of Marker genes n (see Table 3), and the calculation formula is shown in Equation (2).
[0069] Table 3. Dictionary Voting Matrix
[0070]
[0071]
[0072] Then, calculate the ratio of each gene in each cell cluster in the expression matrix, that is, the expression value of each gene divided by the sum of the expression values of all genes in its cell cluster. The calculation formula is shown in Equation (3).
[0073]
[0074] Where: The relative expression ratio of gene i in cluster j. When calculating, take A ij (The expression value of a certain gene i in cluster j) and divide it by the sum of the expression values of this gene in all clusters (sum over all j). Compare the contributions of different genes in the same cluster.
[0075] Calculate the weight of each Marker gene in each cell type, that is, the voting number of each Marker gene in each cell type divided by the sum of the voting numbers of all genes in each cell type. The calculation formula is shown in Equation (4).
[0076]
[0077] Where, this formula is for the weight of each marker gene, V m is the "vote number" or weight value of each marker gene in cell type m (obtained from the percent column of the dictionary). V mn represents the expression weight of the marker gene in each cluster. M mn is the marker gene weight of cell type m in cluster n. The output weight M mn represents the contribution size of each cell type in a certain cluster. Marker genes with high weights can better represent the corresponding cell types and contribute to the accuracy of cell annotation.
[0078] After that, multiply each ratio obtained from the expression matrix by each weight calculated from the dictionary voting matrix to obtain the similarity score matrix. The calculation formula is shown in Equation (5):
[0079] Sim=∑ jn Eij×Mmn(5)
[0080] This is the similarity score calculated based on marker genes and gene expression ratios. Multiply the expression E of each gene ij by the weight M of the marker gene mn and sum over all relevant genes to obtain the similarity score of the cluster to a specific cell type. These scores help identify the degree of association of each cluster with a known cell type. By analyzing the similarity matrix, the most likely cell type for each cluster can be determined. This is very important in cell annotation for identifying the biological characteristics of different cell populations.
[0081] First, for each cell cluster i, find the maximum value of the similarity scores among all cell types m. The calculation formula is shown in Equation (6):
[0082]
[0083] Determine the first threshold λ by calculating the difference between the maximum value Smax max of the similarity scores of all cell clusters and the median S median . The calculation formula is shown in Equation (7): 1, The calculation formula is shown in Equation (7):
[0084] λ1 = Smax - Smedian (7)
[0085] If the difference between S max and the second-largest similarity score S max - S median is greater than 0.00001, there is a second threshold λ2. The calculation formula is shown in Equation (8):
[0086] λ2 = Ssec-max - Smedian (8)
[0087] Finally, the cells are annotated according to the cell type corresponding to the λ1 value. If there is a second-highest value λ2, an additional cell type corresponding to the second-highest value λ2 will be added to the original annotation. Note that the λ value may be 0 or negative, in which case the λ value is meaningless.
[0088] S6. FPCAM database construction:
[0089] The data management system - FPCAM based on the R shiny and Python Django frameworks mainly collects, organizes, and stores the original sequencing data of transcriptomics of human and mouse lung fibrosis tissues and various data after bioinformatics analysis. Users can submit transcriptome data, perform single-cell annotation of lung fibrosis, query differentially expressed genes, Marker genes, and cell heterogeneity through the human-computer interaction interface, providing guidance and assistance for the screening of early biomarkers for lung fibrosis diseases. The present invention intends to adopt the B / S (Browser / Server) website framework development mode, that is, the mode of combining the browser with the database. Based on the concept of separating the three terminals, the security, stability, and migratability of the database are maximized (see Figure 4 ).
[0090] (1) Front-end construction
[0091] In the design of the web page front-end, HTML, CSS, and JS are the cornerstones for enriching page elements and enhancing page interaction. This database intends to use a large number of mature HTML architectures and CSS styles, and use a flat design to make the interface more coordinated and beautiful. We intend to use the third-generation VUE as the framework for the front-end application, and use Echarts and HighCharts to highly optimize the rendering efficiency and interaction system, improving the operability and stability while ensuring the display of the database content. We intend to use SVG and Canvas elements to further enrich the visualization means of data charts, and use Ajax to asynchronously and locally obtain data, eliminating the time consumed by the overall page refresh.
[0092] (2) Back-end construction
[0093] The back-end development of the database intends to be built on the Shiny and Django frameworks, with R and Python languages as the analysis and calculation scripts, and connect the database and the server through pymysql. Deeply customize the UI style and response events of Jbrowse and use it as the best means for the visualization of multi-omics data.
[0094] (3) Database construction
[0095] We will use Linux as the development environment for the analysis platform, use MySQL to build the database, and use MongoDB as a supplement. In addition, we will use multiple-layer verification measures to ensure data security, avoid data leakage and damage, and use Apache2 to deploy the back-end server to ensure the high-concurrency use of the database.
[0096] S7. Annotation results
[0097] In the process of data analysis, it is first necessary to construct a dictionary file to integrate the vote counts, total vote counts, and voting ratios of the top 10 marker genes in basal resting cells. This file provides fundamental support for subsequent data processing and visualization. For example, in Table 4, the vote count for the TP63 gene is 4, the total vote count is 21, and the voting ratio is 4 / 21.
[0098] Table 4. Construction of the dictionary file
[0099]
[0100] As shown in Table 5, we can use Formula 1 to calculate the marker gene expression data for cell clusters 0 and 1.
[0101] Table 5. Construction of the marker gene expression matrix for cell clusters
[0102]
[0103]
[0104] The similarity score matrices for clusters 0 to 5 are calculated using Formulas (2, 3, 4, 5). According to Formula 6, the maximum similarity score for cluster 1 is 0.00119, while the maximum similarity score for cluster 4 is 0.000532. These scores reflect the similarity degrees between different cell clusters and provide important references for subsequent clustering analysis.
[0105] Table 6. Construction of the similarity score matrix
[0106]
[0107] λ1 and λ2 are calculated using Formulas (7, 8), and the final cell annotation results are generated based on these values. Subsequently, we compare and analyze the annotation results of our model with those of other models (SCSA, SingleR, SciBet). As shown in Table 7, we list the accuracies of all models and conduct a comparison to evaluate the performance of different methods for lung fibrosis cell annotation.
[0108] Table 7. Annotation results of each model
[0109]
[0110]
[0111] By comparing the cell annotation results of different models, FPCAM achieved an accuracy of 85.7%, outperforming other models. The accuracy of the SCSA model was 82.1%, and both the SingleR model based on the ImmGenData file and the SciBet model based on the 20mouse_organs data were 78.6%. In addition, the SingleR annotation model based on the MouseRNAseqData data had uncertainties in the identification of some cell types, possibly due to the heterogeneity of the dataset or insufficient specific marker genes, resulting in relatively low confidence in the classification of some cells. Overall, FPCAM has high accuracy and stability and has advantages in cell annotation tasks.
Claims
1. A cell annotation method based on pulmonary fibrosis scRNA-Seq data, characterized in that: Follow these steps to achieve this: S1. Data source: scRNA-Seq data related to lung fibrosis; S2. Data preprocessing: Data preprocessing, including removing low-quality data and removing batch effects; S3. Construction of an annotation dictionary for pulmonary fibrosis: sorting and correcting marker genes according to the importance of literature, Marker genes were screened, including duplicates, classified into four levels: cell function, cell line, cell type, and cell subtype, and the annotation dictionary file related to pulmonary fibrosis was compiled; S4. Develop algorithms: Build a set of lung fibrosis-related cell annotation processes based on scRNA-seq data, optimize the algorithm through similarity matrix, manually curated lung fibrosis-related cell-gene association dictionary, and evaluation index; S5.FPCAM database construction: Build an online analysis platform - FPCAM based on R shiny.
2. The method according to claim 1, characterized in that S1. Data source: Construction of an active dust-exposed mouse model to obtain scRNA-Seq data related to lung fibrosis; scRNA-Seq data related to lung fibrosis; S2. Data preprocessing: Obtaining cluster-related gene expression data: The first step is to perform quality control on the scRNA-seq data obtained in step S1, remove low-quality reads and adapter sequences, and remove cells with low quality; The second step is to align the read sequence to the reference genome. After the alignment is completed, the expression of each gene in each cell is extracted by counting the number of reads of each gene in each cell and normalizing it. In addition, the batch effect is removed. Through the above steps, the gene expression matrix is finally generated, and the expression information of each cell is normalized and converted into a format that can be used for high-level analysis. S3. Construction of Pulmonary Fibrosis Annotation Dictionary: Relevant gene expression data were extracted from public databases, and marker genes were sorted and corrected according to the importance of the literature in a custom database. The 879 marker genes screened out, including duplicates, were classified into four levels: cell function, cell line, cell type, and cell subtype, covering 79 cell subtypes. Finally, the top 10 marker genes were selected through voting to construct a dictionary database. Based on the frequency ranking of marker genes, a weighted score was calculated for each gene, and this database was used as a custom database for subsequent cell type annotation. The database ultimately contained 631 marker genes. S4. Development of algorithms: First, an expression matrix was constructed through scRNA-seq analysis. The horizontal axis is the cell cluster and the vertical axis is the marker gene. Each element in the matrix is the expression value of different marker genes in different cell clusters. The number of marker genes i corresponding to each cell cluster is different. The calculation formula is shown in formula (1): Second, construct a dictionary voting matrix, with the horizontal axis representing the marker gene and the vertical axis representing the cell category. Each element in the matrix represents the number of votes for each marker gene in different cell categories. Different cell categories have different numbers of marker genes n. The calculation formula is shown in formula (2): Third, calculate the ratio of each gene in each cell cluster in the expression matrix, that is, the expression value of each gene divided by the sum of the expression values of all genes in its cell cluster. The calculation formula is shown in formula (3): Fourth, calculate the weight of each marker gene in each cell type, that is, the number of votes for each marker gene in each cell type divided by the sum of the number of votes for all genes in each cell type. The calculation formula is shown in formula (4); Fifth, multiply each ratio obtained in the expression matrix by each weight calculated in the dictionary voting matrix to obtain the similarity score matrix. The calculation formula is shown in formula (5): Try = ∑ jn Eij×Mnn (5) Sixth, for each cell cluster i, find the value with the largest similarity score among all cell types m. The calculation formula is shown in formula (6): Seventh, by calculating the maximum value S of the similarity scores of all cell clusters max With the median S median The difference between the two determines the first threshold λ 1, The calculation formula is shown in formula (7): λ1=S max - S median (7) 8. If S max The similarity score S with the second largest max -S median If the difference between them is greater than 0.00001, there is a second threshold λ2, and the calculation formula is shown in formula (8): λ2=S sec-max-S median (8) Ninth, cells will be annotated according to the cell type corresponding to the λ1 value. If there is a second highest value λ2, another cell type corresponding to the second highest value λ2 will be added to the original one for annotation. It should be noted that the λ value may be 0 or negative, in which case the λ value is meaningless. S5.FPCAM database construction: an online analysis platform for single-cell annotation of pulmonary fibrosis based on the RShiny framework; the platform integrates a variety of visualization and data analysis functions, including: Cell annotation calculation: Based on the above voting matrix and similarity score matrix, efficient cell type identification is achieved; UMAP dimensionality reduction visualization: showing the spatial distribution of different cell types; Score heat map: intuitively display the marker gene scores of each cell type; Delta distribution analysis: further analyze the score differences between different cell types and improve the biological explanatory power of annotations.