A biological sample information processing method, system and storage medium
By performing whole genome analysis and feature analysis on biological samples, combined with non-negative least squares regression, the problem of inaccurate cell correlation judgment in existing technologies is solved, the precise positioning of cell lineage stages and states is achieved, and the accuracy of sample comparison is improved.
Patent Information
- Application Number
- CN202310414041.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-18
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2043-04-18
AI Technical Summary
Existing technologies make it difficult to accurately determine the correlation between cells at different stages or states of the same lineage, especially when using single specific markers or global transcriptome correlation coefficients, and lack precision when comparing samples between different animal models or species.
By obtaining transcriptome information of the test and reference biological samples, global analysis of the entire genome, local analysis of differentially expressed genes, and non-negative least squares regression analysis were performed. Combined with feature analysis and biological identification, the sva and limma packages were used to process the data, batch effect elimination and cluster analysis were performed to determine the cell lineage stage and status.
It achieves precise positioning of specific cell types and subpopulations in cells of different lineages, providing more accurate cell similarity judgment and sample comparison results.
Smart Images

Figure CN116434842B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of biological sample sequencing, and in particular to a method for processing biological sample information, a system for processing biological sample information, and a computer-readable storage medium. Background Art
[0002] With the development of sequencing technology, people have found that subpopulations often exist in cells at different stages / states of the same lineage, or even in cells that are exactly the same. However, it is often difficult to determine the correlation (i.e., similarity and difference) between the target cells to be tested and the known reference cells based on only one analysis or a single specific marker. At the same time, for tissue-level sequencing, when we want to compare the correlation between the same or similar samples from different animal models and different species, it is not accurate enough to use only correlation coefficients such as Pearson and Spearman that are beneficial to the global transcriptome.
[0003] To overcome the above-mentioned defects of the existing technology, the field urgently needs a method for comparing the similarities between biological samples such as cells and tissues to determine the cell lineage stage and state, thereby providing a reference for the specific positioning of specific cell types and subpopulations in different lineage cells. Summary of the Invention
[0004] The following is a brief summary of one or more aspects to provide a basic understanding of these aspects. This summary is not an exhaustive overview of all conceivable aspects and is neither intended to identify key or critical elements of all aspects nor to define the scope of any or all aspects. Its sole purpose is to present some concepts of one or more aspects in a simplified form as a prelude to the more detailed description that will be provided later.
[0005] To overcome the above-mentioned defects of the prior art, the present invention provides a method for processing biological sample information, a system for processing biological sample information, and a computer-readable storage medium, which can determine the cell lineage stage and state based on high-throughput sequencing data, thereby providing a reference for the specific location of specific cell types and subpopulations in different lineage cells.
[0006] Specifically, the above-mentioned biological sample information processing method provided according to the first aspect of the present invention includes the following steps: obtaining transcriptome information of the biological sample to be tested and multiple known reference biological samples; based on the transcriptome information, performing global analysis based on whole genes, local analysis based on differential genes, feature analysis based on gene differential expression analysis, and non-negative least squares regression analysis on the biological sample to be tested and the multiple reference biological samples; and determining the overall characteristics of the biological sample to be tested based on the global analysis results, local analysis results, feature analysis results, and non-negative least squares regression analysis results.
[0007] Preferably, in one embodiment of the present invention, before obtaining the transcriptome information of the biological sample to be tested and the multiple reference biological samples, the processing method further includes the following steps: performing biological identification on the biological sample to be tested based on known specific markers, functional characteristics and / or morphological characteristics; and determining a reference biological sample similar to the biological sample to be tested based on the results of the biological identification.
[0008] Preferably, in one embodiment of the present invention, the biological identification based on the specific marker includes at least one of qRT-PCR detection, Western Blot detection, and immunofluorescence detection, and / or the biological identification based on the functional characteristics includes a functional comparison experiment between the biological sample to be tested and a known reference biological sample.
[0009] Preferably, in one embodiment of the present invention, the step of obtaining transcriptome information of the biological sample to be tested and multiple known reference biological samples includes: in response to the result of the biological identification being unable to determine a reference biological sample similar to the biological sample to be tested, performing high-throughput sequencing on the biological sample to be tested and multiple reference biological samples to obtain their transcriptome information respectively.
[0010] Preferably, in one embodiment of the present invention, before performing the global analysis, the local analysis, the feature analysis and the non-negative least squares regression analysis, the processing method further includes the following steps: organizing the transcriptome information from multiple different sources into standardized data in a unified format; and using the combat function of the sva package or the removeBatchEffect function of the limma package to eliminate batch effects on the standardized data.
[0011] Preferably, in one embodiment of the present invention, the step of performing the global analysis includes: based on the transcriptome information, performing correlation analysis on the biological sample to be tested and multiple reference biological samples to obtain a correlation heat map that clusters the biological sample to be tested and multiple reference biological samples that are globally similar to it; based on the transcriptome information, performing clustering and expression heat map analysis on all genes detected in the biological sample to be tested and multiple reference biological samples to obtain an expression heat map that clusters the biological sample to be tested and multiple reference biological samples that are globally similar to it; and determining the global analysis result based on the correlation heat map and the expression heat map.
[0012] Preferably, in one embodiment of the present invention, the step of performing the local analysis includes: based on the transcriptome information, extracting multiple TOP SD genes with the largest standard deviation (SD) from the biological sample to be tested and multiple reference biological samples, and performing local cluster analysis to obtain a TOP SD gene heat map that clusters the biological sample to be tested and multiple reference biological samples that are locally similar to it; performing pairwise differential expression analysis on the transcriptome information of the biological sample to be tested and multiple reference biological samples, and displaying the differential expression analysis results using a volcano plot; and determining the local analysis results based on the TOP SD gene heat map and the volcano plot.
[0013] Preferably, in one embodiment of the present invention, the step of performing pairwise differential expression analysis on the transcriptome information of the biological sample to be tested and the multiple reference biological samples includes: using the limma package to perform pairwise differential expression analysis on the transcriptome information of the biological sample to be tested and the multiple reference biological samples to respectively determine the differential genes of the biological sample to be tested and each of the reference biological samples, wherein a smaller number of differential genes indicates a higher similarity of the biological samples.
[0014] Preferably, in one embodiment of the present invention, the step of performing the feature analysis includes: performing clustering and heat map analysis on the differential genes of the biological sample to be tested and the multiple reference biological samples to obtain a cluster tree heat map that clusters the biological sample to be tested and the multiple reference biological samples with similar features, wherein the differential genes indicate the positioning of the biological sample to be tested in the lineage development; performing principal component analysis on the cluster tree heat map to determine the lineage stage of the biological sample to be tested in the lineage of the reference biological samples with similar features; performing differential expression analysis and full-gene differential fold scatter plot correlation analysis on the whole genes to form a A mature reference biological sample is used as a control to obtain the difference fold value of the whole gene to simulate the biological distance between the biological sample to be tested and the mature reference biological sample, and to avoid batch effects between the reference biological samples; differential expression analysis and difference fold scatter plot correlation analysis of the differential genes are performed on the whole gene to determine the impact of common differentially expressed genes on sample similarity; correlation analysis is performed on the characteristic gene set that has common characteristics with the characteristics of the biological sample to be tested; and the characteristic analysis result is determined based on the cluster tree heat map, the pedigree stage, the biological distance, the impact and the correlation analysis result.
[0015] Preferably, in one embodiment of the present invention, the step of performing the non-negative least squares regression analysis comprises: using the gene expression R of all reference samples (Reference) in the second data set B through the non-negative least squares regression analysis. B Predict the gene expression Q of the target (Query) sample type in the first dataset A A (Query), where Q A =β 0A +β 1A (R B ); Switch the order of the first data set A and the second data set B, that is, data set A is the Reference sample, and data set B is the Query sample, through the gene expression R of all cell types in the first data set A A , predict the gene expression Q of the target sample in the second dataset B B , where Q B =β 0B +β 1B (R A );According to the gene expression Q A , the gene expression R B , the gene expression R A And the gene expression Q B , the coefficient of determination β=2(β AB +0.01)(β BA +0.01), β0A It means when R B The value of the β coefficient when the expression level is 0, that is, the intersection of the regression line predicted by NNLS and the Y axis; β 1A refers to the β coefficient obtained after NNLS analysis; β 0B It means when R A The value of the β coefficient when the expression level is 0, that is, the intersection of the regression line predicted by NNLS and the Y axis; β 1B refers to the β coefficient obtained after NNLS analysis; β AB It refers to the β coefficient value when predicting the target sample A based on the reference sample B. BA It refers to the β coefficient value when predicting the target sample B based on the reference sample A, where a higher coefficient β indicates a more similar sample type; and based on the coefficient β, the non-negative least squares regression analysis result is determined.
[0016] Preferably, in one embodiment of the present invention, the biological sample to be tested and the reference biological sample are cell samples. After determining the overall characteristics of the biological sample to be tested, the processing method further includes the following steps: identifying the cell type of the cell sample to be tested and / or the specific positioning of the cell sample to be tested in its cell lineage based on the overall characteristics of the cell sample to be tested.
[0017] Preferably, in one embodiment of the present invention, the biological sample to be tested and the reference biological sample are tissue samples. After determining the overall characteristics of the biological sample to be tested, the processing method further includes the following steps: based on the overall characteristics of the tissue sample to be tested, identifying experimental-related treatment methods (drug treatment, gene knockout, viral transfection, etc.) and phenotypic changes at the pathological and histological levels (such as aging, apoptosis, endoplasmic reticulum stress, autophagy phenotype, etc.) in the tissue sample to be tested, and / or at least one pathological change and experimental treatment method in the tissue sample to be tested that is similar to that of the reference biological sample.
[0018] Furthermore, the biological sample information processing system provided in accordance with the second aspect of the present invention includes a memory and a processor. The memory stores computer instructions. The processor is connected to the memory and configured to execute the computer instructions stored in the memory to implement the biological sample information processing method provided in any of the above embodiments.
[0019] Furthermore, the computer-readable storage medium provided according to the third aspect of the present invention stores computer instructions, which, when executed by a processor, implement the biological sample information processing method provided in any one of the above embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The above features and advantages of the present invention will be better understood after reading the detailed description of the embodiments of the present disclosure in conjunction with the following drawings. In the drawings, the components are not necessarily drawn to scale, and components with similar related properties or characteristics may have the same or similar reference numerals.
[0021] Figure 1 A flowchart of a method for processing biological sample information provided according to some embodiments of the present invention is shown.
[0022] Figures 2A to 2G A schematic diagram showing the cell morphological characteristics and specific marker expression of murine biliary tree stem cells (mBTSCs) provided according to some embodiments of the present invention is shown.
[0023] Figure 3 A correlation heat map is shown in which a biological sample to be tested is clustered together with multiple reference biological samples for global comparison therewith according to some embodiments of the present invention.
[0024] Figure 4 A schematic diagram of cluster analysis of an expression heat map of a biological sample to be tested and a plurality of reference biological samples for global comparison thereof is shown according to some embodiments of the present invention.
[0025] Figures 5A-5B A TOP SD gene heat map is shown in which a test biological sample provided according to some embodiments of the present invention and multiple reference biological samples for local comparison are clustered together.
[0026] Figures 6A to 6E A volcano plot is shown for differential expression analysis between transcriptome information of a test biological sample and multiple reference biological samples according to some embodiments of the present invention.
[0027] Figures 7A-7B A cluster tree heat map is shown in which a plurality of reference biological samples are clustered together to compare characteristics of a biological sample to be tested with provided in some embodiments of the present invention.
[0028] Figures 8A to 8D A schematic diagram of the lineage stages of a biological sample to be tested in the lineage of a reference biological sample with which its characteristics are compared is shown according to some embodiments of the present invention.
[0029] Figures 9A to 9G A scatter plot of the fold differences of all genes provided according to some embodiments of the present invention is shown.
[0030] Figures 10A to 10C A schematic diagram of differential expression analysis of all genes and correlation analysis of differential fold scatter plots of differential genes provided according to some embodiments of the present invention is shown.
[0031] Figure 11 A regression diagram of non-negative least squares regression analysis provided according to some embodiments of the present invention is shown.
[0032] Figures 12A-12B Shown are a correlation heat map and a correlation analysis diagram of cell lineages provided according to some embodiments of the present invention.
[0033] Figure 13 A flow chart illustrating information processing of tissue samples according to some embodiments of the present invention is shown.
[0034] Figure 14 A correlation heat map is shown in which a biological sample to be tested is clustered together with multiple reference biological samples for global comparison therewith according to some embodiments of the present invention.
[0035] Figure 15 A schematic diagram of cluster analysis of an expression heat map of a biological sample to be tested and a plurality of reference biological samples for global comparison thereof is shown according to some embodiments of the present invention.
[0036] Figures 16A to 16C A TOP SD gene heat map is shown in which a test biological sample provided according to some embodiments of the present invention and multiple reference biological samples for local comparison are clustered together.
[0037] Figures 17A to 17F A volcano plot is shown for differential expression analysis between transcriptome information of a test biological sample and multiple reference biological samples according to some embodiments of the present invention.
[0038] Figure 18 A cluster tree heat map is shown in which a plurality of reference biological samples are clustered together to compare characteristics of a biological sample to be tested with provided in some embodiments of the present invention.
[0039] Figure 19 A principal component analysis diagram provided according to some embodiments of the present invention is shown.
[0040] Figures 20A to 20D A scatter plot of the fold differences of all genes provided according to some embodiments of the present invention is shown.
[0041] Figures 21A to 21D A schematic diagram of differential expression analysis of all genes and correlation analysis of differential fold scatter plots of differential genes provided according to some embodiments of the present invention is shown. DETAILED DESCRIPTION
[0042] The following specific embodiments illustrate the embodiments of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. Although the description of the present invention will be introduced in conjunction with the preferred embodiment, this does not mean that the features of this invention are limited to this embodiment. On the contrary, the purpose of introducing the invention in conjunction with the embodiment is to cover other options or modifications that may be extended based on the claims of the present invention. In order to provide a deep understanding of the present invention, the following description will include many specific details. The present invention can also be implemented without using these details. In addition, in order to avoid confusion or blurring the focus of the present invention, some specific details will be omitted in the description.
[0043] In the description of the present invention, it should be noted that, unless otherwise expressly specified or limited, the terms "mounted," "connected," and "connected" should be understood in a broad sense. For example, they may refer to fixed, detachable, or integral connections; mechanical or electrical connections; direct or indirect connections through an intermediate medium; and internal communication between two components. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on the specific circumstances.
[0044] Furthermore, the terms "upper," "lower," "left," "right," "top," "bottom," "horizontal," and "vertical" used in the following description should be understood to refer to the orientations depicted in that section and the accompanying drawings. These relative terms are used solely for convenience of description and do not necessarily imply that the devices described herein must be manufactured or operated in a specific orientation. Therefore, they should not be construed as limiting the present invention.
[0045] It will be understood that although the terms "first," "second," "third," etc. may be used herein to describe various components, regions, layers, and / or portions, these components, regions, layers, and / or portions should not be limited by these terms, and these terms are merely used to distinguish different components, regions, layers, and / or portions. Thus, a first component, region, layer, and / or portion discussed below may be referred to as a second component, region, layer, and / or portion without departing from some embodiments of the present invention.
[0046] As mentioned above, existing technologies often struggle to determine the correlation (i.e., similarity and difference) between target cells and known reference cells using only a single analysis or specific marker. Furthermore, for tissue-level sequencing, when comparing the correlation between identical or similar samples from different animal models or species, using only correlation coefficients like Pearson and Spearman, which are beneficial for the global transcriptome, is not sufficiently accurate.
[0047] To overcome the above-mentioned defects of the prior art, the present invention provides a method for processing biological sample information, a system for processing biological sample information, and a computer-readable storage medium, which can determine the cell lineage stage and state based on high-throughput sequencing data, thereby providing a reference for the specific location of specific cell types and subpopulations in different lineage cells.
[0048] In some non-limiting embodiments, the biological sample information processing method provided in the first aspect of the present invention can be implemented via the biological sample information processing system provided in the second aspect of the present invention. Specifically, the biological sample information processing system can be configured with a memory and a processor. The memory includes, but is not limited to, the computer-readable storage medium provided in the third aspect of the present invention, which stores computer instructions. The processor is connected to the memory and configured to execute the computer instructions stored in the memory to implement the biological sample information processing method provided in the first aspect of the present invention.
[0049] The following describes the operating principles of the aforementioned biological sample information processing system in conjunction with several examples of biological sample information processing methods. Those skilled in the art will appreciate that these examples of biological sample information processing methods are merely non-limiting implementations of the present invention, intended to clearly illustrate the main concepts of the present invention and provide specific solutions that facilitate implementation by the public. They are not intended to limit the full functionality or operating methods of the biological sample information processing system. Similarly, the biological sample information processing system is merely a non-limiting implementation of the present invention and does not limit the execution entities or execution order of the steps in these biological sample information processing methods.
[0050] Please refer to Figure 1 , Figure 1 A flowchart of a method for processing biological sample information provided according to some embodiments of the present invention is shown.
[0051] like Figure 1 As shown, the present invention can use mouse biliary tree stem cells (mBTSCs) as the biological sample to be tested, and human biliary tree stem cells (hBTSCs), human hepatic stem cells (hHpSCs), human hepatic progenitor cells (hHBs), and human mature hepatocytes (hAHeps) as reference biological samples.
[0052] Biliary tree stem cells (BTSCs) are multipotent stem cells residing in the peribiliary glands of bile duct tissue. They have been shown to express markers of hepatic, biliary, and pancreatic stem and progenitor cells and have the potential to differentiate into functional hepatocytes, biliary epithelial cells, and pancreatic islet cells. Research on human biliary tree stem cells (hBTSCs) is well established. BSCs have been isolated and purified from biliary tissue using a serum-free culture system. Gene expression profiles compared in vitro with those of hepatic stem cells, hepatic progenitor cells, and mature hepatocytes have revealed that the human biliary tree exhibits a lineage-like maturation process, progressing through hepatic stem cells (hHpSCs), hepatic progenitor cells (hHBs), and mature hepatocytes (hAHeps). However, this lineage-like maturation process remains to be determined in mammals. Due to the similarities in organ size and metabolic profile between humans and mice, mice are often used as small animal models for translational and preclinical research. Human and mouse stem / progenitor cell networks in other systems, such as hematopoietic stem / progenitor cell networks, are highly similar. By establishing a method for parallel comparison of cell lineages across species, we can help define unidentified cell lineages in other species (e.g., commonly used small and large animal models such as mice or pigs). From a gross anatomical perspective, the mouse biliary system is more complex than that of humans. However, the mouse's major branches, such as the hepatopancreatic bile duct, common bile duct, and main pancreatic duct, are highly similar to those of the human biliary system. Therefore, using human biliary tree stem cell isolation and culture methods, we can also obtain mouse biliary tree stem cells. However, after obtaining mouse biliary tree stem cells (mBTSCs) from adjacent sites, how to evaluate the similarity between mouse biliary tree stem cells (mBTSCs) and human biliary tree stem cells (hBTSCs) remains an unresolved issue in this field.
[0053] During the processing of biological sample information, the present invention can first perform biological identification of murine biliary tree stem cells (mBTSCs) based on known specific markers, functional characteristics, and morphological characteristics, and then determine reference biological samples similar to murine biliary tree stem cells (mBTSCs) based on the results of the biological identification.
[0054] Please refer to further Figures 2A to 2G , Figures 2A to 2G A schematic diagram showing the cell morphological characteristics and specific marker expression of murine biliary tree stem cells (mBTSCs) provided according to some embodiments of the present invention is shown.
[0055] like Figures 2A to 2GAs shown, biliary tree stem cells (BTSCs) highly express EpCAM, PDX1, and SOX17 and have the potential to differentiate into hepatic, biliary, and pancreatic tracts. This means that BTSCs can be easily distinguished from mesenchymal cells and immune cells.
[0056] Furthermore, in this embodiment, the biological sample to be tested and the reference biological sample are cells of the same lineage and type, and the results of the above biological identification still cannot give satisfactory results. High-throughput sequencing (RNA-seq) is required for the biological sample to be tested (target cell) and the reference biological sample (known cell). It is best to have more than three repeated high-throughput sequencing (RNA-seq) samples for each cell to avoid intra-group errors.
[0057] Before analyzing mouse biliary tree stem cells (mBTSCs) and multiple reference biological samples, the acquired transcriptome information needed to be standardized into a unified format. The biological sample information processing system formatted high-throughput sequencing (RNA-seq) data into FPKM or TPM formats, and then further standardized data into log2(FPKM+1) or log2(TPM+1) formats. Batch effects were then removed from the standardized data using the combat function in the sva package or the removeBatchEffect function in the limma package.
[0058] Afterwards, a global analysis of the test biological sample and multiple reference biological samples is initiated. The present invention can first perform a correlation analysis on the test biological sample and multiple reference biological samples based on transcriptome information to obtain a correlation heatmap that clusters the test biological sample and multiple globally similar reference biological samples. Optionally, the correlation analysis method includes, but is not limited to, existing correlation analysis methods such as Pearson, Kendall, and Spearman.
[0059] Please refer to Figure 3 , Figure 3 A correlation heat map is shown in which a biological sample to be tested is clustered together with multiple reference biological samples for global comparison therewith according to some embodiments of the present invention.
[0060] like Figure 3 As shown, various cells in biological samples tend to cluster together, and similar cells also cluster together. Figure 3Shown is a correlation heatmap of murine biliary tree stem cells (mBTSCs, three samples), murine mature hepatocytes (mAHeps, three samples), and four human cell subsets (hBTSCs, hHpSCs, hHBs, and hAHeps, three samples each). High-throughput sequencing (RNA-seq) data from murine biliary tree stem cells and murine mature hepatocytes were combined with high-throughput sequencing (RNA-seq) data from the four human cell subsets. Correlation analysis was then performed on all five cell types to determine the lineage stages of the murine cell subsets. As shown in the figure, the correlation heatmap of the six cell types reveals two major clusters: mBTSCs, hBTSCs, and hHpSCs, while mAHeps, hHBs, and hAHeps cluster together. This suggests that mBTSCs are similar to hBTSCs and hHpSCs, but it is impossible to further distinguish which of mBTSCs and hBTSCs is more similar to each other. At the same time, mAHeps and hHBs are similar to hAHeps, but it is impossible to further distinguish which of mAHeps and hAHeps is more similar to each other than hHBs and hAHeps. Therefore, further similarity analysis is needed between the test biological samples and multiple reference biological samples.
[0061] At this point, the present invention can further perform clustering and expression heat map analysis on all genes detected in the test biological sample and multiple reference biological samples based on transcriptome information, to obtain an expression heat map that clusters the test biological sample and multiple globally similar reference biological samples. Optionally, clustering methods include but are not limited to existing clustering methods such as ward.D, ward.D2, single, complete, average, mcquitty, median, and centroid.
[0062] Please refer to Figure 4 , Figure 4 A schematic diagram of cluster analysis of an expression heat map of a biological sample to be tested and a plurality of reference biological samples for global comparison thereof is shown according to some embodiments of the present invention.
[0063] pass Figure 4 The similarities between mBTSCs and hBTSCs, hHpSCs, hHBs, and hAHeps can be seen. Furthermore, mouse mAHeps and human hAHeps can be clustered together, demonstrating the similarity between mAHeps and hAHeps. However, this does not prove that mouse mBTSCs are more similar to human hBTSCs. Instead, it leads to the conclusion that mBTSCs are more similar to hHpSCs.
[0064] In order to obtain more refined and accurate analysis and reduce the impact of data noise on the global analysis of transcriptome information, the present invention can extract multiple TOP SD genes with the largest standard deviation (SD) from the biological sample to be tested and multiple reference biological samples based on transcriptome information, and perform local cluster analysis to obtain a TOP SD gene heat map that clusters the biological sample to be tested and multiple reference biological samples that are locally similar to it.
[0065] Please refer to Figures 5A-5B , Figures 5A-5B A TOP SD gene heat map is shown in which a test biological sample provided according to some embodiments of the present invention and multiple reference biological samples for local comparison are clustered together.
[0066] like Figures 5A-5B As shown, the present invention can Figures 5A-5B The cluster analysis diagram of the expression heat map shown in the figure was optimized, and the 1000 and 3000 top SD genes with the largest standard deviation (SD) were analyzed, respectively. Here, the number of top SD genes includes but is not limited to 1000, 2000, 3000, 5000, and 10000. Using the number of top SD genes, the true differences in biological sample information between cells are amplified. As shown in the figure, the top 3000 and 1000 genes with the standard deviation (SD) in the six cell types (mBTSCs, mAHeps, hBTSCs, hHpSCs, hHBs, and hAHeps) show the same pattern as the correlation heat map. mBTSCs have similar expression patterns with hBTSCs and hHpSCs, while having more distinct expression patterns with precursor or mature lineages. From the results, it can be seen that mBTSCs are more similar to hBTSCs and hHpSCs. Here, the similarity between cells can be basically determined, but the correlation between cells and two of the cell types cannot be determined. Because biliary tree stem cells and hepatic stem cells are different stages of the same liver lineage and are very similar, further analysis is needed to clarify the positioning of mBTSCs.
[0067] The present invention can then further focus on local analysis, performing pairwise differential expression analysis on the transcriptome information of the test biological sample and multiple reference biological samples to obtain differences between various cell types and display the differential expression analysis results using a volcano plot. Here, the fewer the number of differentially expressed genes, the higher the similarity. Conversely, the more differentially expressed genes, the lower the similarity.
[0068] Please refer to Figures 6A to 6E , Figures 6A to 6E A volcano plot is shown for differential expression analysis between transcriptome information of a test biological sample and multiple reference biological samples according to some embodiments of the present invention.
[0069] like Figures 6A to 6E As shown, differential expression analysis can be performed using the limma package, with adjusted p-values all below 0.05. The spots on the right side of the figure represent upregulated genes in the cell type, while the spots on the left represent downregulated genes. Spots farther from the center of the figure are genes with the most significant differential expression in a cell type. The differentially expressed genes between hBTSCs and mBTSCs are the least abundant when comparing mBTSCs with cells of other lineages, including 852 downregulated and 468 upregulated genes (logFC>1, P<0.05). This suggests a high degree of similarity between the two cell types. However, the high similarity between hBTSCs and hHpSCs still leaves some uncertainty about the localization of mBTSCs relative to hBTSCs and hHpSCs, making it difficult to determine the localization of mBTSCs within the two.
[0070] To further clarify which cell type mBTSCs are more similar to, the present invention can use refined computational biology analysis based on differential genes to perform clustering and heat map analysis on the differential genes of the test biological sample and multiple reference biological samples, so as to obtain a cluster tree heat map that clusters the test biological sample and multiple reference biological samples with similar characteristics.
[0071] Please refer to Figures 7A-7B , Figures 7A-7B A cluster tree heat map is shown in which a plurality of reference biological samples are clustered together to compare characteristics of a biological sample to be tested with provided in some embodiments of the present invention.
[0072] like Figures 7A-7B As shown, the present invention can generate a cluster tree heat map based on cluster analysis of all differentially expressed genes and upregulated genes in hBTSCs and comparison with mBTSCs. When clustering and heat map analysis were performed based on differentially expressed genes between hBTSCs and hHpSCs (logFC>2, P<0.05), the cluster tree heat map suggested that mBTSCs and hHpSCs appear more similar than hBTSCs.
[0073] Furthermore, judging similarity based solely on genes upregulated in hBTSCs, understood as a panel / signature specific to hBTSCs, reveals that mBTSCs also express the hBTSC-specific signature, indicating that mBTSCs and hBTSCs are more similar. Therefore, when global or TOP SD gene panels alone are insufficient to identify the degree of similarity between a cell type and two other reference biological sample cell types, the present invention can utilize differential expression analysis to further identify differentially expressed genes between the two reference biological sample cell types, thereby better determining the positioning of the test biological sample cells in lineage cell development.
[0074] Principal Component Analysis (PCA) was then performed on the cluster tree heat map to determine the lineage stage of the tested biological sample in the lineage of the reference biological samples with similar characteristics. The results obtained by the above process were further visualized in combination with differential expression analysis.
[0075] Please refer to Figures 8A to 8D , Figures 8A to 8D A schematic diagram of the lineage stages of a biological sample to be tested in the lineage of a reference biological sample with which its characteristics are compared is shown according to some embodiments of the present invention.
[0076] like Figures 8A to 8D As shown, the biospecimen information processing system performed PCA analysis to determine the lineage stage of mBTSCs within the human hepatobiliary stem cell lineage. As shown in the figure, mBTSCs and hBTSCs are more similar in the context of all differentially expressed genes. Therefore, the cluster tree generated by clustering and heat map analysis of differentially expressed genes between the test biospecimen and multiple reference biospecimens may require additional support from principal component analysis (PCA). mBTSCs and hBTSCs are also more similar in the context of upregulated genes. Therefore, principal component analysis of the cluster tree heat map indicates that mBTSCs are closer to hBTSCs than hHpSCs. In summary, all of the above analyses indicate that mBTSCs are genetically closest to hBTSCs.
[0077] The present invention then performs differential expression analysis and fold change scatter plot correlation analysis on the entire genome, using a mature reference biological sample as a control. Fold change (FC) values for the entire genome are then obtained to simulate the biological distance between the test biological sample stem cells and the mature reference biological sample, thereby avoiding batch effects between the reference biological samples. Here, the mature reference biological sample is hAHeps.
[0078] Please refer to Figures 9A to 9G , Figures 9A to 9G A scatter plot of the fold differences of all genes provided according to some embodiments of the present invention is shown.
[0079] like Figures 9A to 9GAs shown, in all scatter plots, the present invention can draw a graph of a single gene based on the logarithmic (fold change) value. Among them, R is the Pearson coefficient of the significant gene. As shown in the figure, mBTSCs, hBTSCs, hHpSCs, and hHBs were compared with hAHeps, and the distances from the target cells (mBTSCs) to hAHeps were calculated, and the difference fold scatter plots were drawn. It can be seen that the distances of mBTSCs and hBTSCs to mature hepatocytes are the most similar (R = 0.68, with 1987 common differentially expressed genes).
[0080] Afterwards, the present invention can perform differential expression analysis on all genes and correlation analysis of the differential fold scatter plots of the differential genes to determine the impact of the common differentially expressed genes on the sample similarity.
[0081] Please refer to Figures 10A to 10C , Figures 10A to 10C A schematic diagram of differential expression analysis of all genes and correlation analysis of differential fold scatter plots of differential genes provided according to some embodiments of the present invention is shown.
[0082] like Figures 10A to 10C As shown, Figures 10A to 10C The shared DEGs of hBTSCs, mBTSCs, and hHpSCs are shown, along with a comparison with hAHeps. DEGs shared by hBTSCs, mBTSCs, and hHpSCs are designated intersected DEGs. R represents the correlation coefficient of the DEGs. Rall represents the correlation coefficient of all DEGs. As shown, relative to mature hAHeps cells, mBTSCs and hBTSCs share 1987 differentially expressed genes, while hHpSCs and mBTSCs share only 1746 differentially expressed genes. This suggests that the distance between mBTSCs and hBTSCs in differentiating into hAHeps is closer than that between hHpSCs. Further examining the effect of shared differentially expressed genes on similarity reveals that 1395 genes in mBTSCs and hBTSCs are more highly correlated, while the 351 intersecting differentially expressed genes are less correlated, with the lowest correlation found between mBTSCs and hBTSCs at 301 genes. Therefore, mBTSCs and hBTSCs are more similar.
[0083] Afterwards, the present invention can perform non-negative least squares regression analysis (NNLS), and the biological sample information processing system applies non-negative least squares (NNLS) regression based on all reference cell types in the human dataset B (reference cell types from dataset B, R B) gene expression to predict the target cell type in mouse dataset A (Query cell type from dataset A, Q A ) gene expression. A =β 0A +β 1A (R B ), Q A and R B Respectively represent the filtered gene expression of the target cell type from dataset A and all cell types from dataset B. Afterwards, the present invention can switch the order of dataset A and dataset B, that is, dataset A is the Reference sample, and dataset B is the Query sample, and the gene expression of all cell types in dataset A is used as the A Predict the gene expression Q of the target cell type in dataset B B , where Q B =β 0B +β 1B (R A ). Further, the present invention can express the gene Q A , the gene expression R B , the gene expression R A And the gene expression Q B , the coefficient of determination β=2(β AB +0.01)(β BA +0.01). β 0A It means when R B The value of the β coefficient when the expression level is 0, that is, the intersection of the regression line predicted by NNLS and the Y axis; β 1A refers to the β coefficient obtained after NNLS analysis; β 0B It means when R A The value of the β coefficient when the expression level is 0, that is, the intersection of the regression line predicted by NNLS and the Y axis; β 1B refers to the β coefficient obtained after NNLS analysis; β AB It refers to the β coefficient value when predicting the target sample A based on the reference sample B. BA This refers to the β coefficient value when predicting target sample B based on reference sample A. A higher β coefficient indicates more similar sample types. The present invention can then determine the non-negative least squares regression analysis results based on the β coefficient.
[0084] Please refer to Figure 11 , Figure 11 A schematic diagram showing the results of non-negative least squares regression analysis provided according to some embodiments of the present invention is shown.
[0085] like Figure 11As shown, mBTSCs and human hBTSCs have the highest beta coefficient of 0.73, making them the most similar. Existing techniques for annotating cell types in single-cell sequencing data already use this method, but it has not been used for sample similarity analysis and lineage mapping, so we will not elaborate on this here.
[0086] Please refer to Figures 12A-12B , Figures 12A-12B Shown are a correlation heat map and a correlation analysis diagram of cell lineages provided according to some embodiments of the present invention.
[0087] like Figures 12A-12B As shown in the figure, because biliary tree stem cells share common characteristics with small bile duct cells and mature bile duct cells, we further collected bile duct-related genes. As shown in the figure, mouse and human biliary tree stem cells can be clustered together. Correlation analysis based on these genes suggests that mouse and human biliary tree stem cells are more similar.
[0088] In summary, the bioinformatics system uses genetic characterization studies to identify key biomarkers, which are then used to localize the niches of these cells in the mouse liver, biliary tree, duodenum, and pancreas. This is achieved by isolating stem / progenitor cell subsets and assessing their ability to generate cells with hepatic or pancreatic fates ex vivo under fully defined culture conditions. Murine models have a highly heterogeneous genetic background, similar to the heterogeneous genetic makeup of the human population. Studies to date support the conclusion that both mice and humans possess a lineage maturation network within the biliary tract that differentiates biliary tree stem cells toward the liver and pancreas, making mice an ideal model for translational / preclinical research related to the potential of stem / progenitor cells, including biliary tree stem cells, for cell therapy in both normal and diseased conditions.
[0089] Through the above-mentioned information processing of the test and reference biological samples, it can be concluded that: 1. hBTSCs and hHpSCs are relatively similar in lineage. 2. The cells isolated from the mouse bile duct are most similar to human hBTSCs in terms of biological identification, functional identification, and genetic identification. Therefore, the cells isolated from the mouse bile duct are named mBTSCs. 3. The biological sample information processing method provided by the present invention can effectively identify the cell type and the specific location of the cell in the lineage, providing a set of methodologies for determining the type and status of unknown cells.
[0090] In addition, in another preferred embodiment of the present invention, the above-mentioned biological sample information processing system provided by the present invention can also perform analysis on high-throughput sequencing (RNA-seq) of different liver cancer models (Fah- / - mice and Ncoa5+ / - mice), explore the phenomenon that liver tissues tend to be similar after metformin treatment, evaluate the tissue similarity after metformin treatment of different liver cancer models, and prove that liver cancer induced by different gene knockout mice is in a similar microenvironment after metformin treatment, so as to suggest that metformin has similar therapeutic effects in different liver cancer models and that metformin may have similar rules in treating liver cancer of different causes, thereby once again verifying the reliability and wide applicability of the biological sample information processing method provided by the present invention.
[0091] Metformin (Met) is a commonly prescribed drug for the treatment of type 2 diabetes. In addition to its glucose-lowering effects, metformin directly inhibits complex I (CI) of the electron transport chain, leading to decreased cellular complex I activity and oxidative phosphorylation (OXPHOS) levels. This in turn activates the adenosine monophosphate-activated protein kinase (AMPK) signaling pathway, promoting cell cycle arrest and inhibiting tumor cell proliferation. Preclinical and clinical studies have also shown that long-term metformin use in diabetic patients may be associated with a reduced incidence of HCC. Metformin has also been shown to inhibit non-diabetes-induced HCC in mouse models of HCC. Williams et al. demonstrated that metformin may prevent HCC development by ameliorating the tumor-promoting microenvironment. Treatment with metformin significantly reduced HCC development in the Fah- / - mouse HCC model. Williams et al. also demonstrated that metformin treatment significantly reduced HCC development in the Ncoa5+ / - mouse HCC model. This suggests that metformin may act through similar mechanisms in different HCC models. However, there has been no report on how to use bioinformatics tools to prove that metformin may play a similar anti-cancer effect in different liver cancer models.
[0092] Please refer to Figure 13 , Figure 13 A flow chart illustrating information processing of tissue samples according to some embodiments of the present invention is shown.
[0093] like Figure 13As shown, both the test biological sample and the reference biological sample are liver tissues from liver cancer model mice, with experimental and control groups treated with or without metformin. Biological verification suggests that metformin may act through similar mechanisms in different liver cancer model mice, limiting its function to specific genes in specific pathways. Therefore, the biological sample information processing system performs high-throughput sequencing (RNA-seq) on Fah- / - mice, with at least three replicate RNA-seq samples per cell. High-throughput sequencing data for the Ncoa5+ / - mouse liver cancer model by Williams et al. were also obtained. After obtaining transcriptome information, in-depth exploration was conducted at the transcriptome level. Here, through the processing of the above biological sample information, the biological sample information processing system determines that the Fah- / - mouse liver cancer model is the test biological sample and the Ncoa5+ / - mouse liver cancer model is the reference biological sample for analysis.
[0094] Before conducting transcriptome-level analysis, high-throughput sequencing (RNA-seq) data were formatted as FPKM or TPM, and then further normalized to log2(FPKM+1) or log2(TPM+1) formats. Batch effects were then removed from the normalized data using the combat function in the sva package or the removebatcheffect function in the limma package.
[0095] The present invention can then begin a global analysis of the test biological sample and multiple reference biological samples. Here, based on the transcriptome information, a correlation analysis is first performed on the test biological sample and multiple reference biological samples to obtain a correlation heatmap that clusters the test biological sample and its globally similar reference biological samples. Optionally, correlation analysis methods include, but are not limited to, Pearson, Kendall, and Spearman.
[0096] Please refer to Figure 14 , Figure 14 A correlation heat map is shown in which a biological sample to be tested is clustered together with multiple reference biological samples for global comparison therewith according to some embodiments of the present invention.
[0097] like Figure 14 As shown, after batch correction using the sva package, the biological sample information processing system performed correlation analysis on all samples of the Fah- / - and Ncoa5+ / - liver cancer mouse models. As shown in the figure, metformin treatment was similar in both mouse models (Fah- / - mice, Ncoa5+ / -), but due to the differences in the models, it was impossible to further distinguish whether the metformin-treated group was more similar.
[0098] Based on transcriptome information, the biological sample information processing system performs clustering and expression heat map analysis on all genes detected in the test biological sample and multiple reference biological samples, thereby generating an expression heat map that clusters the test biological sample with its globally similar reference biological samples. Optionally, clustering methods include, but are not limited to, "ward.D," "ward.D2," "single," "complete," "average," "mcquitty," "median," and "centroid."
[0099] Please refer to Figure 15 , Figure 15 A schematic diagram of cluster analysis of an expression heat map of a biological sample to be tested and a plurality of reference biological samples for global comparison thereof is shown according to some embodiments of the present invention.
[0100] like Figure 15 As shown, the biological sample information processing system performed clustering and expression heat map analysis on all genes of all samples in the Fah- / - and Ncoa5+ / - liver cancer mouse models. During the global analysis, the transcriptome information was obviously interfered with by the noise data. Therefore, although the control group (wild type) can be clearly distinguished overall, the other five groups cannot be particularly clearly distinguished.
[0101] The results of the global analysis based on transcriptome information are often affected by model differences and data noise, so more refined analysis is often required. Based on the transcriptome information, the biological sample information processing system extracts multiple top SD genes with the largest standard deviations from the test biological sample and multiple reference biological samples, and performs local cluster analysis to obtain a top SD gene heat map that clusters the test biological sample with multiple reference biological samples that are locally similar.
[0102] Please refer to Figures 16A to 16C , Figures 16A to 16C A TOP SD gene heat map is shown in which a test biological sample provided according to some embodiments of the present invention and multiple reference biological samples for local comparison are clustered together.
[0103] like Figures 16A to 16C As shown, the biological sample information processing system Figure 15The cluster analysis diagram of the expression heat map shown in the figure was optimized, and the 1000 and 3000 top SD genes with the largest standard deviation (SD) were analyzed, respectively. Here, the number of top SD genes includes but is not limited to 1000, 2000, 3000, 5000, and 10000. Using the top SD gene count amplifies the true differences in biological sample information between cells, thereby reducing the impact of model and data noise. As shown in the figure, SD represents standard deviation. The biological sample information processing system compared the transcriptome information of all samples from the Fah- / - and Ncoa5+ / - liver cancer mouse models based on the 1000, 3000, and 5000 top SD genes with the largest standard deviation. The results show that the two metformin-treated groups (12W Met and 39W Ncoa5+ / -met) are more similar.
[0104] Afterwards, we further focused on local analysis and performed pairwise differential expression analysis on the transcriptome information of the tested biological samples and multiple reference biological samples to obtain the differences between various cell types. The results of the differential expression analysis were displayed using a volcano plot. The fewer the number of differential genes, the higher the similarity. Conversely, the more differential genes, the lower the similarity.
[0105] Please refer to Figures 17A to 17F , Figures 17A to 17F A volcano plot is shown for differential expression analysis between transcriptome information of a test biological sample and multiple reference biological samples according to some embodiments of the present invention.
[0106] like Figures 17A to 17F As shown, differential expression analysis can be performed using the limma package. Figure 18 The total number of differentially expressed genes between the two groups is shown, where the spots on the right side of the figure represent highly expressed genes, and the spots on the left side represent lowly expressed genes. Using the differentially expressed genes, it can be found that the two metformin-treated groups are most similar.
[0107] Afterwards, clustering and heat map analysis are performed on the differential genes of the test biological sample and multiple reference biological samples to obtain a cluster tree heat map that clusters the test biological sample and multiple reference biological samples with similar characteristics.
[0108] Please refer to Figure 18 , Figure 18 A cluster tree heat map is shown in which a plurality of reference biological samples are clustered together to compare characteristics of a biological sample to be tested with provided in some embodiments of the present invention.
[0109] like Figure 18As shown in the figure, the biological sample information processing system performs heat map and cluster analysis on all samples from the Fah- / - and Ncoa5+ / - liver cancer mouse models based on all differentially expressed genes. Among them, differentially expressed genes (DEGs) are the differentially expressed genes after all pairwise comparisons. As can be seen, the heat map created using differentially expressed genes based on all pairwise comparisons reveals that the two metformin-treated groups are most similar using all differentially expressed genes.
[0110] The cluster tree heat map was then subjected to principal component analysis to determine the lineage stage of the tested biological sample in the lineage of the reference biological samples with similar characteristics, and the results obtained by the above process were further visualized in combination with differential expression analysis.
[0111] Please refer to Figure 19 , Figure 19 A principal component analysis diagram provided according to some embodiments of the present invention is shown.
[0112] like Figure 19 As shown, the biological sample information processing system further obtains differentially expressed genes and performs principal component analysis (PCA) drawing. Figure 19 The PCA analysis of DEGs was performed for all samples from the Fah- / - and Ncoa5+ / - HCC mouse models compared to the Fah- / - HCC mouse model treated with metformin. Furthermore, to determine the distance between all samples and the 12-week Fah- / - met-treated group, the biological sample information processing system extracted all differentially expressed genes compared to the 12-week Fah- / - met-treated group for PCA analysis. This revealed that the two metformin-treated groups were more similar.
[0113] Subsequently, differential expression analysis and fold-change scatter plot correlation analysis of all genes were performed. 12WFah- / - mice not treated with metformin served as the control group, and the fold-change values of other groups compared with this group were obtained to simulate the distance between other groups and the untreated group. Distance was used instead of differential expression value to further avoid batch effects between reference biological samples.
[0114] Please refer to Figures 20A to 20D , Figures 20A to 20D A scatter plot of the fold differences of all genes provided according to some embodiments of the present invention is shown.
[0115] like Figures 20A to 20DAs shown, the biological sample information processing system calculates the distance of the biological sample to be tested (target group: metformin-treated group), draws the scatter plot, and obtains the correlation analysis of different groups in the Fah- / - and Ncoa5+ / - liver cancer mouse models, among which the correlation analysis of all genes between the logFC of "12W Fah- / -met-treated / 12W Fah-untreated" and all other four groups, FC represents "difference fold value", and all genes are displayed as points.
[0116] Afterwards, differential expression analysis of all genes and correlation analysis of the fold difference scatter plots of differentially expressed genes were performed to determine the impact of common differentially expressed genes on sample similarity.
[0117] Please refer to Figures 21A to 21D , Figures 21A to 21D A schematic diagram of differential expression analysis of all genes and correlation analysis of differential fold scatter plots of differential genes provided according to some embodiments of the present invention is shown.
[0118] like Figures 21A to 21D Figure 2. Correlation analysis of the logFC of the "12W Fah- / -met-treated / 12W Fah-untreated" group with the DEGs of the other four groups. All DEGs are shown as dots. Pink indicates the common DEGs; purple indicates the DEGs of the "12W Fah- / -treated / 12W Fah- / -untreated" group; and green indicates the DEGs associated with the other four Ncoa5+ / - HCC mouse models compared with different control groups.
[0119] The above-described biological sample information processing method leads to the following conclusions: 1. Despite knocking out different genes, both mouse models showed a significant reduction in liver cancer progression after metformin treatment, with high similarity, suggesting that metformin's therapeutic effects on liver cancer share common patterns across different liver cancer models. 2. The biological sample information processing method provided by this invention can effectively identify cell types and their specific location within the lineage, providing a methodology for determining the type and status of unknown cells.
[0120] Furthermore, the biological sample information processing methods, biological sample information processing systems, and computer-readable storage media provided by the present invention can also be used in other non-limiting embodiments. For example, P21 is a known senescence marker, yet p21 is still expressed in normal cells. The biological sample information processing methods provided by the present invention can be used to assess the similarity between P21-positive cells and known senescent cells or induced senescent cells.
[0121] For example, AFP-positive cells often express a specific high level of expression in transplanted hepatocytes, suggesting that AFP-positive hepatocytes may be hepatic progenitor cells. The biological sample information processing method provided by the present invention can be used to assess whether AFP-positive cells during embryonic development are more similar to hepatocytes from the fetal liver at a certain week or postpartum day.
[0122] In addition, cell types often differ between species. The biological sample information processing method provided by the present invention can be used to evaluate the similarity of cells from other species to human lineage cells, such as the similarity between mouse mesenchymal stem cells (MSCs) and human MSCs.
[0123] Treatment with metformin significantly reduces the incidence of liver cancer in the Fah- / - mouse liver cancer model and the Ncoa5+ / - mouse liver cancer model. However, because liver cancer has a large heterogeneity, the biological sample information processing method provided by the present invention can analyze whether the liver cancer tissue after metformin treatment acts through similar pathways, thereby alleviating the occurrence of liver cancer.
[0124] In addition to the above embodiments, the biological sample information processing method provided by the present invention can also be used to compare and analyze the old mice after gene knockout with normal old mice when specific mouse genes (such as p16, p21, etc.) are knocked out, to evaluate whether the knockout of specific genes will affect the normal phenotype of p21 and other genes, thereby causing the problem of excessive differences between gene knockout mice and normal physiological mice.
[0125] In summary, the present invention combines bioinformatics analysis methods such as cluster analysis, difference analysis, NNLS analysis, correlation analysis, and principal component analysis, as well as evaluation methods such as the correlation coefficient based on global genes and TOP SD genes, the clustering degree based on global genes and TOP SD genes, the difference multiple, the number of differential genes, the correlation coefficient based on the difference multiple, the correlation coefficient based on the differential genes, the clustering degree based on the differential genes, the β regression coefficient based on NNLS analysis, and the correlation coefficient based on characteristic genes. Therefore, it is possible to accurately judge the cell lineage stage and state based on high-throughput sequencing data, thereby providing a reference for the specific positioning of specific cell types and subpopulations in different lineage cells, and can also judge the similarity between different cell samples after culture and between tissue samples obtained after different experimental treatments.
[0126] Those skilled in the art will understand that these embodiments are merely some non-limiting implementation methods provided by the present invention, and are intended to clearly illustrate the main concepts of the present invention and provide some specific solutions that are convenient for the public to implement, rather than to limit all functions or all working modes that can be adopted by the biological sample information processing system.
[0127] Although the above methods are illustrated and described as a series of acts for simplicity of explanation, it is to be understood and appreciated that these methods are not limited by the order of the acts, as some acts may occur in a different order and / or concurrently with other acts from those illustrated and described herein or not illustrated and described herein but understandable to those skilled in the art according to one or more embodiments.
[0128] The previous description of the disclosure is provided to enable any person skilled in the art to make or use the disclosure. Various modifications to the disclosure will be apparent to those skilled in the art, and the general principles defined herein may be applied to other variations without departing from the spirit or scope of the disclosure. Thus, the disclosure is not intended to be limited to the examples and designs described herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for processing biological sample information, characterized in that: The following steps are involved: Obtaining transcriptome information of a biological sample to be tested and multiple known reference biological samples; Based on the transcriptome information, a global analysis based on whole genes, a local analysis based on differential genes, a feature analysis based on gene differential expression analysis, and a non-negative least squares regression analysis are performed on the biological sample to be tested and the multiple reference biological samples, wherein the step of performing the local analysis comprises: based on the transcriptome information, extracting multiple TOP SD genes with the largest standard deviations from the biological sample to be tested and the multiple reference biological samples to perform local clustering analysis to obtain a TOP SD gene heat map that clusters the biological sample to be tested and the multiple reference biological samples that are locally similar to it; performing pairwise differential expression analysis on the transcriptome information of the biological sample to be tested and the multiple reference biological samples, and displaying the differential expression analysis results using a volcano plot; and determining the local analysis results based on the TOP SD gene heat map and the volcano plot; and The overall characteristics of the biological sample to be tested are determined based on the global analysis results, the local analysis results, the feature analysis results and the non-negative least squares regression analysis results.
2. The processing method according to claim 1, wherein Before obtaining the transcriptome information of the biological sample to be tested and the plurality of reference biological samples, the processing method further includes the following steps: Performing biological identification on the biological sample to be tested based on known specific markers, functional characteristics and / or morphological characteristics; and Based on the results of the biological identification, a reference biological sample similar to the biological sample to be tested is determined.
3. The processing method according to claim 2, wherein Biological identification based on the specific marker includes at least one of qRT-PCR detection, Western Blot detection, immunofluorescence detection, and / or The biological identification based on the functional characteristics includes a functional comparison experiment between the biological sample to be tested and a known reference biological sample.
4. The processing method according to claim 2, characterized in that The step of obtaining transcriptome information of the biological sample to be tested and a plurality of known reference biological samples includes: In response to the result of the biological identification failing to identify a reference biological sample similar to the biological sample to be tested, high-throughput sequencing is performed on the biological sample to be tested and the plurality of reference biological samples to obtain transcriptome information thereof.
5. The processing method according to claim 1, wherein Before performing the global analysis, the local analysis, the feature analysis, and the non-negative least squares regression analysis, the processing method further includes the following steps: Organize transcriptome information from multiple sources into standardized data in a unified format; and The batch effect of the standardized data is removed using the combat function of the sva package or the removebatcheffect function of the limma package.
6. The processing method according to claim 1, wherein The steps of performing the global analysis include: Based on the transcriptome information, performing a correlation analysis on the biological sample to be tested and the plurality of reference biological samples to obtain a correlation heat map that clusters the biological sample to be tested and the plurality of globally similar reference biological samples; Based on the transcriptome information, performing clustering and expression heat map analysis on all genes detected in the biological sample to be tested and the plurality of reference biological samples to obtain an expression heat map that clusters the biological sample to be tested and the plurality of globally similar reference biological samples; and Based on the correlation heat map and the expression heat map, the global analysis result is determined.
7. The processing method according to claim 1, characterized in that The step of performing pairwise differential expression analysis on the transcriptome information of the biological sample to be tested and the plurality of reference biological samples comprises: The limma package is used to perform differential expression analysis on the transcriptome information of the test biological sample and multiple reference biological samples, so as to respectively determine the differential genes of the test biological sample and each of the reference biological samples, wherein the fewer the number of differential genes, the higher the similarity of the biological samples.
8. The processing method according to claim 7, characterized in that The steps of performing the feature analysis include: performing clustering and heat map analysis on the differential genes between the test biological sample and the plurality of reference biological samples to obtain a cluster tree heat map that clusters the test biological sample and the plurality of reference biological samples with similar characteristics, wherein the differential genes indicate the positioning of the test biological sample in lineage development; performing principal component analysis on the cluster tree heat map to determine the lineage stage of the biological sample to be tested in the lineage of reference biological samples with similar characteristics; Performing differential expression analysis and scatter plot correlation analysis of the whole gene, using mature reference biological samples as controls, to obtain the fold difference value of the whole gene, so as to simulate the biological distance between the biological sample to be tested and the mature reference biological sample, and avoid batch effects between the reference biological samples; Perform differential expression analysis on the whole genes and correlation analysis of the differential fold difference scatter plot of the differentially expressed genes to determine the impact of the common differentially expressed genes on the sample similarity; Performing correlation analysis on a set of characteristic genes that have common characteristics with the characteristics of the biological sample to be tested; and The feature analysis result is determined based on the cluster tree heat map, the lineage stage, the biological distance, the impact and the correlation analysis result.
9. The processing method according to claim 7, characterized in that The steps of performing the non-negative least squares regression analysis include: Through the non-negative least squares regression analysis, the gene expression R of all reference samples in the second data set B was used. B Predict the gene expression Q of the target sample type in the first dataset A A , where Q A = β 0A + β 1A (R B ), β 0A It means when R B The value of the β coefficient when the expression is 0, β 1A refers to the β coefficient obtained after the non-negative least squares regression analysis; Switch the order of the first dataset A and the second dataset B, take the dataset A as the reference sample, and take the dataset B as the target sample, and analyze the gene expression R of all cell types in the first dataset A. A , predict the gene expression Q of the target sample in the second dataset B B , where Q B =β 0B +β 1B (R A ), β 0B It means when R A The value of the β coefficient when the expression is 0, β 1B refers to the β coefficient obtained after the non-negative least squares regression analysis; According to the gene expression Q A , the gene expression R B , the gene expression R A And the gene expression Q B , the coefficient of determination β=2(β AB + 0.01)(β BA + 0.01), where β AB It refers to the β coefficient value when predicting the target sample A based on the reference sample B. BA Refers to the β coefficient value when predicting target sample B based on reference sample A; and The non-negative least squares regression analysis result is determined according to the coefficient β.
10. The processing method according to claim 1, wherein The biological sample to be tested and the reference biological sample are cell samples. After determining the overall characteristics of the biological sample to be tested, the processing method further includes the following steps: According to the overall characteristics of the cell sample to be tested, the cell type of the cell sample to be tested and / or the specific location of the cell sample to be tested in its cell lineage are identified.
11. The processing method according to claim 1, wherein The biological sample to be tested and the reference biological sample are tissue samples. After determining the overall characteristics of the biological sample to be tested, the processing method further includes the following steps: Based on the overall characteristics of the tissue sample to be tested, the experimental treatment methods and phenotypic changes at the pathological and histological levels of the tissue sample to be tested, and / or the pathological changes and experimental treatment methods of at least one of the reference biological samples in the tissue sample to be tested are identified.
12. A biological sample information processing system, characterized in that: include: a memory having computer instructions stored thereon; as well as A processor is connected to the memory and configured to execute computer instructions stored in the memory to implement the biological sample information processing method according to any one of claims 1 to 11.
13. A computer-readable storage medium having computer instructions stored thereon, characterized in that: When the computer instructions are executed by a processor, the biological sample information processing method according to any one of claims 1 to 11 is implemented.
Citation Information
Patent Citations
RNA-seq online analysis report system and generation method thereof
CN109584962A
Single cell cross-species cell type identification method
CN115064220A