Classified storage method and system for biological cell data

By standardizing and feature matrix conversion of cellomic data, building molecular feature fingerprints and classification systems, the problem of difficulty in integrating and classifying biological cell data in the existing technology is solved, and efficient data storage and dynamic updates are achieved.

CN120089206AInactive Publication Date: 2025-06-03JINING KESHUN BIOTECHNOLOGY CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510156457.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-12
Publication Date
2025-06-03
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing technology is difficult to effectively integrate and classify cell data from different cancer types, which makes researchers need to spend a lot of time and effort to screen and organize data when analyzing data.

Method used

By obtaining multi-source cellomics data, performing data standardization and feature matrix conversion, constructing cell molecular feature fingerprint data, and performing hierarchical classification and clustering analysis, establishing cell data association matrix and search tree to achieve dynamic updates and classification matching.

Benefits of technology

It realizes refined classification, dynamic update and efficient storage of biological cell data, reduces the time and energy of data screening and sorting, and improves the efficiency of data analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120089206A_ABST
    Figure CN120089206A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of classified storage of biological data, in particular to a classified storage method and system for biological cell data. The method comprises the following steps: acquiring multisource cytomics data; performing data standardization on the multisource cytomics data to obtain standard cell characteristic data; performing characteristic matrix conversion on the standard cell characteristic data to obtain a cell characteristic digital matrix; performing molecular fingerprint construction on the biological cells based on the cell characteristic digital matrix to obtain cell molecular characteristic fingerprint data; performing hierarchical classification on the standard cell characteristic data to obtain a cell phenotype classification system; performing hierarchical labeling on the standard cell characteristic data according to the cell phenotype classification system to obtain cell phenotype labeling data; and performing ontology mapping on the cell phenotype labeling data to obtain a cell phenotype relationship network. According to the method, refined classification, dynamic updating and efficient storage of the biological cell data are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of biological data classification and storage, and particularly to a method and system for classifying and storing biological cell data. Background Art

[0002] In biomedical research and clinical applications, the storage and management of biological cell data have always been key links. Biological cell data not only contains genomic information of cells, but also covers multi-dimensional complex information such as transcriptome, proteome, and metabolome. However, when dealing with such a large and complex biological cell data, the existing data storage methods face the following technical problems. Among them, the problem of data classification and storage is particularly prominent. In the traditional storage mode, biological cell data is often simply classified according to sample source or sequencing time. This extensive storage method cannot fully explore the internal relevance and hierarchy of the data. For example, in cancer research, there are subtle differences in the gene expression patterns of cell data of different cancer types, and these differences are crucial for revealing the pathogenesis of cancer and finding new therapeutic targets. However, the existing storage methods are difficult to effectively integrate and classify these potentially related data, resulting in researchers spending a lot of time and effort in data screening and sorting when analyzing the data. Summary of the Invention

[0003] Based on this, it is necessary for the present invention to provide a method and system for classifying and storing biological cell data to solve at least one of the above technical problems.

[0004] To achieve the above object, a method for classifying and storing biological cell data includes the following steps:

[0005] Step S1: Obtain multi-source cell omics data; perform data standardization on the multi-source cell omics data to obtain standard cell feature data; perform feature matrix transformation on the standard cell feature data to obtain a cell feature digital matrix; construct a molecular fingerprint of biological cells based on the cell feature digital matrix to obtain cell molecular feature fingerprint data;

[0006] Step S2: Perform hierarchical classification on the standard cell feature data to obtain a cell phenotype classification system; perform hierarchical annotation on the standard cell feature data according to the cell phenotype classification system to obtain cell phenotype annotation data; perform ontology mapping on the cell phenotype annotation data to obtain a cell phenotype relationship network;

[0007] Step S3: Calculate the similarity of the cell molecular feature fingerprint data to obtain cell similarity score data; perform hierarchical clustering on the cell molecular feature fingerprint data according to the cell similarity score data to obtain an initial cell clustering result; perform dynamic threshold optimization on the initial cell clustering result to obtain optimized cell classification data;

[0008] Step S4: Construct an association matrix for the cell phenotype relationship network and the optimized cell classification data to obtain a cell data association matrix; perform compressed storage on the cell data association matrix to obtain compressed cell association data; construct an index structure based on the compressed cell association data to obtain a cell data retrieval tree;

[0009] Step S5: Obtain new biological cell data; calculate the classification matching degree of the new biological cell data according to the cell feature digital matrix to obtain classification matching score data; perform an update judgment according to the classification matching score data to obtain a classification update trigger signal; perform recursive optimization on the cell data retrieval tree according to the classification update trigger signal to obtain dynamically updated cell classification data.

[0010] Through standardization processing and feature matrix conversion, the present invention can unify multi-dimensional data of genomics, transcriptomics, proteomics, and metabolomics into a standardized cell feature digital matrix. The construction of molecular fingerprints further digitalizes cell features, enabling cell data from different sources and types to be compared and classified within the same framework. The construction of a hierarchical classification system can fully exploit the hierarchy of cell data, systematically classify cells according to phenotypic characteristics, and enable a clearer understanding of the internal relationships between cells. Through similarity calculation and hierarchical clustering, based on cell molecular feature fingerprint data, the similarities and differences between cells can be accurately identified, thereby achieving more accurate classification. Through dynamic threshold optimization, the reliability of classification is further improved. By dynamically adjusting the clustering threshold, it can better adapt to the feature distributions of different cell populations and avoid classification errors caused by fixed thresholds. The construction of the cell phenotype relationship network enables cell classification not only based on a single molecular feature but also in combination with the phenotypic information of cells, further improving the accuracy of classification. Through the construction of an association matrix and compressed storage, the cell phenotype relationship network and the optimized cell classification data can be efficiently integrated, reducing waste of storage space, and at the same time further improving the storage efficiency through compression technology. The cell data retrieval tree generated through the construction of an index structure can quickly locate and retrieve target cell data, greatly improving the data retrieval speed and enabling more efficient acquisition of the required information. Through the dynamic update mechanism, new biological cell data can be processed in real time. Through the calculation of classification matching degree and dynamic update trigger signal, the classification attribution of new data can be quickly determined, and the existing classification system can be recursively optimized, enabling better exploration of the subtle differences and potential associations in cell data. In summary, the present invention realizes the refined classification, dynamic update, and efficient storage of biological cell data.

[0011] Preferably, the present invention also provides a classification storage system for biological cell data for executing the classification storage method of biological cell data as described above. The classification storage system for biological cell data includes:

[0012] A data preprocessing module, which is used to obtain multi-source omics data of cells; perform data standardization on the multi-source omics data of cells to obtain standard cell feature data; perform feature matrix conversion on the standard cell feature data to obtain a cell feature digital matrix; construct molecular fingerprints for biological cells based on the cell feature digital matrix to obtain cell molecular feature fingerprint data;

[0013] A classification system construction module, which is used to hierarchically classify the standard cell feature data to obtain a cell phenotype classification system; hierarchically annotate the standard cell feature data according to the cell phenotype classification system to obtain cell phenotype annotation data; perform ontology mapping on the cell phenotype annotation data to obtain a cell phenotype relationship network;

[0014] A data classification module, which is used to calculate the similarity of the cell molecular feature fingerprint data to obtain cell similarity score data; perform hierarchical clustering on the cell molecular feature fingerprint data according to the cell similarity score data to obtain an initial cell clustering result; perform dynamic threshold optimization on the initial cell clustering result to obtain optimized cell classification data;

[0015] An index construction module, which is used to construct an association matrix for the cell phenotype relationship network and the optimized cell classification data to obtain a cell data association matrix; perform compressed storage on the cell data association matrix to obtain cell compressed association data; construct an index structure according to the cell compressed association data to obtain a cell data retrieval tree;

[0016] A data dynamic update module, which is used to obtain new biological cell data; calculate the classification matching degree of the new biological cell data according to the cell feature digital matrix to obtain classification matching score data; perform update judgment according to the classification matching score data to obtain a classification update trigger signal; recursively optimize the cell data retrieval tree according to the classification update trigger signal to obtain dynamically updated cell classification data.

[0017] Through the close cooperation among various modules, the present invention forms a complete process for classifying, storing and managing biological cell data. From data preprocessing to classification system construction, then to data classification, index construction and dynamic update, each module provides support for the final classification result and data management. This not only improves the performance and efficiency of the system, but also provides a more convenient and efficient data management experience for users. Technical personnel can quickly obtain high-quality and refined classified cell data through the system, greatly reducing the time and effort for data screening and sorting. Description of the Drawings

[0018] By reading the following detailed description with reference to the accompanying drawings, other features, objects and advantages of the present invention will become more obvious:

[0019] Figure 1The figure shows a schematic flow chart of the steps of a method for classifying and storing biological cell data according to an embodiment.

[0020] Figure 2 The figure shows a detailed schematic flow chart of step S3 according to an embodiment.

[0021] Figure 3 The figure shows a detailed schematic flow chart of step S36 according to an embodiment. Detailed implementation manners

[0022] The technical method of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0023] In addition, the accompanying drawings are only schematic diagrams of the present invention and are not necessarily drawn to scale. The same reference numerals in the figures denote the same or similar parts, and thus repeated descriptions thereof will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. The functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor methods and / or microcontroller methods.

[0024] It should be understood that although the terms "first", "second", etc. may be used here to describe various units, these units should not be limited by these terms. These terms are only used to distinguish one unit from another. For example, without departing from the scope of the exemplary embodiments, the first unit can be called the second unit, and similarly the second unit can be called the first unit. The term "and / or" used here includes any and all combinations of one or more of the listed related items.

[0025] To achieve the above object, please refer to Figures 1 to 3 , the present invention provides a method for classifying and storing biological cell data, including the following steps:

[0026] Step S1: Obtain multi-source cell omics data; perform data standardization on the multi-source cell omics data to obtain standard cell feature data; perform feature matrix conversion on the standard cell feature data to obtain a cell feature digital matrix; construct a molecular fingerprint of biological cells based on the cell feature digital matrix to obtain cell molecular feature fingerprint data;

[0027] Step S2: Hierarchically classify the standard cell feature data to obtain a cell phenotype classification system; hierarchically annotate the standard cell feature data according to the cell phenotype classification system to obtain cell phenotype annotation data; perform ontology mapping on the cell phenotype annotation data to obtain a cell phenotype relationship network;

[0028] Step S3: Calculate the similarity of the cell molecular feature fingerprint data to obtain cell similarity score data; hierarchically cluster the cell molecular feature fingerprint data according to the cell similarity score data to obtain an initial cell clustering result; perform dynamic threshold optimization on the initial cell clustering result to obtain optimized cell classification data;

[0029] Step S4: Construct an association matrix for the cell phenotype relationship network and the optimized cell classification data to obtain a cell data association matrix; compress and store the cell data association matrix to obtain compressed cell association data; construct an index structure based on the compressed cell association data to obtain a cell data retrieval tree;

[0030] Step S5: Obtain new biological cell data; calculate the classification matching degree of the new biological cell data according to the cell feature digital matrix to obtain classification matching score data; perform an update judgment according to the classification matching score data to obtain a classification update trigger signal; recursively optimize the cell data retrieval tree according to the classification update trigger signal to obtain dynamically updated cell classification data.

[0031] In this embodiment, multi-source omics data are first obtained, including cell transcriptome sequencing data, proteomic mass spectrometry data, and metabolomic data. These data are obtained through the Illumina NovaSeq sequencing platform, the Orbitrap mass spectrometer of Thermo Fisher Scientific, and the UPLC-MS system of Waters, respectively. For transcriptome data, expression quantification is performed using the HTSeq tool, and the data are normalized using Pandas; for proteomic data, peptide identification and background noise elimination are performed using the MaxQuant software; for metabolomic data, peak calibration and metabolite quantification are performed using the XCMS software. Standard cell feature data are finally obtained through these steps. Next, the standard cell feature data are converted into a feature matrix. The data are matrixed using the NumPy library in Python, and gene expression data, protein expression data, and metabolite expression data are integrated into a unified digital matrix. Based on this feature matrix, the t-SNE algorithm is used to construct molecular fingerprints for cells, and cell molecular feature fingerprint data are obtained. The hierarchical clustering method is used, and the AgglomerativeClustering class in the Scikit-learn library is used to construct a cell phenotype classification system according to the cell molecular feature fingerprint data. The standard cell feature data are hierarchically labeled according to this classification system, and the labeling information is added to the data using the Pandas library to obtain cell phenotype labeling data. Ontology mapping is performed on the cell phenotype labeling data, and the Cytoscape software is used to construct a cell phenotype relationship network to integrate the gene function, protein function, and metabolic pathway information of cells. Two methods, cosine similarity and Euclidean distance, are used to calculate the similarity of the cell molecular feature fingerprint data to obtain cell similarity score data. Based on these score data, hierarchical clustering is used to perform clustering analysis on cells to obtain an initial cell clustering result. Through the dynamic threshold optimization method, the clustering result is adjusted to obtain optimized cell classification data. An association matrix is constructed between the cell phenotype relationship network and the optimized cell classification data, and the two are merged into an association matrix using the Pandas library. The sparse matrix tool in the SciPy library is used to compress and store the association matrix to obtain cell compressed association data. Based on these compressed data, a cell data retrieval tree is constructed, and the Balanced Tree tool is used to optimize the balance of the tree structure. When new biological cell data are obtained, using the same feature extraction rules, the classification matching degree of the new data is calculated according to the cell feature digital matrix to obtain classification matching score data. According to these score data, an update judgment is made. If the score is lower than the set dynamic threshold, a classification update is triggered. Finally, according to the classification update trigger signal, the cell data retrieval tree is recursively optimized to update the tree structure to obtain dynamically updated cell classification data.

[0032] Preferably, step S1 includes the following steps:

[0033] Step S11: Obtain multi-source cell omics data, where the multi-source cell omics data includes raw cell transcriptome sequencing data, raw cell proteome mass spectrometry data, and raw cell metabolome data;

[0034] Specifically, the Illumina NovaSeq 6000 sequencing platform can be used to perform high-throughput transcriptome sequencing on biological cell samples to obtain raw cell transcriptome sequencing data. These data contain the expression profiles of all genes in the cells and are stored in the database of the local server in the form of a large number of short sequence reads. Each sample generates approximately 10 GB of raw sequencing files. The Orbitrap Fusion Lumos mass spectrometer from Thermo Fisher Scientific is used to perform proteome analysis on the same set of biological cell samples. After lysing the cells and performing enzymatic digestion, the peptide segments are injected into the mass spectrometer for detection to obtain raw cell proteome mass spectrometry data. These data record the types and abundance information of proteins in the cells. The ACQUITY UPLC-Xevo TQ-S micro-column liquid chromatography-tandem mass spectrometer from Waters is used to analyze the cell metabolome. By separating and detecting the cell extracts, raw cell metabolome data are obtained. These data reflect the types and contents of metabolites in the cells.

[0035] Step S12: Perform quality control filtering on the raw cell transcriptome sequencing data to obtain cell transcriptome sequencing data;

[0036] Specifically, the raw cell transcriptome sequencing data can be imported into the FastQC software, which will generate a detailed quality report, including information on base quality distribution, sequence length distribution, base content deviation, adapter contamination, and other aspects. According to the quality report, it is found that some of the sequencing data have problems such as low-quality bases (Q value less than 20) and residual adapter sequences. The Trimmomatic tool is used for data filtering. In Trimmomatic, the following parameters are set: ILLUMINACLIP:TruSeq3-PE.fa:2:30:10, which is used to remove adapter sequences; LEADING:20 and TRAILING:20, which are used to remove bases with a quality lower than 20 at both ends of the sequence; SLIDINGWINDOW:4:20, which is used to perform a sliding window quality check on the middle part of the sequence. The window size is 4 bases. When the average quality within the window is lower than 20, the sequence is truncated starting from this window; MINLEN:36, which is used to remove sequence fragments with a length less than 36 bases after filtering. After the above operations, cell transcriptome sequencing data are obtained, and the average quality value of these data is increased to above Q30.

[0037] Step S13: Eliminate the background noise from the original cell proteome mass spectrometry data to obtain the cell proteome data;

[0038] Specifically, the original cell proteome mass spectrometry data can be imported into the MaxQuant software. In the parameter settings of the software, select the "background noise elimination" function and set an appropriate noise threshold. For example, set the noise threshold to 1% of the mass spectrometry signal intensity (i.e., only the data points with a signal intensity higher than 1% of the total signal intensity are retained). MaxQuant analyzes the mass spectrometry spectrum through a built-in algorithm, identifies and removes those noise signals below the threshold. After the above processing, the cell proteome data is obtained.

[0039] Step S14: Perform peak calibration on the original cell metabolome data to obtain the cell metabolome data, and record the cell transcriptome sequencing data, cell proteome data, and cell metabolome data as standard cell characteristic data;

[0040] Specifically, the original cell metabolome data can be imported into the MassLynx software. In the "data processing" module of the software, select the "peak calibration" function. According to the experimental conditions and instrument performance, set the calibration parameters. For example, select the "internal standard calibration" mode, use an internal standard with a known concentration (such as deuterated metabolites) as a reference, and by adjusting the calibration parameters of the retention time and mass-to-charge ratio (m / z), control the deviation between the theoretical value and the actual measured value of the internal standard within 0.01 m / z. The software calibrates the peaks of each metabolite through an automatic fitting algorithm to ensure that the peak positions of all metabolites are accurate. After peak calibration, the cell metabolome data is obtained, and these data can more accurately reflect the types and contents of metabolites in the cells. Integrate the cell metabolome data with the cell transcriptome sequencing data and cell proteome data, and record it as standard cell characteristic data.

[0041] Step S15: Perform a feature matrix transformation on the standard cell characteristic data to obtain a cell characteristic digital matrix, and construct a molecular fingerprint for the biological cells based on the cell characteristic digital matrix to obtain cell molecular characteristic fingerprint data.

[0042] Specifically, for the detailed implementation process of this embodiment, please refer to the sub-steps of Step S15.

[0043] By integrating the original cell transcriptome sequencing data, proteome mass spectrometry data, and metabolome data, the present invention can comprehensively describe the characteristics of cells from multiple dimensions of gene expression, protein expression, and metabolite levels. This makes the description of cell characteristics richer and more comprehensive, avoiding the information loss caused by a single data type. Through quality control filtering, low-quality data can be effectively removed, improving the accuracy and reliability of the data. Through background noise elimination, interference signals can be removed, enhancing the accuracy and sensitivity of protein identification. Through peak calibration, the deviation in the data can be corrected, improving the accuracy of metabolite quantification. Through feature matrix transformation, complex multi-source data can be converted into a unified digital matrix form, making the data more conducive to subsequent mathematical analysis and calculation. Through the cell molecular fingerprint constructed based on the cell feature digital matrix, the characteristics of cells can be described in a concise and intuitive manner, facilitating comparison and classification among different cells.

[0044] Preferably, step S15 includes the following steps:

[0045] Step S151: Quantify the expression level of the cell transcriptome sequencing data to obtain cell gene expression data, and fill in the missing values of the cell gene expression data to obtain complete cell gene expression data;

[0046] Specifically, the cell transcriptome sequencing data can be aligned with a reference genome sequence (such as the human genome hg38), and an efficient alignment can be performed using an alignment tool such as HISAT2 to generate an alignment result file (in.bam format). The HTSeq tool is used to read the alignment result file, and the read counts of each gene are statistically calculated according to a gene annotation file (such as a GTF-format genome annotation file). The TPM (Transcripts Per Million) method is selected for normalization processing, and the TPM value of each gene is calculated through the formula TPM = (Gene Length × Read Counts) × 10^6. After obtaining the cell gene expression data, it is found that some genes have missing values, that is, the expression levels of some genes in some samples are zero or missing. The KNN (K-Nearest Neighbors) algorithm is used to select the 5 genes most similar to the target gene as references, and the missing values are filled based on the expression levels of these genes. In this way, complete cell gene expression data is obtained.

[0047] Step S152: Identify peptide segments in the cell proteome data to obtain cell protein expression data, and correct the outliers of the cell protein expression data to obtain corrected cell protein expression data;

[0048] Specifically, cell proteome mass spectrometry data (.raw format) can be imported into MaxQuant software. In the parameter settings of the software, select the "peptide identification" function and set the following key parameters: the digestion rule is Trypsin / P (trypsin digestion rule), allowing a maximum of 2 missed cleavage sites; the mass tolerance is set to 10 ppm (parts per million) for precursor ions and 0.02 Da (Dalton) for fragment ions; the database is selected as the human protein sequence database (UniProt human protein library). MaxQuant matches the mass spectrometry data through its built-in search engine, identifies the corresponding peptide sequences, and generates protein identification results. After obtaining the cell protein expression data, it is found that some data have outliers. For example, the expression levels of some proteins are too high or too low, which are caused by mass spectrometry detection errors or sample contamination. The Z-score method is used for outlier detection, and data points with an absolute Z-score greater than 3 are regarded as outliers. For these outliers, the median replacement method is used for correction, that is, the median expression level of the protein in other samples is used to replace the outliers. In this way, the corrected cell protein expression data is obtained.

[0049] Step S153: Quantify metabolites in the cell metabolome data to obtain cell metabolite expression data, and perform data normalization on the cell metabolite expression data to obtain standard cell metabolite expression data;

[0050] Specifically, cell metabolome data (.raw format) can be imported into XCMS software. In the parameter settings of the software, select the "metabolite quantification" function and set the following key parameters: the retention time correction method is "obiwarp", which is used to correct the retention time deviation between different samples; the mass accuracy is set to 0.01 Da; the minimum peak intensity for peak detection is set to 1000. With these parameter settings, XCMS can automatically identify and quantify the peaks of metabolites, generating the quantitative data of metabolites. After obtaining the cell metabolite expression data, it is necessary to perform data normalization on it. Select the "total ion intensity normalization" method. By dividing the quantitative values of all metabolites in each sample by the total ion intensity of the sample, the total ion intensity of each sample is normalized to 1, and finally the standard cell metabolite expression data is obtained.

[0051] Step S154: Perform dimensionality reduction on the cell complete gene expression data to obtain cell gene feature vectors, and perform feature selection on the corrected cell protein expression data to obtain cell protein feature vectors;

[0052] Specifically, the complete gene expression data of cells (represented by TPM values) can be imported into the R language environment, and the prcomp function can be used for PCA analysis. During the analysis, the first 10 principal components are selected, which can explain approximately 80% of the variation in the data. Through PCA analysis, the high-dimensional gene expression data is reduced to a 10-dimensional principal component space, obtaining the cell gene feature vector. Use a variance-based filtering method for feature selection of the corrected protein expression data of cells. The specific operation is as follows: First, calculate the variance of the expression level of each protein in all samples, and then select the proteins with a variance greater than 1 as the feature proteins. This can remove those proteins with small changes in expression levels between samples and retain those proteins with significant expression differences as the feature vectors. Through the above operations, the cell protein feature vector is obtained.

[0053] Step S155: Perform pattern extraction on the standard cell metabolite expression data to obtain the cell metabolite feature vector, and perform feature fusion on the cell gene feature vector, the cell protein feature vector, and the cell metabolite feature vector to obtain the cell feature digital matrix;

[0054] Specifically, the standard cell metabolite expression data can be imported into the Python environment, and the sklearn.ensemble.RandomForestClassifier module can be used. In the random forest model, set the following parameters: n_estimators = 100 (indicating the number of trees in the forest), max_depth = 10 (the maximum depth of the tree), min_samples_split = 2 (the minimum number of samples required to split an internal node). By training the random forest model, extract the importance score of each metabolite, and select the metabolites with an importance score higher than 0.05 as the feature vectors. Perform feature fusion on the cell gene feature vector (obtained by PCA dimensionality reduction), the cell protein feature vector (obtained by variance filtering), and the cell metabolite feature vector. The specific operation is to merge these three feature vectors by column to form a comprehensive cell feature digital matrix. For example, if the gene feature vector has 10 dimensions, the protein feature vector has 20 dimensions, and the metabolite feature vector has 15 dimensions, then the final cell feature digital matrix is 45 dimensions. Through the above operations, the cell feature digital matrix is obtained.

[0055] Step S156: Calculate the similarity of the cell feature digital matrix to obtain the cell population feature similarity matrix;

[0056] Specifically, the cell feature digital matrix can be imported into the R language environment, and the cosine function (from the lsa package) can be used to calculate the cosine similarity matrix. The calculation formula for cosine similarity is A and B represent the feature vectors of two cells respectively. Use the dist function to calculate the Euclidean distance matrix. The calculation formula of the Euclidean distance is: A i The letters A and B represent the i-th element of vector A and vector B respectively. The cosine similarity matrix and the Euclidean distance matrix are weighted and fused, with weights of 0.7 and 0.3 respectively. Through the above operations, the cell population feature similarity matrix is ​​obtained.

[0057] Step S157: constructing molecular fingerprints of biological cells according to the cell population feature similarity matrix to obtain cell molecular feature fingerprint data.

[0058] Specifically, the cell population feature similarity matrix can be imported into the Python environment and the sklearn.manifold.TSNE module can be used. In the t-SNE model, the following parameters are set: n_components = 2 (indicating that the data is reduced to two-dimensional space for visualization), perplexity = 30 (indicating the number of neighbors for each point), and learning_rate = 200 (learning rate). Through the t-SNE algorithm, the high-dimensional cell feature similarity matrix is ​​mapped to a two-dimensional space to generate the two-dimensional coordinates of each cell. These two-dimensional coordinates can intuitively reflect the similarities and differences between cells and form the molecular feature fingerprints of cells. For example, similar cells will cluster together in two-dimensional space, while cells with large differences will be distributed in different areas.

[0059] The present invention effectively solves the missing problem that occurs during data collection by filling missing values, and ensures the integrity of the data. By correcting outliers, abnormal fluctuations and errors in the data are eliminated, and the accuracy and reliability of the data are improved. By data standardization, the deviations caused by different experimental conditions or measurement methods are eliminated, so that the metabolome data are comparable between different samples. Through standardization, transcriptome, proteome and metabolome data can be fused and analyzed under the same framework, solving the problem of heterogeneity of multi-source data. Through dimensionality reduction and feature selection, redundant information can be removed, and the most representative and discriminative feature vectors can be extracted, which significantly improves the efficiency of data processing and the accuracy of analysis. Through pattern extraction, and fusing gene, protein and metabolite feature vectors, the feature information of cells can be extracted from different levels to form a comprehensive cell feature digital matrix. By calculating the cell population feature similarity matrix, the similarity between cells can be quantified.

[0060] Preferably, step S2 comprises the following steps:

[0061] Step S21: annotating the cell transcriptome sequencing data to obtain a cell gene ontology thesaurus;

[0062] Specifically, cell gene expression data can be imported into a local server. The BLAST tool is used to align these gene sequences with the human genome sequences in the Ensembl database, and the alignment parameters are selected as follows: the e-value threshold is set to 1 e-10 ; the word size is set to 11. Through the alignment, each gene is assigned corresponding Gene Ontology (GO) annotation information, including annotations in three aspects: molecular function, biological process, and cellular component. For example, a certain gene is annotated as "involved in cell cycle regulation (GO:0007049)" and "has protein kinase activity (GO:0004672)". Finally, the annotation information of all genes is integrated into a cell gene ontology thesaurus and stored in tabular form, with each row representing a gene and its corresponding GO annotation information.

[0063] Step S22: Identify proteins from the cell proteome data to obtain a cell proteome profile, and extract functional annotations from the cell proteome profile to obtain a cell protein function thesaurus;

[0064] Specifically, cell proteome data can be imported into the MaxQuant software. In the parameter settings of the software, select the "protein identification" function and set the following parameters: the digestion rule is Trypsin / P, allowing a maximum of 2 missed cleavage sites; the mass tolerance is set to 10 ppm for precursor ions and 0.02 Da for fragment ions; the database selection is the human protein sequence database (UniProt human protein library). With these settings, MaxQuant can accurately identify proteins and generate protein identification results. The identified protein sequences are aligned with the UniProt database, and the functional annotation information of each protein is extracted. For example, a certain protein is annotated as "involved in apoptosis (GO:0006915)" and "has protease activity (GO:0004175)". Finally, the functional annotation information of all proteins is integrated into a cell protein function thesaurus and stored in tabular form, with each row representing a protein and its corresponding functional annotation information.

[0065] Step S23: Cluster the metabolite profiles of the cell metabolome data to obtain a cell metabolome molecular map, and extract pathway information from the cell metabolome molecular map to obtain a cell metabolic pathway thesaurus;

[0066] Specifically, standard cell metabolite expression data can be imported into the MetaboAnalyst platform, and the "metabolite clustering analysis" function can be selected. In the clustering analysis, the hierarchical clustering method is selected, the distance metric is set to Euclidean distance, and the clustering algorithm is the Ward method. Through these settings, MetaboAnalyst generates a clustering heatmap of metabolites, showing the similarities and differences between different metabolites, and forming a molecular map of the cell metabolome. Using the pathway analysis function of MetaboAnalyst, the metabolites are matched with known metabolic pathways. The specific operation is to compare the metabolite list with the metabolic pathways in the KEGG (Kyoto Encyclopedia of Genes and Genomes) database, and the comparison parameters are set as follows: the p-value threshold is 0.05 to screen out significantly relevant metabolic pathways. For example, certain metabolites are significantly related to the "tricarboxylic acid cycle (TCA Cycle)" or "glycolysis" pathways. Finally, this pathway information is integrated into a cell metabolic pathway thesaurus and stored in tabular form, with each row representing a metabolite and its corresponding metabolic pathway information.

[0067] Step S24: Perform hierarchical association clustering on the cell gene ontology thesaurus to obtain a cell gene ontology hierarchical tree, and perform functional stratification on the cell protein function thesaurus to obtain a cell protein function hierarchical tree;

[0068] Specifically, the cell gene ontology thesaurus can be imported into the Cytoscape software, and the "Hierarchical Clustering" function in the "Network Analysis" module can be selected. In the clustering analysis, a similarity metric method based on GO annotation is selected. For example, "Semantic Similarity" is used to calculate the similarity between genes. The clustering parameters are set as follows: the similarity threshold is 0.7. Through these settings, Cytoscape generates a cell gene ontology hierarchical tree, showing the hierarchical relationship and functional association between genes. For example, certain genes are further subdivided into sub-functions such as "G1 phase regulation" and "S phase regulation" under the function of "cell cycle regulation". Next, using the DAVID bioinformatics resource platform, the protein function annotation information is imported into DAVID, and the "Functional Annotation Clustering" function is selected. The parameters are set as follows: the similarity threshold is 0.5 to screen out protein groups with similar functions. For example, certain proteins are stratified into "apoptosis-related proteins" and "cell proliferation-related proteins". Finally, this stratification information is integrated into a cell protein function hierarchical tree and stored in a tree structure diagram, with each layer representing a functional level.

[0069] Step S25: Classify the cell metabolic pathway thesaurus to obtain a cell metabolic pathway hierarchy tree, and integrate the cell gene ontology hierarchy tree and the cell metabolic pathway hierarchy tree to obtain a cell molecular function classification system;

[0070] Specifically, the cell metabolic pathway thesaurus can be imported into the KEGG Mapper tool, and the "Pathway Mapping" function is selected. In this function, metabolites are matched with metabolic pathways in the KEGG database, and the pathways are classified according to the distribution of metabolites. For example, metabolites mainly distributed in pathways such as "glycolysis", "tricarboxylic acid cycle", and "fatty acid metabolism" are classified into the "energy metabolism" category; metabolites mainly distributed in pathways such as "nucleotide metabolism" and "amino acid metabolism" are classified into the "biosynthesis" category. Through the above operations, a cell metabolic pathway hierarchy tree is obtained, where each level represents a different metabolic function category. Next, the gene ontology hierarchy tree and the metabolic pathway hierarchy tree are imported into Cytoscape, and the "Merge Networks" function is selected to associate the two based on common biological functions. For example, the "cellular energy metabolism" function in the gene ontology hierarchy tree is associated with the "energy metabolism" category in the metabolic pathway hierarchy tree to form a unified cell molecular function classification system. Finally, this classification system is displayed in the form of a network diagram, clearly reflecting the associations between genes and metabolic pathways at different functional levels.

[0071] Step S26: Integrate the cell molecular function classification system and the cell metabolic pathway hierarchy tree to obtain a cell phenotype classification system;

[0072] Specifically, the data of the cell molecular function classification system and the cell metabolic pathway hierarchy tree can be imported into the R language environment. The igraph package is used to process and visualize network data. By writing an R script, the nodes (genes and metabolic pathways) in the gene ontology hierarchy tree and the metabolic pathway hierarchy tree are merged according to functional similarity. For example, the "apoptosis" function in the gene ontology hierarchy tree is merged with the "apoptosis-related metabolic pathway" in the metabolic pathway hierarchy tree to form a new functional node "apoptosis phenotype". During the merging process, the functional similarity threshold is set to 0.6, that is, only nodes with a functional similarity greater than 0.6 will be merged. Through the above steps, a comprehensive cell phenotype classification system is constructed, where each phenotype node represents the comprehensive functional characteristics of cells at the gene and metabolic levels.

[0073] Step S27: Hierarchically annotate the standard cell feature data according to the cell phenotype classification system to obtain cell phenotype annotation data; perform ontology mapping on the cell phenotype annotation data to obtain a cell phenotype relationship network.

[0074] Specifically, for the detailed implementation process of this embodiment, please refer to the sub-steps of step S27.

[0075] The present invention extracts the functional information of cells from three levels of genes, proteins, and metabolites through transcriptome annotation, protein function annotation, and metabolic pathway annotation, forming a gene ontology thesaurus, a protein function thesaurus, and a metabolic pathway thesaurus. This provides rich semantic information for the description of cell characteristics, enabling cell characteristics to be not limited to numerical data but also including the description of biological functions. Through hierarchical association clustering and functional stratification, the functional information of genes, proteins, and metabolic pathways is integrated to construct a cell gene ontology hierarchical tree, a protein function hierarchical tree, and a metabolic pathway hierarchical tree. This can reveal the internal associations and hierarchical structures of cell functions. By integrating the gene ontology hierarchical tree, the protein function hierarchical tree, and the metabolic pathway hierarchical tree, a cell molecular function classification system is constructed, and further fused to obtain a cell phenotype classification system. This can comprehensively and systematically describe the phenotypic characteristics of cells. Through hierarchical annotation, a clear hierarchical structure and functional labels can be assigned to cell data. This makes the classification of cell data clearer and easier to understand. Through ontology mapping, a cell phenotype relationship network is constructed. This can reveal the functional associations and hierarchical relationships between cells. For example, researchers can quickly identify the similarities and differences between cells, as well as their interactions in biological processes, through the relationship network.

[0076] Preferably, step S27 includes the following steps:

[0077] Step S271: Perform gene function annotation on the standard cell characteristic data according to the cell phenotype classification system to obtain cell gene function annotation data;

[0078] Specifically, the gene function classifications defined in the cell phenotype classification system can be imported into the DAVID platform, and the "Functional Annotation Tool" function module is selected. The gene expression data in the standard cell characteristic data is matched with the cell phenotype classification system, and the "Gene Ontology (GO)" and "Pathway Analysis" functions of the DAVID platform are used to perform function annotation on each gene. For example, for a gene involved in "cell cycle regulation", the DAVID platform will provide detailed GO annotations, such as "GO:0007049" (cell cycle). In the R language environment, the AnnotationDbi package is used to compare the gene expression data with the human genome annotation database (such as org.Hs.eg.db) to extract the detailed function information of each gene. Through the above operations, cell gene function annotation data is obtained, where each gene is given a cell phenotype-related function annotation, such as "Gene A is involved in cell cycle regulation and has protein kinase activity". These annotation data are stored in tabular form.

[0079] Step S272: Perform protein function annotation on the standard cell feature data according to the cell phenotype classification system to obtain cell protein function annotation data;

[0080] Specifically, the protein function classifications defined in the cell phenotype classification system can be imported into the Cytoscape software and its "Network Annotation" function module can be utilized. Match the protein expression data in the standard cell feature data with the cell phenotype classification system. Through the "Annotation Importer" plug-in of Cytoscape, compare the protein data with the functional annotation information in the UniProt database. For example, for proteins involved in "apoptosis", the UniProt database will provide detailed GO annotations such as "GO:0006915" (apoptosis). During the comparison process, set the similarity threshold to 0.7. Through the above operations, cell protein function annotation data is obtained, where each protein is assigned a cell phenotype-related functional annotation, such as "Protein X is involved in apoptosis and has protease activity". These annotation data are stored in the form of a network graph, with each node representing a protein and its functional annotation.

[0081] Step S273: Perform metabolic pathway annotation on the standard cell feature data according to the cell phenotype classification system to obtain cell metabolic pathway annotation data;

[0082] Specifically, the metabolic pathway classifications defined in the cell phenotype classification system can be imported into the KEGG Mapper tool. Match the metabolite expression data in the standard cell feature data with the cell phenotype classification system. Utilize the "Pathway Enrichment Analysis" function of KEGG Mapper to compare the metabolites with the metabolic pathways in the KEGG database, and set the comparison parameters as: the p-value threshold is 0.05, and the q-value threshold is 0.1 to screen out significantly relevant metabolic pathways. For example, certain metabolites are significantly related to the "Glycolysis" or "TCA Cycle" pathways. Then, in the R language environment, use the clusterProfiler package to further analyze the enrichment of metabolic pathways. Through the enrichKEGG function, compare the metabolite data with the KEGG database and generate a detailed metabolic pathway annotation report. Finally, cell metabolic pathway annotation data is obtained, where each metabolite is assigned a detailed metabolic pathway annotation, such as "Metabolite A is involved in the glycolysis pathway (map00010)". These annotation data are stored in tabular form.

[0083] Step S274: Integrate the cell gene function annotation data and the cell protein function annotation data to obtain molecular function annotation data;

[0084] Specifically, the cell gene function annotation data and the cell protein function annotation data can be imported into the R language environment respectively. Use the inner_join function of the dplyr package to merge the gene function annotation data and the protein function annotation data based on cell phenotype classification. For example, merge the functional annotations of genes and proteins involved in "cell cycle regulation" into the same phenotype classification. During the merging process, set the matching parameter as: the phenotype classifications of genes and proteins must be exactly the same. Through the above operations, molecular function annotation data is obtained, where each phenotype classification contains the comprehensive functional annotations of genes and proteins. For example, under the "cell cycle regulation" phenotype classification, there are both functional annotations of genes (such as "Gene A is involved in cell cycle regulation and has protein kinase activity") and functional annotations of proteins (such as "Protein X is involved in cell cycle regulation and has protease activity").

[0085] Step S275: Perform correlation mapping on the cell gene function annotation data and the cell metabolic pathway annotation data to obtain cell phenotype annotation data;

[0086] Specifically, the cell gene function annotation data and the cell metabolic pathway annotation data can be imported into the R language environment. Use the left_join function of the dplyr package to perform correlation mapping on the gene function annotation data and the metabolic pathway annotation data based on cell phenotype classification. For example, correlate the functional annotation of genes involved in "cell cycle regulation" with the relevant metabolic pathway annotation (such as the "tricarboxylic acid cycle" pathway). During the correlation process, set the matching parameter as: the phenotype classifications of genes and metabolic pathways must be the same. Through the above processing, cell phenotype annotation data is obtained, where each phenotype classification contains the comprehensive annotations of gene functions and metabolic pathways. For example, under the "cell cycle regulation" phenotype classification, there are both functional annotations of genes (such as "Gene A is involved in cell cycle regulation and has protein kinase activity") and metabolic pathway annotations (such as "Metabolite A is involved in the tricarboxylic acid cycle pathway").

[0087] Step S276: Perform ontology mapping on the cell phenotype annotation data to obtain a cell phenotype relationship matrix, and construct a network topology structure based on the cell phenotype relationship matrix to obtain a cell phenotype relationship network.

[0088] Specifically, the cell phenotype annotation data can be imported into the Cytoscape software, and its "Network Merge" function can be used to integrate the gene function annotation and metabolic pathway annotation data into a network. In Cytoscape, each gene and metabolite is represented as a node, while the association between gene functions and metabolic pathways is represented as an edge. The "Network Analyzer" plugin of Cytoscape is used to perform topological analysis on the network, and the characteristic parameters of the network, such as the degree of nodes, betweenness centrality, and clustering coefficient, are extracted. The igraph package of the R language is used to process the network. Through the graph_from_data_frame function of the igraph package, the cell phenotype annotation data is converted into a graph object, and the layout_nicely function is used to optimize the layout of the network. Finally, a cell phenotype relationship network is obtained, where each node represents a gene or metabolite, and the edge represents the functional association between them. For example, in the network, there is an edge between gene A (involved in cell cycle regulation) and metabolite A (involved in the tricarboxylic acid cycle pathway), indicating their functional association in the cell cycle regulation phenotype.

[0089] Through gene function annotation, protein function annotation, and metabolic pathway annotation, the present invention annotates cell characteristic data from multiple dimensions, enriching the semantic information of cell data. This not only provides the functional information of cells at the gene and protein levels but also combines the dynamic changes of metabolic pathways, making the description of cell characteristics more comprehensive. By integrating gene function annotation data and protein function annotation data, the functional state of cells can be systematically described at the molecular level. This helps to reveal the synergistic effects between genes and proteins. Through association mapping, the internal connections between gene expression, protein function, and metabolic pathways can be revealed. This enables the description of cell functions not to be limited to a single level but to reflect the dynamic changes of cells at multiple levels. Through ontology mapping, the functional annotations of cells can be associated with standardized ontology terms, making the description of cell data more standardized and unified. This not only improves the comparability and integratability of data but also facilitates data sharing across laboratories and studies. By constructing a cell phenotype relationship network, the functional associations and hierarchical structures between cells can be intuitively displayed.

[0090] Preferably, step S3 includes the following steps:

[0091] Step S31: Extract features from the cell molecular feature fingerprint data to obtain a cell fingerprint feature description vector, and perform Euclidean distance evaluation on the cell fingerprint feature description vector to obtain cell fingerprint feature distance data;

[0092] Specifically, the cell molecular feature fingerprint data (stored in the form of a digital matrix) can be imported into the Python environment. The NumPy library is used to preprocess the data, for example, perform normalization processing to ensure that the value of each feature is between 0 and 1. Through the above operations, each feature value in the original data matrix is converted into a normalized value. The Euclidean distance between the cell fingerprint feature description vectors is calculated using the distance module in the SciPy library. Specifically, the cdist function is used, which can calculate the Euclidean distance between each row of two matrices. Finally, a Euclidean distance matrix is obtained, where each element represents the Euclidean distance between two cell fingerprint feature description vectors. For example, distance_matrix[i][j] represents the distance between the i-th cell and the j-th cell.

[0093] Step S32: Evaluate the cosine similarity of the cell fingerprint feature description vectors to obtain cell feature similarity data;

[0094] Specifically, the cosine similarity between the cell fingerprint feature description vectors can be calculated using the cosine function in the SciPy library. The calculation formula for cosine similarity is where E and F are two fingerprint feature description vectors. Through the above steps, the cosine similarity between the normalized feature vectors is calculated. Specifically, the cdist function is used and the distance metric is specified as cosine, thereby generating a cosine similarity matrix. Finally, a cosine similarity matrix is obtained, where each element represents the cosine similarity between two cell fingerprint feature description vectors. For example, cosine_similarity_matrix[i][j] represents the similarity between the i-th cell and the j-th cell. The value range of cosine similarity is between 0 and 1, and the closer the value is to 1, the more similar the two vectors are.

[0095] Step S33: Perform weighted fusion on the cell fingerprint feature distance data and the cell feature similarity data to obtain cell feature comprehensive score data;

[0096] Specifically, the cell fingerprint feature distance data and the cell feature similarity data can be imported into the Python environment. These two distance data respectively represent the distance and similarity information between cells. Different weights are assigned to the cell fingerprint feature distance data and the cell feature similarity data. According to the experimental design and data characteristics, the weight of the cell fingerprint feature distance data is selected as 0.4, and the weight of the cell feature similarity data is 0.6, because the cosine similarity can better reflect the inherent similarity between cell features in some cases. Through the weighted average formula comprehensive score = (0.4 × Euclidean distance) + (0.6 × cosine similarity), calculate the comprehensive score between each pair of cells. Finally, a cell feature comprehensive score matrix is obtained, where each element represents the comprehensive similarity score between two cells.

[0097] Step S34: Perform distribution feature statistics on the cell feature comprehensive score data to obtain cell population score distribution parameters;

[0098] Specifically, the cell feature comprehensive score matrix can be imported into the Python environment. Use the statistical functions in the SciPy library to calculate the distribution parameters of the scores. Specifically, calculate the mean, standard deviation, skewness, and kurtosis statistical parameters of the comprehensive scores. For example, by calculating the mean and standard deviation, it is found that the average value of the comprehensive scores is 0.75, and the standard deviation is 0.12, which indicates that the comprehensive similarity scores of most cell pairs are concentrated between 0.63 and 0.87. Through statistical analysis, finally, the cell population score distribution parameters are obtained.

[0099] Step S35: Perform cross-modal heterogeneous distribution feature similarity correction on the cell feature comprehensive score data based on the cell population score distribution parameters to obtain cell similarity score data;

[0100] Specifically, the comprehensive score data can be analyzed based on the cell population score distribution parameters (such as mean, standard deviation, skewness, etc.). It is found that the comprehensive score data has a certain skewed distribution, that is, the scores of some cell pairs are significantly higher or lower than the average level. The Z-score normalization method is adopted to convert each score into a normalized value relative to the population distribution. The specific formula is Z-score = score - mean. Through the above operations, the comprehensive score data is converted into normalized data centered on 0 with a standard deviation of 1. Finally, the corrected cell similarity score data is obtained, and these data are closer to the normal distribution in terms of distribution, reducing the bias caused by heterogeneous distribution.

[0101] Step S36: Perform hierarchical clustering on the cell molecular feature fingerprint data according to the cell similarity score data to obtain the initial cell clustering result; perform dynamic threshold optimization on the initial cell clustering result to obtain optimized cell classification data.

[0102] Specifically, for the detailed implementation process of this embodiment, please refer to the sub-steps of step S36.

[0103] The present invention combines two evaluation methods, Euclidean distance and cosine similarity, to measure the similarity between cell molecular feature fingerprints from different perspectives. The Euclidean distance can quantify the absolute difference between feature vectors, while the cosine similarity can measure the directional similarity between feature vectors. This can more comprehensively reflect the similarity between cells and avoid the bias caused by a single similarity measure. Through weighted fusion, corresponding weights can be assigned according to the importance of different evaluation methods, so as to obtain a more reliable similarity score. Through distribution feature statistics and cross-modal heterogeneous distribution feature similarity correction. This can adjust the distribution differences between different modal data, making the similarity score more accurate and consistent, and reducing the classification error caused by data distribution heterogeneity. Through classification, a lineage tree of cells can be generated to reveal the hierarchical relationship between cells. This can not only identify the main classifications of cells but also further subdivide cell subpopulations. Through dynamic threshold optimization, the clustering threshold can be dynamically adjusted according to the feature distribution of the cell population, which can avoid classification errors caused by a fixed threshold and improve the adaptability and flexibility of cell classification.

[0104] Preferably, step S36 includes the following steps:

[0105] Step S361: Construct a hierarchical tree for the cell molecular feature fingerprint data according to the cell similarity score data to obtain an initial cell lineage clustering tree;

[0106] Specifically, the cell similarity score data can be imported into the Python environment. Using the linkage function in the SciPy library, hierarchical clustering analysis is performed based on the similarity scores between cells to construct a cell lineage tree. In hierarchical clustering, the "Ward" method is selected, which optimizes the clustering result by minimizing the within-class variance. The specific code is linkage_matrix = linkage(similarity_matrix, method='ward'), where similarity_matrix is the cell similarity score matrix. Use the dendrogram function of the Matplotlib library to visualize the hierarchical clustering result and generate a tree diagram of the cell lineage tree. By adjusting the parameters of the tree diagram, such as the color and label of the branches, the similarity relationship between cells can be intuitively observed.

[0107] Step S362: Prune and optimize the initial cell lineage clustering tree to obtain an optimized clustering tree of cell subpopulations, and merge the nodes of the optimized clustering tree of cell subpopulations to obtain an initial cell clustering result;

[0108] Specifically, the initial cell lineage clustering tree can be analyzed first to determine the nodes that need to be pruned. By setting the clustering threshold, for example, choosing the distance threshold as 1.5, the fcluster function is used to extract the clustering labels from the hierarchical clustering results. The code is cluster_labels = fcluster(linkage_matrix, t = 1.5, criterion = 'distance'). This step divides the cells into multiple subpopulations. Next, calculate the within-cluster density of each subpopulation. The formula is within-cluster density = number of cells within the subpopulation / average similarity of cell pairs within the subpopulation. For subpopulations with a within-cluster density lower than 0.6, merge them with adjacent subpopulations. Finally, obtain the optimized clustering tree of cell subpopulations and generate the initial cell clustering results.

[0109] Step S363: Calculate the within-cluster density of the initial cell clustering results to obtain the within-cluster density distribution data of cell subpopulations;

[0110] Specifically, the initial cell clustering results can be imported into the Python environment. These results contain the clustering labels to which each cell belongs. For each cell subpopulation (i.e., each cluster), calculate the within-cluster density. The formula for within-cluster density is within-cluster density = number of cells within the subpopulation / average similarity of cell pairs within the subpopulation. Extract the similarity score of cell pairs within each subpopulation and calculate the average of these scores. Then, count the number of cells within each subpopulation and calculate the within-cluster density according to the above formula. For example, for a subpopulation containing 100 cells, the average of all similarity scores between these 100 cells is 0.85. Therefore, the within-cluster density of this subpopulation is 0.85. Through the above operations, the within-cluster density distribution data of each cell subpopulation is obtained.

[0111] Step S364: Adjust the threshold of the within-cluster density distribution data of cell subpopulations to obtain the dynamic threshold parameter of cell subpopulations;

[0112] Specifically, the density distribution data within the clusters can be analyzed first to determine appropriate dynamic threshold parameters. By plotting the histogram of the density within the clusters, the distribution of the density within the clusters can be observed. For example, the density within the clusters of most subpopulations is concentrated between 0.7 and 0.9, but there are also a few subpopulations with a density within the clusters lower than 0.6. The median of the density within the clusters is selected as the initial threshold and fine-tuned according to the distribution. The initial threshold is set to 0.75, and the subpopulations with a density within the clusters lower than 0.75 are re-evaluated according to the density distribution of the subpopulations. For example, for a subpopulation with a density within the clusters lower than 0.75, the distribution of the cell similarity within it is further examined. If an obvious subgroup structure is found, the threshold is appropriately lowered to include these subgroups; conversely, if the cell similarity within the subpopulation is low and the distribution is relatively dispersed, the threshold is raised to exclude these subpopulations. Through the above operations, the dynamic threshold parameters of the cell subpopulations are finally obtained.

[0113] Step S365: Based on the dynamic threshold parameters of the cell subpopulations, perform subpopulation differentiation adjustment on the initial cell clustering result to obtain optimized cell classification data.

[0114] Specifically, the initial cell clustering result and the dynamic threshold parameters can be imported into the Python environment. For each cell subpopulation, the cells within the subpopulation are re-evaluated and adjusted according to the dynamic threshold parameters. For example, for a subpopulation with a density within the clusters lower than the dynamic threshold (such as 0.75), the distribution of the cell similarity within it is further analyzed. If an obvious subgroup structure is found within the subpopulation, that is, the similarity between some cells is significantly higher than that of other cells, these cells are divided into new subpopulations. The specific operation is to calculate the similarity scores of cell pairs within the subpopulation and use the AgglomerativeClustering function in the Scikit-learn library to perform hierarchical clustering on the cells within the subpopulation, setting the n_clusters parameter to 2 or 3 to identify the subgroup structure within the subpopulation. For a subpopulation with a density within the clusters higher than the dynamic threshold, it is further confirmed whether the cell similarity within it is evenly distributed. If it is found that the similarity of some cells is significantly lower than that of other cells, these cells are separated from the current subpopulation and re-assigned to other subpopulations or form new subpopulations. Through the above operations, the optimized cell classification data is finally obtained.

[0115] By constructing an initial clustering tree for cell lineages, the hierarchical relationships between cells can be visually displayed, revealing the lineage structure of cell populations. Through pruning optimization and node merging, redundant information can be removed and the structure of the clustering tree can be optimized, making the division of cell subpopulations more reasonable and clear. By calculating the density distribution data within clusters, the tightness and distribution characteristics of cell subpopulations can be quantified. The calculation of the density within clusters provides a data basis for dynamic threshold adjustment, enabling the clustering process to be optimized according to the actual distribution characteristics of cell subpopulations. By adjusting the clustering threshold based on the density distribution data within clusters, it is possible to more flexibly adapt to the characteristic differences of different cell subpopulations and avoid classification errors caused by fixed thresholds. Through subgroup differentiation adjustment, the cell classification results can be further optimized.

[0116] Preferably, step S4 includes the following steps:

[0117] Step S41: Extract the topological structure features of the cell phenotype relationship network to obtain cell network topological feature data;

[0118] Specifically, the cell phenotype relationship network can be imported into the Cytoscape software, and the "Network Analyzer" plugin of Cytoscape can be used to extract the topological structure features of the network, including the degree, betweenness centrality, closeness centrality, and clustering coefficient of the nodes. For example, the degree of a node represents the number of edges directly connected to that node, reflecting the connectivity of the cell phenotype in the functional network. The betweenness centrality represents the mediating role of the node in the network, and the higher the value, the more important the node is in the network. Further, the igraph package in the R language is used to quantitatively analyze these features. Through the degree, betweenness, closeness, and clustering_coefficient functions of the igraph package, the topological feature values of each node are calculated, and these values are exported as cell network topological feature data. These data are stored in tabular form, with each row representing a cell phenotype node and its corresponding topological feature values.

[0119] Step S42: Quantify the class distribution characteristics of the optimized cell classification data to obtain cell class distribution feature data;

[0120] Specifically, the optimized cell classification data can be imported into the Python environment, and the Pandas library can be used to process the data and count the distribution of each cell category. Calculate the number of cells, proportion, and distribution differences between categories for each category. For example, by using the value_counts function, the distribution of the number of cells in each category can be obtained. Then, statistical functions in the SciPy library, such as chi2_contingency, are used to perform a chi-square test on the category distribution to evaluate whether the distribution differences between different categories are statistically significant. In addition, the average intra-cluster density of each category is calculated through the groupby and mean functions to obtain the intra-cluster density distribution characteristics of each category. Finally, the cell category distribution characteristic data are obtained, which are stored in tabular form, with each row representing a cell category and its corresponding distribution characteristic values.

[0121] Step S43: Perform non-linear feature mapping on the cell network topology feature data to obtain a cell topology mapping matrix, and perform high-dimensional feature mapping on the cell category distribution feature data to obtain a cell classification mapping matrix;

[0122] Specifically, for the cell network topology feature data, the t-SNE (t-Distributed Stochastic Neighbor Embedding) algorithm can be used for non-linear feature mapping. Import the topology feature data into the Python environment and use the TSNE class in the Scikit-learn library for mapping. During the mapping process, set the parameters: n_components = 2 (map the data to a two-dimensional space), perplexity = 30 (perplexity, representing the number of neighbors for each point), learning_rate = 200 (learning rate). Through the above operations, a cell topology mapping matrix is obtained, where each cell phenotype node is mapped to a point in the two-dimensional space. Then, for the cell category distribution feature data, PCA (Principal Component Analysis) is used for high-dimensional feature mapping. Use the PCA class in the Scikit-learn library for mapping. During the mapping process, set the parameter: n_components = 3 (extract the first three principal components). Through the above steps, a cell classification mapping matrix is obtained, where each cell category is represented as a combination of the three principal components. Finally, two mapping matrices are obtained: the cell topology mapping matrix and the cell classification mapping matrix.

[0123] Step S44: Correlate the cell topology mapping matrix and the cell classification mapping matrix to obtain a cell feature correlation matrix;

[0124] Specifically, the cell topology mapping matrix and the cell classification mapping matrix can be imported into the Python environment. The two matrices are merged into a data frame according to the unique identifier of the cell phenotype node (such as the node ID). For example, assume that the column names of the cell topology mapping matrix are ['node_id', 'x', 'y'], and the column names of the cell classification mapping matrix are ['node_id', 'pc1', 'pc2', 'pc3']. Through the merge function, the two data frames are merged according to node_id to generate a comprehensive data frame containing the cell topology coordinates and classification features. The merged data is matrixized using the NumPy library to generate the cell feature correlation matrix. Finally, the cell feature correlation matrix is stored in the form of a two-dimensional array, where each row represents a cell phenotype node and its corresponding topology and classification feature vectors.

[0125] Step S45: Evaluate the sparsity of the cell feature correlation matrix to obtain the cell correlation sparsity parameter;

[0126] Specifically, the cell feature correlation matrix can be imported into the Python environment. Calculate the proportion of non-zero elements in the matrix. Specifically, use the count_nonzero function in the NumPy library to count the number of non-zero elements in the matrix, and use the size attribute to obtain the total number of elements in the matrix. By calculating the sparsity of non-zero element proportion = total number of elementsnumber of non-zero elements, the cell correlation sparsity parameter is obtained. For example, assume that the size of the matrix is 1000×1000 and the number of non-zero elements is 10000, then the sparsity is 0.01. Through the above operations, the cell correlation sparsity parameter is obtained.

[0127] Step S46: Sparsify the cell feature correlation matrix based on the cell correlation sparsity parameter to obtain the cell sparse correlation matrix;

[0128] Specifically, the threshold for sparsification processing can be determined according to the sparsity parameter (such as the sparsity is 0.01). In order to remove the unimportant elements in the matrix, select to retain the elements in the matrix whose absolute value is greater than a certain threshold. For example, assume that the threshold is 0.1. Use the where function and boolean indexing in the NumPy library to set the elements in the matrix whose absolute value is less than 0.1 to 0. The specific operation is: matrix[np.abs(matrix)<0.1] = 0. Through the above operations, the unimportant elements in the correlation matrix are set to zero, realizing the sparsification of the matrix. Then, use the csr_matrix function in the SciPy library to convert the sparsified matrix into the compressed sparse row format. Finally, the cell sparse correlation matrix is obtained, and this matrix is stored in the sparse format.

[0129] Step S47: Encode the data of the cell sparse association matrix to obtain cell-encoded association data, and extract metadata from the cell-encoded association data to obtain cell index feature data;

[0130] Specifically, the cell sparse association matrix can be imported into the Python environment. Using the binary encoding method, convert the non-zero elements in the matrix into binary form. For example, assume that the non-zero element value in the sparse matrix is 1.5. Convert it to the binary representation 1.5 -> 1 (integer part) and 0.5 -> 1 (fractional part). Through the above operations, each element in the sparse matrix is converted into binary encoding to generate cell-encoded association data. Then, extract the key feature information of each cell phenotype node, such as node ID, topological coordinates, and classification features. Use the DataFrame and describe functions in the Pandas library to extract the statistical information of each feature (such as mean, standard deviation, maximum value, minimum value), and store these statistical information as metadata. Finally, obtain the cell index feature data, which are stored in tabular form, and each row represents a cell phenotype node and its corresponding index feature information.

[0131] Step S48: Structure the multi-source cell omics data into a tree based on the cell index feature data to obtain cell index structure data, and perform balance optimization on the cell index structure data to obtain a cell data retrieval tree.

[0132] Specifically, the cell index feature data can be imported into the Python environment. Use the DecisionTreeClassifier or DecisionTreeRegressor class in the Scikit-learn library to construct a decision tree based on the cell index feature data. When constructing the decision tree, select appropriate parameters, such as setting max_depth (the maximum depth of the tree) to 10 and min_samples_split (the minimum number of samples required to split an internal node) to 2. Through the above operations, obtain the cell index structure data, where each node represents a cell phenotype node and its corresponding feature information. Then, use the BalancedTree tool to perform balance optimization on the decision tree. By adjusting the structure of the tree, ensure that the depth difference of each branch does not exceed 1. Finally, obtain the cell data retrieval tree.

[0133] Through the extraction of topological structure features, the present invention can transform complex network relationships into quantitative features, making the association structure between cells clearer. Through the quantification of category distribution features, the distribution of cell classification can be quantified, providing a basis for the statistical analysis and visualization of cell classification. Through non-linear feature mapping and high-dimensional feature mapping, the topological features of the cell network and the category distribution features of cells can be mapped into the same feature space, further integrating the network structure and classification information of cells. Through matrix association, the network structure and classification information of cells can be systematically integrated. This provides a basis for the compressed storage and index construction of cell data, making data storage and retrieval more efficient. Through sparsity evaluation and sparsification processing, redundant information can be removed, reducing the occupancy of data storage space. Through data encoding and extraction of metadata, the storage format of data can be further optimized, making the storage of cell data more compact and efficient. Through tree structuring, an efficient cell data retrieval tree can be constructed. Through balance optimization, the performance of the retrieval tree can be further improved, reducing the retrieval time. The retrieval tree after balance optimization can distribute data more evenly, avoiding low retrieval efficiency caused by uneven data distribution.

[0134] Preferably, step S5 includes the following steps:

[0135] Step S51: Obtain newly added biological cell data, and construct a feature extraction template according to the cell feature digital matrix to obtain cell feature extraction rules;

[0136] Specifically, newly added biological cell data can be collected, including new cell transcriptome, proteome, and metabolome data. The cell feature digital matrix is a matrix containing standardized feature values, with each row representing a cell sample and each column representing a feature. Analyze the structure and feature distribution of this matrix to determine the extraction rules for key features. For example, for transcriptome data, genes with an expression level higher than a certain threshold (such as transcripts per million TPM > 1) are selected as features; for proteome data, proteins with an abundance higher than a certain threshold (such as normalized intensity > 0.05) are selected as features; for metabolome data, metabolites with a concentration higher than a certain threshold (such as normalized concentration > 0.1) are selected as features. Through the above operations, a set of cell feature extraction rules are constructed, which define which features are important.

[0137] Step S52: Extract features from the newly added biological cell data according to the cell feature extraction rules to obtain a cell newly added feature vector, and calculate the similarity between the cell newly added feature vector and the cell feature digital matrix to obtain cell feature similarity data;

[0138] Specifically, the newly added biological cell data can be imported into the Python environment. According to the feature extraction rules, the corresponding features are extracted from the newly added data. For example, for the newly added transcriptome data, the Pandas library is used to extract gene features with an expression level higher than TPM > 1; for proteome data, protein features with an abundance higher than the normalized intensity > 0.05 are extracted; for metabolome data, metabolite features with a concentration higher than the normalized concentration > 0.1 are extracted. Through the above operations, the new cell feature vectors are obtained. Next, the similarity between these new feature vectors and the existing cell feature digital matrix is calculated. The dot function in the NumPy library is used to calculate the dot product between vectors, and the norm function is used to calculate the norm of the vectors, thereby calculating the cosine similarity. The specific formula is: Through the above operations, cell feature similarity data are obtained, which reflect the similarity between the newly added cells and the existing cells.

[0139] Step S53: Perform cell population nearest neighbor feature quantization based on the cell feature similarity data to obtain cell neighborhood distribution data;

[0140] Specifically, the cell feature similarity data can be imported into the Python environment. Through the KNeighborsClassifier class in the Scikit-learn library, the parameter n_neighbors = 5 is set, indicating that each cell considers its 5 nearest neighbors. Then, the fit method is used to train the existing cell data, and the predict method is used to classify and predict the newly added cell data. During the prediction process, the KNN algorithm will find the 5 nearest neighbor cells according to the similarity between the newly added cells and the existing cells, and vote according to the categories of these neighbor cells to determine the category of the newly added cells. Through the above operations, cell neighborhood distribution data are obtained, which reflect the distribution of the newly added cells among their nearest neighbor cells.

[0141] Step S54: Perform Bayesian probability evaluation on the cell neighborhood distribution data to obtain cell class probability distribution data, and perform cell population classification confidence evaluation based on the cell class probability distribution data to obtain classification matching score data;

[0142] Specifically, the cell neighborhood distribution data can be imported into the Python environment. Calculate the posterior probability of each newly added cell belonging to each category. Suppose there are 3 cells belonging to category C and 2 cells belonging to category D among the nearest neighbor cells of a certain newly added cell. According to the category distribution of these neighbor cells, calculate the posterior probabilities of this newly added cell belonging to category C and category D. For example, use the formula P(category C|data) = P(data)P(data|category C)×P(category A), where P(data|category C) is the conditional probability, P(category C) is the prior probability, and P(data) is the total probability of the data. Through the above operations, the cell category probability distribution data is obtained. Next, calculate the classification matching score for each newly added cell. For example, if the highest value in the posterior probability of a newly added cell is 0.75, this value can be used as the classification matching score, indicating that the confidence of this cell belonging to its predicted category is 75%. Through the above operations, the classification matching score data is obtained.

[0143] Step S55: Extract the cell population threshold features based on the classification matching score data to obtain the cell dynamic threshold parameters, and update and judge the classification matching score data according to the cell dynamic threshold parameters to obtain the classification update trigger signal;

[0144] Specifically, the classification matching score data can be imported into the Python environment. Analyze the distribution of the classification matching scores. By calculating the mean and standard deviation of the scores, a dynamic threshold is determined. For example, assume that the mean of the classification matching scores is 0.7 and the standard deviation is 0.1, then the dynamic threshold can be set as the mean minus one standard deviation, that is, 0.6. This threshold is used to judge whether the classification confidence of the newly added cell is high enough. Next, update and judge the classification matching score data according to the dynamic threshold parameters. For each newly added cell, if its classification matching score is higher than the dynamic threshold (such as 0.6), it is considered that the classification result of this cell is reliable and does not need to be updated; if it is lower than this threshold, it is considered that the classification result is not reliable enough and a classification update needs to be triggered. For example, if the classification matching score of a newly added cell is 0.55, which is lower than the dynamic threshold of 0.6, a classification update trigger signal will be generated. Through the above steps, the cell dynamic threshold parameters can be obtained, and the classification update trigger signal is generated based on these parameters.

[0145] Step S56: Dynamically reconstruct the cell data retrieval tree according to the classification update trigger signal to obtain the dynamically updated cell classification data.

[0146] Specifically, the classification update trigger signal can be imported into the Python environment. For each newly added cell that triggers an update, its classification result needs to be re-evaluated, and the cell data retrieval tree is dynamically reconstructed. Specifically, the DecisionTreeClassifier class in the Scikit-learn library is used to retrain the decision tree model, incorporating the newly added cell data into the training set and adjusting the structure of the tree according to the latest data distribution. During the reconstruction process, the Balanced Tree tool is used to optimize the balance of the decision tree, ensuring that the depth difference of each branch does not exceed 1. For example, assume that after the classification update of the newly added cell, it is reassigned to a new category, the node information of the decision tree is updated, and the branch structure of the tree is readjusted. Through the above operations, the dynamically updated cell classification data is obtained.

[0147] By obtaining newly added biological cell data and constructing feature extraction rules, the present invention can incorporate new data into the classification system in real time, ensuring the timeliness of cell data classification. This can quickly respond to the addition of new data and avoid inaccurate classification caused by data lag. Through feature extraction and similarity calculation, the cell classification results can be continuously optimized. Through similarity calculation and quantization of the nearest neighbor features of cell populations, the similarity between newly added cell data and existing cell data can be accurately evaluated, and the subtle differences between cells can be identified. By analyzing the cell neighborhood distribution data, the local characteristics and distribution laws of cell populations can be revealed. Through Bayesian probability evaluation, the probability distribution of cell category attribution can be quantified. Through classification confidence evaluation, the classification results with high confidence can be further screened to improve the accuracy and credibility of classification. Through extraction of cell population threshold features and dynamic adjustment of threshold parameters, the classification threshold can be dynamically adjusted according to the actual distribution characteristics of cell populations, avoiding classification errors caused by fixed thresholds. Through update judgment, the classification results that need to be updated can be intelligently identified and the corresponding update operations can be triggered. Through dynamic reconstruction, the cell classification data can be updated in real time, ensuring the accuracy and timeliness of the retrieval tree. The optimized retrieval tree can support fast data retrieval and update operations, significantly improving the response speed and performance of the system.

[0148] Preferably, the present invention also provides a classification storage system for biological cell data, which is used to execute the classification storage method of biological cell data as described above. The classification storage system for biological cell data includes:

[0149] A data preprocessing module, which is used to obtain multi-source cell omics data; perform data standardization on the multi-source cell omics data to obtain standard cell feature data; perform feature matrix conversion on the standard cell feature data to obtain a cell feature digital matrix; and construct a molecular fingerprint of biological cells based on the cell feature digital matrix to obtain cell molecular feature fingerprint data.

[0150] A classification system construction module, which is used to hierarchically classify standard cell feature data to obtain a cell phenotype classification system; hierarchically label the standard cell feature data according to the cell phenotype classification system to obtain cell phenotype labeled data; perform ontology mapping on the cell phenotype labeled data to obtain a cell phenotype relationship network;

[0151] A data classification module, which is used to calculate the similarity of cell molecular feature fingerprint data to obtain cell similarity score data; hierarchically cluster the cell molecular feature fingerprint data according to the cell similarity score data to obtain an initial cell clustering result; perform dynamic threshold optimization on the initial cell clustering result to obtain optimized cell classification data;

[0152] An index construction module, which is used to construct an association matrix for the cell phenotype relationship network and the optimized cell classification data to obtain a cell data association matrix; perform compressed storage on the cell data association matrix to obtain cell compressed association data; construct an index structure according to the cell compressed association data to obtain a cell data retrieval tree;

[0153] A data dynamic update module, which is used to obtain newly added biological cell data; calculate the classification matching degree of the newly added biological cell data according to the cell feature digital matrix to obtain classification matching score data; perform update judgment according to the classification matching score data to obtain a classification update trigger signal; recursively optimize the cell data retrieval tree according to the classification update trigger signal to obtain dynamically updated cell classification data.

[0154] Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, it is intended to cover all changes falling within the meaning and scope of the equivalent elements of the application documents within the present invention.

[0155] The above are only specific embodiments of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather to the widest scope consistent with the principles and novel features invented herein.

Claims

1. A classification storage method for biological cell data, characterized in that: The following steps are involved: Step S1: acquiring multi-source cell omics data; performing data standardization on the multi-source cell omics data to obtain standard cell feature data; performing feature matrix conversion on the standard cell feature data to obtain a cell feature digital matrix; constructing a molecular fingerprint of biological cells based on the cell feature digital matrix to obtain cell molecular feature fingerprint data; Step S2: hierarchically classify the standard cell feature data to obtain a cell phenotype classification system; hierarchically annotate the standard cell feature data according to the cell phenotype classification system to obtain cell phenotype annotated data; and ontology map the cell phenotype annotated data to obtain a cell phenotype relationship network; Step S3: performing similarity calculation on the cell molecular feature fingerprint data to obtain cell similarity score data; performing hierarchical clustering on the cell molecular feature fingerprint data according to the cell similarity score data to obtain initial cell clustering results; performing dynamic threshold optimization on the initial cell clustering results to obtain optimized cell classification data; Step S4: constructing an association matrix between the cell phenotype relationship network and the optimized cell classification data to obtain a cell data association matrix; compressing and storing the cell data association matrix to obtain cell compressed association data; constructing an index structure based on the cell compressed association data to obtain a cell data retrieval tree; Step S5: Acquire newly added biological cell data; calculate the classification matching degree of the newly added biological cell data according to the cell feature digital matrix to obtain classification matching score data; make an update judgment according to the classification matching score data to obtain a classification update trigger signal; recursively optimize the cell data retrieval tree according to the classification update trigger signal to obtain dynamically updated cell classification data.

2. The method for classifying and storing biological cell data according to claim 1, characterized in that: Step S1 includes the following steps: Step S11: acquiring multi-source cell omics data, wherein the multi-source cell omics data includes original cell transcriptome sequencing data, original cell proteome mass spectrometry data and original cell metabolome data; Step S12: performing quality control filtering on the original cell transcriptome sequencing data to obtain cell transcriptome sequencing data; Step S13: Eliminate background noise from the original cell proteome mass spectrometry data to obtain cell proteome data; Step S14: peak calibration is performed on the original cell metabolome data to obtain cell metabolome data, and the cell transcriptome sequencing data, cell proteome data and cell metabolome data are recorded as standard cell feature data; Step S15: Perform feature matrix conversion on the standard cell feature data to obtain a cell feature digital matrix, and construct a molecular fingerprint of the biological cell based on the cell feature digital matrix to obtain cell molecular feature fingerprint data.

3. The method for classifying and storing biological cell data according to claim 2, characterized in that: Step S15 includes the following steps: Step S151: quantifying the expression of the cell transcriptome sequencing data to obtain cell gene expression data, and filling the missing values ​​of the cell gene expression data to obtain complete cell gene expression data; Step S152: performing peptide identification on the cell proteome data to obtain cell protein expression data, and performing outlier correction on the cell protein expression data to obtain cell corrected protein expression data; Step S153: performing metabolite quantification on the cell metabolome data to obtain cell metabolite expression data, and performing data standardization on the cell metabolite expression data to obtain standard cell metabolite expression data; Step S154: performing dimensionality reduction on the complete gene expression data of the cell to obtain a cell gene feature vector, and performing feature selection on the corrected protein expression data of the cell to obtain a cell protein feature vector; Step S155: performing pattern extraction on the standard cell metabolite expression data to obtain a cell metabolite feature vector, and performing feature fusion on the cell gene feature vector, the cell protein feature vector and the cell metabolite feature vector to obtain a cell feature digital matrix; Step S156: performing similarity calculation on the cell feature digital matrix to obtain a cell population feature similarity matrix; Step S157: constructing molecular fingerprints of biological cells according to the cell population feature similarity matrix to obtain cell molecular feature fingerprint data.

4. The method for classifying and storing biological cell data according to claim 1, characterized in that: Step S2 includes the following steps: Step S21: annotating the cell transcriptome sequencing data to obtain a cell gene ontology thesaurus; Step S22: performing protein identification on the cell proteome data to obtain a cell proteome spectrum, and performing functional annotation extraction on the cell proteome spectrum to obtain a cell protein function vocabulary; Step S23: performing metabolic spectrum clustering on the cell metabolome data to obtain a cell metabolome molecular map, and extracting pathway information from the cell metabolome molecular map to obtain a cell metabolic pathway vocabulary; Step S24: performing hierarchical association clustering on the cell gene ontology vocabulary to obtain a cell gene ontology hierarchical tree, and performing functional stratification on the cell protein function vocabulary to obtain a cell protein function hierarchical tree; Step S25: performing pathway classification on the cell metabolic pathway vocabulary to obtain a cell metabolic pathway hierarchical tree, and integrating the cell gene ontology hierarchical tree and the cell metabolic pathway hierarchical tree to obtain a cell molecular function classification system; Step S26: integrating the cell molecular function classification system and the cell metabolic pathway hierarchy tree to obtain a cell phenotype classification system; Step S27: hierarchically annotating the standard cell feature data according to the cell phenotype classification system to obtain cell phenotype annotated data; and ontology mapping the cell phenotype annotated data to obtain a cell phenotype relationship network.

5. The method for classifying and storing biological cell data according to claim 4, characterized in that: Step S27 includes the following steps: Step S271: annotating the gene function of the standard cell feature data according to the cell phenotype classification system to obtain cell gene function annotation data; Step S272: annotating the protein function of the standard cell feature data according to the cell phenotype classification system to obtain cell protein function annotation data; Step S273: annotating the standard cell characteristic data with metabolic pathways according to the cell phenotype classification system to obtain cell metabolic pathway annotation data; Step S274: Integrate the cell gene function annotation data and the cell protein function annotation data to obtain molecular function annotation data; Step S275: performing association mapping on the cell gene function annotation data and the cell metabolic pathway annotation data to obtain cell phenotype annotation data; Step S276: ontology mapping is performed on the cell phenotype annotation data to obtain a cell phenotype relationship matrix, and a network topology structure is constructed according to the cell phenotype relationship matrix to obtain a cell phenotype relationship network.

6. The method for classifying and storing biological cell data according to claim 1, characterized in that: Step S3 includes the following steps: Step S31: extracting features from the cell molecular feature fingerprint data to obtain a cell fingerprint feature description vector, and performing Euclidean distance evaluation on the cell fingerprint feature description vector to obtain cell fingerprint feature distance data; Step S32: performing cosine similarity evaluation on the cell fingerprint feature description vector to obtain cell feature similarity data; Step S33: performing weighted fusion on the cell fingerprint feature distance data and the cell feature similarity data to obtain cell feature comprehensive score data; Step S34: performing distribution characteristic statistics on the cell characteristic comprehensive score data to obtain the cell population score distribution parameter; Step S35: performing cross-modal heterogeneous distribution feature similarity correction on the cell feature comprehensive score data based on the cell population score distribution parameter to obtain cell similarity score data; Step S36: hierarchically clustering the cell molecular feature fingerprint data according to the cell similarity score data to obtain an initial cell clustering result; and dynamically optimizing the initial cell clustering result to obtain optimized cell classification data.

7. The method for classifying and storing biological cell data according to claim 6, characterized in that: Step S36 includes the following steps: Step S361: constructing a hierarchical tree of the cell molecular feature fingerprint data according to the cell similarity score data to obtain an initial cell lineage clustering tree; Step S362: pruning and optimizing the initial cell lineage clustering tree to obtain a cell subpopulation optimized clustering tree, and merging nodes of the cell subpopulation optimized clustering tree to obtain an initial cell clustering result; Step S363: performing intra-cluster density calculation on the initial cell clustering results to obtain cell subpopulation cluster density distribution data; Step S364: adjusting the threshold of the cell subpopulation cluster density distribution data to obtain a cell subpopulation dynamic threshold parameter; Step S365: performing subpopulation differentiation adjustment on the initial cell clustering result based on the cell subpopulation dynamic threshold parameter to obtain optimized cell classification data.

8. The method for classifying and storing biological cell data according to claim 1, characterized in that: Step S4 includes the following steps: Step S41: extracting topological structure features of the cell phenotype relationship network to obtain cell network topological feature data; Step S42: quantifying the category distribution characteristics of the optimized cell classification data to obtain cell category distribution characteristic data; Step S43: performing nonlinear feature mapping on the cell network topology feature data to obtain a cell topology mapping matrix, and performing high-dimensional feature mapping on the cell category distribution feature data to obtain a cell classification mapping matrix; Step S44: Correlating the cell topology mapping matrix and the cell classification mapping matrix to obtain a cell feature correlation matrix; Step S45: performing sparsity evaluation on the cell feature association matrix to obtain a cell association sparsity parameter; Step S46: performing sparse processing on the cell feature association matrix based on the cell association sparsity parameter to obtain a cell sparse association matrix; Step S47: performing data encoding on the cell sparse association matrix to obtain cell encoding association data, and performing metadata extraction on the cell encoding association data to obtain cell index feature data; Step S48: Based on the cell index feature data, the multi-source cell omics data is structured into a tree to obtain cell index structure data, and the cell index structure data is balanced and optimized to obtain a cell data retrieval tree.

9. The method for classifying and storing biological cell data according to claim 1, characterized in that: Step S5 includes the following steps: Step S51: acquiring newly added biological cell data, and constructing a feature extraction template according to the cell feature digital matrix to obtain a cell feature extraction rule; Step S52: extracting features from the newly added biological cell data according to the cell feature extraction rule to obtain a newly added cell feature vector, and calculating the similarity between the newly added cell feature vector and the cell feature digital matrix to obtain cell feature similarity data; Step S53: quantifying the nearest neighbor features of the cell population according to the cell feature similarity data to obtain cell neighborhood distribution data; Step S54: performing Bayesian probability evaluation on the cell neighborhood distribution data to obtain cell category probability distribution data, and performing cell population classification confidence evaluation based on the cell category probability distribution data to obtain classification matching score data; Step S55: extracting cell population threshold features according to the classification matching score data to obtain a cell dynamic threshold parameter, and updating and judging the classification matching score data according to the cell dynamic threshold parameter to obtain a classification update trigger signal; Step S56: dynamically reconstructing the cell data retrieval tree according to the classification update trigger signal to obtain dynamically updated cell classification data.

10. A classification storage system for biological cell data, characterized in that: For executing the classification and storage method of biological cell data as claimed in claim 1, the classification and storage system of biological cell data comprises: The data preprocessing module is used to obtain multi-source cell omics data; standardize the multi-source cell omics data to obtain standard cell feature data; perform feature matrix conversion on the standard cell feature data to obtain a cell feature digital matrix; construct molecular fingerprints of biological cells based on the cell feature digital matrix to obtain cell molecular feature fingerprint data; The classification system construction module is used to hierarchically classify the standard cell feature data to obtain a cell phenotype classification system; hierarchically annotate the standard cell feature data according to the cell phenotype classification system to obtain cell phenotype annotation data; and ontology map the cell phenotype annotation data to obtain a cell phenotype relationship network; The data classification module is used to calculate the similarity of the cell molecular feature fingerprint data to obtain the cell similarity score data; hierarchically cluster the cell molecular feature fingerprint data according to the cell similarity score data to obtain the initial cell clustering result; dynamically optimize the initial cell clustering result by threshold value to obtain the optimized cell classification data; An index construction module is used to construct an association matrix between the cell phenotype relationship network and the optimized cell classification data to obtain a cell data association matrix; compress and store the cell data association matrix to obtain cell compressed association data; and construct an index structure based on the cell compressed association data to obtain a cell data retrieval tree; The data dynamic update module is used to obtain newly added biological cell data; calculate the classification matching degree of the newly added biological cell data according to the cell feature digital matrix to obtain classification matching score data; make update judgments according to the classification matching score data to obtain a classification update trigger signal; recursively optimize the cell data retrieval tree according to the classification update trigger signal to obtain dynamically updated cell classification data.

Citation Information

Cited By

  • Single cell transcriptome cell type automatic annotation method and device based on consensus voting, electronic equipment and storage medium

    CN120954520A

  • Method, device, electronic equipment and storage medium for single-cell transcriptome cell type automatic annotation based on consensus voting

    CN120954520B

  • Analysis method and device based on multi-version cell annotation results and electronic equipment

    CN121662173A