Disease subtype classification method and device, electronic equipment and computer program product
By constructing and fusing a similarity matrix of multi-omics data, and combining it with multi-stage clustering analysis, the problem of unstable disease subtype identification in existing technologies was solved, and more stable and high-resolution disease subtype classification was achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-01
- Publication Date
- 2026-04-07
AI Technical Summary
Existing disease classification methods rely on clinical symptoms and artificially defined features, which are difficult to stabilize and have insufficient resolution, making them unable to effectively identify unknown subtypes of complex diseases.
By acquiring multi-omics data, constructing a similarity matrix and fusing it, and combining multi-stage clustering analysis of spectral clustering, unsupervised clustering and consensus clustering, disease subtypes can be identified.
It achieves objective and stable identification of disease subtypes, improves the stability and resolution of typing, and avoids the influence of noise and parameter selection in traditional methods.
Smart Images

Figure CN121808429A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of medical data processing technology, and in particular relates to a disease subtype classification method, a disease subtype classification device, an electronic device, and a computer program product. Background Technology
[0002] In the field of complex disease research, the identification of disease subtypes is crucial for revealing disease heterogeneity, understanding differences in pathological mechanisms, and improving precision diagnosis and treatment capabilities. However, existing disease classification methods still mainly rely on clinical symptoms, imaging observations, or a limited number of biomarkers, and depend on artificially set characteristics or prior assumptions. They have limited ability to capture unknown structures or potential subtypes, and are prone to problems such as classification instability and / or insufficient classification resolution when processing complex medical data. Summary of the Invention
[0003] This application provides a disease subtype classification method, a disease subtype classification device, an electronic device, and a computer program product, which can help improve the stability and resolution of classification.
[0004] Firstly, this application provides a method for classifying disease subtypes, including: Obtain at least two types of omics data from multiple patient samples with the target disease; Similarity matrices were constructed for each type of omics data. The similarity matrices were used to describe the feature similarity between any two patient samples under the corresponding omics data. The similarity matrices from various omics datasets are fused to obtain a fused similarity matrix; Multi-stage clustering analysis based on the fusion similarity matrix is used to obtain the subtyping results for the target disease. The multi-stage clustering analysis includes at least two of the following clustering methods: spectral clustering, unsupervised clustering, and consensus clustering.
[0005] Secondly, this application provides a disease subtype classification device, comprising: The acquisition module is used to acquire at least two types of omics data from multiple patient samples suffering from the target disease; The module is used to construct similarity matrices for various types of omics data. The similarity matrix is used to describe the feature similarity between any two patient samples under the corresponding omics data. The fusion module is used to fuse similarity matrices from various omics datasets to obtain a fused similarity matrix. The clustering module is used to perform multi-stage clustering analysis based on the fusion similarity matrix to obtain the subtyping results for the target disease. The multi-stage clustering analysis includes at least two of the following clustering processing methods: spectral clustering, unsupervised clustering, and consensus clustering.
[0006] Thirdly, this application provides an electronic device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the method described in the first aspect.
[0007] Fourthly, this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the method described in the first aspect above.
[0008] Fifthly, this application provides a computer program product comprising a computer program that, when executed by one or more processors, implements the steps of the method described in the first aspect.
[0009] The advantages of this application compared to existing technologies are as follows: This application's solution achieves objective and stable subtype identification of patient groups for target diseases by constructing a similarity matrix for multi-omics data, matrix fusion based on the similarity matrix, and conducting multi-stage clustering analysis. On one hand, by acquiring at least two types of omics data and constructing corresponding similarity matrices for each, the representation of differences between patients based on different data types can be expressed in a unified similarity form. This ensures that different omics data are comparable within the same mathematical space, providing a structured foundation for the fusion of similarity matrices. On the other hand, the multi-stage clustering strategy overcomes the technical shortcomings of traditional single-clustering methods, which are susceptible to noise, initialization issues, and parameter selection, leading to unstable results. In summary, this application's solution does not rely on prior assumptions, making the disease subtype identification process more comprehensive and robust, and improving the stability and resolution of the final subtype classification.
[0010] It is understood that the beneficial effects of the second to fifth aspects mentioned above can be found in the relevant descriptions in the first aspect mentioned above, and will not be repeated here. Attached Figure Description
[0011] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0012] Figure 1 This is a schematic diagram illustrating the implementation process of the disease subtype classification method provided in the embodiments of this application; Figure 2 This is a structural block diagram of the disease subtype classification device provided in the embodiments of this application; Figure 3 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0013] The embodiments of the technical solution of this application will now be described in detail with reference to the accompanying drawings. These embodiments are only used to more clearly illustrate the technical solution of this application and are therefore merely examples, and should not be used to limit the scope of protection of this application.
[0014] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the application; the terms “comprising” and “having”, and any variations thereof, in the specification, claims, and foregoing description of the drawings are intended to cover non-exclusive inclusion.
[0015] In the description of the embodiments of this application, technical terms such as "first" and "second" are used only to distinguish different objects and should not be construed as indicating or implying relative importance or implicitly indicating the number, specific order, or primary and secondary relationship of the indicated technical features.
[0016] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0017] In the description of the embodiments in this application, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " in this document generally indicates that the preceding and following related objects have an "or" relationship.
[0018] In the description of the embodiments of this application, the term "multiple" refers to two or more (including two), unless otherwise expressly and specifically defined.
[0019] This application proposes a method for classifying disease subtypes. This method can be applied to electronic devices with data processing capabilities, such as servers, personal computers, or smartphones; however, this application does not limit the specific type of such electronic device. Please refer to... Figure 1 , Figure 1 The implementation process of the disease subtype classification method applied to this electronic device is presented, and detailed below: Step 101: Obtain at least two types of omics data from multiple patient samples with the target disease.
[0020] The electronic device can first acquire raw data from multiple patient samples suffering from the target disease, specifically omics data. Each patient sample typically corresponds to one individual patient. Omics data refers to high-dimensional medical data acquired from different levels, including but not limited to: clinical laboratory data, imaging data, genetic data, protein data, metabolic data, and microbial data, etc., which are not limited in this embodiment. To improve the accuracy of subtype classification, the electronic device can focus on at least two types of omics data, that is, select and acquire at least two types of data from the available omics data to provide a comprehensive description of the multidimensional biological state of the patient samples.
[0021] In some examples, each type of omics data includes multiple features. For instance, clinical laboratory data may include features of multiple indicators such as complete blood count and biochemical parameters; imaging data may include volume features of 109 brain regions from structural magnetic resonance imaging (MRI); genetic data may include features of 20,000 single nucleotide polymorphism (SNP) sites from whole-exome sequencing; protein data may include 1,000 protein abundance features from plasma proteomics; metabolic data may include 300 small molecule compound features from serum metabolomics; and microbial data may include 500 bacterial community abundance features from 16S rRNA sequencing.
[0022] In some examples, the target disease can be a complex neurodegenerative disease that can be classified, such as Parkinson's disease or Alzheimer's disease, without obvious structural changes on imaging. This application does not limit this.
[0023] In some examples, electronic devices can obtain omics data through various methods such as file import or database query. This application does not limit the method of obtaining omics data.
[0024] This step provides a multi-source raw data foundation for subsequent feature similarity calculations and fusion, ensuring that subsequent processing can integrate information from different biological levels to obtain stable and interpretable disease subtypes.
[0025] Step 102: Construct similarity matrices for each type of omics data.
[0026] For each type of omics data, the electronic device can construct a similarity matrix describing the degree of similarity between any two patient samples; that is, the similarity matrix is specifically a two-dimensional matrix formed by the feature similarity between any two patient samples as its matrix elements. It can be understood that the rows and columns of the similarity matrix each correspond to a specific patient sample, and its matrix elements are used to quantify the degree of similarity between two patient samples. When constructing the similarity matrix, the electronic device can first calculate the feature distance between patient samples, and then calculate the feature similarity between patient samples based on this feature distance, thereby constructing the similarity matrix.
[0027] This step can convert the various features of omics data into a matrix form that reflects the relationship structure between patient samples, so that different types of omics data can be fused and compared in a unified form of similarity.
[0028] Step 103: Fuse the similarity matrices of various omics data to obtain a fused similarity matrix.
[0029] Electronic devices can unify and fuse similarity matrices of various omics data to obtain a fused similarity matrix that comprehensively reflects the structural information in different omics data. In some examples, the various similarity matrices can be fused with the same weight; alternatively, based on a preset weighting strategy, the corresponding weights of each similarity matrix can be calculated separately before fusion. This application does not limit this approach.
[0030] This step integrates data from multiple heterogeneous omics sources into a unified structure, thereby overcoming the limitations of single-omics information and providing a more robust data foundation for subsequent clustering processing.
[0031] Step 104: Perform multi-stage clustering analysis based on the fusion similarity matrix to obtain the subtyping results for the target disease.
[0032] Electronic devices can perform multi-stage clustering analysis on patient samples based on a fusion similarity matrix. This multi-stage clustering analysis includes at least two of the following clustering methods: spectral clustering, unsupervised clustering, and consensus clustering.
[0033] Spectral clustering refers to constructing a Laplacian matrix using a fused similarity matrix, and obtaining the spectral embedding space through eigenvalue decomposition, enabling samples to exhibit good cluster separability in this low-dimensional space. In some examples, the constructed Laplacian matrix can be a normalized Laplacian matrix, an unnormalized general Laplacian matrix, a normalized Laplacian matrix from a random walk, a symmetric normalized Laplacian matrix, etc., and this application does not limit this.
[0034] Unsupervised clustering refers to classifying data in the spectral embedding space using unsupervised learning methods such as hierarchical clustering, LDA (Latent Dirichlet Allocation), or adaptive K-means clustering. This application does not limit the specific clustering algorithm used in this unsupervised clustering process. In some examples, adaptive K-means clustering can be used, which performs K-means clustering based on multiple candidate K values and selects the optimal number of clusters based on the silhouette coefficient.
[0035] Consensus clustering refers to repeatedly sampling and clustering each candidate K value to construct a consensus matrix, and evaluating the clustering stability based on its cumulative distribution function to determine the optimal number of clusters.
[0036] Ultimately, the electronic device can obtain a classification result for the target disease, that is, a classification label for different disease subtypes under the target disease. It is important to note, however, that this classification result is not a diagnostic result for each patient sample, but only a subtype classification result obtained after analyzing the patient samples.
[0037] This step can improve the accuracy and stability of subtype classification through a multi-stage clustering strategy, making the obtained disease subtypes have good interpretability and reproducibility.
[0038] In some embodiments, the electronic device may use a Gaussian kernel to perform distance-similarity transformation, thereby constructing a similarity matrix. Based on this, step 102 may include: Step 1021: For each type of omics data, calculate the feature distance between any two patient samples under the omics data.
[0039] For each type of omics data, electronic devices can calculate the feature distance between any two patient samples. The feature distance is a mathematical quantity used to measure the degree of difference between two patient samples in the omics data space; generally speaking, the larger the value, the less similar the corresponding two patient samples are.
[0040] In this embodiment, different distance measurement methods can be selected according to the characteristics of different omics data; that is, the distance type between features is determined based on the omics type of the omics data. For example, Euclidean distance can be used to calculate the distance between features for clinical laboratory data, cosine distance can be used for imaging data, Hamming distance can be used for gene data, correlation distance can be used for protein data, Manhattan distance can be used for metabolic data, and Brectis distance can be used for microbial data. It can be understood that selecting different distance measurement methods for different omics data can maximize the preservation of biological differences between different omics, thereby ensuring that the constructed similarity matrix can truly reflect the biological similarity between patient samples under the corresponding omics data. Of course, the same distance metric can be used for each type of omics data; or, other possible distance metrics can be introduced according to actual application needs. For example, geometric distance (including but not limited to Minkowski distance, Mahalanobis distance and Chebyshev distance) can be used to calculate the distance between features for clinical test data. The embodiments of this application do not limit the types of distances that can be used for various types of omics data, and can make adaptive choices according to actual conditions.
[0041] This step can obtain the degree of difference between different patient samples under a certain type of omics data, laying the foundation for the subsequent construction of a similarity matrix that can reflect patient similarity; and since the distance calculation method between different omics data is matched with their biological characteristics, it can avoid measurement bias caused by improper selection of distance type, thereby improving the overall reliability of this scheme.
[0042] Step 1022: Determine the median distance based on the distances between all features obtained from the omics data.
[0043] After obtaining the inter-feature distances for all patient samples in a given omics dataset, statistical analysis of the overall distribution of these distances is needed to determine the median of the inter-feature distances. This median distance reflects the typical scale of the distance distribution under this type of omics dataset. Since the value ranges and distribution patterns of different omics datasets vary significantly, this median distance can serve as a scale parameter when constructing a similarity matrix, enabling comparisons of distances across different omics datasets at the same scale.
[0044] This step enables adaptive evaluation of distance distributions across different omics datasets, providing a reasonable scale basis for converting feature distances into similarity. Furthermore, this step avoids using fixed parameters, which can mitigate potential scaling issues and thus help improve the stability and generalization performance of the subsequent similarity matrix.
[0045] Step 1023: Construct a similarity matrix for the omics data based on the feature distance between any two patient samples and the median distance.
[0046] After obtaining the distances and median distances between all features of a class of omics data, a similarity matrix corresponding to that class of omics data can be constructed based on these two values. In some examples, the similarity matrices corresponding to different omics data can be represented as follows: S_clinical = exp(-D 2 euclidean / (2σ1 2 )) S_imaging = exp(-D 2 cosine / (2σ2 2 )) S_genetic = exp(-D 2 hamming / (2σ3 2 )) S_protein = exp(-D 2 correlation / (2σ4 2 )) S_metabolic = exp(-D 2 manhattan / (2σ5 2 )) S_microbiome = exp(-D 2 bray_curtis / (2σ6 2 )) Where S_clinical is the similarity matrix for clinical laboratory data, S_imaging is the similarity matrix for imaging data, S_genetic is the similarity matrix for gene data, S_protein is the similarity matrix for protein data, S_metabolic is the similarity matrix for metabolic data, and S_microbiome is the similarity matrix for microbial data; D is the distance between features, with subscripts euclidean for Euclidean distance, cosine for cosine distance, hamming for Hanning distance, correlation for correlation distance, Manhattan for Manhattan distance, and bray_curtis for Bray-Curtis distance; σ i This represents the adaptive bandwidth parameter under different omics data, and its specific value is the median distance under the corresponding omics data.
[0047] This step can convert distances, which are difficult to use directly for fusion analysis, into similarities that can be used to represent network structures, thereby constructing a similarity matrix.
[0048] In some embodiments, the electronic device can dynamically adjust the weight contribution of each omics data to avoid the curse of dimensionality; based on this, step 103 may include: Step 1031: Standardize the similarity matrices of various omics data into probability matrices.
[0049] In multi-omics data, similarity matrices generated from different omics types may differ in numerical range, distribution characteristics, and dimensional scale. Directly fusing similarity matrices from various omics datasets could lead to a situation where similarity matrices with larger numerical ranges from certain omics datasets receive unreasonable weightings during the fusion process. Therefore, electronic devices can first standardize each similarity matrix into a probability matrix by normalizing all elements in each row of the similarity matrix so that the sum of the elements in that row is 1; that is, the final probability matrix has a row sum of 1. This processing method allows similarity matrices from various omics datasets to be converted to a uniform numerical scale, making subsequent matrix fusion more fair and stable.
[0050] This step eliminates the dimensional differences between similarity matrices of different omics data by transforming the probability matrix, so that the subsequent fusion process can be carried out on a uniform scale, thereby improving the stability and rationality of the fusion results.
[0051] Step 1032: Calculate the weights of each type of omics data based on the probability matrices corresponding to each type of omics data.
[0052] The electronic device can then determine the importance of different omics data for disease subtyping. Generally, directly assigning weights manually is highly subjective, while simple averaging cannot reflect the differences in the contribution of different omics to subtyping. Based on this, the embodiments of this application can use the largest eigenvalue (denoted as λ1) as the basis for determining the weights. Specifically, the largest eigenvalue refers to the eigenvalue with the largest value obtained after eigenvalue decomposition of the probability matrix.
[0053] In some examples, the weights of various omics data can be calculated using the following formula: w i = λ 1(S_i) / Σλ 1(S_ j) Where S represents omics data, w represents weights, and λ1 represents the largest eigenvalue in the probability matrix; specifically, w i λ1 represents the weight of the i-th omics data. S_i ) represents the largest eigenvalue in the probability matrix corresponding to the i-th type of omics data, Σλ1(S_ j The expression represents the summation of the largest eigenvalues in the probability matrices corresponding to each type of omics data. In other words, the denominator is the sum of the largest eigenvalues in the probability matrices corresponding to each type of omics data, and the numerator is the largest eigenvalue in the probability matrix corresponding to a particular type of omics data. This allows the calculation of the weight of that type of omics data, which represents the contribution of the similarity matrix of that omics data to subsequent fusion.
[0054] This step dynamically calculates the weights of various omics data by using the maximum eigenvalue, so that omics data with high contribution account for a higher proportion in the fusion, thereby improving the ability of the fusion results to express the actual disease classification structure. This can avoid the bias caused by human intervention.
[0055] Step 1033: Perform matrix fusion based on the similarity matrices corresponding to various omics data and the weights of various omics data to obtain a fused similarity matrix.
[0056] After obtaining the weights of various omics data, the electronic device can perform a weighted summation with the corresponding similarity matrix. Furthermore, the electronic device can incorporate an identity matrix during the fusion process to strengthen the diagonal elements of the matrix, ensuring high stability of each patient sample's similarity to itself.
[0057] In some examples, electronic devices may specifically employ the following fusion formula for matrix fusion: S fused = Σ(w i × S i ) +α× I Among them, S fused S represents the fusion similarity matrix; i w represents the similarity matrix corresponding to the i-th type of omics data; i α represents the weight of the i-th omics data; α represents the preset regularization parameter, which is usually a non-zero positive number and can be used to control the degree of diagonal enhancement and improve the numerical stability of the fused matrix; I represents the identity matrix.
[0058] This step improves the stability and robustness of the fusion matrix by weighted fusion of similarity matrices from different omics and by adding an identity matrix and regularization coefficients during the fusion process. This allows the fusion matrix to more accurately reflect the comprehensive similarity relationships of multidimensional biological information between different patient samples, providing a reliable foundation for subsequent clustering.
[0059] In some embodiments, the multi-stage clustering analysis proposed in step 104 may include spectral clustering and unsupervised clustering, wherein the unsupervised clustering may specifically be adaptive K-means clustering. The specific implementation flow of step 104 can then be as follows: Step 1041: Generate a Laplacian matrix for spectral clustering based on the fusion similarity matrix, and perform eigenvalue decomposition on the Laplacian matrix to obtain the spectral embedding representation.
[0060] First, the electronic device can convert the fused similarity matrix into a Laplacian matrix, which can be used in spectral clustering to represent the connection structure between different patient samples. In some examples, this conversion process can be specifically described as follows: First, construct a degree matrix D based on the fused similarity matrix, where D(i,i) = Σ j S fused (i, j) is used to express the connection strength of the i-th sample; then, the Laplacian matrix L = D - S can be calculated based on the fused similarity matrix and degree matrix. fused .
[0061] Then, the electronic device can perform eigenvalue decomposition on the Laplacian matrix to calculate the smallest n+1 eigenvalues and the corresponding eigenvectors. Eigenvalue decomposition can be used to extract key directions describing the graph structure, allowing high-dimensional similarity relationships to be expressed in a low-dimensional space. Thus, the smallest n+1 eigenvalues and their corresponding eigenvectors can be calculated. The smallest n+1 eigenvalues refer to the first n+1 eigenvalues selected from all eigenvalues in ascending order, where n represents the dimension of the subsequent spectral embedding.
[0062] Finally, the electronic device can map the eigenvectors corresponding to the n eigenvalues (excluding the first eigenvalue) out of the n+1 eigenvalues to an n-dimensional spectral embedding space. The electronic device can construct this n-dimensional spectral embedding space for subsequent spectral embedding. This operation is a crucial step in spectral clustering, transforming the original similarity matrix into a lower-dimensional and more easily separable geometric space. Specifically, in spectral theory, the eigenvector corresponding to the first minimum eigenvalue is usually a constant vector, mapping all samples to the same value and lacking the ability to distinguish samples; therefore, it needs to be discarded. The remaining n eigenvectors can serve as new data representations, which, after being mapped to the n-dimensional spectral embedding space, form the sample points in the spectral embedding space.
[0063] In some embodiments, the electronic device may first perform L2 normalization on the remaining n feature vectors to ensure that each sample can be located on a unit sphere in the spectral embedding space, thereby creating an ideal geometric structure for subsequent K-means clustering.
[0064] Step 1042: Under the spectral embedding representation, K-means clustering is performed based on each preset candidate K value, and the first candidate optimal solution is determined based on the preset clustering evaluation index.
[0065] The electronic device can perform K-means clustering on multiple candidate cluster numbers K (i.e., candidate K values) in the spectral embedding space based on the sample representation after spectral embedding. The candidate K values are typically set within a preset range, such as 2 to 10. In some examples, considering the sensitivity of the K-means algorithm to initialization, the electronic device can perform 10 independent algorithm runs for each candidate K value and employ a K-means++ initialization strategy to improve convergence quality.
[0066] Based on the clustering results corresponding to each candidate K value, the electronic device can calculate the average silhouette coefficient for each candidate K value and determine the candidate K value with the highest average silhouette coefficient as the first candidate optimal solution. This average silhouette coefficient can be understood as being used to evaluate the clustering quality under different candidate K values, comprehensively considering both intra-cluster compactness and inter-cluster separation, with a value range of [-1, 1]. Generally speaking, the larger the average silhouette coefficient, the better the clustering effect. Therefore, the electronic device can determine the candidate K value with the highest average silhouette coefficient as the first candidate optimal solution.
[0067] In some embodiments, the electronic device may also simultaneously record its cluster center and sample label at the candidate K value with the highest average profile coefficient, thereby providing a basis for subsequent verification.
[0068] Step 1043: Determine the subtyping result for the target disease based on the first candidate optimal solution.
[0069] In some examples, it is assumed that the mean silhouette coefficient reaches its maximum value when K=4, indicating that the omics data of the patient samples are naturally divided into four relatively close and well separated clusters, thus corresponding to four different subtypes.
[0070] In some embodiments, to further verify the stability and reliability of the clustering results, multi-stage clustering analysis may also include consensus clustering processing; based on this, before step 1043, the specific implementation process of step 104 may further include: Step 1044: Repeatedly cluster each candidate K value to construct a consensus matrix corresponding to each candidate K value based on the clustering results.
[0071] To evaluate the stability of clustering results with different candidate K values, the electronic device can perform repeated clustering experiments using repeated sampling. The number of repeated clustering experiments corresponding to each candidate K value can be 100 or other values, which is not limited in this embodiment. In each repeated clustering experiment, the electronic device only randomly selects a subset of patient samples and features of omics data to participate in clustering. For example, in at least two classes of omics data from multiple patient samples, only 80% of the sample subset is randomly selected for repeated clustering experiments. In addition, the electronic device can also perform random sampling along the feature dimension, thereby increasing the randomness of repeated clustering experiments. The above methods can evaluate the robustness of clustering results to data perturbation and simulate noise and missing data in real clinical data.
[0072] For each candidate K value, based on the clustering results of repeated clustering, the electronic device can obtain the frequency at which any two patient samples are assigned to the same cluster, thereby constructing the corresponding consensus matrix C. Here, C(i,j) represents the frequency at which patient samples i and j are assigned to the same cluster in all repeated clustering experiments for that candidate K value. It can be understood that the diagonal elements of the consensus matrix are always 1, that is, C(i,i) = 1, because each patient sample is completely identical to itself; the off-diagonal elements of the consensus matrix take values between [0,1]. Ideally, the element values of the consensus matrix corresponding to patient samples belonging to the same true subtype should be close to 1, while the element values of the consensus matrix corresponding to patient samples belonging to different subtypes should be close to 0, and the consensus matrix exhibits a clear block diagonal structure.
[0073] Step 1045: Evaluate the clustering stability of each candidate K value based on the consensus matrix, and determine the second candidate optimal solution based on the clustering stability.
[0074] As described earlier, the diagonal elements of the consensus matrix are always 1, and the consensus matrix exhibits a clear block diagonal structure. Based on these diagonal elements, the consensus matrix is divided into an upper triangular part (above the diagonal elements) and a lower triangular part (below the diagonal elements). Electronic devices can extract the element values of the upper triangular part of the consensus matrix for each candidate K value, thereby calculating the Cumulative Distribution Function (CDF) of that candidate K value. It can be understood that the CDF can be used to quantitatively reflect the concentration and dispersion of the consensus matrix, thereby assessing the clustering stability of the corresponding candidate K values.
[0075] The electronic device can thus analyze the cumulative distribution function (CDF) curves corresponding to each candidate K value to determine the second candidate optimal solution among all candidate K values based on the analysis results. Generally, for the optimal number of clusters, the CDF curve should exhibit a bimodal distribution characteristic, specifically: the consensus values of a large number of sample pairs are clustered around 0 (indicating between different clusters) and around 1 (indicating within the same cluster). Based on this, the electronic device can analyze the CDF curves corresponding to each candidate K value and identify the candidate K value with the most stable cluster structure as the second candidate optimal solution by comparing the area change rate under the CDF curves of different candidate K values. Among them, the candidate K value with the most stable cluster structure generally has the smallest area change rate under its CDF curve.
[0076] Accordingly, step 1043 can be specifically manifested as: determining the classification result for the target disease based on the first candidate optimal solution and the second candidate optimal solution. Specifically, the electronic device can first determine the difference between the first candidate optimal solution and the second candidate optimal solution. If the difference between the two is less than a preset difference threshold, it means that the difference between the two is not significant, and the electronic device can consider the second candidate optimal solution as the target optimal solution (because the second candidate optimal solution has better stability) to obtain the classification result for the target disease; conversely, if the difference between the two is greater than or equal to the difference threshold, it means that the difference between the two is significant, and there may be some problems with clustering at this time. The electronic device can introduce other clustering processing methods for further clustering processing. No limitation is made on other clustering processing methods here.
[0077] In some embodiments, prior to step 102, the electronic device may further perform preprocessing operations on the multi-omics data, and the disease subtype classification method provided in this application embodiment may further include: A1 imputes missing values in various omics data to obtain imputed omics data.
[0078] In clinical and multi-omics data processing, data gaps are common, such as missing blood test results, insufficient gene sequencing coverage, and loss of metabolomics signals. In this embodiment, since the data to be processed is multi-omics data, the imputation processing method to be used can be determined based on the omics type of the omics data.
[0079] In some examples, median imputation can be used for clinical laboratory data; for other omics data (i.e., gene data, protein data, metabolic data, and microbial data), KNN imputation can be used, where K can be 5 or other values, without limitation here.
[0080] Based on this, electronic devices may also consider removing features with a missing rate greater than a preset missing rate threshold to avoid excessive data loss affecting subsequent processing. The preset missing rate threshold can be 30% or other values, which are not limited here.
[0081] A2 standardizes the imputed omics data to obtain standardized omics data.
[0082] Standardization refers to the process of using mathematical transformations to make features with different dimensions and / or distributions comparable. Because omics data have complex feature sources, electronic devices can select different standardization methods based on the statistical type of the omics data.
[0083] In some examples, for continuous features in omics data, electronic devices can perform Z-score normalization; for count features in omics data, electronic devices can perform normalization based on the Centered Log Ratio Transformation (CLR); and for dichotomous features in omics data, electronic devices can maintain the original format.
[0084] A3 filters standardized omics data of various types to obtain filtered omics data.
[0085] Filtering can be used to remove redundant data from various omics datasets. In this embodiment, filtering may include, but is not limited to, at least one of the following: variance-based filtering, correlation-based filtering, mutual information-based filtering, relief algorithm-based filtering, recursive feature elimination-based filtering, univariate / multivariate logistic regression-based filtering, and Lasso (Least Absolute Shrinkage and Selection Operator) regression-based filtering. As examples only, variance-based filtering can retain features with the top 80% variance (or other variance thresholds); correlation-based filtering can remove redundant features with correlation coefficients > 0.95 (or other correlation coefficient thresholds), etc. The processing methods for each possible filtering method will not be elaborated here.
[0086] In some embodiments, after obtaining the subtyping results for the target disease, the electronic device can also perform multi-level statistical analysis and biological interpretation; then the disease subtyping classification method provided in this application embodiment may further include: B1, based on the typing results, performs differential analysis on the omics data of different subtypes to obtain differentially expressed molecules for each subtype.
[0087] The typing results have classified multiple patient samples into several subtypes of the target disease. Electronic devices can group various omics data (such as clinical data, imaging data, gene data, protein data, metabolic data, and microbial data) according to the subtype to which the patient samples belong, and perform differential analysis between the groups. Differential analysis can employ methods such as t-tests; the specific test method can be selected according to different data distributions, and this application does not limit this approach. Through differential analysis, genes, proteins, or metabolites that are particularly abnormal in a certain subtype can be screened from a large number of features in multi-omics data; these screened features are the differentially expressed molecules.
[0088] B2 maps differentially expressed molecules to predefined biological pathways for enrichment analysis.
[0089] Biological pathways refer to functional modules composed of interrelated genes, proteins, or metabolites, such as dopamine pathways, cholinergic pathways, or inflammatory response pathways. To identify subtypes of functional abnormalities from differentially expressed molecules, electronic devices can compare these molecules with pre-constructed biological pathways.
[0090] Based on this, pathway enrichment analysis is performed, and statistical methods (including but not limited to hypergeometric tests) are used to assess whether molecules in a certain biological pathway appear significantly more frequently in the list of differentially expressed molecules, in order to determine whether the biological pathway is significantly affected in a certain subtype.
[0091] B3, based on the enrichment results, quantifies the pathway function to generate explanations of the biological mechanisms of each subtype.
[0092] Electronic devices can quantify pathway function based on enrichment results, that is, quantify the functional impairment or activity changes of biological pathways to obtain clear numerical indicators. Based on the quantitative results, biological mechanism explanations for each disease subtype can be generated. For example, in a subtype, abnormalities in multiple genes of the dopamine pathway can be explained as impaired dopamine functional modules in that subtype; if inflammatory response pathways are significantly enriched, it can be explained as a prominent inflammatory response in that subtype.
[0093] Based on this, electronic devices can also integrate relevant published medical literature to form causal relationships between manifestations and biological mechanisms, thereby further clarifying the rationality of subtype classification.
[0094] As can be seen from the above, the embodiments of this application achieve objective and stable identification of potential subtype structures in complex medical data through the construction of a similarity matrix of multi-omics data, weighted fusion based on eigenvalues, and multi-stage clustering analysis. First, this method can fully utilize the complementary information of different omics data, significantly improving the discriminative power and robustness of subtype classification. Second, the fused similarity matrix, under spectral embedding processing, can form a clearer low-dimensional representation, making subsequent clustering results more objective and reliable. Furthermore, by combining the silhouette coefficient evaluation in K-means clustering with consensus clustering verification, the sensitivity of clustering results to parameter selection and data perturbation can be effectively avoided, thereby obtaining more credible optimal classification results. Based on this, the embodiments of this application further perform differential analysis and biological pathway interpretation after obtaining the classification results, thereby revealing the corresponding pathological mechanism characteristics for different subtypes.
[0095] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0096] Corresponding to the disease subtype classification method provided above, this application also provides a disease subtype classification device. Please refer to... Figure 2 The disease subtype classification device 2 in this embodiment includes: The acquisition module 201 is used to acquire at least two types of omics data from multiple patient samples suffering from the target disease; Module 202 is used to construct similarity matrices for various types of omics data. The similarity matrix is used to describe the feature similarity between any two patient samples under the corresponding omics data. The fusion module 203 is used to fuse similarity matrices from various omics data to obtain a fused similarity matrix. Clustering module 204 is used to perform multi-stage clustering analysis based on the fusion similarity matrix to obtain the subtyping results for the target disease. The multi-stage clustering analysis includes at least two of the following clustering processing methods: spectral clustering, unsupervised clustering, and consensus clustering.
[0097] In some embodiments, the construction module 202 includes: The distance calculation unit is used to calculate the inter-feature distance between any two patient samples under each type of omics data, wherein the distance type of the inter-feature distance is determined based on the omics type of the omics data; The median determination unit is used to determine the median distance among all feature distances calculated in omics data. The matrix construction unit is used to construct a similarity matrix for any two patient samples in the omics data based on the feature distance between them and the median distance.
[0098] In some embodiments, omics data includes at least two of the following categories: clinical laboratory data, imaging data, gene data, protein data, metabolic data, and microbial data; wherein the feature distances corresponding to clinical laboratory data are Euclidean distances, the feature distances corresponding to imaging data are cosine distances, the feature distances corresponding to gene data are Hamming distances, the feature distances corresponding to protein data are correlation distances, the feature distances corresponding to metabolic data are Manhattan distances, and the feature distances corresponding to microbial data are Brectis distances.
[0099] In some embodiments, the fusion module 203 includes: Matrix standardization unit is used to standardize the similarity matrix of various omics data into probability matrices respectively; The weight calculation unit is used to calculate the weight of each type of omics data based on the probability matrix corresponding to each type of omics data. The weight is used to represent the contribution of the corresponding omics data in the fusion. The matrix fusion unit is used to perform matrix fusion based on the similarity matrix corresponding to various omics data and the weights of various omics data to obtain a fused similarity matrix.
[0100] In some embodiments, where the multi-stage clustering analysis includes spectral clustering and unsupervised clustering, the unsupervised clustering is specifically adaptive K-means clustering. The clustering module 204 includes: The spectral clustering unit is used to generate a Laplacian matrix for spectral clustering based on the fusion similarity matrix, and to perform eigenvalue decomposition on the Laplacian matrix to obtain the spectral embedding representation. K-means clustering unit is used to perform K-means clustering based on each preset candidate K value under spectral embedding representation, and to determine the first candidate optimal solution based on the preset clustering evaluation index. The typing result determination unit is used to determine the typing result for the target disease based on the first candidate optimal solution.
[0101] In some embodiments, where the multi-stage clustering analysis further includes consensus clustering processing, the clustering module 204 further includes: The consensus matrix construction unit is used to perform repeated clustering on each candidate K value to construct the consensus matrix corresponding to each candidate K value based on the clustering results. The consensus matrix is used to represent the frequency with which any two patient samples are assigned to the same cluster in the repeated clustering of candidate K values. The evaluation unit is used to evaluate the clustering stability of each candidate K value based on the consensus matrix, and to determine the second candidate optimal solution based on the clustering stability. Accordingly, the typing result determination unit is specifically used to determine the typing result for the target disease based on the first candidate optimal solution and the second candidate optimal solution.
[0102] In some embodiments, the disease subtype classification device 2 further includes: The imputation processing module is used to impute missing values in various types of omics data to obtain imputed omics data. The imputation processing method is determined based on the omics type of the omics data. The standardization processing module is used to standardize the imputed omics data of various types to obtain standardized omics data of various types. The standardization processing method is determined based on the data statistical type of each feature data under the omics data. The filtering module is used to filter standardized omics data of various types to obtain filtered omics data. The filtering process is used to remove redundant data from the various omics data.
[0103] In some embodiments, the disease subtype classification device 2 further includes: The first analysis module is used to perform differential analysis on omics data of different subtypes based on the typing results, and to obtain differential molecules of each subtype; The second analysis module is used to map differentially expressed molecules to preset biological pathways to perform enrichment analysis. An explanation generation module is used to quantitatively assess pathway function based on enrichment results, and to generate biological mechanism explanations for each subtype based on the assessment.
[0104] This application's embodiments achieve objective and stable identification of potential subtype structures in complex medical data through the construction of a similarity matrix from multi-omics data, weighted fusion based on eigenvalues, and multi-stage clustering analysis. First, this method fully utilizes complementary information from different omics data, significantly improving the discriminative power and robustness of subtype classification. Second, the fused similarity matrix, under spectral embedding processing, can form a clearer low-dimensional representation, making subsequent clustering results more objective and reliable. Furthermore, by combining silhouette coefficient evaluation in K-means clustering with consensus clustering verification, the sensitivity of clustering results to parameter selection and data perturbation can be effectively avoided, thus obtaining more credible optimal classification results. Based on this, this application's embodiments further perform differential analysis and biological pathway interpretation after obtaining the classification results, thereby revealing the corresponding pathological mechanism characteristics for different subtypes.
[0105] Corresponding to the disease subtype classification method provided above, this application also provides an electronic device. Please refer to... Figure 3 The electronic device 3 in this application embodiment includes: a memory 301, and one or more processors 302. Figure 3 (Only one is shown in the image) and a computer program stored in memory 301 and executable on the processor. Specifically, the processor 302 performs the following steps by running the aforementioned computer program stored in memory 301: Obtain at least two types of omics data from multiple patient samples with the target disease; Similarity matrices were constructed for each type of omics data. The similarity matrices were used to describe the feature similarity between any two patient samples under the corresponding omics data. The similarity matrices from various omics datasets are fused to obtain a fused similarity matrix; Multi-stage clustering analysis based on the fusion similarity matrix is used to obtain the subtyping results for the target disease. The multi-stage clustering analysis includes at least two of the following clustering methods: spectral clustering, unsupervised clustering, and consensus clustering.
[0106] Assuming the above is the first possible implementation, then the second possible implementation provided based on the first possible implementation is as follows: For each type of omics data, the feature distance between any two patient samples is calculated under the omics data, wherein the distance type of the feature distance is determined based on the omics type of the omics data; Determine the median distance among all feature distances calculated using omics data; Based on the feature distance and median distance between any two patient samples in the omics data, a similarity matrix is constructed in the omics data.
[0107] In the third possible implementation method provided based on the first possible implementation method described above, the similarity matrices of various omics data are fused to obtain a fused similarity matrix, including: The similarity matrices for various omics data are standardized into probability matrices respectively; Based on the probability matrices corresponding to various omics data, the weights of various omics data are calculated respectively. The weights are used to represent the contribution of the corresponding omics data in the fusion. A fused similarity matrix is obtained by fusing the similarity matrices corresponding to various omics data and the weights of various omics data.
[0108] In the fourth possible implementation provided based on the first possible implementation described above, where the multi-stage clustering analysis includes spectral clustering and unsupervised clustering, the unsupervised clustering can specifically be adaptive K-means clustering. Multi-stage clustering analysis is performed based on a fused similarity matrix to obtain a classification result for the target disease, including: A Laplacian matrix for spectral clustering is generated based on the fusion similarity matrix, and eigenvalue decomposition is performed on the Laplacian matrix to obtain the spectral embedding representation. Under the spectral embedding representation, K-means clustering is performed based on each preset candidate K value, and the first candidate optimal solution is determined based on the preset clustering evaluation index. The subtyping result for the target disease is determined based on the first candidate optimal solution.
[0109] In the fifth possible implementation provided based on the fourth possible implementation described above, where the multi-stage clustering analysis further includes consensus clustering processing, before determining the subtyping result for the target disease based on the first candidate optimal solution, the processor 302 further performs the following steps when running the aforementioned computer program stored in the memory 301: Each candidate K value is repeatedly clustered to construct a consensus matrix corresponding to each candidate K value based on the clustering results. The consensus matrix is used to represent the frequency with which any two patient samples are assigned to the same cluster in the repeated clustering of candidate K values. The clustering stability of each candidate K value is evaluated based on the consensus matrix, and the second candidate optimal solution is determined based on the clustering stability. Accordingly, the subtyping result for the target disease is determined based on the first candidate optimal solution, including: The classification results for the target disease are determined based on the first and second candidate optimal solutions.
[0110] In a sixth possible implementation provided based on the first, second, third, fourth, or fifth possible implementations described above, before constructing similarity matrices for various types of omics data, the processor 302 further performs the following steps when running the computer program stored in the memory 301: Missing values in various omics data are imputed to obtain imputed omics data. The imputation method is determined based on the omics type of the omics data. The imputed omics data are standardized to obtain standardized omics data. The standardization process is determined based on the statistical data type of each feature data under the omics data. Standardized omics data of various types are filtered to obtain filtered omics data of various types. The filtering process is used to remove redundant data in various omics data.
[0111] In a seventh possible implementation provided based on the first, second, third, fourth, or fifth possible implementations described above, after obtaining the subtyping results for the target disease through multi-stage clustering analysis based on the fusion similarity matrix, the processor 302 further performs the following steps when running the computer program stored in the memory 301: Based on the typing results, differential analysis was performed on the omics data of different subtypes to obtain the differentially expressed molecules of each subtype; Differential molecules are mapped to predefined biological pathways for enrichment analysis; The enrichment results are used to quantitatively assess pathway function in order to generate biological mechanism explanations for each subtype based on the assessment results.
[0112] It should be understood that, in the embodiments of this application, the processor 302 may be a central processing unit (CPU), but it may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.
[0113] Memory 301 may include read-only memory and random access memory, and provides instructions and data to processor 302. Some or all of memory 301 may also include non-volatile random access memory. For example, memory 301 may also store device type information.
[0114] As can be seen from the above, the embodiments of this application achieve objective and stable identification of potential subtype structures in complex medical data through the construction of a similarity matrix of multi-omics data, weighted fusion based on eigenvalues, and multi-stage clustering analysis. First, this method can fully utilize the complementary information of different omics data, significantly improving the discriminative power and robustness of subtype classification. Second, the fused similarity matrix, under spectral embedding processing, can form a clearer low-dimensional representation, making subsequent clustering results more objective and reliable. Furthermore, by combining the silhouette coefficient evaluation in K-means clustering with consensus clustering verification, the sensitivity of clustering results to parameter selection and data perturbation can be effectively avoided, thereby obtaining more credible optimal classification results. Based on this, the embodiments of this application further perform differential analysis and biological pathway interpretation after obtaining the classification results, thereby revealing the corresponding pathological mechanism characteristics for different subtypes.
[0115] This application also provides a computer program product that, when run on an electronic device, enables the electronic device to perform the steps described in the various method embodiments above.
[0116] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the above device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0117] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0118] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of external device software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0119] In the embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For instance, the division of modules or units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between devices or units through some interfaces, and may be electrical, mechanical, or other forms.
[0120] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0121] If the integrated units described above are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can also be implemented by a computer program instructing associated hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable storage medium can include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer-readable storage device, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc. It should be noted that the contents of the aforementioned computer-readable storage media may be appropriately added to or subtracted from the contents according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable storage media may not include electrical carrier signals and telecommunication signals.
[0122] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A method for classifying disease subtypes, characterized in that, include: Obtain at least two types of omics data from multiple patient samples with the target disease; Similarity matrices are constructed for each type of omics data, and the similarity matrices are used to describe the feature similarity between any two patient samples under the corresponding omics data. The similarity matrices of the various omics data are fused to obtain a fused similarity matrix; Multi-stage clustering analysis is performed based on the fusion similarity matrix to obtain the subtyping results for the target disease. The multi-stage clustering analysis includes at least two of the following clustering methods: spectral clustering, unsupervised clustering, and consensus clustering.
2. The disease subtype classification method as described in claim 1, characterized in that, The construction of similarity matrices for each type of omics data includes: For each type of omics data, the feature distance between any two patient samples is calculated under the omics data, wherein the distance type of the feature distance is determined based on the omics type of the omics data; Determine the median distance among all the distances between the features calculated using the omics data; Based on the feature distance between any two patient samples in the omics data, and the median distance, a similarity matrix is constructed in the omics data.
3. The disease subtype classification method as described in claim 1, characterized in that, The process of fusing similarity matrices across various types of omics data to obtain a fused similarity matrix includes: The similarity matrices for each type of omics data are standardized into probability matrices. Based on the probability matrix corresponding to each type of omics data, the weights of each type of omics data are calculated, and the weights are used to represent the contribution of the corresponding omics data in the fusion. The fused similarity matrix is obtained by performing matrix fusion based on the similarity matrices corresponding to the various types of omics data and the weights of the various types of omics data.
4. The disease subtype classification method as described in claim 1, characterized in that, In the case where the multi-stage clustering analysis includes spectral clustering and unsupervised clustering, the unsupervised clustering specifically refers to adaptive K-means clustering. The multi-stage clustering analysis based on the fused similarity matrix to obtain the subtyping results for the target disease includes: Based on the fused similarity matrix, a Laplacian matrix for spectral clustering is generated, and the Laplacian matrix is subjected to eigenvalue decomposition to obtain a spectral embedding representation. Under the spectral embedding representation, K-means clustering is performed based on each preset candidate K value, and the first candidate optimal solution is determined based on the preset clustering evaluation index. The subtyping result for the target disease is determined based on the first candidate optimal solution.
5. The disease subtype classification method as described in claim 4, characterized in that, In the case where the multi-stage clustering analysis further includes consensus clustering processing, the method further includes the following steps before determining the subtyping result for the target disease based on the first candidate optimal solution: Each candidate K value is subjected to repeated clustering to construct a consensus matrix corresponding to each candidate K value based on the clustering results. The consensus matrix is used to represent the frequency at which any two patient samples are assigned to the same cluster in the repeated clustering of the candidate K values. The clustering stability of each candidate K value is evaluated based on the consensus matrix, and a second candidate optimal solution is determined based on the clustering stability. Accordingly, determining the typing result for the target disease based on the first candidate optimal solution includes: The classification result for the target disease is determined based on the first candidate optimal solution and the second candidate optimal solution.
6. The disease subtype classification method according to any one of claims 1 to 5, characterized in that, Before constructing similarity matrices for each type of omics data, the disease subtype classification method further includes: Missing values in various types of omics data are imputed to obtain imputed omics data for each type, wherein the imputation processing method is determined based on the omics type of the omics data; The imputed omics data of various types are standardized to obtain standardized omics data of various types, wherein the standardization processing method is determined based on the data statistical type of each feature data under the omics data; The standardized omics data of various types are filtered to obtain filtered omics data of various types, wherein the filtering process is used to remove redundant data in the omics data of various types.
7. The disease subtype classification method according to any one of claims 1 to 5, characterized in that, After performing multi-stage clustering analysis based on the fusion similarity matrix to obtain the subtyping results for the target disease, the disease subtype classification method further includes: Based on the typing results, differential analysis was performed on the omics data of different subtypes to obtain the differential molecules of each subtype; The differentially expressed molecules are mapped to predefined biological pathways for enrichment analysis. The pathway function is quantitatively assessed based on the enrichment results, so as to generate biological mechanism explanations for each subtype based on the assessment results.
8. A disease subtype classification device, characterized in that, include: The acquisition module is used to acquire at least two types of omics data from multiple patient samples suffering from the target disease; A construction module is used to construct similarity matrices for each type of omics data, wherein the similarity matrices are used to describe the feature similarity between any two patient samples under the corresponding omics data; The fusion module is used to fuse the similarity matrices of various types of omics data to obtain a fused similarity matrix; The clustering module is used to perform multi-stage clustering analysis based on the fusion similarity matrix to obtain the typing results for the target disease. The multi-stage clustering analysis includes at least two of the following clustering processing methods: spectral clustering, unsupervised clustering, and consensus clustering.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method as described in any one of claims 1 to 7.
10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by one or more processors, implements the method as described in any one of claims 1 to 7.