A high-dimensional data clustering and feature structure analysis method

By employing high-dimensional data clustering and feature structure analysis methods, and utilizing Gaussian graphical mixture models and sparse feature association networks, the problem of insufficient mining of feature variable interaction networks in high-dimensional data is solved. This enables accurate hierarchical analysis and in-depth revelation of feature relationships in high-dimensional data, and can be applied to fields such as medical health, big data analysis, and industrial fault diagnosis.

CN121561508BActive Publication Date: 2026-04-17NANJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
NANJING UNIV OF POSTS & TELECOMM
Filing Date
2026-01-22
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing high-dimensional data clustering methods fail to fully explore the dynamic interaction network between feature variables, making it difficult for clustering results to explain the heterogeneity of the system's internal operating modes. Furthermore, they lack a unified framework to automatically determine the optimal number of states and generate sparse feature networks.

Method used

We employ high-dimensional data clustering and feature structure analysis methods, using Gaussian graphical mixture model (GGMM) for unsupervised clustering to construct a sparse feature association network. By combining graph theory index calculation and statistical testing, we automatically determine the optimal number of clusters and reveal the topological differences of feature variables.

Benefits of technology

It achieves accurate hierarchical management of high-dimensional data and in-depth mining of feature relationships, and can intuitively display the dynamic interaction mechanism of feature variables under different states. It is widely used in fields such as medical health, big data analysis, financial risk stratification and industrial IoT equipment fault diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121561508B_ABST
    Figure CN121561508B_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of data mining and artificial intelligence, and discloses a high-dimensional data clustering and feature structure analysis method, which comprises the following steps: 1, high-dimensional data preprocessing and output of standardized data, to obtain a preprocessed standardized data set; 2, construction of a Gaussian graphical mixture model for unsupervised clustering, to output a data sub-population containing a category label; 3, independent construction of a feature correlation network for each data sub-population, to output a topological graph of the feature correlation network of each sub-population; and 4, multi-dimensional graph theory index calculation and statistical test on the feature correlation network of each sub-population, to output a final analysis result. The application solves the problem of difficult processing of high-dimensional complex manifold distribution data, can intuitively display core features and dynamic interaction mechanisms behind different categories, realizes a leap from sample division to mechanism revelation, and can be widely applied in the fields of medical subtype discovery, financial risk transmission analysis and industrial fault diagnosis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of data mining and artificial intelligence technology, specifically involving a method for high-dimensional data clustering and feature structure analysis. Background Technology

[0002] With the explosive growth of the Internet of Things (IoT), edge computing, cloud computing, and next-generation big data technologies, the global digitalization process is advancing at an unprecedented pace. In key areas such as smart healthcare, FinTech, intelligent manufacturing, and social governance, the presentation of data has evolved from low-dimensional, simple structures to high-dimensional, dense, and internally interconnected complex forms. This massive amount of high-dimensional data is not merely a collection of numbers; it often implicitly contains extremely complex nonlinear manifold structures and dynamic and close interaction mechanisms between feature variables. Traditional clustering algorithms such as K-means, hierarchical clustering, and DBSCAN have shown some effectiveness in processing low-dimensional, simply distributed data, but their theoretical shortcomings become increasingly apparent when facing the aforementioned high-dimensional complex systems, making them difficult to adapt to the complex manifold distributions that are widely present in high-dimensional spaces.

[0003] In recent years, probabilistic graphical modeling methods, represented by Gaussian Graphical Mixture Models (GGMMs), have made significant progress in the field of state recognition for high-dimensional data due to their ability to simultaneously model data distribution and covariance structure. However, most existing GGMM applications remain at the level of static feature mean-driven approaches or only focus on the overall sparsity of the covariance matrix, failing to fully delve into the dynamic interaction network between feature variables. This lack of structural information mining means that while clustering results achieve sample partitioning, they are insufficient to explain the heterogeneity of the system's internal operating modes under different states from a mechanistic perspective.

[0004] However, existing methods often separate clustering from network construction, lacking a unified framework that can automatically determine the optimal number of states in an unsupervised environment, generate highly interpretable sparse feature networks for each state, and quantify the statistical differences in topology between different states. Heterogeneous feature mining based on network topology has not been fully explored, yet this is a crucial step in achieving refined hierarchical classification of complex objects, fault mechanism tracing, and personalized intervention decisions. Summary of the Invention

[0005] To overcome the shortcomings of existing data analysis methods that ignore the dynamic interaction relationships between features, this application provides a high-dimensional data clustering and feature structure analysis method. This method can automatically determine the optimal number of clusters, achieve accurate stratification of samples, and reveal the topological differences of feature variables under different states by constructing a sparse feature association network. It can be widely applied to scenarios that require fine-grained segmentation and feature relationship mining of multimodal data, such as medical and health big data analysis, financial risk customer segmentation, industrial IoT equipment fault diagnosis, and complex social network analysis.

[0006] To achieve the above objectives, this application employs the following technical solution:

[0007] This application presents a method for high-dimensional data clustering and feature structure analysis, specifically including the following steps:

[0008] Step 1: High-dimensional data preprocessing and outputting standardized data: Collect high-dimensional data and preprocess it to obtain a preprocessed standardized dataset. ;

[0009] Step 2: Standardize the dataset obtained in Step 1. As input, a Gaussian graphical mixture model is constructed for unsupervised clustering, and the output is a data subgroup containing category labels;

[0010] Step 3: Based on the data subgroups containing category labels output after clustering in Step 2, construct a feature association network independently for each data subgroup and output the topology graph of the feature association network for each subgroup;

[0011] Step 4: Analyze the feature association networks of each subgroup constructed in Step 3. Perform multi-dimensional graph theory index calculations and statistical tests to output the final analysis results.

[0012] A further improvement to this application is that step 1 specifically includes the following steps:

[0013] Step 1.1: Collect high-dimensional data. For the collected high-dimensional data, establish an adaptive cleaning and standardization pipeline to eliminate data noise and dimensional differences.

[0014] Step 1.2: Perform distribution tests on the missing data in the high-dimensional data. For continuous variables that conform to a normal distribution, use the mean imputation method. For variables that are skewed, use the median imputation method to ensure that the statistical characteristics of the high-dimensional data are not destroyed.

[0015] Step 1.3: Automatically remove high-missing dimensions with low information content based on information entropy, and perform sample deduplication based on unique identifiers. Then, perform binning. For the first bin, which is mapped to a discrete state after binning... One characteristic variable Information entropy The calculation formula is:

[0016]

[0017] in, This represents the total number of values ​​that the feature variable can take. For characteristic variables The A specific value, Representing characteristic variables Values The probability of;

[0018] Step 1.4: Transform continuous feature variables into ordered rank variables according to different numerical intervals, by calculating the mean of each feature dimension. with standard deviation Map the data to a unified feature space:

[0019]

[0020] in, For the original dataset, This is the standardized dataset.

[0021] A further improvement in this application is that step 2 specifically includes the following steps:

[0022] Step 2.1: Standardize the dataset obtained in Step 1. As input, a Gaussian graph mixture model is constructed for unsupervised clustering. Graph structure modeling is embedded during the unsupervised clustering process to maximize the penalized log-likelihood function. The optimization objective is to find a set of parameters for an optimal Gaussian graphic mixture model. This makes the parameter set Maximize the value:

[0023]

[0024] in, For Gaussian graphic mixture model parameter set, For standardized datasets The total number, This represents the current number of clusters. For the preprocessed standardized dataset The first in A sample vector, For the first The mixed weights of each cluster, Let be the probability density function of a Gaussian distribution. For precision matrix The sparse penalty term, For the first The mean vector of each cluster. For the first The covariance matrix of each cluster;

[0025] Step 2.2: Standardize the dataset from Step 1. Each feature introduces a dynamic weight allocation mechanism. In the expectation-maximization algorithm, the first feature... In the cluster, the first Weights of each feature The calculation formula is:

[0026]

[0027] in, For the first The feature in the first Local distribution within a cluster, For the first One feature in the standardized dataset Global distribution on For the first The feature in the first Local distribution within a cluster, For the first One feature in the standardized dataset Global distribution on For standardized datasets The total dimension of features, For temperature coefficient, express Divergence;

[0028] Step 2.3, Setting The search targets each element within the search range. The values ​​are then iterated through in step 2.2 until convergence, resulting in a series of Gaussian graphical mixture models with varying parameters. The Bayesian information criterion is then applied. The Gaussian mixture model is evaluated and traversed for search. When the parameter set of a certain Gaussian mixture model... When the value tends to stabilize, calculate the Gaussian mixture model. value:

[0029]

[0030] in, This represents the maximum value of the likelihood function for the Gaussian Graphical Mixture Model (GGMM). For standardized datasets The total number, This is a penalty term used to prevent model overfitting. It iterates through a preset interval of cluster numbers and compares all... Value, identify the global minimum value, and lock it. The model state with the minimum value;

[0031] Step 2.4: Using the Bayesian posterior probability formula, calculate the probability of each... The sample vector belongs to the th sample vector. The posterior probability of each cluster :

[0032]

[0033] in, For the first The mixed weights of each cluster, For the first The mean vector of each cluster. For the first Cluster covariance matrices;

[0034] Based on the maximum a posteriori probability principle, determine the first... Final category label for each sample :

[0035]

[0036] Step 2.5: Output includes category labels. The dataset will be standardized. Divided into Cluster-independent subgroups of data containing category labels are used as input data for constructing the feature association network in step 3.

[0037] A further improvement in this application is that step 3 specifically includes the following steps:

[0038] Step 3.1, based on the first step in step 2 The standardized dataset of the nth data subgroup is used to calculate the nth... The validation covariance matrix of the standardized dataset for each data subgroup , , build with The log-likelihood objective function of norm penalty is estimated to be the first... The sparse precision matrix of each cluster :

[0039]

[0040] in, Represents the matrix trace operation. For non-negative regularization hyperparameters, Represents the off-diagonal elements of a sparse precision matrix Norm sum;

[0041] Step 3.2: Using the preset non-negative regularization hyperparameters... Candidate models for generating regularized paths within a value range are used, utilizing the extended Bayesian information criterion. Quality evaluation was performed on each candidate model:

[0042]

[0043] in, For the first The number of standardized datasets for each data subgroup The feature dimension, This represents the estimated number of edges in the network. As a hyperparameter, iterate through and calculate the scores of all candidate Gaussian graphical mixture models along the regularization path, and select... The value corresponding to the minimum As the optimal regularization parameter;

[0044] Step 3.3: Calculate the final sparse precision matrix using the optimal regularization parameters selected in Step 3.2. and the final sparse precision matrix The partial correlation coefficient matrix is ​​transformed into a feature association network based on the partial correlation coefficient matrix. ,in, For nodes, For connecting edges, the final sparse precision matrix middle Then at node With nodes Establish connections between them, with the weights of the connections determined by the partial correlation coefficients. Decide:

[0045]

[0046] in, represents the element values ​​in a sparse precision matrix;

[0047] Step 3.4: Output the feature association network of each subgroup. The topology diagram.

[0048] A further improvement in this application is that step 4 specifically includes the following steps:

[0049] Step 4.1: Develop the feature association network in Step 3. The system is divided to simulate particles in a feature correlation network. The random walk process in the model utilizes the structural similarity between nodes to determine the structural distance. The smallest feature aggregation forms a community with the same function, defining nodes. and nodes Distance between :

[0050]

[0051] in, The number of steps in the random walk. For the node go through Step to Node The probability, For nodes The degree;

[0052] Step 4.2: Calculate the feature association network The centrality indices of each node's strength, closeness, and betweenness are used to introduce a stability verification mechanism for the standardized dataset. Perform sampling with replacement to construct data subpopulations of different sizes and calculate the standardized dataset. centrality vector With the centrality vector of the data subgroup The correlation between levels is quantitatively evaluated using the correlation stability coefficient to find those that meet the requirements. Maximum sample rejection ratio under the given conditions:

[0053]

[0054] in, For standardized datasets The total number, To maximize the size of the removable subsample, if If the centrality ranking of a feature is greater than the preset stability threshold, it is determined that the centrality ranking of the feature is statistically robust, thus eliminating the interference of random errors.

[0055] Step 4.3: Calculate the feature association networks of the two sets of networks. Original observed differences in global strength, network structure dissimilarity, and edge weight differences. A nonparametric permutation test strategy is adopted to randomly shuffle sample labels and reconstruct the feature association network. Constructing the empirical zero distribution of difference statistics, statistics The value is:

[0056]

[0057] in, For the first Difference statistics under random permutations For indicator functions, For the preset number of permutation iterations, when The value is 1 when it is active and 0 otherwise. Multiple checks and corrections;

[0058] Step 4.4: Summarize and Quantify the Assessment ,distance and statistics The analysis results are used to generate a final analysis report.

[0059] A further improvement of this application is that the high-dimensional data clustering and feature structure analysis method is implemented through a high-dimensional data clustering and feature structure analysis system, which includes a data preprocessing and standardization module, a GGMM adaptive clustering module, a feature association network construction module, and a network topology difference analysis module, wherein:

[0060] The data preprocessing and standardization module collects high-dimensional data and preprocesses it to obtain a standardized dataset.

[0061] The GGMM adaptive clustering module performs unsupervised clustering on the standardized dataset obtained by the data preprocessing and standardization module, and outputs a dataset containing category labels as input to the feature association network construction module.

[0062] The feature association network construction module performs sparse structure learning and outputs a feature association network topology graph;

[0063] The network topology difference analysis module takes the feature-related network topology graph as input and performs multi-dimensional graph theory index calculations and statistical tests.

[0064] The beneficial effects of this application are:

[0065] This application is the first to deeply integrate dynamically weighted GGMM clustering with feature network topology analysis, by introducing KL divergence weight updates and... Sparse constraints not only solve the problem that traditional methods struggle to handle high-dimensional, complex manifold distribution data, but also intuitively demonstrate the core features and dynamic interaction mechanisms behind different categories.

[0066] This application represents a leap from sample segmentation to mechanism revelation, and can be widely applied in fields such as medical subtype discovery, financial risk transmission analysis, and industrial fault diagnosis. Attached Figure Description

[0067] Figure 1 This is the overall flowchart of this application.

[0068] Figure 2 This is the chord diagram showing the significant edge strength difference in this application. Detailed Implementation

[0069] The embodiments of the present invention will be disclosed below with reference to the drawings. For clarity, many practical details will be described in the following description. However, it should be understood that these practical details are not intended to limit the present invention. That is, in some embodiments of the present invention, these practical details are not essential. In addition, for the sake of simplicity, some conventional structures and components will be shown in the drawings in a simple schematic manner.

[0070] like Figure 1 As shown, this application discloses a high-dimensional data clustering and feature structure analysis method. This method is implemented through a high-dimensional data clustering and feature structure analysis system, which includes a data preprocessing and standardization module, a GGMM adaptive clustering module, a feature association network construction module, and a network topology difference analysis module. Specifically: the data preprocessing and standardization module collects high-dimensional data and preprocesses it to obtain a standardized dataset; the GGMM adaptive clustering module performs unsupervised clustering on the standardized dataset obtained from the data preprocessing and standardization module, outputting a dataset containing category labels as input to the feature association network construction module; the feature association network construction module performs sparse structure learning and outputs a feature association network topology graph; and the network topology difference analysis module takes the feature association network topology graph as input and performs multi-dimensional graph theory index calculations and statistical tests.

[0071] Specifically, the high-dimensional data clustering and feature structure analysis method includes the following steps:

[0072] Step 1: High-dimensional data preprocessing and outputting standardized data: Collect high-dimensional data and preprocess it to obtain a preprocessed standardized dataset. First, the system fully integrates raw high-dimensional data from various data sources, including but not limited to high-frequency financial transaction databases, industrial SCADA monitoring systems, and medical electronic medical record systems, via standard interfaces. Given the inevitable presence of missing data in the raw high-dimensional data, the high-dimensional data clustering and feature structure analysis system employs an adaptive imputation strategy: performing a distribution pattern test on each feature variable. For variables conforming to a normal distribution, the system automatically uses mean imputation to maintain the central tendency of the data; while for variables exhibiting a skewed distribution, it switches to median imputation to minimize the distortion of the overall distribution statistics by outliers and ensure the statistical integrity of the existing data.

[0073] To address the sparsity and redundancy issues in high-dimensional data, the high-dimensional data clustering and feature structure analysis system introduces a feature selection mechanism based on information entropy. By calculating the information entropy of each dimension, the amount of information it contains is quantified. For dimensions with information entropy below a preset threshold—i.e., quasi-constant features with minimal variation and low discriminative power—the system classifies them as low-information dimensions and automatically removes them. This process effectively achieves preliminary dimensionality reduction of the feature space and filters out invalid noise. Simultaneously, a strict deduplication operation is performed based on the unique identifier of each sample, eliminating completely duplicated redundant records and ensuring the independence of the samples.

[0074] To mitigate the impact of minute numerical fluctuations on the model for continuous variables in the data, the high-dimensional data clustering and feature structure analysis system maps them into several ordered discrete levels based on their numerical ranges. Subsequently, for features with different physical meanings and dimensions, such as blood glucose concentration in medical data... Indices, blood pressure values, or industrial data such as temperature, pressure, and rotational speed, are analyzed by high-dimensional data clustering and feature structure analysis systems. Standardization. This process maps all feature values ​​to the mean. Standard deviation is The standard normal distribution space is used to completely eliminate the problem of convergence failure or uneven weight distribution caused by huge differences in the scale of the dependent variable.

[0075] Specifically, step 1 of this application includes the following steps:

[0076] Step 1.1: Collect high-dimensional data. For the collected high-dimensional data, establish an adaptive cleaning and standardization pipeline to eliminate data noise and dimensional differences.

[0077] Step 1.2: Perform distribution tests on the missing data in the high-dimensional data. For continuous variables that conform to a normal distribution, use the mean imputation method. For variables that are skewed, use the median imputation method to ensure that the statistical characteristics of the high-dimensional data are not destroyed.

[0078] Step 1.3: Automatically remove high-missing dimensions with low information content based on information entropy, and perform sample deduplication based on unique identifiers. Then, perform binning. For the first bin, which is mapped to a discrete state after binning... One characteristic variable Assume its set of values ​​is Its information entropy The calculation formula is:

[0079]

[0080] in, This represents the total number of possible values ​​for this feature variable. For characteristic variables The A specific value, Representing characteristic variables Values The probability is usually estimated using sample frequency.

[0081] Step 1.4: Continuous feature variables are transformed into ordered rank variables according to different numerical intervals. Finally, to eliminate the influence of scale differences in physical quantities on the convergence speed of the optimization algorithm, this application adopts... Standardization strategy. This involves calculating the mean of each feature dimension. with standard deviation Map the data to a unified feature space:

[0082]

[0083] in, For the original dataset, The standardized dataset;

[0084] The data is mapped to a unified feature space. This process ensures that each feature dimension has equal initial weights in subsequent Gaussian Graphical Mixture Model (GGMM) iterations, thereby improving the stability and accuracy of the clustering results.

[0085] Step 2: Standardize the dataset obtained in Step 1. As input, a Gaussian Graphical Mixture Model (GGMM) is constructed for unsupervised clustering, outputting data subgroups containing category labels. In this application, based on... Preferred GGMM adaptive clustering: This involves using the preprocessed standardized dataset... This refers to the Gaussian Graphical Mixture Model (GGMM) with a standardized feature matrix as input. This model not only simulates the probability distribution of samples in space but also deeply integrates the conditional independence structure between variables. To address the issue of the number of clusters (…), To address the pain point of difficulty in determining the value, this system initiates an adaptive search procedure, traversing preset candidate clustering intervals. In each candidate state, the system calculates the Bayesian information criterion (...). This metric introduces a penalty term between the fitting accuracy and parameter complexity of the Gaussian Graphical Mixture Model (GGMM), and the system automatically locks the parameter complexity. The model state with the smallest value is taken as the global optimal solution, thus objectively determining the optimal number of clusters for the data without human intervention.

[0086] After determining the optimal number of clusters, the system initiates expectation maximization. The algorithm performs parameter estimation. During the iteration process, this application introduces an innovative dynamic weighting mechanism. The system utilizes... Divergence measures the difference between each feature's current cluster distribution and the global distribution. A greater difference indicates a higher contribution of the feature to distinguishing the current category. The system dynamically updates feature weights accordingly, automatically amplifying features with significant class-discriminating power while suppressing irrelevant background noise features. This process is repeated until the model's log-likelihood function converges, ultimately outputting labeled clustering results, achieving accurate hierarchical partitioning of complex samples.

[0087] Construction and optimization of sparse feature association networks: For each subgroup after clustering, a graph model is used to represent its internal feature interaction mechanism, transforming the black-box clustering results into an interpretable topological network. Graph minimum absolute shrinkage and selection algorithms are employed. The algorithm, in the accuracy matrix In the solution, introduce Norm penalty terms are used to achieve sparsity in the network structure, where non-zero elements... Directly corresponding features and Partial correlation, i.e., the edges in a network, given other variables.

[0088] Specifically, step 2 includes the following steps:

[0089] Step 2.1: Standardize the dataset obtained in Step 1. As input, a Gaussian graphical mixture model (GGMM) is constructed for unsupervised clustering. To achieve probabilistic modeling of high-dimensional data, graph structure modeling is embedded in the unsupervised clustering process to maximize the penalized log-likelihood function. The optimization objective is to find a set of parameters for an optimal Gaussian graphic mixture model. This makes the parameter set Maximize the value:

[0090]

[0091] in, Let be the parameter set of the Gaussian Graphical Mixture Model (GGMM), representing the independent variance of each dimension. For standardized datasets The total number, This represents the current number of clusters. For the preprocessed standardized dataset The first in A sample vector, For the first The mixed weights of each cluster, Let be the probability density function of a Gaussian distribution. For precision matrix The sparse penalty term, For the first The mean vector of each cluster. For the first The covariance matrix of each cluster, specifically the precision matrix. Defined as covariance matrix The inverse matrix of is mathematically expressed as:

[0092]

[0093] Step 2.2: To solve the above objective function, this invention uses the Expectation-Maximization (EM) algorithm for iterative processing. During this process, the standardized dataset from Step 1 is processed... Each feature is assigned a dynamic weight to quantify its contribution to clustering, aiming to maximize the expected value. In the algorithm, the first In the cluster, the first Weights of each feature The calculation formula is:

[0094]

[0095] in, For the first The feature in the first Local distribution within a cluster, For the first One feature in the standardized dataset Global distribution on For the first The feature in the first Local distribution within a cluster, For the first One feature in the standardized dataset Global distribution on For standardized datasets The total dimension of features, The temperature coefficient controls the sharpness of the weight distribution. express Divergence is used to measure local distribution. With global distribution The difference; the introduction of this weight makes the model maximize In the process, it can adaptively amplify the influence of key features.

[0096] Step 2.3, Setting The search targets each element within the search range. The values ​​are then iterated through in step 2.2 until convergence, resulting in a series of Gaussian graphical mixture models with varying parameters. The Bayesian information criterion is then applied. The Gaussian mixture model is evaluated and traversed for search. When the parameter set of a certain Gaussian mixture model... When the value tends to stabilize, calculate the Gaussian mixture model. value:

[0097]

[0098] in, This represents the maximum value of the likelihood function for the Gaussian Graphical Mixture Model (GGMM). For standardized datasets The total number, This is a penalty term used to prevent model overfitting. It iterates through a preset interval of cluster numbers and compares all... Value, identify the global minimum value, and lock it. The model state with the minimum value;

[0099] At this point, the model achieves an optimal balance between fitting accuracy and model complexity, thus determining the final number of clusters. And the corresponding clustering results.

[0100] Step 2.4, Optimal Model Establishment and Result Output: Once the system has locked onto the model with the minimum... The optimal number of clusters K and its corresponding set of convergence parameters for the values The system will then perform the final sample labeling. Using the Bayesian posterior probability formula, the probability of each sample is calculated. The sample vector belongs to the th sample vector. The posterior probability of each cluster :

[0101]

[0102] in, For the first The mixed weights of each cluster, For the first The mean vector of each cluster. For the first Cluster covariance matrices;

[0103] Based on the Maximum A Posteriori (MAP) principle, determine the first... Final category label for each sample :

[0104]

[0105] Step 2.5: Output includes category labels. The dataset will be standardized. Divided into Cluster-independent subgroups of data containing category labels are used as input data for constructing the feature association network in step 3.

[0106] The system will automatically lock The model state with the minimum value represents the optimal balance between complexity and goodness of fit, thus maximizing the number of clusters. The system features fully automated identification, requiring no manual pre-setting. Furthermore, it improves the model's generalization ability on non-convex datasets by dynamically adjusting the covariance matrix type and initialization parameters, and combines cross-validation to evaluate the stability of clustering results, outputting statistically significant data category labels.

[0107] Step 3: Based on the data subgroups containing category labels output after clustering in Step 2, construct a feature association network independently for each data subgroup and output the topology graph of the feature association network for each subgroup.

[0108] In this step, for scenarios where the feature dimension is much larger than the sample size, the system introduces the Extended Bayesian Information Criterion to prevent overfitting. As a guide for hyperparameter tuning. By adjusting the penalty parameter, The criteria seek an optimal balance between the complexity and sparsity of the network model, ensuring that the generated interconnected network can both accurately capture the core interaction logic and effectively eliminate spurious associations (false edges) caused by noise. Finally, the system maps the optimized non-zero partial correlation coefficients to a weighted undirected network, intuitively representing the interaction mechanism between features in this state mode.

[0109] Specifically, step 3 includes the following steps:

[0110] Step 3.1, based on the first step in step 2 The standardized dataset of the nth data subgroup is used to calculate the nth... The validation covariance matrix of the standardized dataset for each data subgroup , , build with The log-likelihood objective function of norm penalty is estimated to be the first... The sparse precision matrix of each cluster This function aims to solve for the accuracy matrix. (i.e., the inverse of the covariance matrix), whose non-zero elements correspond to the conditional dependencies between features:

[0111]

[0112] in, Represents the matrix trace operation. is a non-negative regularization hyperparameter used to control the sparsity of the feature network (i.e., the number of edges in the network). Represents the off-diagonal elements of a sparse precision matrix Norm sum;

[0113] Step 3.2: Using the preset non-negative regularization hyperparameters... Candidate models for generating regularized paths within a value range are used, utilizing the extended Bayesian information criterion. Quality evaluation was performed on each candidate model:

[0114]

[0115] in, For the first The number of standardized datasets for each data subgroup The feature dimension, This represents the estimated number of edges in the network (i.e., the number of non-zero elements). This is a hyperparameter, typically set to 0.5. It iterates through and calculates the scores of all candidate Gaussian graphical mixture models along the regularization path, then selects... The value corresponding to the minimum As the optimal regularization parameter, it achieves the best balance between network sparsity and goodness of fit, eliminating spurious indirect correlations. This controls the sparsity of the feature network, ensuring that the generated network structure is both interpretable and avoids overfitting.

[0116] Step 3.3: Calculate the final sparse precision matrix using the optimal regularization parameters selected in Step 3.2. and the final sparse precision matrix The partial correlation coefficient matrix is ​​transformed into a feature association network based on the partial correlation coefficient matrix. ,in, For nodes, For connecting edges, the final sparse precision matrix middle Then at node With nodes Establish connections between them, with the weights of the connections determined by the partial correlation coefficients. Decide:

[0117]

[0118] in, represents the element values ​​in a sparse precision matrix;

[0119] Step 3.4: Output the feature association network of each subgroup. The topology diagram is used as the input for the network structure difference analysis in step 4.

[0120] Step 4: Analyze the feature association networks of each subgroup constructed in Step 3. Perform multi-dimensional graph theory index calculations and statistical tests, and output the final analysis results. Graph theory index calculation refers to calculating standardized datasets. centrality vector With the centrality vector of the data subgroup .

[0121] Exploratory Graph Analysis (EGA) combined with the Walktrap algorithm is applied to modularly analyze the feature network. The Walktrap algorithm simulates a random walk process and uses the structural distance between nodes to identify closely related features. The smallest feature aggregations form functional communities, thereby revealing modular collaborative patterns among variables. Simultaneously, centrality indices such as strength, betweenness, and density are calculated to precisely locate the core feature nodes that play a crucial regulatory role in the network.

[0122] To scientifically quantify the differences in network structure among different subtypes, the system performed a network comparison test (NCT). This test employed a non-parametric statistical strategy, executing 1000 random permutation tests. By randomly shuffling sample labels and reconstructing the network, an empirical null distribution of the difference statistics is constructed. Under strict control of the false positive rate, the positions of observed differences (such as differences in global connectivity strength and differences in specific edge weights) within the random distribution are calculated, thereby obtaining the statistical... Value. Matching Multiple test corrections ensure that the differences found have rigorous statistical significance evidence, rather than random error.

[0123] Finally, the system is built based on The stability verification unit for resampling technology. By performing multiple random samplings with replacement on the original data and recalculating the centrality index on each subset, the ranking fluctuations of core feature nodes are observed. The system calculates the relevant stability coefficients (…). If the coefficient exceeds a preset threshold, such as This confirms that the identified core features and network structure have extremely high robustness in a statistical sense and can withstand the test of sample perturbation.

[0124] Specifically, step 4 includes the following steps:

[0125] Step 4.1: Develop the feature association network in Step 3. Divide into parts and in the feature association network A random walk process is performed, utilizing the structural similarity between nodes to calculate the structural distance. The smallest feature aggregation forms a community with the same function, defining nodes. and nodes Distance between :

[0126]

[0127] in, The number of steps in the random walk. For the node go through Step to Node The probability, For nodes Based on this degree, the system divides the feature association network into several non-overlapping communities, revealing the differences in feature grouping structure among different clustering subtypes.

[0128] Step 4.2: Calculate the feature association network The strength of each node ( Proximity ) and median ( To assess the centrality of the dataset, a stability verification mechanism based on Bootstrap resampling is introduced for standardized datasets. Perform sampling with replacement to construct data subpopulations of different sizes and calculate the standardized dataset. centrality vector With the centrality vector of the data subgroup To ensure the robustness of the analysis results, this application introduces a correlation stability coefficient (CSC) to assess the hierarchical correlation between the data points. The centrality metric is evaluated and defined as the maximum sample removal rate that meets the following conditions:

[0129]

[0130] Seeking satisfaction The maximum sample removal ratio under the given conditions. The sensitivity of the quantified feature ranking to sample fluctuations. Among them, For standardized datasets The total number, To maintain the maximum relevance threshold, if If the centrality ranking of this feature is statistically robust, then the interference of random errors is excluded.

[0131] Step 4.3: Calculate the feature association networks of the two sets of networks. In global strength ( Network structural dissimilarity ) and edge weight differences ( The original observed difference values ​​on ) A nonparametric permutation test strategy is adopted. Randomly shuffle sample labels and reconstruct the feature association network. , Next, construct the empirical zero distribution of the difference statistic, in statistics. The value is:

[0132]

[0133] in, For the first Difference statistics under random permutations For indicator functions, For the preset number of permutation iterations, when The value is 1 when it is active and 0 otherwise. Multiple test corrections, if the corrected This confirms that there are significant statistical differences between the two subtypes in terms of network mechanisms.

[0134] Step 4.4, Summarizing the Feature Structure Analysis Results: Summarizing and Quantitatively Evaluating ,distance and statistics The analysis results are used to generate a final analysis report.

[0135] Example

[0136] This application provides a high-dimensional data clustering and feature structure analysis method, aiming to explore its in-depth application and value in the field of medical and health big data. This embodiment selects type 2 diabetes mellitus (T2DM) patient data with typical high-dimensional and complex correlation characteristics. Through refined subtyping of the patient group and pathological feature network analysis, the feasibility, effectiveness, and universality of this application in complex disease phenotypic identification, feature interaction mechanism mining, and result robustness assessment are systematically verified. Specific implementation steps and detailed analysis are as follows:

[0137] (1) This application first establishes a data interface with the hospital's electronic medical record system (EMR) and extracts all the sample data to be analyzed from patients diagnosed with type 2 diabetes, as shown in Table 1.

[0138] Table 1

[0139]

[0140] The sample data in this application encompasses comprehensive demographic information such as age and gender distribution, key clinical indicators such as body mass index (BMI) and systolic / diastolic blood pressure, and detailed laboratory biochemical indicators such as glycated hemoglobin (HbA1c), fasting blood glucose, lipid profiles, and renal function indicators, forming a high-dimensional feature space. To eliminate the noise and quality issues commonly found in raw medical data, a rigorous data governance pipeline was constructed.

[0141] Adaptive imputation of missing values: All feature variables are scanned for missing rates. For variables with a missing rate of 15% or less, they are handled according to their data distribution characteristics: those conforming to a normal distribution are imputed using the mean, and those with a skewed distribution are imputed using the median, thus preserving the original statistical characteristics of the data to the greatest extent possible; for variables with a missing rate exceeding 15%, the imputation is performed accordingly. Variables that are considered low-quality dimensions are removed to avoid introducing excessive interpolation errors.

[0142] Sample deduplication and screening: Deduplication based on unique identifiers was performed to remove completely duplicate redundant records, resulting in 7,053 high-quality, independent and complete valid samples, providing a solid data foundation for subsequent analysis.

[0143] Unified Feature Mapping: Considering the coexistence of continuous and discrete variables in medical data, all continuous clinical variables were divided into five ordered levels (labels), achieving feature discretization and reducing the interference of minor fluctuations in the data. Subsequently, a mapping process was performed on all numerical variables. Standardization maps indicators with different physical dimensions (such as millimoles per liter of blood glucose and millimeters of mercury of blood pressure) to a unified standard feature space (mean of 0 and standard deviation of 1), thereby completely eliminating the potential impact of variable scale differences on the convergence speed and weight allocation of subsequent clustering algorithms.

[0144] (2) The dynamic weighted Gaussian graphic mixture model (GGMM) proposed in this invention is used to perform unsupervised clustering analysis on the standardized high-dimensional feature data, in order to discover hidden patient subtypes.

[0145] This application abandons the traditional method of manually specifying the number of clusters, and instead uses the Bayesian Information Criterion (BIC) as an objective criterion for model selection. The program automatically traverses the preset range of cluster numbers. Calculate the model state. Value. The results show that when When the BIC value reaches its minimum, it indicates that the model achieves the best balance between goodness of fit and complexity when the samples are divided into two classes.

[0146] Based on a determined optimal number of clusters Initialize the GGMM model parameters. During the iteration of the Expectation Maximization (EM) algorithm, the model dynamically updates the mean, covariance matrix, and mixing weights of each Gaussian component until the log-likelihood function tends to converge stably.

[0147] The stability of the clustering results was rigorously evaluated through cross-validation. The system ultimately output two subtypes with statistically significant differences. Based on their clinical characteristics, these subtypes were named "non-obese-metabolic stable type" and "obese-insulin-resistant type," achieving precise stratification of the T2DM patient population. The obese-insulin-resistant type showed a significantly elevated... , (waistline), (Hip circumference), and extremely high liver enzyme levels (ALT, GGT), accompanied by typical high TG (triglycerides), low... The dyslipidemia pattern of high-density lipoprotein cholesterol (HDL-C) and the overall elevation of blood glucose levels suggest that this subtype is a group of patients with obesity-driven, severe insulin resistance and hepatic metabolic stress. Conversely, the non-obese-metabolic stable subtype, although diagnosed with... However, its Liver and kidney function and blood lipid levels were mostly at average or low levels, suggesting that the pathogenesis of this subtype may be mainly related to pancreatic islet function. It is related to relative cellular insufficiency, rather than metabolic overload.

[0148] (3) In order to reveal the differences in pathological mechanisms behind different subtypes, this step constructs a feature association network for each subtype based on the clustering results.

[0149] For the symptom feature data of each subtype, the Graph Minimum Absolute Shrinkage and Selection Operator (GLASSO) algorithm was used for modeling. To eliminate spurious indirect associations in complex feature relationships, the Extended Bayesian Information Criterion (EBIC) was introduced to automatically search for the optimal regularization parameter. This process accurately generated the partial correlation coefficient matrix between feature variables, retaining only edges with significant direct dependencies, successfully constructing a highly interpretable sparse feature network.

[0150] (4) Use graph theory analysis techniques to deeply quantify network topology characteristics and verify the significance of differences through rigorous statistical tests.

[0151] Introducing exploratory graph analysis ( Combined with random walk strategy ( The algorithm performs topological analysis on the generated feature network. By simulating the flow of information within the network, the algorithm automatically identifies closely connected functional modules (communities). The analysis results show significant structural heterogeneity: the non-obese-metabolic stable feature network is divided into 7 functional communities, while the obese-insulin-resistant feature network aggregates into 6 functional communities. This difference in the number and structure of communities suggests a fundamental difference in the complexity of the pathophysiological mechanisms and the modular synergistic patterns between the two subtypes.

[0152] Centrality indices such as strength, proximity, and median strength of each node in the network were calculated. The results showed that in the "obesity-insulin resistance" network, the glomerular filtration rate (GFR)... The centrality of serum creatinine (SCr) ranked first, indicating that pathological changes related to renal function play a core regulatory role in this subtype; while in the "non-obese-metabolic stable type", the centrality of serum creatinine (SCr) was significantly higher than that of other characteristics.

[0153] Finally, 1000 permutation tests were performed using the Network Comparison Test (NCT). Statistical results showed that although there was no significant difference in global network connectivity between the two subtypes ( and , This indicates that the overall symptom burden is similar, but there are highly significant statistical differences in network topology (i.e., edge distribution patterns). ).

[0154] This confirms that although both groups of patients had diabetes, the interaction patterns between their symptoms underwent a fundamental reorganization. Further edge strength analysis identified 38 significantly different edges ( This result is illustrated by a chord diagram (see...). Figure 2 This is visually illustrated. The thickness of the connecting lines in the diagram directly reflects the intensity of the difference: for example, and The connecting lines between them are relatively thick, visually reflecting the significant differences in the relationship between the two subtypes; and and The thinner lines between the symptoms indicate a relatively stable relationship. This visualization method clearly reveals the reorganization of interaction pathways between symptoms (such as the association between blood glucose and kidney function) during the progression from mild to severe diabetes. This finding provides new data evidence for understanding the pathological evolution pathways of different subtypes of diabetes.

[0155] To eliminate the interference of random errors, a correlation stability coefficient is also introduced. The centrality indicators mentioned above were evaluated. The results showed that the key centrality indicators for both subtypes exhibited extremely high stability (coefficients both exceeding [percentage missing]). Specifically, in non-obese metabolically stable individuals... The CS coefficient of the node reaches The CS coefficient of the eGFR node in obesity-insulin resistance is as high as The results far exceed the threshold required by conventional statistics, demonstrating that the identification of these core features has extremely strong robustness in cross-sample resampling, and the conclusion is reliable.

[0156] This application, through four modules—data preprocessing, GGMM clustering, feature network construction, and network difference analysis—not only achieves accurate "state profiling" of samples but also reveals the "feature game" relationships hidden behind the data. Compared to traditional clustering methods, this method exhibits stronger interpretability, noise resistance, and statistical rigor when processing high-dimensional medical and financial data, providing scientific quantitative support for precise intervention and decision-making in complex systems. Furthermore, the specific steps of each module are clearly defined, ensuring the operability of the method and the reliability of the results.

[0157] The above description is merely an embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of the present invention should be included within the scope of the claims of the present invention.

Claims

1. A method for high-dimensional data clustering and feature structure analysis, characterized in that: The method is used in the field of medical and health big data. Specifically, the high-dimensional data clustering and feature structure analysis method includes the following steps: Step 1: High-dimensional data preprocessing and outputting standardized data: High-dimensional data of patients with type 2 diabetes are collected and preprocessed to obtain a preprocessed standardized dataset. ; Step 2: Standardize the dataset obtained in Step 1. As input, a Gaussian graphical mixture model is constructed for unsupervised clustering, aiming to discover latent patient subtypes, and the output is a data subgroup containing category labels; Step 3: Based on the data subgroups containing category labels output after clustering in Step 2, construct a feature association network independently for each data subgroup and output the topology graph of the feature association network for each subgroup. Step 4: Perform multi-dimensional graph theory index calculations and statistical tests on the feature association networks of each subgroup constructed in Step 3, and output the final analysis results, specifically: Step 2 includes the following steps: Step 2.1: Standardize the dataset obtained in Step 1. As input, a Gaussian graph mixture model is constructed for unsupervised clustering. Graph structure modeling is embedded during the unsupervised clustering process to maximize the penalized log-likelihood function. The optimization objective is to find a set of parameters for an optimal Gaussian graphic mixture model. This makes the parameter set Maximize the value: in, This is the parameter set for the Gaussian graphic mixture model. For standardized datasets The total number, This represents the current number of clusters. For the preprocessed standardized dataset The first in A sample vector, For the first The mixed weights of each cluster, Let Gaussian distribution probability density function be used. For precision matrix The sparse penalty term, For the first The mean vector of each cluster. For the first The covariance matrix of each cluster; Step 2.2: Standardize the dataset from Step 1. Each feature introduces a dynamic weight allocation mechanism. In the expectation-maximization algorithm, the first feature... In the cluster, the first Weights of each feature for: in, For the first The feature in the first Local distribution within a cluster, For the first One feature in the standardized dataset Global distribution on For the first The feature in the first Local distribution within a cluster, For the first One feature in the standardized dataset Global distribution on For standardized datasets The total dimension of features, For temperature coefficient, express Divergence; Step 2.3, Setting The search targets each element within the search range. The values ​​are then iterated through in step 2.2 until convergence, yielding Gaussian graphical mixture models with varying parameters. The Bayesian information criterion is then applied. The Gaussian mixture model is evaluated and traversed for search. When the parameter set of a certain Gaussian mixture model... When stable, calculate the Gaussian mixture model. value: in, This represents the maximum value of the likelihood function for the Gaussian graphical mixture model. For standardized datasets The total number, As a penalty term, iterate through the preset interval of cluster numbers and compare all... Value, identify the global minimum value, and lock it. The model state with the minimum value; Step 2.4: Using the Bayesian posterior probability formula, calculate the probability of each... The sample vector belongs to the th sample vector. Posterior probabilities of clusters : in, For the first The mixed weights of each cluster, For the first The mean vector of each cluster. For the first Each cluster covariance matrix; Based on the maximum a posteriori probability principle, determine the first... Final category label for each sample : Step 2.5: Output includes category labels. The dataset will be standardized. Divided into Cluster-independent subgroups of data containing category labels are used as input data for constructing the feature association network in step 3.

2. The high-dimensional data clustering and feature structure analysis method according to claim 1, characterized in that: Step 1 specifically includes the following steps: Step 1.1: Collect high-dimensional data. For the collected high-dimensional data, establish an adaptive cleaning and standardization pipeline to eliminate data noise and dimensional differences. Step 1.2: Perform distribution tests on the missing data in the high-dimensional data. For continuous variables that conform to a normal distribution, use the mean imputation method. For variables that are skewed, use the median imputation method to ensure that the statistical characteristics of the high-dimensional data are not destroyed. Step 1.3: Automatically remove high-missing dimensions with low information content based on information entropy, and perform sample deduplication based on unique identifiers. Then, perform binning. For the first bin, which is mapped to a discrete state after binning... One characteristic variable Information entropy The calculation formula is: in, This represents the total number of values ​​that the feature variable can take. For characteristic variables The Each possible value Representing characteristic variables Values The probability of; Step 1.4: Transform continuous clinical characteristic variables into ordered rank variables according to different numerical intervals, and calculate the mean of each characteristic dimension. with standard deviation Map the data to a unified feature space: in, For the original dataset, This is the standardized dataset.

3. The high-dimensional data clustering and feature structure analysis method according to claim 1, characterized in that: Step 3 specifically includes the following steps: Step 3.1, based on the first step in step 2 The standardized dataset of the nth data subgroup is used to calculate the nth... The validation covariance matrix of the standardized dataset for each data subgroup , , build with The log-likelihood objective function of norm penalty is estimated to be the first... The sparse precision matrix of each cluster : in, Represents the matrix trace operation. For non-negative regularization hyperparameters, Represents the off-diagonal elements of a sparse precision matrix Norm sum; Step 3.2: Using the preset non-negative regularization hyperparameters... Candidate models for generating regularized paths within a value range are used, utilizing the extended Bayesian information criterion. Quality evaluation was performed on each candidate model: in, For the first The number of standardized datasets for each data subgroup For the feature dimension, This represents the estimated number of edges in the network. As a hyperparameter, iterate through and calculate the scores of all candidate Gaussian graphical mixture models along the regularization path, and select... The value corresponding to the minimum As the optimal regularization parameter; Step 3.3: Calculate the final sparse precision matrix using the optimal regularization parameters selected in Step 3.

2. and the final sparse precision matrix The partial correlation coefficient matrix is ​​transformed into a feature association network based on the partial correlation coefficient matrix. ,in, For nodes, For connecting edges, the final sparse precision matrix middle Then at node With nodes Establish connections between them, with the weights of the connections determined by the partial correlation coefficients. Decide: in, represents the element values ​​in a sparse precision matrix; Step 3.4: Output the feature association network of each subgroup. The topology diagram.

4. The high-dimensional data clustering and feature structure analysis method according to claim 3, characterized in that: Step 4 specifically includes the following steps: Step 4.1: Develop the feature association network in Step 3. The system is divided to simulate particles in a feature correlation network. The random walk process in the model utilizes the structural similarity between nodes to determine the structural distance. The smallest closely related features are aggregated into the same functional community, defining nodes. and nodes Structural distance between : in, The number of steps in the random walk. For the node go through Step to Node The probability, For nodes The degree; Step 4.2: Calculate the feature association network The centrality indices of strength, proximity, and betweenness of each node in the standardized dataset are used to introduce a stability verification mechanism. Perform sampling with replacement to construct data subpopulations of different sizes and calculate the standardized dataset. centrality vector With the centrality vector of the data subgroup The correlation between levels is quantitatively evaluated using the correlation stability coefficient to find those that meet the requirements. Maximum sample rejection ratio under the given conditions: in, For standardized datasets The total number, To maximize the size of the removable subsample, if If the centrality ranking of the feature is greater than the preset stability judgment threshold, it is determined that the centrality ranking of the feature is statistically robust, thus eliminating the interference of random error. Step 4.3: Calculate the feature association networks of the two sets of networks. Original observed differences in global strength, network structure dissimilarity, and edge weight differences. A nonparametric permutation test strategy is adopted to randomly shuffle sample labels and reconstruct the feature association network. Constructing the empirical zero distribution of difference statistics, statistics The value is: in, For the first Difference statistics under random permutations For indicator functions, For the preset number of permutation iterations, when The value is 1 when it is active and 0 otherwise. Multiple checks and corrections; Step 4.4: Summarize and Quantify the Assessment ,distance and statistics The analysis results are used to generate a final analysis report.

5. The high-dimensional data clustering and feature structure analysis method according to claim 1, characterized in that: The high-dimensional data clustering and feature structure analysis method is implemented through a high-dimensional data clustering and feature structure analysis system. This system includes a data preprocessing and standardization module, a GGMM adaptive clustering module, a feature association network construction module, and a network topology difference analysis module, wherein: The data preprocessing and standardization module collects high-dimensional data and preprocesses it to obtain a standardized dataset. The GGMM adaptive clustering module performs unsupervised clustering on the standardized dataset obtained by the data preprocessing and standardization module, and outputs a dataset containing category labels as input to the feature association network construction module. The feature association network construction module performs sparse structure learning and outputs a feature association network topology graph; The network topology difference analysis module takes the feature-related network topology graph as input and performs multi-dimensional graph theory index calculations and statistical tests.

Citation Information

Patent Citations

  • Large-scale data mining method based on multi-tuple data optimization

    CN119089404A

  • Attention deficit hyperactivity disorder subtype identification method based on brain network topology hub deviation

    CN120727247A