Data analysis method and device, electronic equipment and storage medium

By performing data distribution simulation and clustering on high-dimensional data, personalized thresholds are determined, and multiple rounds of feature screening are carried out to achieve efficient data analysis, which solves the problem of key features being filtered in traditional methods and improves the accuracy and reliability of the analysis.

CN120086560APending Publication Date: 2025-06-03INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510238681.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

Traditional high-dimensional data analysis techniques rely on unified thresholds or standards to filter differential features, resulting in some key features being filtered out, affecting the accuracy and reliability of the analysis results.

Method used

By simulating the data distribution of the features to be analyzed in multiple samples, the first distribution data is obtained; then the data clustering process is performed based on the first distribution data, and the first threshold and the second threshold of each sample are determined; then, the feature screening of the features to be analyzed is obtained, the first feature is obtained, and the secondary screening is performed to obtain the target feature; finally data analysis is performed based on the target feature.

Benefits of technology

This method can capture individualized features more accurately, improve the sensitivity and accuracy of feature screening, improve the accuracy and reliability of data analysis, and avoid the problem of key features being filtered out.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120086560A_ABST
    Figure CN120086560A_ABST
Patent Text Reader

Abstract

The invention discloses a data analysis method and device, electronic equipment and a storage medium, and relates to the technical field of data processing, data distribution simulation processing is carried out on a sample to adapt to characteristics of various non-normal distribution data, and then two thresholds of a data state of a to-be-analyzed feature in the sample can be determined to determine the data state of the to-be-analyzed feature in the sample. According to the method, the to-be-analyzed features are screened through the first feature and the second feature, the data state of the sample can be fully considered, so that the individualized features can be captured more accurately, furthermore, the volatility of the to-be-analyzed features can be further concerned by performing secondary screening on the screened to-be-analyzed features, namely the first feature, the sensitivity and accuracy of feature screening are improved, and the accuracy of feature screening is improved. And then data analysis is performed through the screened target features, so that the accuracy and reliability of data analysis are improved, and the technical effects of accurately identifying and screening key features and ensuring heterogeneity between samples in the data analysis process, thereby improving the accuracy and reliability of data analysis are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, and in particular, to a method and device for data analysis, an electronic device, and a storage medium. Background Art

[0002] In the field of artificial intelligence, especially in the application of large models, processing high-dimensional, non-normal distribution, sparse, and noisy data is a common challenge. These data usually come from diverse data sources and have complex distribution characteristics. Their core characteristics include high dimensionality, non-normal distribution, sparsity, and noise interference. High dimensionality means that each sample contains a large number of feature dimensions, which not only increases the complexity of data processing but also poses higher requirements on the storage and computing capabilities of the model. The characteristic of non-normal distribution makes traditional statistical methods based on the assumption of normal distribution difficult to apply, while sparsity is manifested as extremely low or near-zero values of certain features in specific samples, resulting in a large amount of invalid or redundant information in the data matrix. The heterogeneity between sample individuals, especially in the case of complex and diverse data distributions, may lead to the omission or misjudgment of key features. In addition, the interference of noise may mask the true signal of the data, thereby affecting the judgment and prediction accuracy of the model.

[0003] Although existing high-dimensional data analysis technologies have made certain progress, they still have limitations in processing non-normal distribution, sparse, and noisy data. Traditional high-dimensional data analysis usually relies on a unified threshold or standard to screen for differential features, which may filter out some key features, thus affecting the accuracy and reliability of the analysis results. Summary of the Invention

[0004] This application provides a method and device for data analysis, an electronic device, and a storage medium, so as to at least solve the problem that traditional high-dimensional data analysis in related technologies usually relies on a unified threshold or standard to screen for differential features, which may filter out some key features, thereby affecting the accuracy and reliability of the analysis results.

[0005] This application provides a method for data analysis, including:

[0006] Performing data distribution simulation processing on multiple features to be analyzed in multiple samples according to a preset shape parameter to obtain first distribution data; wherein, the preset shape parameter is a positive parameter used to define the distribution shape of the first distribution data, and the preset shape parameter includes at least two positive parameters;

[0007] Performing data clustering processing on the first distribution data to obtain a first threshold and a second threshold corresponding to each of the multiple samples; wherein, the first threshold and the second threshold are at least used to determine the data states of the multiple features to be analyzed in the multiple samples, and the second threshold is greater than the first threshold;

[0008] Feature screening processing is performed on multiple features to be analyzed according to the first threshold and the second threshold corresponding to each of the multiple samples, obtaining multiple first features, and secondary screening processing is performed on the multiple first features to obtain target features;

[0009] Data analysis processing is performed according to the target features to obtain a data analysis result.

[0010] This application also provides a data analysis device, including:

[0011] A simulation unit, configured to perform data distribution simulation processing on multiple features to be analyzed in multiple samples according to preset shape parameters, obtaining first distribution data; wherein, the preset shape parameters are positive parameters used to define the distribution shape of the first distribution data, and the preset shape parameters include at least two positive parameters;

[0012] A clustering unit, configured to perform data clustering processing according to the first distribution data, obtaining the first threshold and the second threshold corresponding to each of the multiple samples; wherein, the first threshold and the second threshold are at least used to determine the data status of multiple features to be analyzed in multiple samples, and the second threshold is greater than the first threshold;

[0013] A first screening unit, configured to perform feature screening processing on multiple features to be analyzed according to the first threshold and the second threshold corresponding to each of the multiple samples, obtaining multiple first features;

[0014] A second screening unit, configured to perform secondary screening processing on the multiple first features to obtain target features;

[0015] An analysis unit, configured to perform data analysis processing according to the target features to obtain a data analysis result.

[0016] This application also provides an electronic device, including: a memory, configured to store a computer program; a processor, configured to implement the steps of any of the above data analysis methods when executing the computer program.

[0017] This application also provides a computer-readable storage medium, in which a computer program is stored, wherein the computer program implements the steps of any of the above data analysis methods when executed by a processor.

[0018] This application also provides a computer program product, including a computer program, and the computer program implements the steps of any of the above data analysis methods when executed by a processor.

[0019] The method and device for data analysis, electronic device, and storage medium of the present application can adapt to the characteristics of various non-normal distribution data by performing data distribution simulation processing on samples. Then, by determining two thresholds that can determine the data state of the feature to be analyzed in the samples, the feature to be analyzed is screened. This can fully consider the data state of the samples, thereby being able to capture individual characteristics more precisely. Further, by performing secondary screening on the screened feature to be analyzed, i.e., the first feature, the volatility of the feature to be analyzed can be further concerned, improving the sensitivity and accuracy of feature screening. Then, data analysis is performed on the screened target features to improve the accuracy and reliability of data analysis. Therefore, it can solve the technical problem that traditional high-dimensional data analysis usually relies on unified thresholds or standards to screen differential features, which will filter out some key features, thus affecting the accuracy and reliability of the analysis results, and achieve the technical effect of accurately identifying and screening key features during the data analysis process, ensuring the heterogeneity between samples, and thereby improving the accuracy and reliability of data analysis. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] To more clearly illustrate the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0021] Figure 1 It is a schematic flowchart of a method for data analysis provided by an embodiment of the present application;

[0022] Figure 2 It is a schematic flowchart of another method for data analysis provided by an embodiment of the present application;

[0023] Figure 3 It is a schematic structural diagram of a device for data analysis provided by an embodiment of the present application;

[0024] Figure 4 It is a schematic structural diagram of another device for data analysis provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0025] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.

[0026] It should be noted that in the description of this application, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. The terms "first", "second", etc. in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0027] To enable those skilled in the art of this technology to better understand the solution of this application, the following further details this application in conjunction with the accompanying drawings and specific embodiments.

[0028] Figure 1 A flowchart of a method for data analysis provided for an embodiment of this application. In combination with the execution process of the method for data analysis, the method is described in detail.

[0029] As Figure 1 shown, the method for data analysis includes:

[0030] Step 101, perform data distribution simulation processing on multiple to-be-analyzed features in multiple samples according to preset shape parameters to obtain first distribution data; wherein, the preset shape parameters are positive parameters used to define the distribution shape of the first distribution data, and the preset shape parameters include at least two positive parameters.

[0031] In the embodiments of this application, the ways of performing data distribution simulation processing include but are not limited to: Beta distribution model, etc. The preset shape parameters are parameters pre-customized and selected, for example: the (alpha) parameter and (beta) parameter in the Beta distribution, etc. Specifically, this application does not limit the ways of data distribution simulation processing and the preset shape parameters.

[0032] Furthermore, the two positive parameters included in the preset shape parameters can jointly determine the distribution shape of the first distribution data. Data distribution simulation processing refers to using a statistical distribution (such as: Beta distribution) to simulate the distribution of each to-be-analyzed feature. The purpose is to generate data that conforms to specific statistical characteristics and can be used for further analysis, modeling or testing.

[0033] At the same time, it should be noted that before performing data distribution simulation processing, normalization processing needs to be performed on the to-be-analyzed features to facilitate improving the efficiency and accuracy of data processing. Therefore, before performing data distribution simulation processing on multiple to-be-analyzed features in multiple samples, the following methods can also be adopted but are not limited to: perform normalization processing on multiple initial features of multiple initial samples to obtain multiple to-be-analyzed features.

[0034] Specifically, the simulation processing of the data distribution of multiple features to be analyzed can be carried out by, but not limited to, formula (1):

[0035]

[0036] where f(x cnr |α,β)~Beta(x cnr |α,β) is the first distribution data, and x cnr represents the c-th (c = 1, …, c) feature in the r-th (r = 1, …, r) group of data of the n-th (n = 1, …, n) sample, that is, multiple features to be analyzed. α and β are two positive parameters in the preset shape parameters. α determines the size of the left tail of the Beta distribution curve in the interval [0, 1], and β determines the size of the right tail of the Beta distribution curve in the interval [0, 1]. Γ(·) is the gamma function, and the gamma function is also known as the second Euler integral, which is an extension of the factorial function to real and complex numbers.

[0037] Step 102: Perform data clustering processing according to the first distribution data to obtain the first threshold and the second threshold corresponding to each of the multiple samples; wherein, the first threshold and the second threshold are at least used to determine the data status of the multiple features to be analyzed in the multiple samples, and the second threshold is greater than the first threshold.

[0038] In the embodiments of the present application, the methods for performing data clustering processing include, but are not limited to: k-means clustering method, etc. Among them, each sample can be regarded as a heterogeneous population composed of the features to be analyzed in K groups or clusters, and each group or cluster represents a data status. The data status includes, but is not limited to: low-state features, medium-state features, high-state features, etc.

[0039] The first threshold and the second threshold are thresholds for distinguishing the data status of multiple features to be analyzed. For example: in a certain sample, the features to be analyzed less than the first threshold are low-state features, the features to be analyzed greater than or equal to the first threshold and less than or equal to the second threshold are medium-state features, and the features to be analyzed greater than the second threshold are high-state features.

[0040] The determination of the first threshold and the second threshold can adopt, but is not limited to: first, use the log-likelihood function of the complete data and use the Expectation-Maximization (EM) algorithm to determine the target mixing ratio of the features to be analyzed, and then determine the first threshold and the second threshold through the target mixing ratio. The mixing ratio refers to the probability that multiple features to be analyzed belong to different data statuses.

[0041] Step 103: Perform feature screening on multiple features to be analyzed according to the first threshold and the second threshold corresponding to each of the multiple samples, obtain multiple first features, and perform secondary screening on the multiple first features to obtain target features.

[0042] In the embodiments of the present application, regarding the screening of multiple features to be analyzed, first, screening is performed through the first threshold and the second threshold, and screening can be performed according to the data status of the features to be analyzed, so as to eliminate the features to be analyzed with relatively small differences and obtain the first features.

[0043] When performing secondary screening, methods such as the coefficient of variation and non-parametric rank sum test can be used but are not limited to these. In secondary screening, considering that the values of features in different samples may show certain fluctuations with changes in sample conditions, in order to measure the magnitude of data sequence fluctuations, the coefficient of variation can be used for screening, and then the non-parametric rank sum test is used to screen out the features to be analyzed with relatively large differences to obtain the target features.

[0044] Step 104: Perform data analysis processing based on the target features to obtain a data analysis result.

[0045] In the embodiments of the present application, the method of performing data analysis processing based on the target features is a custom-selected method. Performing data analysis based on the target features can improve the accuracy of data analysis. Specifically, regarding the process of data analysis processing, the following methods can be used but are not limited to these:

[0046] a. Data preparation: Ensure the data quality of the target features, including handling missing values, outliers, standardizing or normalizing data, etc. b. Model selection: Select a suitable statistical model or machine learning algorithm according to the purpose of data analysis. For example: If it is a classification problem, models such as logistic regression, support vector machine, decision tree, etc. may be selected; if it is a regression problem, models such as linear regression, ridge regression, etc. may be selected. c. Data analysis: Use the target features and the selected model to perform data analysis processing.

[0047] The method for data analysis of the present application can adapt to the characteristics of various non-normal distribution data through data distribution simulation processing of samples. Then, by determining two thresholds that can determine the data state of the feature to be analyzed in the sample, the feature to be analyzed is screened. This can fully consider the data state of the sample, so as to more accurately capture individual characteristics. Further, by performing secondary screening on the feature to be analyzed after screening, that is, the first feature, the volatility of the feature to be analyzed can be further concerned, and the sensitivity and accuracy of feature screening can be improved. Then, data analysis is performed on the target feature selected through screening to improve the accuracy and reliability of data analysis. Therefore, it can solve the technical problem that traditional high-dimensional data analysis usually relies on a unified threshold or standard to screen differential features, which will filter out some key features, thus affecting the accuracy and reliability of the analysis result, and achieve the technical effect of accurately identifying and screening key features during the data analysis process, ensuring the heterogeneity between samples, and thus improving the accuracy and reliability of data analysis.

[0048] In an implementable embodiment of the present application, as a refinement of the above step 101, when performing data distribution simulation processing, the following methods can also be used but are not limited to: performing data expansion processing on the preset shape parameter through a first preset function to obtain a random data set; performing data calculation processing according to multiple features to be analyzed and the preset shape parameter to obtain a feature data set; performing data distribution simulation processing according to the random data set and the feature data set to obtain first distribution data.

[0049] In the embodiment of the present application, the first preset function is a user-defined function, for example: the distribution function of the above formula (1), etc. Specifically, the present application does not limit the setting of the first preset function.

[0050] The purpose of data expansion processing is to generate a random data set, which means using the first preset function to transform the preset shape parameter to create a random sample with certain statistical characteristics, such as: B(α,β) in the above formula (1). The feature data set includes the interaction or combination between features, such as: x in the above formula (1) cnr α-1 (1 - x cnr ) β-1 。

[0051] Data distribution simulation processing refers to using the generated random data set and feature data set to perform data distribution simulation processing. It includes but is not limited to the following steps: a. Model fitting, for example: Beta model; b. Distribution feature calculation: calculating the statistical characteristics of the fitted distribution, such as: mean, variance, skewness, kurtosis, etc.; c. Simulation generation: based on the fitted distribution model, generating simulation data to simulate the possible distribution of real data.

[0052] After the data distribution simulation process, the first distribution data is obtained, which can reflect the distribution characteristics of the feature to be analyzed after considering the random parameters, providing a basis for understanding the behavior of the data characteristics under the influence of randomness.

[0053] In an implementable embodiment of the present application, as a refinement of the above step 102, when performing data clustering processing, the following method can be used but is not limited to: The first distribution data includes the initial shape parameters and the initial mixing ratios corresponding to each of the multiple samples. Through the second preset function, the initial shape parameters and the initial mixing ratios corresponding to each of the multiple samples are subjected to the expectation maximization process to obtain the first threshold and the second threshold corresponding to each of the multiple samples; wherein, the initial shape parameter is a parameter describing the data state of the multiple features to be analyzed, and the initial mixing ratio is the probability that the multiple features to be analyzed belong to different data states.

[0054] In the embodiment of the present application, the initial shape parameters and the initial mixing ratios are the data carried in the first distribution data. After the data distribution simulation process is performed according to two positive parameters α and β in the preset shape parameters, the initial shape parameters and the initial mixing ratios can be obtained. It can be understood that the positive parameters α and β are two randomly set shape parameters.

[0055] The second preset function is a user-defined function, for example: the log-likelihood function of the complete data, etc. Specifically, the present application does not limit the setting of the second preset function.

[0056] Regarding the determination principle of the first threshold and the second threshold, the following method can be used but is not limited to: Since each sample can be regarded as a heterogeneous population composed of the features to be analyzed in K groups or clusters, and the data states include but are not limited to: low-state features, medium-state features, and high-state features.

[0057] Therefore, it can be considered that each feature to be analyzed is in one of the K = 3 data states, and the data state is characterized by each K cluster. At the same time, to determine the thresholds between the states (i.e., the first threshold and the second threshold), it is necessary to determine through the initial shape parameters and the initial mixing ratios. Among them, θ represents the initial shape parameter, θ = (α 1 , β 1 , …, α K , β K ), α k and β k represent the shape parameters of the kth cluster. ρ represents the initial mixing ratio, ρ = (ρ 1 , …, ρ K ), between 0 and 1, and ρ kRepresents the mixing ratio belonging to the k-th cluster. At this time, the first threshold and the second threshold can be determined by calculating the probability density of the feature to be analyzed. Regarding the calculation of the probability density, formula (2) can be used but is not limited to:

[0058]

[0059] Among them, c is each feature to be analyzed, k represents the k-th cluster of multiple features to be analyzed, α knr , β knr respectively represent the two shape parameters, i.e., the initial shape parameters, of the k-th cluster feature in the r-th (r = 1, …, r) group of data of the n-th (n = 1, …, n) sample, ρ k is the mixing ratio of the k-th cluster, i.e., the initial mixing ratio, f(X|τ,θ) is the probability density of the feature to be analyzed, f(X|α k , β k ) represents the distribution data corresponding to the k-th cluster of multiple features to be analyzed, Beta(x cnr |α knr , β knr ) can also represent the first distribution data. At this time, the probability density that demarcates the low-state feature and the middle-state feature can be used as the first threshold, and the probability density that demarcates the middle-state feature and the high-state feature can be used as the second threshold.

[0060] Furthermore, in addition to using the probability density, the first threshold and the second threshold can also be determined by calculating the maximum likelihood values of ρ and θ. To calculate the maximum likelihood of ρ and θ, a latent vector y c =(y c1 , …, y cK ) needs to be introduced for each feature c to be analyzed. Among them, if the feature c to be analyzed belongs to the k-th group, then y ck is 1, otherwise y c is 0. The C×K matrix Y and the first distribution data are combined to form the complete data (X,Y). Regarding the calculation of the maximum likelihood values of ρ and θ, formula (3) can be used but is not limited to:

[0061]

[0062] Among them, y ck is the latent vector that the feature c to be analyzed belongs to the k-th cluster. The maximum likelihood values of ρ and θ can be determined through l C (ρ,θ,Y|X) and Furthermore, at convergence, the probability clustering solution can also be obtained from the expected value of y ck , that is, the posterior probability that the feature c to be analyzed belongs to the clustering cluster k, ι C(ρ,θ,Y|θ) represents the likelihood values of the shape parameter and the mixing ratio obtained after performing likelihood logarithmic estimation on the feature c to be analyzed, that is, the maximum likelihood values of ρ and θ and

[0063] Take the minimum value among all the maximum likelihood values of ρ as the first threshold, and take the maximum value among all the maximum likelihood values of ρ as the second threshold.

[0064] Optimize the initial shape parameter and the initial mixing ratio through a second preset function to obtain the first threshold and the second threshold for feature screening, which helps to improve the fitting degree and prediction ability of the data. The data status of the samples can be fully considered through the first threshold and the second threshold, so as to be able to capture individual features more precisely.

[0065] In an implementable embodiment of the present application, when performing expectation maximization processing on the initial shape parameter and the initial mixing ratio, the following methods can also be used but are not limited to: perform likelihood estimation processing according to the initial shape parameter and the initial mixing ratio corresponding to each of multiple samples to obtain the first shape parameter and the first mixing ratio corresponding to each of the multiple samples, and perform fitting density calculation processing according to the first shape parameter and the first mixing ratio corresponding to each of the multiple samples to obtain the fitting density estimation values corresponding to each of the multiple samples; in the case where any of the fitting density estimation values corresponding to the multiple samples is less than the density threshold, perform likelihood estimation processing according to the first shape parameter and the first mixing ratio corresponding to each of the multiple samples to obtain the second shape parameter and the second mixing ratio corresponding to each of the multiple samples; perform fitting density calculation processing according to the second shape parameter and the second mixing ratio corresponding to each of the multiple samples to obtain the fitting density estimation values corresponding to each of the multiple samples; repeat the above steps until the fitting density estimation values corresponding to each of the multiple samples are all greater than or equal to the density threshold, and obtain the first target mixing ratio corresponding to each of the multiple samples that meets the first preset condition, and the second target mixing ratio corresponding to each of the multiple samples that meets the second preset condition; determine the first target mixing ratio corresponding to each of the multiple samples as the first threshold corresponding to each of the multiple samples, and determine the second target mixing ratio corresponding to each of the multiple samples as the second threshold corresponding to each of the multiple samples.

[0066] In the embodiment of the present application, the density threshold is a threshold set by the user, for example: 1, etc. Specifically, the present application does not limit the setting of the density threshold.

[0067] When performing expectation maximization, it includes two steps: the expectation step and the maximization step. First, in the expectation step, the expected value of the complete-data log-likelihood function is obtained based on the initial shape parameters and the initial mixture proportion estimates. Then, in the maximization step, the parameter estimates are optimized by maximizing the expected complete-data log-likelihood function to obtain the maximum likelihood value. and Iterate the expectation and maximization steps until the log-likelihood function converges to at least one local optimum.

[0068] Specifically, regarding the calculation of the expected value of the complete-data log-likelihood function, that is, the method of calculating the first shape parameter and the first mixture proportion based on the initial shape parameters and the initial mixture proportion, it can be carried out by, but not limited to, formula (4):

[0069]

[0070] where (t) represents the number of iterations, α in· , β in· represent the shape parameters corresponding to the features to be analyzed in the i-th cluster among n samples, x cn· represents the feature c to be analyzed among n samples, that is, multiple features to be analyzed, ρ i represents the mixture proportion corresponding to the feature to be analyzed in the i-th cluster, α kn. , β kn. represent the shape parameters corresponding to the features to be analyzed in the k-th cluster among n samples.

[0071] ρ k represents the mixture proportion corresponding to the feature to be analyzed in the k-th cluster, E[y ck |X, ρ (t) , θ (t) includes performing likelihood logarithmic estimation on the corresponding feature c to be analyzed and, after iteration, obtaining the shape parameters and mixture proportions, represents the posterior probability that the feature c to be analyzed belongs to the k-th cluster.

[0072] Regarding the fitting density calculation process, it can be carried out by, but not limited to, formula (5):

[0073]

[0074] where ω i represents the fitting density estimate value, and the threshold for separating low-state data and medium-state data is the minimum value of ρ 1 when ω 1 ≥1. Similarly, the threshold for dividing medium-state data and high-state data is the maximum value of ρ 2 when ω 2 ≥1. f(X|α k , βk ) represents the distribution data corresponding to the k-th cluster of multiple features to be analyzed, f(X|α i ,β i ) represents the distribution data corresponding to the i-th cluster of multiple features to be analyzed. Specifically, ρ 1 represents the mixing ratio corresponding to the features to be analyzed in the first cluster, that is, the mixing ratio corresponding to the low-state data, ρ 2 represents the mixing ratio corresponding to the features to be analyzed in the second cluster, that is, the mixing ratio corresponding to the medium-state data.

[0075] Further, it should be noted that the first preset condition is some conditions set by the user. For example: when the fitting density estimation values corresponding to multiple samples are all greater than or equal to the density threshold, likelihood estimation processing is performed through the initial shape parameters and initial mixing ratios corresponding to multiple samples, and the minimum mixing ratio obtained is the first target mixing ratio.

[0076] The second preset condition is also some conditions set by the user. Further, it should be noted that there are limit values for performing likelihood estimation processing on the initial shape parameters and initial mixing ratios corresponding to multiple samples, that is, the maximum mixing ratios and maximum shape parameters corresponding to multiple samples will be obtained. The setting of the second preset condition is related to the maximum mixing ratio. For example: when the fitting density estimation values corresponding to multiple samples are all greater than or equal to the density threshold, likelihood estimation processing is performed through the initial shape parameters and initial mixing ratios corresponding to multiple samples, and the maximum mixing ratio obtained is the second target mixing ratio.

[0077] By continuously adjusting the shape parameters and mixing ratios, the fitting degree of the model to the data distribution is improved until a mixing ratio that meets specific conditions is found. The mixing ratio of the specific conditions is then used as the threshold for feature screening, which helps to improve the performance and accuracy of feature screening.

[0078] In an implementable embodiment of the present application, when performing feature screening processing on multiple features to be analyzed according to the first threshold and the second threshold, the following methods can also be used but are not limited to: obtaining the initial feature values corresponding to multiple features to be analyzed; among multiple features to be analyzed, removing the features to be analyzed whose initial feature values are greater than the first threshold and less than the second threshold to obtain multiple first features.

[0079] In the embodiment of the present application, for multiple features to be analyzed, first, the initial feature values corresponding to multiple features to be analyzed need to be obtained. The initial feature values can be the values in the original data or the values calculated through a certain preprocessing method.

[0080] Intermediate state data of the feature to be analyzed with the initial eigenvalue greater than the first threshold and less than the second threshold. Since the feature differences of the intermediate state data are small, the intermediate state data needs to be removed to obtain the first feature.

[0081] By removing features with eigenvalues within a specific range to reduce the size of the feature set, the data state of the samples can be fully considered, so that individual features can be captured more accurately.

[0082] In an implementable embodiment of the present application, when performing a secondary screening process on multiple first features, the following methods can be used but are not limited to: obtaining the coefficient of variation corresponding to each of the multiple first features; among the multiple first features, removing the first features with a coefficient of variation less than the preset coefficient threshold to obtain multiple second features; performing a screening process on the multiple second features to obtain the target feature.

[0083] In the embodiment of the present application, the preset coefficient threshold is a threshold set by the user. For example: 0.15, etc. Specifically, the present application does not limit the preset coefficient threshold.

[0084] Considering that the values of features in different samples may show certain fluctuations with the change of sample conditions, in order to measure the magnitude of the data sequence fluctuations, the coefficient of variation is introduced. If the coefficient of variation of the first feature corresponding to a feature is greater than 0.15, it indicates that the first feature has significant fluctuations, and this feature information helps to distinguish samples under different conditions.

[0085] Through a multi-stage feature selection process, first remove features with insufficient fluctuations based on the coefficient of variation, and then obtain the final target feature through further screening. This helps to focus on the fluctuations of the features to be analyzed and improve the sensitivity and accuracy of feature screening.

[0086] In an implementable embodiment of the present application, when performing a screening process on multiple second features to obtain the target feature, the following methods can be used but are not limited to: determining different types of samples among multiple samples; and performing a non-parametric test process on the multiple second features according to different types of samples to obtain the test results; among the multiple second features, removing the second features with test results greater than or equal to the preset test threshold to obtain the target feature.

[0087] In the embodiment of the present application, the preset test threshold is a threshold set by the user. For example: 0.05, etc. Specifically, the present application does not limit the preset test threshold.

[0088] Further, regarding the implementation process of this embodiment, for example: perform non-parametric tests on the features to be analyzed of samples under different conditions to obtain the test result p-value. If the p-value is less than 0.05, it can be considered that there are significant differences in this feature among samples under different conditions. The smaller the p-value, the more significant the difference, and the greater the possibility of this feature being a differential feature. Arrange the calculated p-values in ascending order and select the first part of the features as candidate differential features. Conduct multiple tests to determine how many features should be finally selected, thereby optimizing the number of selected features.

[0089] Identify and retain those features with significant differences between different sample types through non-parametric tests, improving the sensitivity and accuracy of feature screening.

[0090] Further, to facilitate the understanding of the implementation process of the data analysis method of this application, a flowchart of another data analysis method provided by the embodiment of this application is as Figure 2 shown, where beta distribution modeling is used to simulate high-dimensional data distribution (corresponding to the above step 101), which can flexibly adapt to the data characteristics of various non-normal distributions, thereby improving the adaptability and accuracy of the model.

[0091] Construct a KN clustering method to determine the differential feature threshold (corresponding to the above step 102). The KN clustering method based on Beta data distribution can cluster the data into different states (such as: low state, medium state, and high state), and objectively infer the threshold between these states. Calculate a set of specific differential feature thresholds for each sample. The differential feature threshold can more accurately reflect the data characteristics of individual samples, thereby improving the accuracy and reliability of the model in individual analysis.

[0092] According to the obtained threshold, combine the coefficient of variation and the rank sum test method to construct a screening model to screen differential features (corresponding to the above step 103), and identify the features with significant differences by combining the coefficient of variation and non-parametric rank sum test. Use the coefficient of variation to measure the volatility of feature values; use the non-parametric test method to screen out the differential features closely related to the sample conditions. This application has high sensitivity and accuracy in identifying differential features, optimizes data quality, and can provide a reference for the analysis and research in the field of high-dimensional data.

[0093] Scenario application and method effect verification (corresponding to the above step 104) can accurately identify and screen out key features during the data analysis process, ensuring the heterogeneity between samples, thereby improving the accuracy and reliability of data analysis.

[0094] Regarding the method of data analysis of the present application, the present application also provides an application scenario embodiment for illustration. Specifically, for example, taking the genomic data of tumor samples as an example, the differential methylation characteristic sites are extracted by the present application to assist subsequent analysis, so as to help predict disease risks, assist in diagnosis and formulate personalized treatment plans. Each NR column of this data contains the methylation levels of C cytosine-phosphate-guanine (CpG) sites in one of R sample types from N patients. By analyzing the methylation distribution status of these data, the present application can screen out methylation sites with significant differences. These differential sites may serve as potential biomarkers, providing important bases for disease prediction, diagnosis and treatment, thereby improving the efficiency and accuracy of data analysis. The specific embodiments are as follows:

[0095] 1) Obtain the differential methylation thresholds of different samples by the KN clustering method:

[0096] In order to cluster the CpG sites in the sample type into three clusters representing three methylation states and infer the thresholds between these states, the KN model is applied to fit the dataset modeled by the beta distribution. The methylation state thresholds for different samples are inferred, so that more hypomethylated and hypermethylated CpG sites are identified as differentially methylated sites, and there are individualized methylation state thresholds for different samples.

[0097] 2) Construct a screening model to screen differentially methylated sites and predict the tumor purity estimate:

[0098] Taking into account the two conditions of "the mean methylation level of CpG sites in tumor samples needs to be within the differential methylation threshold range" and "the coefficient of variation is greater than 0.15", the sample data is screened to retain the data that meets the conditions. Next, the non-parametric Wilcoxon rank-sum test is performed on the methylation levels of the tumor sample set and the normal sample set of each site screened out to evaluate the significance of the difference. Experiments found that when the p-value of the top 1100 sites is selected as the final differentially methylated sites, the Pearson correlation coefficient between the obtained tumor sample purity results and the Bayesian Segmentation Model for Tumor Mutation Burden and Hypermutation (ABSOLUTE) is the highest, thus verifying the effectiveness of this method.

[0099] In this embodiment, not only the characteristic that the methylation levels of differentially methylated sites fluctuate greatly in tumor samples is considered, but also the property that the means of these sites should mostly be in the intermediate methylation region is considered. Subsequently, the non-parametric rank sum test method is used to further screen out the differentially methylated sites closely related to tumor purity. The experimental results verify the effectiveness and reliability of the data analysis method, and further prove the potential application value of this method in the study of high-dimensional data characteristics.

[0100] In summary, this application can achieve the following technical effects:

[0101] The data analysis method of this application can adapt to the characteristics of various non-normal distribution data by performing data distribution simulation processing on samples. Then, by determining two thresholds that can determine the data state of the feature to be analyzed in the sample, the feature to be analyzed is screened. This can fully consider the data state of the sample, so as to more accurately capture individual characteristics. Further, by performing secondary screening on the screened feature to be analyzed, that is, the first feature, the volatility of the feature to be analyzed can be further concerned, improving the sensitivity and accuracy of feature screening. Then, data analysis is performed on the screened target feature to improve the accuracy and reliability of data analysis. Therefore, it can solve the technical problem that traditional high-dimensional data analysis usually relies on unified thresholds or standards to screen differential features, which will filter out some key features, thus affecting the accuracy and reliability of the analysis results, and achieve the technical effect of accurately identifying and screening key features during the data analysis process, ensuring the heterogeneity between samples, and thus improving the accuracy and reliability of data analysis.

[0102] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method.

[0103] The embodiment of this application also provides a data analysis device, Figure 3 which is a schematic structural diagram of a data analysis device provided by this application, as Figure 3 shown, including:

[0104] A simulation unit 31, configured to perform data distribution simulation processing on multiple features to be analyzed in multiple samples according to preset shape parameters to obtain first distribution data; wherein, the preset shape parameters are positive parameters used to define the distribution shape of the first distribution data, and the preset shape parameters include at least two positive parameters;

[0105] A clustering unit 32, configured to perform data clustering processing according to first distribution data to obtain a first threshold and a second threshold corresponding to each of multiple samples; wherein, the first threshold and the second threshold are at least used to determine data states of multiple to-be-analyzed features in the multiple samples, and the second threshold is greater than the first threshold.

[0106] A first screening unit 33, configured to perform feature screening processing on multiple to-be-analyzed features according to the first threshold and the second threshold corresponding to each of the multiple samples to obtain multiple first features.

[0107] A second screening unit 34, configured to perform secondary screening processing on the multiple first features to obtain target features.

[0108] An analysis unit 35, configured to perform data analysis processing according to the target features to obtain a data analysis result.

[0109] In an embodiment of the present application, as Figure 4 shown, the simulation unit 31 includes:

[0110] An expansion module 311, configured to perform data expansion processing on a preset shape parameter through a first preset function to obtain a random data set.

[0111] A calculation module 312, configured to perform data calculation processing according to multiple to-be-analyzed features and a preset shape parameter to obtain a feature data set.

[0112] A simulation module 313, configured to perform data distribution simulation processing according to the random data set and the feature data set to obtain first distribution data.

[0113] In an embodiment of the present application, the clustering unit 32 is further configured to perform expectation maximization processing on the initial shape parameter and the initial mixing ratio corresponding to each of the multiple samples through a second preset function to obtain the first threshold and the second threshold corresponding to each of the multiple samples.

[0114] Wherein, the initial shape parameter is a parameter describing the data states of multiple to-be-analyzed features, and the initial mixing ratio is the probability that multiple to-be-analyzed features belong to different data states.

[0115] In an embodiment of the present application, as Figure 4 shown, the clustering unit 32 includes:

[0116] An estimation module 321, configured to perform likelihood estimation processing according to the initial shape parameter and the initial mixing ratio corresponding to each of the multiple samples to obtain a first shape parameter and a first mixing ratio corresponding to each of the multiple samples.

[0117] A calculation module 322, configured to perform fitting density calculation processing according to the first shape parameters and the first mixing ratios corresponding to multiple samples respectively, so as to obtain fitting density estimation values corresponding to the multiple samples respectively;

[0118] The estimation module 321 is further configured to, when the fitting density estimation values corresponding to any multiple samples are less than a density threshold, perform likelihood estimation processing according to the first shape parameters and the first mixing ratios corresponding to the multiple samples respectively, so as to obtain second shape parameters and second mixing ratios corresponding to the multiple samples respectively;

[0119] The calculation module 322 is further configured to perform fitting density calculation processing according to the second shape parameters and the second mixing ratios corresponding to the multiple samples respectively, so as to obtain fitting density estimation values corresponding to the multiple samples respectively;

[0120] A determination module 323, configured to repeat the above steps until the fitting density estimation values corresponding to the multiple samples are all greater than or equal to the density threshold, so as to obtain a first target mixing ratio corresponding to the multiple samples that meets a first preset condition, and a second target mixing ratio corresponding to the multiple samples that meets a second preset condition;

[0121] The determination module 323 is further configured to determine the first target mixing ratio corresponding to the multiple samples as the first threshold corresponding to the multiple samples respectively, and determine the second target mixing ratio corresponding to the multiple samples as the second threshold corresponding to the multiple samples respectively.

[0122] In an embodiment of the present application, as Figure 4 shown, the first screening unit 33 includes:

[0123] An acquisition module 331, configured to acquire initial feature values corresponding to multiple features to be analyzed respectively;

[0124] An elimination module 332, configured to eliminate the features to be analyzed with initial feature values greater than the first threshold and less than the second threshold among the multiple features to be analyzed, so as to obtain multiple first features.

[0125] In an embodiment of the present application, as Figure 4 shown, the second screening unit 34 includes:

[0126] An acquisition module 341, configured to acquire the coefficient of variation corresponding to each of the multiple first features;

[0127] An elimination module 342, configured to eliminate the first features with a coefficient of variation less than a preset coefficient threshold among the multiple first features, so as to obtain multiple second features;

[0128] A screening module 343, configured to perform screening processing on the multiple second features to obtain target features.

[0129] In one embodiment of the present application, the screening module 343 is further configured to:

[0130] Determine different types of samples among a plurality of samples; and perform non-parametric test processing on a plurality of second features according to the different types of samples to obtain a test result;

[0131] Among the plurality of second features, remove the second features whose test results are greater than or equal to a preset test threshold to obtain target features.

[0132] For the description of the features in the corresponding embodiment of the data analysis device, reference may be made to the relevant description in the corresponding embodiment of the data analysis method, which will not be elaborated here one by one.

[0133] An embodiment of the present application further provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any of the above-mentioned data analysis method embodiments.

[0134] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in any of the above-mentioned data analysis method embodiments when running.

[0135] In an exemplary embodiment, the above-mentioned computer-readable storage medium may include, but is not limited to: USB flash drive, read-only memory (ROM for short), random access memory (RAM for short), mobile hard disk, magnetic disk or optical disc and other various media that can store computer programs.

[0136] An embodiment of the present application further provides a computer program product. The above-mentioned computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above-mentioned data analysis method embodiments.

[0137] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above-mentioned data analysis method embodiments.

[0138] Those skilled in the art may further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered as exceeding the scope of this application.

[0139] The above has introduced in detail a method and apparatus for data analysis, an electronic device, and a storage medium provided by this application. Specific examples are used herein to elaborate on the principles and implementation manners of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application. It should be noted that for those of ordinary skill in the art, without departing from the principles of this application, several improvements and modifications can be made to this application, and these improvements and modifications also fall within the protection scope of the claims of this application.

Claims

1. A method for data analysis, characterized in that: include: Performing data distribution simulation processing on a plurality of features to be analyzed in a plurality of samples according to preset shape parameters to obtain first distribution data; wherein the preset shape parameters are positive parameters for defining the distribution shape of the first distribution data, and the preset shape parameters include at least two positive parameters; Performing data clustering processing according to the first distribution data to obtain a first threshold and a second threshold corresponding to each of the multiple samples; wherein the first threshold and the second threshold are at least used to determine the data status of the multiple features to be analyzed in the multiple samples, and the second threshold is greater than the first threshold; Performing feature screening processing on the multiple features to be analyzed according to the first threshold and the second threshold corresponding to each of the multiple samples to obtain multiple first features, and performing secondary screening processing on the multiple first features to obtain target features; Data analysis and processing are performed according to the target characteristics to obtain data analysis results.

2. The data analysis method according to claim 1, characterized in that: The step of performing data distribution simulation processing on a plurality of features to be analyzed in a plurality of samples according to preset shape parameters to obtain first distribution data comprises: Performing data expansion processing on the preset shape parameters by using a first preset function to obtain a random data set; Performing data calculation and processing according to the plurality of features to be analyzed and the preset shape parameters to obtain a feature data set; Data distribution simulation processing is performed according to the random data set and the characteristic data set to obtain the first distribution data.

3. The data analysis method according to claim 1, characterized in that: The first distribution data includes initial shape parameters and initial mixing ratios corresponding to each of the multiple samples, and the data clustering process is performed according to the first distribution data to obtain first thresholds and second thresholds corresponding to each of the multiple samples, including: Performing expectation maximization processing on the initial shape parameters and the initial mixing ratio corresponding to each of the plurality of samples through a second preset function to obtain the first threshold and the second threshold corresponding to each of the plurality of samples; The initial shape parameter is a parameter describing the data state of the multiple features to be analyzed, and the initial mixing ratio is the probability that the multiple features to be analyzed belong to different data states.

4. The data analysis method according to claim 3, characterized in that: The performing expectation maximization processing on the initial shape parameters and the initial mixing ratio corresponding to each of the plurality of samples by using a second preset function to obtain the first threshold and the second threshold corresponding to each of the plurality of samples comprises: Performing likelihood estimation processing according to the initial shape parameters and the initial mixing ratio corresponding to each of the multiple samples to obtain first shape parameters and first mixing ratios corresponding to each of the multiple samples, and performing fitting density calculation processing according to the first shape parameters and the first mixing ratio corresponding to each of the multiple samples to obtain fitting density estimation values ​​corresponding to each of the multiple samples; When the fitting density estimation value corresponding to any of the plurality of samples is less than a density threshold, likelihood estimation processing is performed according to the first shape parameters and the first mixing ratio corresponding to the plurality of samples to obtain second shape parameters and second mixing ratio corresponding to the plurality of samples; Perform fitting density calculation processing according to the second shape parameters and the second mixing ratio corresponding to each of the multiple samples to obtain the fitting density estimation value corresponding to each of the multiple samples; Repeat the above steps until the fitted density estimation values ​​corresponding to the multiple samples are all greater than or equal to the density threshold, thereby obtaining a first target mixing ratio corresponding to the multiple samples that meets the first preset condition and a second target mixing ratio corresponding to the multiple samples that meets the second preset condition; The first target mixture ratio corresponding to each of the plurality of samples is determined as the first threshold corresponding to each of the plurality of samples, and the second target mixture ratio corresponding to each of the plurality of samples is determined as the second threshold corresponding to each of the plurality of samples.

5. The data analysis method according to claim 1, characterized in that: The performing feature screening processing on the plurality of features to be analyzed according to the first threshold and the second threshold to obtain the plurality of first features comprises: Obtaining initial feature values ​​corresponding to each of the plurality of features to be analyzed; Among the multiple features to be analyzed, the features to be analyzed whose initial feature values ​​are greater than the first threshold and less than the second threshold are eliminated to obtain the multiple first features.

6. The data analysis method according to claim 1, characterized in that: The performing secondary screening on the plurality of first features to obtain the target features comprises: Obtaining the coefficient of variation corresponding to each of the plurality of first features; Among the multiple first features, first features whose coefficient of variation is less than a preset coefficient threshold are eliminated to obtain multiple second features; The plurality of second features are screened to obtain the target feature.

7. The data analysis method according to claim 6, characterized in that: The screening process of the plurality of second features to obtain the target feature comprises: Determine different types of samples among the multiple samples; and perform non-parametric test processing on the multiple second features according to the different types of samples to obtain test results; Among the multiple second features, the second features whose inspection results are greater than or equal to a preset inspection threshold are eliminated to obtain the target feature.

8. A data analysis device, characterized in that: include: A simulation unit, used for performing data distribution simulation processing on a plurality of features to be analyzed in a plurality of samples according to preset shape parameters to obtain first distribution data; wherein the preset shape parameters are positive parameters for defining the distribution shape of the first distribution data, and the preset shape parameters include at least two positive parameters; a clustering unit, configured to perform data clustering processing according to the first distribution data, to obtain a first threshold and a second threshold corresponding to each of the plurality of samples; wherein the first threshold and the second threshold are at least used to determine the data states of the plurality of features to be analyzed in the plurality of samples, and the second threshold is greater than the first threshold; A first screening unit, configured to perform feature screening processing on the plurality of features to be analyzed according to the first threshold and the second threshold corresponding to each of the plurality of samples, to obtain a plurality of first features; A second screening unit, configured to perform a secondary screening process on the plurality of first features to obtain a target feature; The analysis unit is used to perform data analysis and processing according to the target characteristics to obtain data analysis results.

9. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the data analysis method according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the data analysis method according to any one of claims 1 to 7.