Semi-supervised feature selection method using cost-entropy information granulation for diabetes detection
By using the Spark framework for parallel processing and weighted granularity construction in diabetes detection, combined with cost entropy to evaluate feature importance, the efficiency and accuracy issues in high-dimensional data processing are solved, enabling more efficient feature selection and early screening.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NANTONG UNIV
- Filing Date
- 2026-01-20
- Publication Date
- 2026-06-05
AI Technical Summary
Existing technologies for diabetes detection suffer from several problems, including low efficiency in processing high-dimensional medical data, insufficient consideration of the differences in clinical misdiagnosis costs, inadequate utilization of unlabeled samples, and difficulty in balancing the correlation between global and local features, leading to a decline in the accuracy of early screening.
We employ the Spark framework for parallel processing, combining covariance matrix analysis and hemispherical partitioning strategies to construct weighted granular compartments, and evaluate feature importance through cost entropy to optimize feature selection.
It improves the classification accuracy of multidimensional feature datasets, reduces computational complexity, enhances the accuracy and efficiency of feature selection, makes better use of unlabeled data, and strengthens the robustness of the model.
Smart Images

Figure CN122158156A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of smart healthcare technology, and in particular to a semi-supervised feature selection method for cost entropy information granulation in diabetes detection. Background Technology
[0002] With the development of smart healthcare technology, diabetes detection faces the critical challenge of processing high-dimensional medical data. Traditional feature selection methods generally suffer from low computational efficiency, insufficient consideration of the differences in clinical misdiagnosis costs, and inadequate utilization of massive amounts of unlabeled samples when processing multi-source heterogeneous medical data such as blood glucose indicators and metabolic parameters. Especially in semi-supervised learning scenarios, existing algorithms struggle to effectively balance the global correlation of diabetes features (such as the interaction between insulin resistance and multiple organs) and local specificity (such as blood glucose time-series fluctuation patterns), leading to a significant decrease in early screening accuracy. This invention proposes a semi-supervised feature selection method based on cost entropy hemispherical partitioning. This method innovatively integrates the quantification of clinical misdiagnosis costs into feature importance assessment, analyzes the nonlinear correlation between features through covariance matrix analysis, optimizes the feature space structure by combining a hemispherical partitioning strategy, and utilizes the Spark framework to achieve efficient parallel computation, providing a more efficient and reliable technical solution for intelligent diabetes diagnosis.
[0003] Current mainstream methods are mostly limited to a single global or local feature analysis perspective, making it difficult to comprehensively consider both the global and local characteristics of the data, leading to a one-sidedness in feature selection. Furthermore, the processing strategies for unlabeled samples are relatively coarse, lacking a refined decision class partitioning mechanism, further affecting the accuracy of classification tasks. In addition, traditional feature selection algorithms often ignore the cost of classification errors when evaluating feature importance, causing the selected feature subset to fail to achieve optimal classification results in practical applications. For example, Shu et al. proposed a particle sphere method based on neighborhood discriminability, generating pseudo-labels and evaluating feature importance through dynamic particle sphere partitioning; however, this method is sensitive to parameter selection during particle sphere construction and has high computational complexity, limiting its application on large-scale datasets. Moreover, while the Zentropy-based method proposed by Yuan et al. achieves efficient feature selection through a novel uncertainty metric, its objective function does not consider the different costs of misclassification, potentially affecting the robustness of the model in practical applications. Balancing global and local information is also a key issue in feature selection. Although the Semi2MNR method proposed by Qian et al. balances feature selection by minimizing neighborhood redundancy and maximizing neighborhood relevance, it only considers the sample's... The analysis of neighbor relationships overemphasizes local neighborhood analysis and fails to effectively capture global feature associations.
[0004] In the field of parallel processing, although some methods have attempted to improve processing efficiency by leveraging distributed computing frameworks, their parallel computing capabilities have not been fully utilized during data partitioning and information granulation construction, resulting in significant room for improvement in overall computational efficiency. More importantly, existing methods generally face excessive computational resource and time consumption when processing large-scale datasets, making it difficult to meet the stringent real-time and efficiency requirements of practical applications. Summary of the Invention
[0005] The purpose of this invention is to provide a semi-supervised feature selection method with cost entropy information granularization for diabetes detection, so as to improve the classification accuracy of multidimensional feature datasets, reduce the computational complexity of classification methods, and improve the accuracy and efficiency of feature selection.
[0006] The inventive concept of this invention is as follows: This invention provides an efficient data processing and analysis method that leverages the Spark framework to achieve parallel processing and improve computational efficiency. It utilizes the covariance matrix for feature correlation analysis and combines a hemispherical partitioning strategy to optimize the feature space. Furthermore, this invention introduces weighted granular compartments, dividing the compartments into core and boundary compartments based on the sample distribution within each compartment. This method aims to improve the classification accuracy of multidimensional feature datasets, reduce the computational complexity of classification methods, and enhance the accuracy and efficiency of feature selection.
[0007] To achieve the above-mentioned objectives, the present invention employs the following technical solution: a semi-supervised feature selection method for cost entropy information granularization in diabetes detection, comprising the following steps: S1. Use Spark's in-memory computing technology to divide the diabetes dataset into different subsets and construct weighted granular compartments; S2. Based on the principal direction analysis of the covariance matrix, each decision class is hemispherically granulated, and pseudo-labels are assigned to unlabeled samples. S3. Introduce cost entropy to evaluate features based on feature differences and obtain feature importance; S4. Rank the features by importance and select a subset of features; S5. Aggregate the results of feature selection from each child node to obtain the optimal feature subset.
[0008] Further, step S1 includes the following steps: S11. Read the diabetes dataset, determine its feature set and decision classification, and the decision information system... ,in It is a labeled sample set. It is an unlabeled sample set. It is a feature set. It's a decision-making class. Using Spark to process the diabetes dataset... Divide into subsets This method fully leverages the advantages of distributed computing and can effectively improve the efficiency of processing large-scale datasets by processing these subsets in parallel on multiple nodes. S12. Construction of Virtual Samples: Based on the feature values in the diabetes dataset, create three virtual samples: samples of the maximum, average, and minimum values; for decision information systems... ,in It is a labeled sample set. It is an unlabeled sample set. It is a feature set. It's a decision-making category, with three virtual samples. (Maximum value) (average value), The minimum value is defined as follows: (1) S13, Regarding decision information systems ,in It is a labeled sample set. It is an unlabeled sample set. It is a feature set. It's a decision-making category. Then the sample Relative to feature subset fuzzy neighborhood radius The definition is as follows: (2) in, Indicates based on feature subset The distance between samples, Sort the eigenvalues in ascending order by the absolute difference, and take the top ones. The mean of the values is used as an auxiliary sample. This is the floor function; S14. Generate weighted granular compartments: For each sample Generate its fuzzy space, and retain only the granular compartments containing samples that satisfy the bidirectional neighborhood relationship as the core granular compartments. For each sample Apart from the samples in the core granular compartment, the granular compartments containing the samples that satisfy a weak single correlation are the boundary granular compartments. ,Right now (3) (4) in:
[0009] S15. A weighted mechanism is introduced to weight the features in the core and boundary granular compartments. The weight allocation formula is as follows: (5) in, This represents the weight of the samples in the core particle compartment. This represents the weight of the samples in the boundary granular compartment. and Representing samples respectively Whether it belongs to the core pod or the boundary pod is indicated by the function. If the sample If it belongs to the core particle compartment, then ,otherwise If the sample If it belongs to the boundary granular compartment, then ,otherwise .
[0010] Further, step S2 includes the following steps: S21, In a semi-supervisory decision-making system In the middle, for each decision class Sample set Each sample Indicates having A sample of features. The sample mean is defined as: (6) Then the covariance matrix The Line number Column elements are defined as: (7) in Indicates sample In the The values of each feature It is the mean of this feature; S22, Covariance matrix for each decision class Perform eigenvalue decomposition. The formula is: (8) in, Represents the covariance matrix The 1 eigenvalue, These are the corresponding eigenvectors. Each eigenvector indicates a direction in the feature space, and its corresponding eigenvalue reflects the magnitude of the variance in that direction.
[0011] S23. To extract the direction representing the main trend of the sample distribution, the eigenvector with the largest eigenvalue is selected as the main direction. Let... Then the corresponding eigenvector satisfy: (9) The vector Defined as the current decision class The main direction.
[0012] S24. To determine the range of hemispherical granulation, the square root of each eigenvalue is calculated based on the eigenvalue decomposition results, and the radius is obtained by calculating the average value according to the designed weighted granulation chamber. : (10) in It is the number of features. It is the weight of the granular compartment. These are the eigenvalues of the covariance matrix decomposition.
[0013] According to decision-making categories main direction and the radius in formula (10) A hemisphere can be constructed. This paper uses unlabeled samples. The process of constructing a hemisphere around the center of the sphere to achieve granulation, with the sample set within the hemisphere as follows: (11) S25, Unlabeled Sample Belongs to the decision-making category The membership degree is: (12) in This represents the labeled samples belonging to the decision class. A set of. The hemisphere represents the decision-making category. The number of labeled samples is calculated, and the membership degrees of all decision classes are normalized to ensure that the sum of the membership degrees is 1. The formula is: (13) Finally, the decision class with the highest membership degree was selected. As unlabeled samples The tag, namely: (14) Further, step S3 includes the following steps: S31, In a semi-supervisory decision-making system In, for any feature Define features On the sample and The distance between them is Then the cost of feature difference It can be defined as: (15) in, It is the base of the natural logarithm. . S32, In a semi-supervisory decision-making system Among them It is a labeled sample set. It is an unlabeled sample set. It is a feature set. It is a decision class, for any feature Cost entropy Defined as: (16) (17) in, It is the total number of samples in the decision-making system. It is a sample and In features The joint probability distribution on, Represents any two samples in a decision-making system and In features The cost of the difference.
[0014] S33, Features Importance Defined as: (18) in, It is the highest cost entropy among all features.
[0015] Further, step S4 includes the following steps: S41. For all features in the decision-making system The importance of its features is calculated according to formula (18); S42. Based on the calculated feature importance The features in the diabetes dataset are ranked, and the top-ranked features are selected to form a locally optimal feature subset.
[0016] Furthermore, S5 includes: under the Spark distributed framework, aggregating the locally optimal feature subsets of each child node through the reduce operation of the master node to form a globally optimal feature subset.
[0017] Meanwhile, the present invention proposes an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the computer program is executed, it implements the steps of the method described in the present invention.
[0018] Furthermore, the present invention proposes a computer-readable storage medium having a computer program stored thereon, the computer program being configured to implement the steps of the method described in the present invention when invoked by a processor.
[0019] Finally, the present invention proposes a computer program product, including a computer program / instructions, characterized in that the computer program / instructions, when executed by a processor, implement the steps of the method described in the present invention.
[0020] Compared with the prior art, the beneficial effects of the present invention are as follows: (1) By utilizing the in-memory computing technology and parallel computing capabilities of the Spark framework, the dataset is divided into multiple subsets and processed in parallel, which can significantly improve data processing efficiency. Compared with traditional single-machine processing methods, this invention can greatly shorten the computation time, especially when processing large-scale datasets, where the advantages are more obvious.
[0021] (2) This invention proposes a semi-supervised feature selection method for diabetes detection using cost entropy information granularization. By constructing weighted granular compartments to process the data, it comprehensively considers both global and local characteristics. The core granular compartment focuses on the local relationships between features, while the boundary granular compartments capture global feature information. By introducing a weighting mechanism, samples in the core and boundary granular compartments are assigned differentiated weights, thereby optimizing the discriminative power and accuracy of feature selection.
[0022] (3) A fuzzy label learning method based on hemispherical granulation is proposed. By mapping unlabeled samples to hemispherical regions in the feature space and using the information of labeled samples within the hemisphere for label learning, the classification accuracy of unlabeled samples can be effectively improved. This method can better utilize unlabeled data and improve the robustness of the model.
[0023] (4) By introducing cost entropy and distance factors to evaluate features, and combining feature difference costs, the most important features for the decision system can be effectively selected, thereby improving the accuracy and robustness of model prediction. Attached Figure Description
[0024] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used together with the embodiments of the invention to explain the invention and do not constitute a limitation thereof.
[0025] Figure 1 This is an overall flowchart of the method proposed in this invention.
[0026] Figure 2 This is a data processing framework diagram of the method proposed in this invention.
[0027] Figure 3 The flowchart shows the algorithm of the method proposed in this invention. Detailed Implementation
[0028] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. Of course, the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0029] Example 1: See Figure 1-3 As shown, this embodiment provides a semi-supervised feature selection method for cost entropy information granularization in diabetes detection, used to predict whether a patient has diabetes, including the following steps: Data preparation stage: Data source: Clinical data of 10,000 patients were collected from the electronic medical record (EMR) system of a large hospital, including age, gender, medical history, symptoms, laboratory test results (such as blood glucose, blood lipids, liver function and other indicators) and imaging characteristics (such as quantitative characteristics of CT and MRI images).
[0030] Data cleaning: Standardizing the data, such as normalizing age to the 0-1 range, handling missing values (e.g., filling with the mean or median), and removing outliers (e.g., blood sugar levels outside the normal physiological range).
[0031] Feature encoding: For categorical features (such as gender, medical history, etc.), one-hot encoding or label encoding is performed so that the model can process them.
[0032] S1. Use Spark's in-memory computing technology to divide the diabetes dataset into different subsets and construct weighted granular compartments; S2. Based on the principal direction analysis of the covariance matrix, each decision class is hemispherically granulated, and pseudo-labels are assigned to unlabeled samples. S3. Introduce cost entropy to evaluate features based on feature differences and obtain feature importance; S4. Rank the features by importance and select a subset of features; S5. Aggregate the results of feature selection from each child node to obtain the optimal feature subset.
[0033] Step S1 includes the following steps: S11. Read the diabetes dataset, determine its feature set and decision classification, and the decision information system... ,in It is a labeled sample set. It is an unlabeled sample set. It is a feature set. It's a decision-making class. Using Spark to process the diabetes dataset... Divide into subsets This method accelerates the construction of information granules by processing these subsets in parallel across multiple nodes, fully leveraging the advantages of distributed computing to effectively improve the efficiency of processing large-scale datasets. Assume a subset of data is obtained at a certain node. As shown in Table 1.
[0034] Table 1
[0035] S12. Construction of Virtual Samples: Based on the feature values in the diabetes dataset, create three virtual samples: samples of the maximum, average, and minimum values; for decision information systems... ,in It is a labeled sample set. It is an unlabeled sample set. It is a feature set. It's a decision-making category, with three virtual samples. (Maximum value) (average value), The minimum value is defined as follows: (1) Data subset In the sample, the three virtual samples are as follows: , , .
[0036] S13, Regarding decision information systems ,in It is a labeled sample set. It is an unlabeled sample set. It is a feature set. It's a decision-making category. Then the sample Relative to feature subset fuzzy neighborhood radius The definition is as follows: (2) in, Indicates based on feature subset The distance between samples, Sort the eigenvalues in ascending order by the absolute difference, and take the top ones. The mean of the values is used as an auxiliary sample. This is the floor function; S14. Generate weighted granular compartments: For each sample Generate its fuzzy space, and retain only the granular compartments containing samples that satisfy the bidirectional neighborhood relationship as the core granular compartments. For each sample Apart from the samples in the core granular compartment, the granular compartments containing the samples that satisfy a weak single correlation are the boundary granular compartments. ,Right now (3) (4) in:
[0037] S15. A weighted mechanism is introduced to weight the features in the core and boundary granular compartments. The weight allocation formula is as follows: (5) in, This represents the weight of the samples in the core particle compartment. This represents the weight of the samples in the boundary granular compartment. and Representing samples respectively Whether it belongs to the core pod or the boundary pod is indicated by the function. If the sample If it belongs to the core particle compartment, then ,otherwise If the sample If it belongs to the boundary granular compartment, then ,otherwise .
[0038] In this example, three virtual samples (maximum, minimum, and average) are used as centers, and the neighborhood radius is used as the radius to form three coverage areas. The intersection of the three areas is called the "core granular compartment" with a weight of 0.8. The intersection of any two areas but not the intersection of the three is called the "boundary granular compartment" with a weight of 0.2.
[0039] Step S2 includes the following steps: S21, In a semi-supervisory decision-making system In the middle, for each decision class Sample set Each sample Indicates having A sample of features. The sample mean is defined as: (6) Then the covariance matrix The Line number Column elements are defined as: (7) in Indicates sample In the The values of each feature It is the mean of this feature; The covariance matrix calculated based on the constructed weighted granular hopper in this example is as follows:
[0040] S22, Covariance matrix for each decision class Perform eigenvalue decomposition. The formula is: (8) in, Represents the covariance matrix The 1 eigenvalue, These are the corresponding eigenvectors. Each eigenvector indicates a direction in the feature space, and its corresponding eigenvalue reflects the magnitude of the variance in that direction.
[0041] In this example, the eigenvalue decomposition results are shown in Table 2: Table 2 Eigenvalue decomposition results
[0042] S23. To extract the direction representing the main trend of the sample distribution, the eigenvector with the largest eigenvalue is selected as the main direction. Let... Then the corresponding eigenvector satisfy: (9) The vector Defined as the current decision class The main direction.
[0043] From the table data, we can see that in this example, the eigenvector with the largest eigenvalue [0.30, 0.55, 0.55, 0.30, 0.10] is selected as the direction with the highest sample density.
[0044] S24. To determine the range of hemispherical granulation, the square root of each eigenvalue is calculated based on the eigenvalue decomposition results, and the radius is obtained by calculating the average value according to the designed weighted granulation chamber. : (10) in It is the number of features. It is the weight of the granular compartment. These are the eigenvalues of the covariance matrix decomposition.
[0045] According to decision-making categories main direction and the radius in formula (10) A hemisphere can be constructed. This paper uses unlabeled samples. The process of constructing a hemisphere around the center of the sphere to achieve granulation, with the sample set within the hemisphere as follows: (11) S25, Unlabeled Sample Belongs to the decision-making category The membership degree is: (12) in This represents the labeled samples belonging to the decision class. A set of. The hemisphere represents the decision-making category. The number of labeled samples is calculated, and the membership degrees of all decision classes are normalized to ensure that the sum of the membership degrees is 1. The formula is: (13) Finally, the decision class with the highest membership degree was selected. As unlabeled samples The tag, namely: (14) In this example, samples S1, S2, S3, S4, S7, and S8 within hemisphere S9 were labeled as diabetic; samples S2, S4, S5, S6, and S8 within hemisphere S10 were labeled as non-diabetic; samples S1, S2, S3, S4, and S7 within hemisphere S11 were labeled as diabetic; and samples S2, S4, S5, S6, and S8 within hemisphere S12 were labeled as non-diabetic.
[0046] Step S3 includes the following steps: S31, In a semi-supervisory decision-making system In, for any feature Define features On the sample and The distance between them is Then the cost of feature difference It can be defined as: (15) in, It is the base of the natural logarithm. . S32, In a semi-supervisory decision-making system Among them It is a labeled sample set. It is an unlabeled sample set. It is a feature set. It is a decision class, for any feature Cost entropy Defined as: (16) (17) in, It is the total number of samples in the decision-making system. It is a sample and In features The joint probability distribution on, Represents any two samples in a decision-making system and In features The cost of the difference.
[0047] S33, Features Importance Defined as: (18) in, It is the highest cost entropy among all features.
[0048] Step S4 includes the following steps: S41. For all features in the decision-making system The importance of its features is calculated according to formula (18); Based on the label learning results, the weighting coefficients can be calculated as shown in Table 3: Table 3. Weighted Feature Importance
[0049] S42. Based on the calculated feature importance The features in the diabetes dataset are ranked, and the top-ranked features are selected to form a locally optimal feature subset.
[0050] The features are ranked according to their weighted importance, and the top-ranked features are selected to form the optimal feature subset. The final selected feature subset is: FBG, HbA1c, and PFC.
[0051] Step S5 includes: aggregating the locally optimal feature subsets selected by each node under the Spark distributed framework to obtain the optimal feature subset, such as... Figure 2 The feature subsets of each node in the framework diagram are aggregated as shown.
[0052] Example 2: This example proposes an electronic system, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method steps of the present invention.
[0053] Example 3: This example proposes a computer-readable storage medium storing a computer program thereon. When the computer program is executed by a processor, it implements the steps of the method described in this invention, which will not be repeated here.
[0054] Example 4: This example proposes a computer program product, including a computer program / instructions. When the computer program / instructions are executed by a processor, they implement the steps of the method described in this invention, which will not be repeated here.
[0055] It should be noted that the processing flow of embodiments 2-4 corresponds to the specific steps of the method provided in embodiment 1 of the present invention, and has the corresponding functional modules and beneficial effects of the method. Technical details not described in detail in this embodiment can be found in the method provided in embodiment 1 of the present invention.
[0056] The program code used to implement the methods of this application may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing device, such that when executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0057] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A semi-supervised feature selection method for cost entropy information granularization in diabetes detection, characterized in that, Includes the following steps: S1. Use Spark's in-memory computing technology to divide the diabetes dataset into different subsets and construct weighted granular compartments; S2. Based on the principal direction analysis of the covariance matrix, each decision class is hemispherically granulated, and pseudo-labels are assigned to unlabeled samples. S3. Introduce cost entropy to evaluate features based on feature differences and obtain feature importance; S4. Rank the features by importance and select a subset of features; S5. Aggregate the results of feature selection from each child node to obtain the optimal feature subset.
2. The semi-supervised feature selection method for cost entropy information granulation in diabetes detection according to claim 1, characterized in that, S1 includes the following steps: S11. Read the diabetes dataset, determine its feature set and decision classification, and the decision information system... , where is the labeled sample set, It is an unlabeled sample set. It is a feature set. It's a decision-making class, using Spark to process the diabetes dataset. Divide into subsets And process these subsets in parallel on multiple nodes; S12. Construction of Virtual Samples: Based on the feature values in the diabetes dataset, create three virtual samples: samples of the maximum, average, and minimum values; for decision information systems... ,in It is a labeled sample set. It is an unlabeled sample set. It is a feature set. It's a decision-making category, with three virtual samples. For the maximum value, For average value, To find the minimum value, the following definitions apply: (1); S13, Regarding decision information systems ,in It is a labeled sample set. It is an unlabeled sample set. It is a feature set. It's a decision-making category. Then the sample Relative to feature subset fuzzy neighborhood radius The definition is as follows: (2); in, Representation based on feature subset The distance between samples, Sort the eigenvalues in ascending order by the absolute difference, and take the top ones. The mean of the values is used as an auxiliary sample. This is the floor function; S14. Generate weighted granular compartments: For each sample Generate its fuzzy space, and retain only the granular compartments containing samples that satisfy the bidirectional neighborhood relationship as the core granular compartments. For each sample Apart from the samples in the core granular compartment, the granular compartments containing the samples that satisfy a weak single correlation are the boundary granular compartments. ,Right now: (3); (4); in: ; S15. Introduce a weighted mechanism to weight the features in the core and boundary granular compartments. The weight allocation formula is as follows: (5); in, This represents the weight of the samples in the core particle compartment. This represents the weight of the samples in the boundary granular compartment. and Representing samples respectively Whether it belongs to the core granular compartment and the boundary granular compartment is an indicator function, if the sample If it belongs to the core particle compartment, then ,otherwise If the sample If it belongs to the boundary granular compartment, then ,otherwise .
3. The semi-supervised feature selection method for cost entropy information granulation in diabetes detection according to claim 1, characterized in that, S2 includes the following steps: S21, In a semi-supervisory decision-making system In the middle, for each decision class Sample set , where each sample Indicates having A sample of features is defined with the following mean: (6); Then the covariance matrix The Line number Column elements are defined as: (7); in Indicates sample In the The values of each feature It is the mean of this feature; S22, Covariance matrix for each decision class The eigenvalue decomposition formula is as follows: (8); in, Represents the covariance matrix The 1 eigenvalue, For each eigenvector, there is a direction in the feature space, and the corresponding eigenvalue reflects the variance in that direction. S23. To extract the direction representing the main trend of the sample distribution, the eigenvector with the largest eigenvalue is selected as the main direction. Let... Then the corresponding eigenvector satisfy: (9); The vector Defined as the current decision class The main direction; S24. To determine the range of hemispherical granulation, the square root of each eigenvalue is calculated based on the eigenvalue decomposition results, and the radius is obtained by calculating the average value according to the designed weighted granulation chamber. : (10); in It is the number of features. It is the weight of the granular compartment. These are the eigenvalues of the covariance matrix decomposition; According to decision-making categories main direction and the radius in formula (10) Construct a hemisphere with unlabeled samples The process of constructing a hemisphere around the center of the sphere to achieve granulation, with the sample set within the hemisphere as follows: (11); S25, Unlabeled Sample Belongs to the decision-making category The membership degree is: (12); in This represents the labeled samples belonging to the decision class. The set, The hemisphere represents the decision-making category. The number of labeled samples is calculated, and the membership degrees of all decision classes are normalized to ensure that the sum of the membership degrees is 1. The formula is: (13); Finally, the decision class with the highest membership degree was selected. As unlabeled samples The tag, namely: (14)。 4. A semi-supervised feature selection method for cost entropy information granulation in diabetes detection according to claim 1, characterized in that, S3 includes the following steps: S31, In a semi-supervisory decision-making system In, for any feature Define features On the sample and The distance between them is Then the cost of feature difference Defined as: (15); in, It is the base of the natural logarithm. ; S32, In a semi-supervisory decision-making system Among them It is a labeled sample set. It is an unlabeled sample set. It is a feature set. It is a decision class, for any feature Cost entropy Defined as: (16); (17); in, It is the total number of samples in the decision-making system. It is a sample and In features The joint probability distribution on, Represents any two samples in a decision-making system and In features The cost of difference; S33, Features Importance Defined as: (18); in, It is the highest cost entropy among all features.
5. A semi-supervised feature selection method for cost entropy information granulation in diabetes detection according to claim 1, characterized in that, S4 includes the following steps: S41. For all features in the decision-making system The importance of its features is calculated according to formula (18); S42. Based on the calculated feature importance The features in the diabetes dataset are ranked, and the top-ranked features are selected to form a locally optimal feature subset.
6. A semi-supervised feature selection method for cost entropy information granulation in diabetes detection according to claim 1, characterized in that, The S5 includes: under the Spark distributed framework, aggregating the locally optimal feature subsets of each child node through the reduce operation of the master node to form a globally optimal feature subset.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the computer program is executed, it implements the steps of the method as described in any one of claims 1 to 6.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, The computer program is configured to implement the steps of the method according to any one of claims 1 to 6 when invoked by a processor.
9. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method according to any one of claims 1 to 6.