A method for reducing medical test indexes based on the combination of rough set positive domain and information entropy

By combining the rough set positive domain and information entropy methods, the comprehensive importance of medical test indicators is calculated, which solves the problem of redundant attributes in high-dimensional data and achieves more accurate simplification and improved diagnostic efficiency.

CN120408127BActive Publication Date: 2025-09-09NANJING UNIV OF INFORMATION SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510919382.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-04
Publication Date
2025-09-09
Estimated Expiration
2045-07-04

AI Technical Summary

Technical Problem

Existing medical test index data processing methods are unable to effectively balance deterministic and non-deterministic elements in high-dimensional data, resulting in increased redundant attributes and reduced model generalization ability and diagnostic efficiency.

Method used

A method based on the combination of rough set positive domain and information entropy is adopted to calculate the comprehensive importance of medical test indicators through weighted average, and iterative simplification is performed to reasonably balance deterministic and non-deterministic elements to generate a more accurate simplified set.

Benefits of technology

It improves the diagnostic accuracy and efficiency of medical tests, reduces unnecessary test items, reduces medical costs, and enhances clinical interpretability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120408127B_ABST
    Figure CN120408127B_ABST
Patent Text Reader

Abstract

The present invention discloses a medical test index simplification method based on the combination of rough set positive domain and information entropy, which is suitable for high-dimensional data attribute simplification and knowledge discovery in scenarios such as medical data analysis. By integrating the two perspectives of rough set theory and information theory methods, a new simplification strategy is proposed, which obtains the comprehensive importance of attributes by weighted average, and its weight is determined by the ratio of the positive domain to the boundary domain. It performs iterative simplification according to the comprehensive importance, reasonably takes into account the deterministic and non-deterministic elements, solves the feasibility problem of attribute simplification in large-scale fuzzy information decision-making systems, and obtains a more accurate simplified set. It is convenient to focus on key diagnostic indicators, improve diagnostic efficiency and accuracy, enhance clinical interpretability, reduce unnecessary test items, and play a role in optimizing data, improving diagnostic efficiency, and saving medical costs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a medical test index simplification method based on the synthesis of rough set positive domain and information entropy, and belongs to the technical field of medical test index data processing. Background Art

[0002] In the current context of explosive growth in medical data, a single examination often generates hundreds of feature parameters, forming a high-dimensional fuzzy information decision system. Attribute reduction, by extracting key features, optimizes data, improves diagnostic efficiency, and reduces medical costs.

[0003] Taking lung CT nodule detection as an example, using the first eight supervised attributes from a preliminary screening of 52 significant attributes achieved 100% accuracy, sensitivity, and specificity. Furthermore, experiments have shown that increasing the number of attributes does not significantly improve performance. In unsupervised scenarios, increasing the number of attributes may even reduce classification effectiveness. This suggests that more features are not necessarily better; redundant attributes can introduce noise and reduce model generalization. This reduction in data dimensions allows doctors to focus on key diagnostic features such as nodule morphology and density gradients, significantly improving clinical decision-making efficiency.

[0004] In information decision-making systems, attribute simplification is necessary to address the complexity and accuracy of high-dimensional data calculations and eliminate the impact of redundant and irrelevant attributes on the calculation process and final results. The commonly used attribute simplification methods are as follows:

[0005] 1. The discernible matrix method based on rough set theory determines the core attributes and gradually removes redundant attributes. This method has huge space complexity and is not suitable for processing high-dimensional data;

[0006] 2. Combining fuzzy mathematics with rough set theory, the positive domain dependency attribute reduction method selects attributes by calculating their importance to the lower approximation of the decision class. This method relies on deterministic elements and ignores non-deterministic elements.

[0007] 3. The information gain attribute reduction method, based on information theory, measures the amount of information in an attribute by calculating entropy reduction, thereby filtering attributes. This method focuses on the information entropy brought by uncertain elements and ignores the deterministic elements in decision making. Summary of the Invention

[0008] The technical problem to be solved by the present invention is to provide a medical test index simplification method based on the combination of rough set positive domain and information entropy, which reasonably takes into account both deterministic and non-deterministic elements, solves the feasibility problem of attribute simplification in large-scale fuzzy information decision systems, and obtains a more accurate simplified set.

[0009] The present invention adopts the following technical solutions to solve the above technical problems:

[0010] A medical test index reduction method based on the combination of rough set positive domain and information entropy is used to reduce all medical test indicators for a certain disease, including the following steps:

[0011] Step 1: Obtain a medical test index data set, where each sample in the medical test index data set includes a patient, all test index data for the patient with a certain disease, and a diagnosis conclusion for the patient with a certain disease; convert the medical test index data set into a fuzzy information decision system. ,in, represents the set of all patients, Indicates the patients, , is the number of all patients, Represents a set of all medical test indicators for a certain disease. Indicates the indicators, , is the number of all medical examination indicators, Indicates that into disjoint diagnostic conclusion categories, Indicates the Diagnostic conclusion categories, , is the number of diagnostic conclusion categories;

[0012] Step 2: Calculate the set of all medical test indicators using the rough set positive domain method The positive domain of Calculate the weighted coefficient in the positive domain And the diagnostic conclusion category set right Dependence ;

[0013] Step 3: Convert the fuzzy information decision system into a deterministic information decision system through clustering method , calculated in a deterministic information decision system Information entropy And all attributes Conditional information entropy under ;

[0014] Step 4: Set the reduction set to , and initialize the reduced set is an empty set;

[0015] Step 5, for the collection Not selected For each index in the current reduced set, compare each index with the current reduced set Combine them to obtain the union corresponding to each indicator, calculate the importance of the union corresponding to each indicator based on the positive domain of the rough set and the importance based on information entropy, and then calculate the comprehensive importance of the union corresponding to each indicator;

[0016] Step 6: The indicator corresponding to the maximum comprehensive importance Add to the current reduction set In, get the collection , as the new reduced set ;

[0017] Step 7: Calculate the diagnostic conclusion category set using the same method as step 2 The new reduced set obtained in step 6 Dependence , and the new reduced set obtained in step 6 is calculated using the same method as step 3 Conditional entropy under ;

[0018] Step 8: Determine whether and , if satisfied, then the new reduced set obtained in step 6 As the final index reduction set output; otherwise return to step 5, in the set Not selected Continue to select the indicator with the largest comprehensive importance from the indicators and add it to the new simplified set obtained in step 6 In the example above, a new reduction set is generated again. , until satisfied and .

[0019] Compared with the prior art, the present invention adopts the above technical solution and has the following technical effects:

[0020] 1. This paper integrates the perspectives of rough set theory and information theory to propose a new reduction strategy. This strategy uses a weighted average to determine the comprehensive importance of test indicators, with the weight determined by the ratio of the positive domain to the boundary domain. Iterative reduction based on the comprehensive importance rationally balances deterministic and non-deterministic elements, addressing the feasibility of attribute reduction in large-scale fuzzy information decision systems. This strategy yields a more accurate reduced set, ensuring the accuracy of diagnostic conclusions from medical tests within the reduced set.

[0021] 2. The method proposed in this invention facilitates focusing on key diagnostic indicators, improving diagnostic efficiency and accuracy, enhancing clinical interpretability, reducing unnecessary testing items, and reducing medical costs. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 This is a simplified flow chart of the medical test indicators of the present invention;

[0023] Figure 2 This is a flow chart of determining attribute importance based on the rough set positive domain method of the present invention;

[0024] Figure 3 is a flow chart of determining attribute importance using the information entropy method of the present invention;

[0025] Figure 4 It is a flowchart of the iterative reduction of the present invention. DETAILED DESCRIPTION

[0026] The embodiments of the present invention are described in detail below, and examples of the embodiments are shown in the accompanying drawings. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be interpreted as limiting the present invention.

[0027] like Figure 1 As shown, the present invention proposes a method for reducing medical laboratory test indicators based on the integration of rough set positive domains and information entropy. First, a distance-based fuzzy equivalence relation is defined within a fuzzy decision system, along with a lower approximation operator and a fuzzy positive domain. Furthermore, the importance of conditional attributes based on the positive domain is defined. Second, the fuzzy decision system is binarized, transforming it into a deterministic information decision system. The importance of conditional attributes based on information entropy is defined. The two importances are weighted averaged to obtain a comprehensive importance. Heuristic reduction is performed based on the comprehensive importance, ultimately yielding a reduced attribute set.

[0028] Data preparation: Preprocess the medical test index data and diagnostic results and convert them into a fuzzy information decision system. is a fuzzy information decision system, in which is the domain, representing the patients undergoing medical examination; is a finite set of non-empty conditional attributes, representing all medical test indicators for a certain disease. For each attribute , there is a mapping , assign the attribute value (fuzzy value) to The individuals in represent the test values ​​of relevant medical tests; Represents the symbolic decision attribute, indicating the diagnostic conclusion. Partition into disjoint decision classes ,individual For each diagnosis The membership degree of is defined as:

[0029] .

[0030] Step 1: Determine the importance of each attribute with respect to the deterministic element by using the rough set positive domain method.

[0031] Step 1.1 In the fuzzy decision information system In the fuzzy equivalence relation, the distance is used as the metric.

[0032] Step 1.1.1 For each conditional attribute , define the distance function:

[0033] ,

[0034] Indicates the differences between different individuals in attributes The degree of difference in medical test data under the current situation.

[0035] Step 1.1.2: Condition attribute subset , define the composite distance function:

[0036] ,

[0037] This means that the comprehensive difference in medical test data between different individuals under the combined effect of multiple attributes is determined by the maximum value of the data differences under all their individual attributes.

[0038] Step 1.2 Define the decision class Metric-based fuzzy approximation operator

[0039] Step 1.2.1 Define the lower approximation operator:

[0040] ,

[0041] The lower approximation operator measures the individual Give the degree of certainty of a diagnosis conclusion, quantified as individual The minimum difference between an individual with this diagnosis and an individual without this diagnosis.

[0042] Step 1.2.2 Define the upper approximation operator:

[0043] ,

[0044] The upper approximation operator evaluates the individual The probability of receiving a certain diagnosis is defined by its maximum similarity to individuals receiving the same diagnosis.

[0045] Step 1.3 Define attribute subsets The fuzzy positive domain of :

[0046] ,

[0047] The fuzzy positive domain reflects the The larger the positive domain, the higher the individual's certainty membership, indicating that the attribute subset The greater the diagnostic certainty.

[0048] Step 1.4 Calculate the positive domain of the entire attribute set .

[0049] Step 1.5 Define decision variables For attribute subsets Dependencies:

[0050] ,

[0051] Dependence is the normalization of the fuzzy positive domain, and the value is in the interval Inside.

[0052] Step 1.6 Calculate the dependency of the decision attribute on the full set of conditional attributes .

[0053] Step 1.7 Define properties Importance based on positive domain:

[0054] ,

[0055] Measuring medical testing How much does it affect the certainty of the diagnosis?

[0056] Step 2: Determine the importance of each attribute with respect to the non-deterministic element using the information entropy method

[0057] Step 2.1 Binarization of fuzzy decision system

[0058] Step 2.1.1 For each conditional attribute , using the fuzzy clustering algorithm FCM to cluster the samples into Clusters: , where each individual For each cluster The membership degree is .

[0059] Step 2.1.2 According to the maximum membership principle, the fuzzy clustering is transformed into a deterministic classification. Belong to the category with the maximum membership, that is, for each medical test indicator , according to the individual detection value, the individuals are divided into kind.

[0060] Step 2.1.3 Generate a deterministic information decision system :

[0061] Here we choose the fuzzy clustering algorithm FCM, which can retain the fuzzy membership attributes and facilitate the docking of fuzzy related algorithms. If there is no need to retain fuzziness, we can choose the k-means algorithm for clustering and directly generate a deterministic information decision system. .

[0062] Step 2.2 The information entropy of decision attributes is defined in:

[0063] ,

[0064] Information entropy measures the degree of uncertainty of diagnostic conclusions in mathematical form.

[0065] Step 2.3 Calculate information entropy ;

[0066] Step 2.4 In the method of dividing and refining, the attribute subset is defined Induced joint partitioning of sample sets:

[0067] ,

[0068] in, Indicates that the sample set has attributes The following categories.

[0069] Step 2.5 Define attribute subsets in Conditional entropy :

[0070] ,

[0071] Conditional entropy measures the The degree of uncertainty in the diagnostic conclusion.

[0072] Step 2.6 Calculate all attributes Information entropy under ;

[0073] Step 2.7 Define properties The importance of based on information entropy:

[0074] ,

[0075] Measured medical testing How big is the impact on diagnostic uncertainty?

[0076] Step 3: Define the weighting coefficient by the proportional relationship between the positive domain and the boundary domain :

[0077] ,

[0078] In a given fuzzy decision system, since each individual There is a clear diagnosis ,The boundary region quantifies the remaining uncertainty after considering all ,unambiguously classifiable objects in the positive region, which is complementary to the ,positive domain value.

[0079] Step 4: Define conditional attributes The overall importance of:

[0080] ,

[0081] Taking into account both deterministic and uncertain decisions, the medical testing Impact on diagnosis.

[0082] Step 5: Perform iterative reduction to obtain the attribute reduction set

[0083] Step 5.1 Initialize the reduced set ;

[0084] Step 5.2 For each unselected attribute , calculate the comprehensive importance;

[0085] Step 5.2.1 Calculate the conditional attribute set The fuzzy positive domain of

[0086] Step 5.2.2 Calculate the decision's dependence on the attribute subset ;

[0087] Step 5.2.3 Calculate properties Importance based on positive domain:

[0088] ,

[0089] Step 5.2.4 Calculate the conditional attribute set Induced division:

[0090] ,

[0091] Step 5.2.5 Calculate conditional entropy ;

[0092] Step 5.2.6 Calculate properties Importance based on information entropy:

[0093] ,

[0094] Step 5.2.7 Calculate conditional attributes The overall importance of:

[0095] ,

[0096] Step 5.3 Select the current optimal attribute: ;

[0097] Step 5.4 Update the reduced set ;

[0098] Step 5.5 Verify termination conditions ,and , if satisfied, output the reduced set , otherwise return to step 5.2.

[0099] Example 1

[0100] It is a fuzzy information decision table based on the lung nodule identification scenario, as shown in Table 1, where: individual sample set Represents 9 patients; conditional attribute set Represents 6 medical test indicators, whose values ​​are all pre-processed data. Indicates the rating of the depth of dyeing, represents the nuclear-cytoplasmic ratio, express Index, that is, the proportion of positive cells (%), Indicates the degree of cell arrangement disorder, Indicates the nodule edge morphology (lobulation index), Represents the nodule growth rate; decision attribute set , respectively represent the two categories of lung nodule diagnosis, .

[0101] Table 1

[0102]

[0103] Step 1: Determine the importance of each attribute with respect to the deterministic element by using the rough set positive domain method.

[0104] 1.1 In fuzzy decision information system For each conditional attribute , calculate the distance function:

[0105] ,

[0106] Only conditional attributes are listed here The distance function of:

[0107] ,

[0108] 1.2 Calculating all attributes The distance function ,

[0109] ,

[0110] 1.3 Calculate the total attributes of each sample The lower approximation of :

[0111] ,

[0112] Calculated:

[0113] , , ,

[0114] , , ,

[0115] , , ,

[0116] 1.4 Calculating decision variables For all attribute subsets Dependencies:

[0117] .

[0118] Step 2: Determine the importance of each attribute with respect to the non-deterministic element using the information entropy method

[0119] 2.1 Binarization of Fuzzy Decision System

[0120] For each conditional attribute , using the fuzzy clustering algorithm FCM to cluster the samples into Clusters (in this embodiment, ), , where each individual The degree of membership in each cluster is According to the maximum membership principle, the fuzzy clustering is transformed into a deterministic classification, and each individual Belong to the category with the maximum membership, and generate a deterministic information decision system , as shown in Table 2.

[0121] Table 2

[0122]

[0123] 2.2 In Calculate the information entropy of decision attributes in

[0124] ,

[0125] 2.3 Calculating the full attribute subset Induced sample partitioning: ,

[0126] 2.4 Calculating a subset of all attributes The conditional entropy of: .

[0127] Step 3: Calculate the weighting coefficient :

[0128] .

[0129] Step 4: Initialize the reduced set , perform iterative reduction to obtain the attribute reduction set

[0130] 4.1 First Iteration

[0131] 4.1.1 For each unselected attribute , calculate the attribute dependency based on the positive domain:

[0132] , , ,

[0133] , , ,

[0134] 4.1.2 Computed Properties Importance based on positive domain:

[0135] because ,so , the results are shown in Table 3.

[0136] Table 3

[0137]

[0138] 4.1.3 In Calculation condition attribute set Induced division :

[0139]

[0140] 4.1.4 Calculating Conditional Entropy of Attributes :

[0141] , , ,

[0142] , , ,

[0143] 4.1.5 Computed Properties Importance based on information entropy: , the results are shown in Table 4.

[0144] Table 4

[0145]

[0146] 4.1.6 Calculating Conditional Attributes The overall importance of: , the results are shown in Table 5.

[0147] Table 5

[0148]

[0149] 4.1.7 Select the current optimal attribute: ;

[0150] 4.1.8 Updating the Reduced Set ;

[0151] 4.1.9 At this time , , the termination condition is not reached, and the second round of iteration is performed.

[0152] 4.2 Second Iteration

[0153] 4.2.1 For each unselected attribute , calculated properties Based on the importance of the positive domain, the results are shown in Table 6;

[0154] ,

[0155] Table 6

[0156]

[0157] 4.2.2 Computed Properties Based on the importance of information entropy, the results are shown in Table 7:

[0158] ,

[0159] exist Calculation condition attribute set Induced division :

[0160]

[0161] Table 7

[0162]

[0163] 4.2.3 Calculating Conditional Attributes The overall importance of: , the results are shown in Table 8.

[0164] Table 8

[0165]

[0166] 4.2.4 Select the current optimal attribute: ;

[0167] 4.2.5 Updating the Reduced Set ;

[0168] 4.2.6 At this time , , the termination condition is not reached, and the third round of iteration is performed.

[0169] 4.3 Third Iteration

[0170] 4.3.1 For each unselected attribute , calculated properties Based on the importance of the positive domain, the results are shown in Table 9:

[0171] ,

[0172] Table 9

[0173]

[0174] 4.3.2 Computed Properties Importance based on information entropy: Since we have obtained ,so ,property The importance based on information entropy is ;

[0175] 4.3.3 Attributes at this time The comprehensive importance of Decide, maximum;

[0176] 4.3.4 Select the current optimal attribute: ;

[0177] 4.3.5 Updating the Reduced Set ;

[0178] 4.3.6 At this time , , the termination condition is not reached, and the fourth round of iteration is performed.

[0179] 4.4 Fourth Iteration

[0180] 4.4.1 For each unselected attribute , calculated properties Based on the importance of the positive domain, the results are shown in Table 10:

[0181] ,

[0182] Table 10

[0183]

[0184] 4.4.2 Select the current optimal attribute: ;

[0185] 4.4.3 Updating the Reduced Set ;

[0186] 4.4.4 At this time , , the termination condition is not reached, and the fifth round of iteration is performed.

[0187] 4.5 Fifth Iteration

[0188] 4.5.1 For each unselected attribute , computed properties Based on the importance of the positive domain, the results are shown in Table 11:

[0189] ,

[0190] Table 11

[0191]

[0192] 4.5.2 Select the current optimal attribute: ;

[0193] 4.5.3 Updating the Reduced Set ;

[0194] 4.5.4 At this time , , the termination condition is reached, the reduction process ends, and the attribute reduction set is obtained .

[0195] This indicates that in the decision-making diagnosis of lung nodules, among the six medical test indicators, the "nuclear-cytoplasmic ratio" has no meaning for the test conclusion and is a redundant test and can be cancelled.

[0196] Example 2

[0197] It is a fuzzy information decision table based on the diabetic nephropathy diagnosis scenario, as shown in Table 12, where the individual sample set Represents 10 diabetic patients; conditional attribute set Represents 4 kinds of medical test indicators, whose values ​​are all pre-processed data, among which Indicates fasting blood sugar, represents glycosylated hemoglobin, Indicates the urine microalbumin / creatinine ratio, Represents systolic blood pressure; decision attribute set , indicating whether the patient was diagnosed with diabetic nephropathy.

[0198] Table 12

[0199]

[0200] According to the method of the present invention, the reduced set is , indicating that "fasting blood glucose" and "urine microalbumin / creatinine ratio" are the core attributes for diagnosing diabetic nephropathy. This simplified result is consistent with the pathological mechanism of diabetic nephropathy "hyperglycemia → renal microvascular damage".

[0201] Based on the same inventive concept, an embodiment of the present application provides a computer device, including a memory, a processor, and a computer program stored in the memory and runnable on the processor. When the processor executes the computer program, the steps of the aforementioned medical test index simplification method based on the combination of rough set positive domain and information entropy are implemented.

[0202] Based on the same inventive concept, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the steps of the aforementioned medical test index simplification method based on the combination of rough set positive domain and information entropy.

[0203] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0204] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0205] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0206] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 The steps for the function specified in one or more boxes.

[0207] The above embodiments are only for illustrating the technical idea of ​​the present invention and cannot be used to limit the protection scope of the present invention. Any changes made on the basis of the technical solution in accordance with the technical idea proposed by the present invention shall fall within the protection scope of the present invention.

Claims

1. A medical test index simplification method based on the combination of rough set positive domain and information entropy is used to simplify all medical test indicators for a certain disease, which is characterized by: The steps include: Step 1: Obtain a medical test index data set, where each sample in the medical test index data set includes a patient, all test index data for the patient with a certain disease, and a diagnosis conclusion for the patient with a certain disease; convert the medical test index data set into a fuzzy information decision system (U, A, D), where U = {x1, x2, ..., x N } represents the set of all patients, x i represents the i-th patient, i=1,…,N, N is the number of all patients, A={a1,a2,…,a M } represents the set of all medical test indicators for a disease, a m Indicates the mth index, m=1,…,M, M is the number of all medical examination indicators, U / D={D1,D2,…,D R } means dividing U into disjoint diagnostic conclusion categories, D r represents the rth diagnostic conclusion category, r = 1,…,R, R is the number of diagnostic conclusion categories; Step 2: Calculate the positive domain of all medical test index set A by using the rough set positive domain method. Calculate the weight coefficient α and the dependence γ of the diagnostic conclusion category set D on A based on the positive domain of set A. A (D); The specific process is as follows: Step 2.1, in the fuzzy information decision system (U, A, D), for each medical test index a m , calculate the distance between any two patients and obtain the distance matrix corresponding to each medical test indicator Among them, for medical examination index a m , any two patients x i ,x j The distance calculation formula is as follows: Among them, a m (x i ), a m (x j ) represent patients x i ,x j Medical test indicators m The value of Step 2.2: Based on the distance matrices corresponding to all medical test indicators, the distance matrix d of set A is obtained. A ;d A The elements in are represented as: Step 2.3, according to the distance matrix d A Calculate the lower approximation of each patient with respect to set A as follows: Step 2.4, based on step 2.3, calculate the positive domain POS of set A A (D), the formula is as follows: Step 2.5, according to the positive domain POS of set A A (D) Calculate the weighting coefficient α, the formula is as follows: Among them, |U| represents the number of all patients, that is, N; Step 2.6, according to the positive domain POS of set A A (D) Calculate the dependence γ of the diagnostic conclusion category set D on A A (D), the formula is as follows: Step 3: The fuzzy information decision system is converted into a deterministic information decision system (U, A′, D) by clustering method. The information entropy H(D) of D and the conditional information entropy H(D|A) under all attributes A′ are calculated in the deterministic information decision system. ′ ); Step 4, set the reduced set to B, and initialize the reduced set B to an empty set; Step 5: For each indicator in set A that is not selected into set B, combine each indicator with the current reduced set B to obtain the corresponding union of each indicator. Calculate the importance of the union of each indicator based on the positive domain of the rough set and the importance based on information entropy, and then calculate the comprehensive importance of the union of each indicator. Step 6: The index a corresponding to the maximum comprehensive importance * Add to the current reduction set B, and get the set B∪{a * }, as the new reduced set B; Step 7: Calculate the dependency γ of the diagnostic conclusion category set D on the new reduced set B obtained in step 6 using the same method as step 2. B (D), and calculate the conditional entropy H(D|B) under the new reduced set B obtained in step 6 using the same method as step 3; Step 8: Determine whether γ is satisfied B (D) = γ A (D) and H(D|B)=H(D|A) ′ ), if it is satisfied, the new reduced set B obtained in step 6 is output as the final index reduced set; otherwise, return to step 5, continue to select the index with the largest comprehensive importance from the indexes in set A that are not selected into B, and add it to the new reduced set B obtained in step 6, and generate a new reduced set B again until γ is satisfied. B (D) = γ A (D) and H(D|B)=H(D|A) ′ ).

2. The medical test index simplification method based on rough set positive domain and information entropy synthesis according to claim 1 is characterized in that: The specific process of step 3 is as follows: Step 3.1: For each medical test indicator a m , using fuzzy clustering algorithm to cluster all patients into k clusters, according to the maximum membership principle, the fuzzy clustering is converted into deterministic classification, each patient is assigned to the classification corresponding to the maximum membership, and a deterministic information decision system (U, A′, D) is generated; Step 3.2, calculate the information entropy H(D) of D in (U,A′,D): Among them, P(D r ) indicates that the patient is classified into the rth diagnostic conclusion category D r The probability of |D r | represents the number of patients classified into the rth diagnostic conclusion category; Step 3.3, calculate the joint partition of all patients induced by the full attribute A′: U / A′, and then calculate the conditional information entropy H(D|A ′ ).

3. The medical test index simplification method based on rough set positive domain and information entropy synthesis according to claim 2 is characterized in that: The specific process of step 5 is as follows: Step 5.1, for each unselected indicator a m ∈A\B, a m Combined with the reduced set B obtained in the previous iteration, we get B∪{a m }, in the first iteration, Calculate each B∪{a m } in, Denotes each patient's relationship with the set B∪{a m } lower approximation; Calculate each B∪{a m Dependence based on positive domain Step 5.2, calculate each B∪{a m Importance of positive domain based on rough set: Among them, γ B (D) is the dependency of the current reduced set B. In the first iteration, γ B (D) = 0; Step 5.3, calculate each B∪{a m }induced partition U / (B∪{a m })={X1,…,X T }, the division in the first iteration is the cluster obtained by using the fuzzy clustering algorithm; Step 5.4, calculate each B∪{a m }'s conditional entropy H(D|B∪{a m }): Among them, |X t | represents the set X t The number of patients in r ∩X t | represents the set X t Diagnosis is D r The number of patients, T represents the number of patients in the condition set B∪{a m } represents the number of categories into which patients are divided, and |U| represents the number of all patients, that is, N; Step 5.5, calculate each B∪{a m } Importance SGF(a based on information entropy m ,B,D): SGF(a m ,B,D)=H(D|B)-H(D|B∪{a m }) Where H(D|B) is the conditional entropy of the reduced set B obtained in the previous iteration. In the first iteration, H(D|B) = H(D); Step 5.6, calculate each B∪{a m The comprehensive importance of CS(a m ,B,D):

4. A computer device comprising a memory, a processor, and a computer program stored in the memory and capable of running on the processor, characterized in that: When the processor executes the computer program, the steps of the medical test index simplification method based on the combination of rough set positive domain and information entropy as described in any one of claims 1 to 3 are implemented.

5. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the medical test index simplification method based on the combination of rough set positive domain and information entropy as described in any one of claims 1 to 3 are implemented.

Citation Information

Patent Citations

  • Medical auxiliary examination system knowledge acquisition and inference method based on rough set

    CN105718726A

  • Diabetes parallel attribute reduction method based on coevolution discrete particle swarm optimization

    CN117059284A