Disease diagnosis method based on fuzzy neighborhood operator and generalized Shapley value
By constructing a disease diagnosis method based on fuzzy neighborhood operators and generalized Shapley values, the problems of global correlation and commutativity of overlapping functions in attribute reduction are solved, achieving more efficient and accurate disease diagnosis, which is applicable to the field of medical data processing.
Patent Information
- Application Number
- CN202511211613.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-28
- Publication Date
- 2025-11-07
AI Technical Summary
Existing attribute reduction methods ignore the global correlation between attributes, and the commutativity of overlap functions limits the application scope of fuzzy neighborhood operators, resulting in insufficient accuracy and efficiency in disease diagnosis.
A disease diagnosis method based on fuzzy neighborhood operators and generalized Shapley values is adopted. By constructing new fuzzy neighborhood operators, generalized Shapley values and Choquet integrals, the global correlation between attributes is considered, attribute reduction is performed, and a smart classifier is combined for disease diagnosis.
It significantly improves the accuracy and efficiency of disease diagnosis, reduces data dimensionality, reduces computational burden, enhances the ability to resolve fuzzy information, and improves the performance and applicability of diagnostic models.
Smart Images

Figure CN120913818A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of medical data processing and disease diagnosis, and particularly relates to a disease diagnosis method based on fuzzy neighborhood operators and generalized Shapley values. BACKGROUND
[0002] In medical diagnosis, accurate disease diagnosis often depends on the analysis of a large amount of medical data. These data contain various disease characteristics (conditional attributes) and corresponding diagnosis results (decision attributes). Attribute reduction is an important step to filter out key features from a large number of features, which can reduce data dimensionality and improve the efficiency and accuracy of the diagnosis model.
[0003] Traditional rough set reduction methods are usually based on the discriminative ability of a single attribute, such as dependency, information entropy, etc., which may ignore the interaction between attributes. If the classification ability of two attributes is weak when they are used individually, but significantly improved when they are used jointly, these seemingly "redundant" attributes may actually be necessary. Therefore, a new attribute reduction method is needed to consider the interaction between attributes.
[0004] In the prior art, some scholars have proposed a new non-additive operator using an overlap function, and introduced an attribute reduction method based on the fuzzy neighborhood measure of the operator, which considers the interaction between attributes. However, the method does not consider the correlation between attributes from a global perspective, and the commutativity of the overlap function limits the application range of the fuzzy neighborhood operator.
[0005] Based on the above problems, the present application proposes a disease diagnosis method based on fuzzy neighborhood operators and generalized Shapley values, which realizes attribute reduction considering global attribute correlation by constructing new fuzzy neighborhood operators, generalized Shapley values and Choquet integral, and applies it to disease diagnosis to improve the accuracy and efficiency of diagnosis. SUMMARY
[0006] The present application aims to solve the problem of ignoring global attribute correlation and the commutativity of the overlap function limiting the application range in the prior art, and proposes a disease diagnosis method based on fuzzy neighborhood operators and generalized Shapley values.
[0007] To achieve the above purpose, the present application adopts the following technology: a disease diagnosis method based on fuzzy neighborhood operators and generalized Shapley values, comprising the following steps:
[0008] Comprising the following steps:
[0009] Constructing fuzzy neighborhood operators and fuzzy neighborhoods, and calculating the upper and lower fuzzy neighborhood measures based on fuzzy neighborhoods;
[0010] determining a generalized Shapley value based on the upper and lower fuzzy neighborhood measures;
[0011] performing attribute reduction on a disease diagnosis decision table using the generalized Shapley value;
[0012] combining the reduced decision table with an intelligent classifier to obtain a disease diagnosis result.
[0013] Further, the constructing fuzzy neighborhood operator comprises:
[0014] In a fuzzy covering approximation space, based on a pseudo-overlapping function and its corresponding residual implication, a plurality of types of fuzzy neighborhood operators are defined.
[0015] Further, the constructing fuzzy neighborhood comprises:
[0016] In a fuzzy covering information table, for any fuzzy covering, a plurality of types of fuzzy neighborhoods are defined by combining the fuzzy neighborhood operator to capture the local neighborhood characteristics of samples under the fuzzy covering.
[0017] Further, the calculating upper and lower fuzzy neighborhood measures comprises: based on the fuzzy neighborhood, integrating neighborhood membership and inclusion membership, defining upper and lower fuzzy neighborhood measures to quantify the inclusion relationship between fuzzy neighborhoods and distinguish the difference between upper approximation and lower approximation.
[0018] Further, the determining generalized Shapley value comprises:
[0019] The upper and lower fuzzy neighborhood measures Shapley values are combined to expand the upper and lower generalized Shapley values to quantify the importance of attributes or attribute subsets in the fuzzy neighborhood environment and consider the interaction between attributes.
[0020] Further, the attribute reduction using generalized Shapley value comprises:
[0021] Based on the generalized Shapley value, the importance of attributes in the decision information table is evaluated, redundant attributes are removed, key diagnostic information is retained, and a reduced decision information table is obtained.
[0022] Further, the Choquet integral is defined based on the upper and lower generalized Shapley values to integrate multi-attribute information.
[0023] Further, attribute reduction of the decision information table is performed based on the Choquet integral, specifically: after integrating attribute information using the Choquet integral, the importance of attributes is evaluated and redundant attributes are removed to obtain a reduced decision information table.
[0024] In summary, due to the adoption of the above-mentioned technical disease diagnosis method based on the fuzzy neighborhood operator and the generalized Shapley value, the beneficial effects of the present application are:
[0025] 1. At the level of fuzzy information processing, the newly defined fuzzy neighborhood operator and fuzzy neighborhood break through the limitations of traditional methods in describing the relationship between samples in a fuzzy environment. These new operators and neighborhoods can more flexibly and subtly capture the fuzzy characteristics and local neighborhood information between samples, fully combine the structural characteristics of fuzzy coverage, and improve the representation ability of complex and fuzzy medical data, laying a solid foundation for subsequent information analysis and processing.
[0026] 2. The fuzzy neighborhood measure realizes the comprehensive quantification of the inclusion relationship of the fuzzy neighborhood, provides a more detailed measurement standard for fuzzy information by distinguishing the difference between the upper and lower approximations, enhances the analysis ability of fuzzy information, and makes the originally difficult-to-quantify fuzzy relationship measurable and analyzable. Meanwhile, the combination of the fuzzy neighborhood measure and the generalized Shapley value expands the application of the Shapley value in the fuzzy context, accurately evaluates the importance of attributes in the fuzzy environment, considers the interaction between attributes, overcomes the shortcomings of traditional methods in processing fuzzy information, and makes the evaluation of attribute importance more accurate and reasonable.
[0027] 3. In terms of attribute reduction, the two new methods based on the generalized Shapley value and Choquet integral can effectively eliminate redundant features in the decision table. They significantly reduce the number of features while retaining key diagnostic information, significantly reduce the data dimension, and reduce the computational burden, creating favorable conditions for the efficient operation of subsequent intelligent classifiers. The Choquet integral based on the upper and lower generalized Shapley values fully integrates the advantages of both, considers the correlation and importance difference between attributes, and realizes the accurate fusion of multi-attribute information. Compared with traditional integral methods, it can more reasonably integrate attribute information in a fuzzy environment, improve the reliability of information fusion, and provide a stronger basis for diagnostic decision-making.
[0028] 4. The proposed fuzzy decision method integrates upper and lower approximation information, effectively deals with the common information fuzziness and uncertainty in medical diagnosis, and can adapt to different decision-making scenarios through parameter adjustment, improving the flexibility and reliability of decision-making and providing a more scientific judgment method for disease diagnosis. In practical medical diagnosis applications, the number of features is significantly reduced through attribute reduction, while key information is retained. Based on this, intelligent classifiers are used for diagnosis, which not only significantly improves the accuracy of diagnosis, but also optimizes the model performance and reduces the complexity of the model, enabling it to function stably in different types of disease diagnosis datasets, with wide applicability and practical value. It provides an efficient and accurate new method for the medical diagnosis field. Attached Figure Description
[0029] Figure 1 A flowchart of the method of the present invention is shown. Detailed Implementation
[0030] The disease diagnosis method based on fuzzy neighborhood operators and generalized Shapley values in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0031] To more clearly and intuitively demonstrate the practical application effects and advantages of this invention in disease diagnosis based on fuzzy neighborhood operators and generalized Shapley values, and to verify its feasibility and effectiveness, the invention is further described below with reference to embodiments. Through specific scenario simulations and data calculations, the method is explained in detail how it plays a role in actual disease diagnosis based on fuzzy neighborhood operators and generalized Shapley values, helping readers better understand the technical details and practical value of the invention. The invention is further described below with reference to embodiments;
[0032] See Figure 1 The present invention provides a disease diagnosis method applicable to fuzzy neighborhood operators and generalized Shapley values; comprising the following steps:
[0033] (1) Propose the (I, PO) fuzzy β neighborhood operator
[0034] Definition 1. Let there be a fuzzy β-covered approximate space β-FCAS, denoted as We present four classes of pseudo-overlapping functions PO and their corresponding residual implications. and Fuzzy beta neighborhood operators, these operators are defined as:
[0035] (1)
[0036] (2)
[0037] (3)
[0038] (4)
[0039] We call these four different types of fuzzy β-neighborhood operators (I, PO) fuzzy β-neighborhood operators.
[0040] Four types of fuzzy neighborhood operators based on pseudo-overlap function and its residual implication are defined, which breaks the limitation of traditional fuzzy neighborhood operator type and enriches the mathematical expression form of neighborhood operator in fuzzy covering approximation space. The four types of operators can more flexibly depict the neighborhood relationship between samples in fuzzy environment, providing diversified basic tools for subsequent fuzzy neighborhood construction and measure calculation, and enhancing the representation ability of complex fuzzy information.
[0041] (2) The (I, PO) fuzzy β-neighborhood is proposed
[0042] Definition 2. Let a β-fuzzy covering information table (β-FCIT) be given, and any is a β-fuzzy covering (β-FC), we define four new types of β-neighborhoods For any x∈U and x∈(U,F), r∈1,2,3,4, including the following four different types:
[0043]
[0044] The four new types of neighborhood defined based on fuzzy covering information table accurately capture the local neighborhood characteristics of samples under fuzzy covering. Compared with traditional fuzzy neighborhood, it fully combines the structural information of fuzzy covering, can more subtly describe the distribution of fuzzy information around the sample, and improves the accuracy of depicting the neighborhood relationship of samples, laying a more reasonable foundation for subsequent quantitative analysis of fuzzy neighborhood measure.
[0045] (3) The (I, PO) fuzzy β-neighborhood measure is proposed
[0046] Definition 3. Let (U,F) be a β-FCIT, define the upper and lower (I, PO) fuzzy β-neighborhood measures of X and For any x∈U:
[0047]
[0048] is a membership degree related to X, r∈1,2,3,4, membership degree (i.e. β-neighborhood), is the membership degree contained in X.
[0049] The definition of upper and lower fuzzy neighborhood measure realizes the comprehensive quantification of fuzzy neighborhood inclusion relationship by integrating "neighborhood membership degree" and "inclusion membership degree". The measure not only reflects the overall characteristics of sample neighborhood, but also distinguishes the difference between "upper approximation" and "lower approximation", providing more detailed quantitative indicators for information measurement in fuzzy environment and enhancing the analysis ability of fuzzy information.
[0050] (4) A generalized Shapley value based on the upper and lower (I, PO) fuzzy β neighborhood measure is proposed.
[0051] Definition 4. Suppose β-FCIT(U,F), P(U) is a function of U={x1,x2,...,x}. n A power set in}, for any POI-β fuzzy measure in U, It is the upper (I, PO)-β fuzzy measure in U. The upper and lower generalized Shapley values of U are respectively and have
[0052]
[0053] Combining fuzzy neighborhood measures with generalized Shapley values expands the application scope of Shapley values in fuzzy scenarios. Generalized Shapley values can quantify the importance of attributes (or subsets) in fuzzy neighborhood environments, while also considering the interactions between attributes. This overcomes the shortcomings of traditional Shapley values in handling fuzzy information, improving the accuracy and rationality of attribute importance assessment.
[0054] (5) Based on the above, a new method for attribute reduction of decision tables based on fuzzy β-neighborhood measure and generalized Shapley value is proposed.
[0055] Input: Input decision information table T = (U, A∪{d}), as shown in Table 1, where U = {x1, x2, ..., x n}, A={a1,a2,…,a m},b ij =a j (x i ),i∈{1,2,…,n},j∈{1,2,…,m}.
[0056] Table 1. Decision Information Table T=(U,A∪{d})
[0057]
[0058]
[0059] Output: r-generalized Shapley value reduction in T (where r∈{1,2,3,4})
[0060] Step 1: Induce a fuzzy-covering approximate space: for any a j ∈A, we define a fuzzy set in Let be the variance of attribute aj. It is a fuzzy beta-cover (β∈[0,1]). Thus, (U, F) is a fuzzy β-covering approximation space. We call T = (U, F, d) a fuzzy β-covering information decision table.
[0061] Step 2: Obtain all
[0062] Step 3: Calculate the minimal lower generalized Shapley value reduction of (U, F, d), the specific steps are as follows:
[0063] (a)
[0064] (b) For any D∈U / d
[0065] (c) Continue
[0066] (d)
[0067] (e)
[0068] (f) If
[0069] (g)
[0070] (h)
[0071] (i) Until
[0072] (g) Return
[0073] Based on the generalized Shapley value and Choquet integral, two attribute reduction methods are proposed to effectively eliminate redundant features in the decision table. As shown in Table 3, the reduction rates of the two methods for different disease data sets are very high, which greatly reduces the number of features while retaining key diagnostic information. This effect significantly reduces the data dimension and reduces the computational complexity, laying a foundation for the efficient operation of subsequent classifiers.
[0074] (6) Propose the Choquet integral of upper and lower generalized Shapley value
[0075] Definition 5. Suppose (U, F) is a β-FCIT, U = {x1, x2, …, x n}, A∈F(U). For any r = 1, 2, 3, 4, define a pair of upper and lower r-generalized Shapley value (U and ) of A about U, respectively represented as and
[0076]
[0077] where {x'1,x'2,…,x' n} is a parameter of {x1,x2,…,x n}, and A(x'1)≤A(x'2)≤…≤A(x' n ). P(U) is a power set of U={x1,x2,…,x n}, X (i) ={x' i ,x' i+1 ,…,x' n}, and
[0078]
[0079] The integral combines the advantages of the generalized Shapley value and Choquet integral, fully considers the correlation and importance difference between attributes, and realizes the accurate fusion of multi-attribute information. Compared with traditional integral methods, it can more reasonably integrate attribute information in fuzzy environment, improve the robustness and accuracy of information fusion, and provide a more reliable quantitative basis for decision-making.
[0080] (7) The (I, PO)-fuzzy decision of x is proposed
[0081] Definition 6. Let T=(U,F,d) be a β-FCITD, U / d={D1,…,D n}. For any x∈U, the β-fuzzy decision of x is denoted as:
[0082]
[0083] The Choquet integral decision-making method based on upper and lower generalized Shapley values integrates the upper and lower approximate information, and can effectively deal with the fuzziness and uncertainty of information in medical diagnosis. The decision-making method can adapt to different decision-making scenarios through parameter adjustment, improving the flexibility and reliability of decision-making, and providing a more scientific basis for disease diagnosis.
[0084] (8) Based on the above content, a new attribute reduction method of decision table based on Choquet integral of generalized Shapley value is proposed
[0085] Input: input decision information table T=(U,A∪{d}), as shown in Table 1, where U={x1,x2,…,x n}, A={a1,a2,…,a m}, b ij= a j (x i ),i∈{1,2,…,n},j∈{1,2,…,m}.
[0086] Output: r -Choquet Integral Reduction (where r ∈ {1,2,3,4})
[0087] Step 1: Induce a fuzzy-cover approximation space: For any a j ∈ A, we define a fuzzy set where is the variance of attribute aj. Then is a fuzzy β-cover (β ∈ [0,1]), Thus, (U,F) is a fuzzy β-cover approximation space. We call T = (U,F,d) a fuzzy β-cover information decision table (β-FCITD).
[0088] Step 2: Obtain all
[0089] Step 3: Calculate where x k ∈ U, U / d = {D1,…,D t ,…,D s} is the partition of universe U with respect to decision attribute d.
[0090] Step 4: Calculate the minimal lower generalized Choquet integral reduction of (U,F,d), with the following steps:
[0091] (a)
[0092] (b) For any D ∈ U / d
[0093] (c) Continue
[0094] (d)
[0095] (e)
[0096] (f) If
[0097] (g)
[0098] (h)
[0099] (i) Until
[0100] (g) Return
[0101] (9) Medical diagnosis method based on "attribute reduction + intelligent classifier"
[0102] Input: Complete medical diagnosis decision information table.
[0103] Output: Medical diagnosis results and decisions.
[0104] Step 1: Use the new method of decision table attribute reduction based on GSV and CHI of fuzzy neighborhood β measure to delete redundant disease characteristics and realize dimension reduction of medical decision information table.
[0105] Step 2: Select intelligent classifier model and establish intelligent classifier diagnosis model. According to the reduction determined in step 1, construct a new medical diagnosis data set. Then, use it to train the diagnosis model;
[0106] Step 3: Use the trained intelligent classifier to classify the test sample symptom set to obtain the medical diagnosis results.
[0107] (10) Application of medical diagnosis method based on "attribute reduction + intelligent classifier" in disease diagnosis problems
[0108] In order to evaluate the efficacy and efficiency of the attribute reduction algorithm in this study in the real world medical diagnosis scene, experiments were conducted using four data sets from the UCI data set, which are heart CT scan image data, diabetic retinopathy image data, lung cancer after large lung resection lesion data, and hepatitis data. The information of the four data sets is shown in Table 2. Then, the two reduction methods proposed (the reduced data information is shown in Table 3) are used to realize the diagnosis of medical diseases.
[0109] Table 2 Disease diagnosis data information
[0110]
[0111]
[0112] Table 3 Reduced data information
[0113]
[0114] Finally, on the one hand, the data in Table 2 is directly processed by medical diagnosis methods (support vector machine, decision tree, random forest, K nearest neighbor algorithm, naive Bayes model and logistic regression); on the other hand, the data is first optimized by the reduction method based on GSV and CHI proposed by us, and then the medical diagnosis method is used to process the data (the data after attribute reduction of Table 3 data), which realizes the role of improving the accuracy and optimizing the model. The results are shown in Tables 4, 5, 6 and 7.
[0115] Table 4 Comparison of medical diagnosis of cardiac CT scan image data (Dataset 1)
[0116] Table 5 Comparison of medical diagnosis of diabetic retinopathy image data (Dataset 2)
[0117]
[0118] Table 6 Comparison of medical diagnosis of post-lung resection lesion data (Dataset 3)
[0119]
[0120]
[0121] Table 7 Comparison of medical diagnosis of hepatitis data (Dataset 4)
[0122]
[0123] The above merely describes the preferred embodiments of the present application, but the protection scope of the present application is not limited thereto, and any person skilled in the art can make equivalent replacements or changes to the disease diagnosis method and the inventive concept thereof based on the fuzzy neighborhood operator and the generalized Shapley value within the technical range disclosed by the present application, which should be covered within the protection scope of the present application.
Claims
1. A disease diagnosis method based on fuzzy neighborhood operators and generalized Shapley values, characterized by, The method comprises the following steps: constructing fuzzy neighborhood operators and fuzzy neighborhoods, calculating upper and lower fuzzy neighborhood measures based on the fuzzy neighborhoods; determining generalized Shapley values based on the upper and lower fuzzy neighborhood measures; performing attribute reduction on a disease diagnosis decision information table using the generalized Shapley values; combining the reduced decision information table with an intelligent classifier to obtain a disease diagnosis result.
2. The disease diagnosis method based on fuzzy neighborhood operator and generalized Shapley value according to claim 1, characterized in that, The constructing of the fuzzy neighborhood operators comprises: In a fuzzy covering approximation space, a plurality of types of fuzzy neighborhood operators are defined based on pseudo-overlapping functions and corresponding residual implications. 3.The disease diagnosis method based on fuzzy neighborhood operator and generalized Shapley value according to claim 1, characterized in that, The constructing of the fuzzy neighborhoods comprises: In a fuzzy covering information table, a plurality of types of fuzzy neighborhoods are defined based on the fuzzy neighborhood operators for any fuzzy covering to capture local neighborhood features of samples under the fuzzy covering. 4.The disease diagnosis method based on fuzzy neighborhood operator and generalized Shapley value according to claim 1, wherein, The calculating of the upper and lower fuzzy neighborhood measures comprises: 5.The disease diagnosis method based on fuzzy neighborhood operator and generalized Shapley value according to claim 1, wherein, Based on the fuzzy neighborhoods, neighborhood membership degrees and inclusion membership degrees are integrated to define upper and lower fuzzy neighborhood measures to quantify the inclusion relationship between fuzzy neighborhoods and distinguish the difference between upper approximation and lower approximation. The determining of the generalized Shapley values comprises: 6.The disease diagnosis method based on fuzzy neighborhood operator and generalized Shapley value according to claim 1, wherein, The upper and lower fuzzy neighborhood measures Shapley values are combined to expand upper and lower generalized Shapley values to quantify the importance of attributes or attribute subsets in the fuzzy neighborhood environment and consider the interaction between attributes. The attribute reduction using the generalized Shapley values comprises: 7.The disease diagnosis method based on fuzzy neighborhood operator and generalized Shapley value according to claim 1, wherein, Based on the generalized Shapley values, the importance of attributes in the decision information table is evaluated, redundant attributes are removed, and key diagnosis information is retained to obtain a reduced decision information table. 8.The disease diagnosis method based on fuzzy neighborhood operator and generalized Shapley value according to claim 1, wherein, Choquet integrals are defined based on the upper and lower generalized Shapley values to fuse multi-attribute information. The attribute reduction of the decision information table based on the Choquet integrals comprises: after the attribute information is fused using the Choquet integrals, the importance of attributes is evaluated and redundant attributes are removed to obtain a reduced decision information table.