Auxiliary medical diagnosis method based on kernel fuzzy rough set

Through the method based on nuclear fuzzy rough set, the problem of multi-grained data feature extraction and fusion in the prior art is solved, and the accurate capture of nonlinear data relationships and unsupervised abnormal detection is achieved, which improves the efficiency and accuracy of auxiliary medical diagnosis.

CN120126729APending Publication Date: 2025-06-10SICHUAN UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202410321572.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-20
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

Existing auxiliary medical diagnostic methods are difficult to comprehensively utilize multi-grained data features, it is difficult to accurately capture the relationship between nonlinear data, and require a large amount of manual labeling of data, which limits the improvement of algorithm performance.

Method used

Using a method based on the kernel fuzzy rough set, the kernel fuzzy relationship matrix is ​​calculated through the Gaussian kernel function, a multi-grained fuzzy information particle set is constructed, the fuzzy approximation accuracy and fuzzy information particle anomaly is calculated, and the anomaly score of the sample is finally calculated by weighted summing to realize unsupervised anomaly detection.

Benefits of technology

Effectively handle the extraction and fusion of multi-grained data features, which can accurately capture the relationship between nonlinear data without labeling data for model training, realizing unsupervised abnormality detection of medical data and assisting medical diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0004750689050000031
    Figure BDA0004750689050000031
  • Figure BDA0004750689050000034
    Figure BDA0004750689050000034
  • Figure BDA0004750689050000035
    Figure BDA0004750689050000035
Patent Text Reader

Abstract

The invention discloses an auxiliary medical diagnosis method based on a kernel fuzzy rough set, and belongs to the technical field of medical data analysis, and the method comprises the following steps: S1, carrying out the standardization processing of medical data, and obtaining the data after the standardization processing; s2, calculating a kernel fuzzy relation matrix by using a Gaussian kernel function according to the normalized data; s3, selecting an attribute subset according to the kernel fuzzy relation matrix, and constructing a multi-granularity fuzzy information particle set; s4, calculating fuzzy approximation precision according to the multi-granularity fuzzy information particle set; s5, calculating the particle anomaly degree of the fuzzy information according to the fuzzy approximation precision; s6, calculating anomaly scores of all samples according to the fuzzy information particle anomaly degree; s7, judging whether the abnormal scores of the samples are greater than a threshold value one by one, if so, outputting medical data abnormal points, and assisting a doctor in diagnosis; otherwise, the data are regarded as normal data until all samples are judged. According to the method, the problems of extraction and fusion of multi-granularity data features of existing medical data anomaly detection are solved, nonlinear medical data with fuzzy uncertainty can be effectively processed, and therefore doctors can be assisted in finding the condition of a patient more quickly and more accurately.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical data analysis, and particularly to an unsupervised medical data anomaly detection method based on kernel fuzzy rough sets. Background Art

[0002] Auxiliary medical diagnosis aims to use anomaly detection methods to discover objects or data that deviate significantly from expected medical behaviors or patterns, thereby assisting medical staff in improving the accuracy and efficiency of diagnosis. In the medical field, abnormal medical data may indicate an abnormal health condition of a patient, requiring timely intervention and treatment to avoid the further development of the disease. Therefore, anomaly detection has important research significance and application value for medical auxiliary diagnosis. It can help medical staff promptly discover and handle patients' health problems, improving the quality of medical services; at the same time, it helps medical institutions more reasonably arrange medical resources and improve the efficiency of medical diagnosis.

[0003] Currently, anomaly detection methods based on machine learning have been widely used in the field of auxiliary medical diagnosis, mainly using classification, regression, and clustering algorithms. Specifically: Classification algorithms can divide patients into normal and abnormal categories based on their medical characteristics to help doctors quickly identify whether a patient has an abnormal condition. Building a classification model requires a large amount of labeled data for training, and manual annotation is a time-consuming, laborious, subjective, and complex behavior, increasing the modeling difficulty. Regression algorithms are mainly used to predict patients' medical data, such as predicting a patient's future health condition or disease risk. Doctors can judge whether a patient has an abnormal condition by comparing the predicted value with the actual situation. However, the regression model may face the problem of insufficient prediction accuracy. The physiological parameters of patients are affected by multiple factors, resulting in uncertainty in the prediction results of the model. Moreover, since each patient's condition is unique, a separate regression model needs to be established for each patient during anomaly detection, resulting in a significant increase in modeling costs. Clustering methods can divide similar patients into different groups through classification. Doctors can determine patients with abnormal conditions by identifying a small number of samples that do not conform to most medical behaviors, thereby promptly discovering potential health problems. However, the clustering model may be affected by parameter selection and is not suitable for detecting a large amount of high-dimensional medical data.

[0004] Generally speaking, the existing auxiliary medical diagnosis methods usually have the following problems: (1) Most of the existing methods are only applicable to data analysis at a single granularity and are difficult to comprehensively utilize data features from different granularities; (2) For non-linear data with fuzzy uncertainty, the existing detection methods are difficult to accurately capture the relationships between data; (3) Most of the anomaly detection methods used in the field of auxiliary medical diagnosis require a large amount of manually labeled data, thus limiting the improvement of the performance of auxiliary medical diagnosis algorithms. Summary of the Invention

[0005] Aiming at the above deficiencies in the prior art, an auxiliary medical diagnosis method based on kernel fuzzy rough sets provided by the present invention solves the problems of extraction and fusion of multi-granularity data features in existing medical data anomaly detection, and can effectively process non-linear data with fuzzy uncertainty, realizing unsupervised anomaly detection of medical data.

[0006] In order to achieve the above invention purpose, the technical solution adopted by the present invention is: an auxiliary medical diagnosis method based on kernel fuzzy rough sets, comprising the following steps:

[0007] S1. Obtain medical data and perform normalization processing on the medical data to obtain normalized data;

[0008] S2. According to the normalized data, use a Gaussian kernel function to calculate a kernel fuzzy relation matrix;

[0009] S3. According to the kernel fuzzy relation matrix, select an attribute subset and construct a multi-granularity fuzzy information granule set;

[0010] S4. According to the multi-granularity fuzzy information granule set, calculate the fuzzy approximation accuracy;

[0011] S5. According to the fuzzy approximation accuracy, calculate the anomaly degree of the fuzzy information granule;

[0012] S6. According to the anomaly degree of the fuzzy information granule, calculate the anomaly scores of all samples;

[0013] S7. Judge one by one whether the anomaly score of the sample is greater than the threshold. If so, output the anomaly data, and repeat step S7 until the judgment of all samples is completed. Otherwise, it is regarded as normal data, and the judgment of the next sample is carried out until the judgment of all samples is completed.

[0014] The beneficial effects of the present invention are as follows: using a Gaussian kernel function to calculate the kernel fuzzy relation matrix between samples, mapping the data to a high-dimensional space to better capture the relationship between data, and solving the problem that non-linear data is difficult to represent; at the same time, constructing fuzzy information granules from multiple granularity levels to obtain a fuzzy information granule set, solving the problem of extraction of multi-granularity data features; characterizing the uncertainty and anomaly characteristics of information granules through fuzzy approximation accuracy and anomaly degree of fuzzy information granules respectively, and then calculating the anomaly scores of samples by weighted summation of the anomaly degree of fuzzy information granules, solving the problem of fusion of multi-granularity data features; the present invention does not require labeled data for model training, and can effectively realize unsupervised anomaly detection of medical data, thereby assisting medical diagnosis.

[0015] Further, the expression of the normalization processing in step S1 is:

[0016]

[0017] Among them, f(·) is a normalization function; c t (x i ) represents the value of the sample x i on the attribute c t . and represent the maximum and minimum values of all samples on the attribute c t .

[0018] The beneficial effect of the above further solution is: normalizing the numerical data in the medical data, unifying the dimension, and reducing the data calculation amount.

[0019] Further, the specific steps of the step S2 are as follows:

[0020] S201. According to the data after normalization processing, obtain an attribute subset for each attribute, and obtain a relative attribute subset for the attribute set after removing one attribute;

[0021] S202. According to expert experience or previous research, set the Gaussian kernel parameter δ. According to the Gaussian kernel parameter, calculate the membership degrees of the kernel fuzzy relations generated by each attribute subset and the relative attribute subset:

[0022]

[0023]

[0024] Among them, K B (x i , x j ) is a membership function, indicating the degree to which the sample x i has a kernel fuzzy relation K j with the sample x B , K B is the kernel fuzzy relation generated by the attribute subset B; δ is the Gaussian kernel parameter, ||B(x i ) - B(x j )|| represents the Euclidean distance between two samples with respect to the attribute subset B; C is the set of all attributes with m attributes. Let B = {b 1 , b 2 , L, b h} (h ∈ [1, m]), b l (x i ), b l (x j ) respectively represent the values of the samples x i , x j on the attribute b l .

[0025] Next, let B = {c t}(t ∈ [1, m]) is a subset of attributes, and P = C - B is a relative subset of attributes. Calculate the membership degrees K of the kernel fuzzy relations generated by each subset of attributes B (x i , x j ), and the membership degree K of the kernel fuzzy relation generated by the relative subset of attributes P (x i , x j ).

[0026] S203. According to the membership degree K of the kernel fuzzy relation of the subset of attributes B (x i , x j ), obtain the kernel fuzzy relation matrix of the subset of attributes According to the membership degree K of the kernel fuzzy relation of the relative subset of attributes P (x i , x j ), obtain the kernel fuzzy relation matrix of the relative subset of attributes

[0027] The beneficial effect of the above further solution is that by using the Gaussian kernel function to construct the kernel fuzzy relation matrix, the problem of difficult to capture the similarity relationship between non-linear data is solved.

[0028] Furthermore, step S3 generates a multi-granularity fuzzy information granule set according to the kernel fuzzy relations generated by each subset of attributes:

[0029]

[0030]

[0031] Among them, G U (K B ) is the multi-granularity fuzzy information granule set; is the fuzzy information granule centered on the sample x B generated by the kernel fuzzy relation K i ; U is the sample set; B is the subset of attributes; K B is the kernel fuzzy relation generated by the subset of attributes B; is the membership degree of the kernel fuzzy relation K B , indicating the degree to which the samples x i and x j have the kernel fuzzy relation K B .

[0032] The beneficial effect of the above further solution is that by constructing fuzzy information granules from multiple granularities simultaneously, a multi-granularity fuzzy information granule set is obtained, solving the problem of extracting multi-granularity data features and making the description of data features more accurate.

[0033] Furthermore, the expression of the fuzzy approximation accuracy in step S4 is as follows:

[0034]

[0035]

[0036]

[0037]

[0038]

[0039] where is the fuzzy approximation accuracy; is the fuzzy set of the lower approximation of the fuzzy information granule, is the fuzzy set of the upper approximation of the fuzzy information granule; is the fuzzy similarity relation K B generating a fuzzy information granule centered on the sample x i ; is the membership function of the lower approximation of the fuzzy information granule, is the membership function of the upper approximation of the fuzzy information granule; K P (x, y) is the degree to which the samples x and y have the fuzzy relation K P under the induction of the attributes in the relative attribute subset P; is the membership function of the fuzzy information granule; inf is the infimum, sup is the supremum; max is the maximum function, min is the minimum function; B is the attribute subset; P is the relative attribute subset; x i is the sample; i is the sample label; y is the sample.

[0040] The beneficial effect of the above further solution is that the uncertainty of the information granule is characterized by the fuzzy approximation accuracy, which is used as a screening factor for the abnormality degree of the fuzzy information granule, improving the accuracy of abnormal point screening.

[0041] Furthermore, the expression of the abnormality degree of the fuzzy information granule in step S5 is as follows:

[0042]

[0043] where GAE(·) is the calculation function of the abnormality degree of the fuzzy information granule; is the fuzzy similarity relation K B generating a fuzzy information granule centered on the sample x i ; B is the attribute subset, P is the relative attribute subset; x i is the sample; i is the sample label; U is the sample set; is the fuzzy approximation accuracy.

[0044] The beneficial effects of the above further solution are as follows: By characterizing the abnormal characteristics of information granules through the abnormality degree of fuzzy information granules, as the screening factor for abnormal points, the accuracy of abnormal point screening is improved.

[0045] Further, the expression of the anomaly score in step S6 is:

[0046]

[0047]

[0048] where KFRAS(·) is the calculation function of the anomaly score; x i is the sample; i is the sample label; m is the maximum number of the attribute set; B t is the t-th attribute subset; GAE(·) is the calculation function of the fuzzy information granule anomaly factor; is the fuzzy information granule centered on the sample x t induced by the attribute subset B i ; is the weight function; U is the sample set.

[0049] The beneficial effects of the above further solution are as follows: By performing weighted summation on the multi-granularity fuzzy information granule abnormality degree to calculate the anomaly score of the sample, the problem of multi-granularity data feature fusion is solved, and the accuracy of abnormal point screening is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 is the flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0051] The following describes the specific embodiments of the present invention to facilitate those skilled in the art of the present technology to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art of the present technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions and creations using the concept of the present invention are within the scope of protection.

[0052] Example 1

[0053] In this embodiment, the kernel fuzzy rough set is an important rough computing model. It inherits the advantages of the kernel function and the fuzzy rough set. By introducing the Gaussian kernel, it maps the data into a high-dimensional space, effectively mining the potential information between samples. At the same time, it constructs a fuzzy similarity space using the concept of the fuzzy set to handle the uncertainty in the data more flexibly. In auxiliary medical diagnosis, medical data is imported into an information system (information table), where each row represents a patient (i.e., a sample), and each column represents an attribute of the patient (i.e., a feature). An information system is represented by the binary tuple (U, C), where U = {x 1 , x 2 ,..., x n} represents the set composed of n patients to be detected, and C = {c 1 , c 2 ,..., c m} represents the set composed of m conditional attributes.

[0054] In this embodiment, first, the Gaussian kernel function is selected to calculate the kernel fuzzy relation matrix of the attribute subset, solving the problem that it is difficult to characterize the similarity between complex non-linear data; then, fuzzy information granules are constructed from multiple granularity levels to solve the problem of extracting multi-granularity data features; the uncertainty and abnormal characteristics of the information granules are characterized by the fuzzy approximation accuracy and the abnormal degree of the fuzzy information granules respectively, and then the abnormal degrees of the multi-granularity fuzzy information granules are weighted and summed to calculate the abnormal score of the sample, solving the problem of fusing multi-granularity data features; this method does not require labeled data for model training and can effectively realize unsupervised anomaly detection of medical data, thus assisting doctors in medical diagnosis.

[0055] The expression for the normalization process in step S1 is:

[0056]

[0057] where f(·) is the normalization function, and c t (x i ) represents the value of the sample x i on the attribute c t , and represent the maximum and minimum values of all samples on the attribute c t .

[0058] In this embodiment, the value range of the numerical data is adjusted to the real number interval from 0 to 1 through the min-max normalization operation.

[0059] Step S2 is specifically as follows:

[0060] S201. For each attribute, obtain an attribute subset based on the normalized data, and for the attribute set with one attribute removed, obtain a relative attribute subset.

[0061] S202. Set the Gaussian kernel parameter δ according to expert experience or previous research. According to the Gaussian kernel parameter, calculate the membership degrees of the kernel fuzzy relations generated by each attribute subset and the relative attribute subset:

[0062]

[0063]

[0064] Among them, K B (x i , x j ) is a membership function, indicating the degree to which the sample x i and the sample x j have the kernel fuzzy relation K B . K B is the kernel fuzzy relation generated by the attribute subset B; δ is the Gaussian kernel parameter, ||B(x i ) - B(x j )|| represents the Euclidean distance between two samples with respect to the attribute subset B; C is the attribute set with m attributes. Let B = {b 1 , b 2 ,..., b h} (h ∈ [1, m]), and b l (x i ), b l (x j ) represent the values of the samples x i , x j on the attribute b l , respectively.

[0065] Next, let B = {c t} (t ∈ [1, m]) be the attribute subset, and P = C - B be the relative attribute subset. Calculate the membership degrees K B (x i , x j ) of the kernel fuzzy relations generated by each attribute subset, and the membership degrees K P (x i , x j ) of the kernel fuzzy relations generated by the relative attribute subset, respectively.

[0066] S203. According to the membership degree K B (x i , x j ) of the kernel fuzzy relation of the attribute subset, obtain the kernel fuzzy relation matrix of the attribute subset Membership degree K of the core fuzzy relation according to the relative attribute subset P (x i , x j ), the core fuzzy relation matrix of the relative attribute subset is obtained

[0067] In this embodiment, the core fuzzy relation is the core concept of the core fuzzy rough set theory, which refers to a fuzzy set K: U×U→[0,1] calculated by the core function defined on U×U. For any (x i , x j ) ∈ U×U, the membership degree represents the degree to which the sample x i has a relationship K j with the sample x B . The core fuzzy relation K B on U can be represented by a core fuzzy matrix .

[0068] According to the membership degrees of the core fuzzy relations generated by the above calculated attribute subsets and relative attribute subsets, the core fuzzy relation matrix M(K B ) of any attribute subset B and the core fuzzy relation matrix M(K P ) of the relative attribute subset P are calculated.

[0069] Furthermore, step S3 sets the attribute subset set, and generates a multi-granularity fuzzy information granule set according to the core fuzzy relations generated by each attribute subset:

[0070]

[0071]

[0072] where G U (K B ) is the multi-granularity fuzzy information granule set; is the fuzzy information granule centered on the sample x B generated by the core fuzzy similarity relation K i ; U is the sample set; B is the attribute subset; K B is the fuzzy relation generated by the attribute subset B; is the membership degree of the fuzzy relation K B , indicating the degree to which the samples x i and x j have the core fuzzy relation K B .

[0073] In this embodiment, in the fuzzy rough set theory, a fuzzy information granule (or "information granule") is a set of a series of objects induced by the core fuzzy relationship between objects. The data features are extracted by constructing fuzzy information granules from multiple granularity levels, which specifically includes the following steps:

[0074] Step 1: Selection of attribute subset

[0075] Any attribute subset can construct a set of core fuzzy relationships, and each core fuzzy relationship can further construct the fuzzy information granules of objects. To simplify the calculation, only one attribute is used to construct the information granules (i.e., basic fuzzy information granules). One attribute c 1 , c 2 , …, c m is successively selected from the set of all attributes C = {c t} to form an attribute subset B t = {c t}, thereby obtaining m attribute subsets

[0076] Step 2: Information granulation

[0077] The process of constructing information granules is also called information granulation. Information granulation is the process of dividing all objects into particles of different granularities. In the fuzzy rough set theory, a group of objects can be aggregated through the fuzzy relationship between objects to form information granules. Using any fuzzy relationship K B to divide all samples U to obtain a set of generalized fuzzy equivalence classes, that is, a set of multi-granularity fuzzy information granules:

[0078]

[0079] where G U (K B ) is the set of multi-granularity fuzzy information granules; is the fuzzy information granule centered on the sample x B generated by the fuzzy similarity relationship K i ; obviously, the information granule is a fuzzy set on U, and its membership function is:

[0080]

[0081] The formula for calculating the base of the information granule is:

[0082]

[0083] Step 3: Construction of multi-granularity fuzzy information granules

[0084] A multi-granularity fuzzy information granule is a set of fuzzy information granules with multiple information granules. For any sample x i , respectively under m attribute subsets B t ∈B, a set of multi-granularity fuzzy information granules of the sample can be obtained through information granulation.

[0085] The expression of the fuzzy approximation accuracy in step S4 is:

[0086]

[0087]

[0088]

[0089]

[0090]

[0091] Among them, is the fuzzy approximation accuracy; is the fuzzy set of the lower approximation of the fuzzy information granule, is the fuzzy set of the upper approximation of the fuzzy information granule; is the core fuzzy similarity relation K B generating a fuzzy information granule centered on the sample x i ; is the membership function of the lower approximation of the fuzzy information granule, is the membership function of the upper approximation of the fuzzy information granule; K P (x, y) is the degree to which the samples x and y have the core fuzzy relation K P under the induction of the attributes in the relative attribute subset P; is the membership function of the fuzzy information granule; inf is the infimum, sup is the supremum; max is the maximum function, min is the minimum function; B is the attribute subset, P is the relative attribute subset; x i is the sample; i is the sample label; y is the sample.

[0092] In this embodiment, the fuzzy approximation accuracy is a measure of the uncertainty of the information contained in the fuzzy information granule. The lower its value, the more discriminative the information granule is. The expression is:

[0093]

[0094] Among them, K P is a relative core fuzzy relation constructed from the attribute subset P = C - B, and is used to evaluate the approximation accuracy of the information granule. and represent A pair of fuzzy sets of the lower approximation and the upper approximation, representing the degree to which an object definitely belongs to an information granule and may belong to the information granule The degree is calculated by the formula:

[0095]

[0096]

[0097] Among them, T and S are two operation operators of fuzzy sets, and the calculation formula is:

[0098]

[0099]

[0100] The expression of the anomaly degree of the fuzzy information granule in step S5 is:

[0101]

[0102] Among them, GAE(·) is the calculation function of the anomaly degree of the fuzzy information granule; is the core fuzzy relation K B generates a fuzzy information granule centered on the sample x i ; B is an attribute subset; x i is a sample; i is the sample label; U is the sample set; is the fuzzy approximation accuracy; P is the relative attribute subset.

[0103] In this embodiment, the anomaly degree of the information granule is an index for measuring the anomaly degree of the fuzzy information granule. The core idea is that if a sample belongs to a more unique information granule (i.e., the lower the fuzzy approximation accuracy), it is more likely to be an outlier. It is necessary to calculate the anomaly degree of the information granules of all samples at multiple granularities.

[0104] The expression of the anomaly degree in step S6 is:

[0105]

[0106]

[0107] Among them, KFRAS(·) is the calculation function of the anomaly score; x i is a sample; i is the sample label; m is the maximum number of the attribute set; B t is the attribute subset number; GAE(·) is the calculation function of the anomaly degree of the fuzzy information granule; is the fuzzy information granule centered on the sample x t induced by the attribute subset B i ; is the weight mapping function; U is the sample set.

[0108] In this embodiment, the sample anomaly score is used to measure the likelihood of a sample belonging to an outlier. The anomaly features from multiple granularity fuzzy information granules are fused by weighted summation of the multi-granularity information granule anomaly degrees.

[0109] In this embodiment, let the anomaly score threshold of the outlier be μ ∈ (0, 1). If the anomaly score KFRAS(x i ) of the sample x i > μ, then x i is determined to be abnormal medical data. By comparing the anomaly score threshold μ of all samples one by one, all abnormal medical data can be calculated.

[0110] Embodiment 2

[0111] In this embodiment, an information table containing 5 samples and 3 attributes is given, as shown on the left side of Table 1 (the right side is the result after normalization).

[0112] Table 1 Original data and normalization results

[0113]

[0114] Step 1: Data normalization processing.

[0115] For the convenience of subsequent calculations, first perform min-max normalization on the input data to convert the numerical data to between 0 and 1. The processing result is shown on the right side of Table 1.

[0116] Step 2: Calculate the kernel fuzzy relation matrix.

[0117] Set the kernel parameter δ = 0.3, let B t = {c t}, and calculate the following kernel fuzzy relation matrix of the attribute subset:

[0118]

[0119]

[0120]

[0121] Let P t = C - B t , and calculate the relative attribute subset kernel fuzzy relation matrix:

[0122]

[0123]

[0124]

[0125] Step 3: Construct multi-granularity fuzzy information granules.

[0126] Describe the uncertainty and abnormal characteristics of data by constructing multi-granularity fuzzy information granules. For the kernel fuzzy relation matrix Construct a fuzzy information granule for each row in the matrix: Taking the sample x 1 as an example, from the fuzzy relation matrix we can get From the fuzzy relation matrix we can get From the fuzzy relation matrix we can get

[0127] Step 4: Calculate the fuzzy approximation accuracy.

[0128] Measure the uncertainty of the information contained in the multi-granularity fuzzy information granules by calculating their approximation accuracy. Taking the relative attribute subset P 1 as an example, calculate the approximation accuracy of the fuzzy information granule. Let Calculate

[0129]

[0130] Similarly, we can obtain:

[0131]

[0132]

[0133] Step 5: Calculate the abnormality degree of the fuzzy information granule.

[0134] Quantitatively describe the abnormality degree of the multi-granularity fuzzy information granule by calculating its abnormality degree. Taking the sample x 1 as an example, calculate the abnormality degree of its information granule 1 under the attribute subset B as:

[0135]

[0136] We can get its corresponding weight coefficient as:

[0137]

[0138] Step 6: Calculate the anomaly score of the sample.

[0139] Calculate the anomaly score of the sample by weighted summation of the abnormality degrees of a group of multi-granularity fuzzy information granules of the sample:

[0140]

[0141] KFRAS(x 2 ) ≈ 0.0758;

[0142] KFRAS(x 3 ) ≈ 0.1059;

[0143] KFRAS(x 4 ) ≈ 0.1263;

[0144] KFRAS(x 5 ) ≈ 0.0442。

[0145] Step 7: Perform anomaly determination through threshold comparison.

[0146] Compare the anomaly scores of all samples. Obviously, the anomaly score of sample x 4 is significantly higher than that of other samples. Let the anomaly score threshold μ be 0.12. Compare the anomaly scores of all samples with the threshold μ, then x 4 is determined as abnormal medical data, output the medical data of patient x 4 , and remind medical staff that there is an anomaly in this patient, and further screening, diagnosis and timely treatment are required.

Claims

1. An auxiliary medical diagnosis method based on kernel fuzzy rough sets, characterized in that: The following steps are involved: S1. Obtain medical data and perform normalization processing on the medical data to obtain normalized data; S2, using the Gaussian kernel function to calculate the kernel fuzzy relationship matrix based on the normalized data; S3, according to the kernel fuzzy relationship matrix, select attribute subsets and construct a multi-granularity fuzzy information particle set; S4, calculating the fuzzy approximation accuracy according to the multi-granularity fuzzy information granule set; S5. Calculate the granularity of the fuzzy information according to the fuzzy approximation accuracy; S6. Calculate the anomaly scores of all samples according to the anomaly degree of the fuzzy information granules; S7. Determine whether the abnormal score of each sample is greater than the threshold value. If so, output the abnormal data and repeat step S7 until all samples are judged. Otherwise, treat it as normal data and judge the next sample until all samples are judged.

2. The auxiliary medical diagnosis method based on kernel fuzzy rough sets according to claim 1 is characterized in that: The step S2 is specifically as follows: S201. According to the normalized data, an attribute subset is obtained for each attribute, and a relative attribute subset is obtained for the attribute set with one attribute removed. The normalized expression is: Among them, f(·) is the normalization function, c t (x i ) represents the sample x i In the attribute c t The value on and Indicates that all samples have attribute c t The maximum and minimum values ​​on ; S202. According to expert experience or previous research, the Gaussian kernel parameter δ is set, and then the membership degree of the kernel fuzzy relationship generated by each attribute subset and the relative attribute subset is calculated according to the Gaussian kernel parameter: Among them, K B (x i ,x j ) is a membership function, which means that sample x i With sample x j With kernel fuzzy relation K B The degree of K B is the kernel fuzzy relationship generated by the attribute subset B, δ is the Gaussian kernel parameter, ||B(x i )-B(x j )|| shows sample x i With x j The Euclidean distance between the attribute subset B, C is the set of all attributes with m attributes, let B = {b1, b2, L, b h }(h∈[1,m]), b l (x i ), b l (x j ) represent the samples x i and x j In attribute b l The value on S203, according to the membership degree K of the kernel fuzzy relationship of the attribute subset B (x i ,x j ), and obtain the kernel fuzzy relationship matrix of the attribute subset Membership degree K of kernel fuzzy relation based on relative attribute subsets P (x i ,x j ), and obtain the kernel fuzzy relationship matrix of the relative attribute subset 3. The auxiliary medical diagnosis method based on kernel fuzzy rough sets according to claim 1 is characterized in that: The expression of the multi-granularity fuzzy information granule set in step S3 is: Among them, G U (K B ) is a set of multi-granularity fuzzy information particles, K is the kernel fuzzy relation B Generated by sample x i is the fuzzy information particle centered, U is the sample set, B is the attribute subset, K B The kernel fuzzy relation generated for attribute subset B, K is the kernel fuzzy relation B The membership degree of sample x i and x j With kernel fuzzy relation K B degree.

4. The auxiliary medical diagnosis method based on kernel fuzzy rough sets according to claim 1 is characterized in that: The expression of the fuzzy approximation accuracy in step S4 is: in, is the fuzzy approximation accuracy, is the fuzzy set approximated by the fuzzy information granule, is the fuzzy set approximated on the fuzzy information granule, is the fuzzy similarity relation K B Generated by sample x i The fuzzy information particle centered at is the membership function approximated under the fuzzy information granule, is the approximate membership function on the fuzzy information granule, K P (x, y) is the attribute induction in the relative attribute subset P, and samples x and y have a fuzzy relationship K P degree, is the membership function of the fuzzy information particle, inf is the infimum, sup is the supremum, max is the maximum function, min is the minimum function, B is the attribute subset, P is the relative attribute subset, x i is a sample, i is the sample number, and y is a sample.

5. The auxiliary medical diagnosis method based on kernel fuzzy rough sets according to claim 1 is characterized in that: The expression of the abnormality of the fuzzy information granularity in step S5 is: Among them, GAE(g) is the calculation function of the abnormality of fuzzy information granularity, is the fuzzy similarity relation K B Generated by sample x i is the fuzzy information particle centered, B is the attribute subset, P is the relative attribute subset, x i is a sample, i is the sample number, U is the sample set, is the fuzzy approximation accuracy.

6. The auxiliary medical diagnosis method based on kernel fuzzy rough sets according to claim 1, characterized in that: The expression of the abnormal score in step S6 is: Among them, KFRAS(g) is the calculation function of the anomaly score, x i is a sample, i is the sample number, m is the maximum number of the attribute set, B t is the t-th attribute subset, GAE(g) is the calculation function of the abnormality degree of fuzzy information particles, For attribute subset B t Induced by sample x i The fuzzy information particle centered at is the weight function, and U is the sample set.

Citation Information

Cited By

  • Unsupervised electricity consumption anomaly detection method fusing multi-scale fuzzy information particles

    CN116304948A

  • Centralized management method for data center cluster

    CN120994145A