Auxiliary medical diagnosis method based on random migration of pellets and balls
Through the Markov random walk method based on particle sphere calculation, the problem that the prior art is difficult to mine multi-grained abnormal information in medical data is solved, and efficient multi-grained representation and auxiliary diagnosis of medical data are achieved.
Patent Information
- Application Number
- CN202411644781.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-18
- Publication Date
- 2025-06-10
AI Technical Summary
Existing auxiliary medical diagnostic methods are difficult to mine multi-grained abnormal information in medical data, and relying on manual labeling data limits the improvement of algorithm performance.
The Markov random walk method based on particle sphere calculation is adopted to cover the sample space by adaptively generating particle spheres of different sizes, calculate the center and radius of each particle sphere, construct a state transfer matrix based on particle spheres, iteratively calculate to generate a steady-state distribution, normalize it to the degree of abnormality of each particle sphere, and associate it with the covered sample to calculate the abnormality score of each sample.
It realizes multi-grained representation of medical data, reduces calculation costs, can quickly and accurately give diagnostic results, and does not require labeled data for model training, effectively realizing auxiliary diagnosis of medical data.
Smart Images

Figure CN120126730A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of medical data analysis, and particularly relates to an auxiliary medical diagnosis method based on the random walk of granulocytes. Background Art
[0002] The auxiliary medical diagnosis technology aims to identify individuals or data points that are significantly different from the conventional medical model through anomaly detection methods, so as to assist doctors to diagnose diseases more accurately and improve the diagnosis and treatment efficiency. In the medical field, abnormal data may indicate problems with the patient's health condition and need to be detected and treated early to prevent the condition from deteriorating. Therefore, anomaly detection technology has important application value for improving the quality of medical services.
[0003] Currently, machine learning technology plays an important role in auxiliary medical diagnosis, mainly including methods based on proximity, classification, and clustering. Proximity methods usually use the distance or density between medical data to identify abnormal patterns. However, in high-dimensional medical data, the distances between samples tend to be the same, and proximity methods cannot accurately detect anomalies. Classification methods analyze the feature information of data and classify the data into normal and abnormal categories to provide clear diagnostic guidance for doctors. However, a large amount of labeled data is required for training these models, and the data labeling process is both time-consuming and subjective, increasing the complexity of model training. Clustering algorithms identify abnormal medical data by dividing similar data into groups and identifying samples that do not conform to the behavior of the majority group. Clustering methods have the potential to reveal the internal structure of data, but their performance may be significantly affected by the selection of clustering algorithm parameters and may not be suitable for processing large-scale high-dimensional medical data.
[0004] The existing auxiliary medical diagnosis methods mainly face the following challenges: (1) Most methods are only applicable to a single granularity and do not fully utilize the multi-granularity information of data; (2) These methods usually rely on manually labeled data for training, which limits the improvement of algorithm performance. Therefore, it is crucial to study an anomaly detection method that can mine multi-granularity information in data to address these challenges and improve the accuracy and efficiency of auxiliary medical diagnosis. Summary of the Invention
[0005] Aiming at the deficiencies in the above-mentioned existing technologies, the auxiliary medical diagnosis method based on granulocyte computing Markov random walk provided by the present invention solves the problem that existing anomaly detection methods are difficult to mine multi-granularity anomaly information in medical data, thereby more effectively realizing auxiliary medical diagnosis.
[0006] To achieve the above invention purpose, the technical solution adopted by the present invention is: an auxiliary medical diagnosis method based on granulocyte computing Markov random walk, including the following steps:
[0007] S1. Obtain medical data, and perform standardization processing on the medical data to obtain the standardized data;
[0008] S2. Adaptively generate granule balls of different sizes to cover the sample space, and calculate the center and radius of each granule ball;
[0009] S3. Calculate the distances between the granule balls, and construct a state transition matrix based on the granule balls;
[0010] S4. Generate a steady-state distribution through iterative calculation, and normalize it to the degree of abnormality of each granule ball;
[0011] S5. Calculate the anomaly score of each sample by associating the degree of abnormality of each granule ball with the samples it covers;
[0012] S6. Judging one by one whether the anomaly score of the sample is greater than the threshold. If so, output the abnormal points of the medical data, and repeat step S6 until all samples are judged. Otherwise, it is regarded as normal data, and the next sample is judged until all samples are judged;
[0013] The beneficial effects of the present invention are as follows: The present invention introduces granule ball calculation, providing an innovative and efficient processing tool for the field of anomaly detection; using granule balls to cover data samples to achieve an adaptive and efficient multi-granularity representation of data, and then constructing a state transition matrix for Markov random walk. Since the number of granule balls is much smaller than the number of samples, the calculation cost is greatly reduced compared with the previous sample-based state transition matrix; the present invention can quickly and accurately give a diagnosis result, thereby reducing the waiting time of patients and improving the medical experience of patients; the present invention does not require labeled data for model training and can effectively realize the auxiliary diagnosis of medical data.
[0014] Further, the expression of the normalization processing in step S1 is:
[0015]
[0016] where f(·) is the normalization processing; x a is the value of the sample x in the medical data on the attribute a; A a is the set composed of the values of all samples in the medical data on the attribute a; min(g) is the minimum value function; max(g) is the maximum value function.
[0017] The beneficial effect of the above further solution is: Performing normalization operation on medical data eliminates the influence of different dimensions and reduces the data calculation amount.
[0018] Further, step S2 is specifically:
[0019] S201. Initialize the entire dataset as a granule ball GB 0 , GB 0 contains all sample points in the medical dataset;
[0020] Define the center and radius of the granule ball:
[0021]
[0022] where c k and r k are the center and radius of the granule ball GB k respectively, x l represents the l-th sample covered by the granule ball GB k , n k is the number of samples covered by the granule ball GB k , and are the values of the sample x l on the 1st dimensional attribute and the b-th dimensional attribute respectively, b is the total number of attributes, is the value of the granule ball center c k on the j-th dimension;
[0023] S202. Execute the 2-Means clustering algorithm on the granule ball GB 0 to divide it into two sub-granule balls and
[0024] S203. Calculate the total distance SD of the original granule ball GB 0 , the sub-granule balls and respectively, that is, the sum of the distances from all samples covered by the granule ball to the granule ball center, and its definition formula is:
[0025]
[0026] where SD k represents the total distance of the granule ball GB k , x l represents the l-th sample covered by the granule ball GB k , n k is the number of samples covered by the granule ball GB k , c k is the center of the granule ball GB k ;
[0027] Set the splitting criterion of the granule ball, that is, if the total distance SD of the original granule ball is greater than the sum of the total distances SD of the two sub-granule balls, the original granule ball will be allowed to be split into two sub-granule balls, otherwise the original granule ball will not be split; the splitting criterion is defined as:
[0028]
[0029] Among them, GB k is the original granule sphere, and are two sub-granule spheres after the original granule sphere GB k is split, SD k is the total distance of the original granule sphere GB k , is the sum of the total distances SD of the two sub-granule spheres, and respectively represent the total distances of the sub-granule spheres and ;
[0030] S204. The granule spheres that meet the splitting standard are split according to the 2-Means clustering algorithm, and S204 is repeated until all granule spheres no longer meet the splitting conditions;
[0031] S205. Obtain the final granule sphere set GBs = {GB 1 , GB 2 , K, GB m}, where m is the total number of granule spheres.
[0032] The beneficial effects of the above further solution are as follows: By accurately calculating the center and radius of each granule sphere, the distribution characteristics of the data can be captured and represented more accurately; compared with the traditional method of processing each sample one by one, the present invention effectively reduces the computational complexity through the hierarchical representation of granule spheres, especially when dealing with high-dimensional data.
[0033] Further, the specific steps of step S3 are as follows:
[0034] S301. Calculate the distances between each granule sphere and construct a distance matrix based on the granule spheres:
[0035] Dis = [dis(GB i , GB j )] m×m
[0036] Among them, Dis is the distance matrix based on the granule spheres, m represents the number of generated granule spheres, and dis(GB i , GB j ) is the element Dis(i, j) in the i-th row and j-th column of Dis, which is the distance between the granule sphere GB i and the granule sphere GB j . The calculation formula is:
[0037] dis(GB i , GB j ) = ||c i - c j || + ri +r j
[0038] where dis(GB i , GB j ) represents the distance between granule balls GB i and granule ball GB j , c i and c j are the centers of granule balls GB i and granule ball GB j respectively, r i and r j are the radii of granule balls GB i and granule ball GB j respectively, and ||g|| represents the 2-norm operation;
[0039] S302. Construct a state transition matrix through the distance matrix:
[0040] P = E -1 Dis
[0041] where P is the state transition matrix, E = diag(e 1 , e 2 , K, e m ) is a diagonal matrix, and the i-th diagonal element in E is m represents the number of generated granule balls, Dis is the distance matrix based on granule balls, Dis(i, j) is the distance between granule balls GB i and granule ball GB j , and E -1 represents the inverse matrix of the diagonal matrix E.
[0042] The beneficial effects of the above further solution are as follows: By representing based on granule balls instead of directly calculating on the original data points, the size of the distance matrix is significantly reduced, thus significantly reducing the storage and calculation complexity; The state transition matrix based on granule balls helps to capture the dynamic changes and patterns of the data in the local neighborhood, thus more effectively discovering the internal structure of the data.
[0043] Furthermore, the specific steps of step S4 are as follows:
[0044] S401. Initialize the time t = 0 and set the probability distribution as m represents the number of generated granule balls;
[0045] S402. Perform Markov random walk on the probability distribution s (t) at time t with probability values.
[0046] S403. Repeat step S402 until s converges to the steady-state distribution s *, that is, s (t+1) -s (t) |≤10 -4 , otherwise, increment the time t by one and proceed to step S402;
[0047] S404. Let each probability value of the steady-state distribution s * represent the degree of abnormality of the corresponding granule ball:
[0048] AD(GB k ) = s * (k)
[0049] where AD(GB k ) represents the degree of abnormality of the granule ball GB k , and s * (k) is the k-th value of the steady-state distribution s * .
[0050] The beneficial effects of the above further solution are as follows: The Markov random walk process can more accurately identify abnormal or unusual data patterns by considering the transition probabilities between granule balls; compared with analyzing each data point one by one, MRW performs calculations at the granule ball level, reducing the computational amount and improving the efficiency of anomaly detection; the Markov random walk process is somewhat robust to noise through iterative calculation of the steady-state distribution and can more stably identify anomalies.
[0051] Furthermore, in the specific step S5, the anomaly score of the sample is defined as:
[0052] AS(x i ) = AD(GB k ) × W(GB k ), x i ∈GB k
[0053] where AS(x i ) represents the anomaly score of the sample x i , AD(GB k ) represents the degree of abnormality of the granule ball GB i to which the sample x k belongs, is the weight of the granule ball GB k , |n k | represents the number of samples covered by the granule ball GB k , and |X| is the total number of samples in the sample set X.
[0054] The beneficial effects of the above further solution are as follows: Characterize the anomaly characteristics of the sample through the degree of abnormality and weight of the granule ball to which the sample belongs, and use them as screening factors for anomaly points to improve the accuracy of anomaly point screening. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] Figure 1 This is the flowchart of the method of the present invention. Detailed implementation manners
[0056] The following describes the detailed implementation manners of the present invention to facilitate those skilled in the art of the present technology to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the detailed implementation manners. For those of ordinary skill in the art of the present technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions created using the concept of the present invention are within the scope of protection.
[0057] Embodiment 1
[0058] As Figure 1 shown, in an embodiment of the present invention, an auxiliary medical diagnosis method based on granular ball computing Markov random walk includes the following steps:
[0059] S1. Obtain medical data and perform standardization processing on the medical data to obtain the data after standardization processing;
[0060] S2. Adaptively generate granular balls of different sizes to cover the sample space, and calculate the center and radius of each granular ball;
[0061] S3. Calculate the distances between the granular balls and construct a state transition matrix based on the granular balls;
[0062] S4. Generate a steady-state distribution through iterative calculation and normalize it to the degree of abnormality of each granular ball;
[0063] S5. Calculate the abnormality score of each sample by associating the degree of abnormality of each granular ball with the samples it covers;
[0064] S6. Judging one by one whether the abnormality score of the sample is greater than the threshold. If so, output the abnormal points of the medical data, and repeat step S6 until all samples are judged. Otherwise, it is regarded as normal data, and the next sample is judged until all samples are judged;
[0065] In this embodiment, the granular ball computing theory provides an effective tool for multi-granularity information fusion and can be directly applied to the analysis model of medical data. In the specific application of granular ball computing, medical data is imported into an information system (or called an information table), where each row represents a patient (or called a sample), and each column represents an attribute (feature) of the sample. The values of the attributes can include numerical types (such as blood pressure, blood sugar, heart rate, etc.). An information system is represented as <X, A>, where X represents the set of all objects and A represents the set of all attributes.
[0066] In this embodiment, first, granular balls with different granularity sizes are constructed. The local characteristics of the samples covered by each granular ball are characterized by the center and radius information of the granular ball. Then, the abnormal characteristics of the granular ball are characterized by Markov random walk. Finally, the degree of abnormality of the granular ball is weighted to calculate the abnormal score of the corresponding sample, solving the problem that the existing method only mines abnormal information from a single granularity. This method does not require labeled data for model training and can effectively achieve unsupervised anomaly detection of medical data, thus achieving the purpose of assisting medical diagnosis.
[0067] The expression for the normalization process in step S1 is:
[0068]
[0069] where f(·) is the normalization process; x a is the value of sample x in medical data on attribute a; A a is the set composed of the values of all samples in medical data on attribute a; min(g) is the minimum value function; max(g) is the maximum value function.
[0070] In this embodiment, the value range of the data is adjusted to the real number interval from 0 to 1 through min-max normalization operation, eliminating the influence of different dimensions on the algorithm performance.
[0071] Step S2 is specifically as follows:
[0072] S201. Initialize the entire data set as a granular ball GB 0 , GB 0 contains all sample points in the medical data set;
[0073] Calculate the center and radius of the granular ball:
[0074]
[0075] where c k and r k are the center and radius of the granular ball GB k respectively, x l represents the l-th sample covered by the granular ball GB k , n k is the number of samples covered by the granular ball GB k , and are the values of sample x l on the first dimension attribute and the b-th dimension attribute respectively, b is the total number of attributes, is the value of the granular ball center c k on the j-th dimension;
[0076] S202. For the granular ball GB0 Execute the 2-Means clustering algorithm to divide it into two sub-grain balls and
[0077] S203. Calculate the total distance SD of the original grain ball GB 0 , sub-grain balls and respectively, that is, the sum of the distances from all samples covered by the grain ball to the center of the grain ball. Its definition formula is:
[0078]
[0079] where SD k represents the total distance of the grain ball GB k , x l represents the l-th sample covered by the grain ball GB k , n k is the number of samples covered by the grain ball GB k , c k is the center of the grain ball GB k ;
[0080] Set the splitting criterion for the grain ball, that is, if the total distance SD of the original grain ball is greater than the sum of the total distances SD of the two sub-grain balls, the original grain ball is allowed to be split into two sub-grain balls, otherwise the original grain ball is not split; the splitting criterion is defined as:
[0081]
[0082] where GB k is the original grain ball, and are the two sub-grain balls after splitting the original grain ball GB k , SD k is the total distance of the original grain ball GB k , is the sum of the total distances SD of the two sub-grain balls, and represent the total distances of the sub-grain balls and respectively;
[0083] S204. The grain balls that meet the splitting criterion are split according to the 2-Means clustering algorithm. Repeat S204 until all grain balls no longer meet the splitting conditions;
[0084] S205. Obtain the final grain ball set GBs = {GB 1 , GB 2 , K, GB m}, where m is the total number of grain balls.
[0085] In this embodiment, sample data is covered with spheres of different sizes, and the local characteristics of the samples covered by each sphere are characterized by the center and radius of the sphere.
[0086] The specific steps of step S3 are as follows:
[0087] S301. Calculate the distance between each sphere to construct a distance matrix based on the spheres:
[0088] Dis = [dis(GB i , GB j )] m×m
[0089] where Dis is the distance matrix based on the spheres, m represents the number of generated spheres, and dis(GB i , GB j ) is the element Dis(i, j) in the i-th row and j-th column of Dis. Dis(i, j) is the distance between sphere GB i and sphere GB j . The calculation formula is:
[0090] dis(GB i , GB j ) = ||c i - c j || + r i + r j
[0091] where dis(GB i , GB j ) represents the distance between sphere GB i and sphere GB j . c i and c j are the centers of sphere GB i and sphere GB j respectively. r i and r j are the radii of sphere GB i and sphere GB j respectively. ||g|| represents the 2-norm operation;
[0092] S302. Construct a state transition matrix through the distance matrix:
[0093] P = E -1 Dis
[0094] where P is the state transition matrix, E = diag(e 1 , e 2 , K, e m ) is a diagonal matrix. The i-th diagonal element in E is m represents the number of generated granular balls, Dis is the distance matrix based on granular balls, and Dis(i,j) is the distance between granular ball GB i and granular ball GB j ; E -1 represents the inverse matrix of the diagonal matrix E.
[0095] In this embodiment, using granular balls to construct the state transition matrix significantly reduces the complexity of data storage and calculation compared with the traditional method of constructing the state transition matrix based on samples.
[0096] The specific steps of step S4 are as follows:
[0097] S401. Initialize the time t = 0 and set the probability distribution as m represents the number of generated granular balls;
[0098] S402. Perform Markov random walk on the probability distribution s (t) at time t with probability values.
[0099] S403. Repeat step S402 until s converges to the steady-state distribution s * , that is, s (t+1) -s (t) |≤10 -4 , otherwise increment the time t by one and enter step S402;
[0100] S404. Each probability value of the steady-state distribution s * represents the degree of abnormality of the corresponding granular ball:
[0101] AD(GB k ) = s * (k)
[0102] where AD(GB k ) represents the degree of abnormality of granular ball GB k , s * (k) is the k-th value of the steady-state distribution s * .
[0103] In this example, using Markov random walk to characterize the abnormal characteristics of granular balls provides an effective mathematical tool for the abnormal detection of granular balls.
[0104] The calculation formula for the abnormal score of the sample in step S5 is defined as:
[0105] AS(x i ) = AD(GB k ) × W(GB k ), x i ∈GB k
[0106] Among them, AS(x i ) represents the anomaly score of sample x i , and AD(GB k ) represents the anomaly degree of the granule ball GB i to which the sample x k belongs. is the weight of the granule ball GB k , |n k | represents the number of samples covered by the granule ball GB k , and |X| is the total number of samples in the sample set X.
[0107] In this embodiment, the sample anomaly score is used to measure the degree to which a sample belongs to an anomaly, and the anomaly score of the sample is obtained by weighting the granule ball anomaly factors containing multi-granularity information.
[0108] In this embodiment, let the anomaly score threshold for judging an anomaly point be μ ∈ (0, 1). If the anomaly score AS(x i ) of the sample x i > μ, then x i is determined to be an anomaly point; by comparing the anomaly scores of all samples with the threshold μ one by one, the anomaly samples in the medical data can be effectively counted. This process can timely send a warning to medical staff, indicating the possible health risks of those patients with higher anomaly scores, so as to prompt them to conduct more in-depth examinations and diagnoses on these patients to ensure that appropriate treatment measures can be taken in a timely manner.
Claims
1. An auxiliary medical diagnosis method based on random walk of a particle sphere, characterized in that: The following steps are involved: S1. Obtain medical data and perform standardized processing on the medical data to obtain standardized data; S2, adaptively generate spheres of different sizes to cover the sample space, and calculate the center and radius of each sphere; S3, calculating the distance between the particles and balls, and constructing a state transfer matrix based on the particles and balls; S4, generating a steady-state distribution through iterative calculation and normalizing it to the abnormality of each sphere; S5. Calculate the anomaly score of each sample by correlating the degree of anomaly of each sphere with the sample it covers; S6. Determine whether the abnormal score of each sample is greater than the threshold value. If so, output the abnormal point of the medical data and repeat step S6 until all samples are judged. Otherwise, treat it as normal data and judge the next sample until all samples are judged.
2. The auxiliary medical diagnosis method based on random walk of particles and balls according to claim 1, characterized in that: The expression for the normalization process in step S1 is: Among them, f(·) is normalized; x a is the value of sample x in the medical data on attribute a; A a is the set of values of all samples in the medical data on attribute a; min(g) is the minimum function; max(g) is the maximum function.
3. The auxiliary medical diagnosis method based on random walk of particles and balls according to claim 1, characterized in that: The step S2 is specifically as follows: S201, initialize the entire data set as a sphere GB0, and calculate the center and radius of the sphere: Among them, c k and r k GB k The center and radius, x l GB k The lth sample covered, n k GB k The number of samples covered, and The samples x are l The values of the first dimension attributes and the bth dimension attributes, where b is the total number of attributes; is the center of the sphere c k The value in the jth dimension; S202: Execute 2-Means clustering algorithm on the ball GB0 to divide it into two sub-balls and S203, respectively calculate the original particle ball GB0, the sub-particle ball and The total distance SD is the sum of the distances from all samples covered by the sphere to the center of the sphere, and its definition formula is: Among them, SD k GB k The total distance, x l GB k The lth sample covered, n k GB k The number of samples covered, c k GB k The center of , ||g|| represents the 2-norm operation; Set the segmentation criteria of the spheres, that is, if the total distance SD of the original sphere is greater than the sum of the total distance SD of the two sub-spheres, the original sphere will be allowed to be segmented into two sub-spheres, otherwise the original sphere will not be segmented; the segmentation criteria are defined as: Among them, GB k For the original pellet, and The original ball GB k The two sub-spheres after segmentation, SD k The original ball GB k The total distance, is the sum of the total distances between the two pellets, and Respectively represent the particle sphere and Total distance; S204, the spheres meeting the segmentation criteria are segmented using the 2-Means clustering algorithm, and S204 is repeated until all spheres no longer meet the segmentation criteria; S205, get the final set of spheres GBs = {GB1, GB2, K, GB m }, m is the total number of spheres.
4. The auxiliary medical diagnosis method based on random walk of particles and balls according to claim 1, characterized in that: The step S3 is specifically as follows: S301, calculate the distance between each particle ball, and construct a distance matrix based on the particle balls: Dis=[dis(GB i ,GB j )] m×m Where Dis is the distance matrix based on the particle sphere, m represents the number of generated particles, dis(GB i ,GB j ) is the element in row i and column j in Dis, Dis(i,j) is the particle ball GB i GB j The distance is calculated as: haze(GB i ,GB j )=||c i -c j ||+r i +r j Among them, dis(GB i ,GB j ) represents the sphere GB i GB j The distance between i and c j GB i GB j The center of i and r j GB i GB j The radius of , ||g|| represents the 2-norm operation; S302. Construct a state transfer matrix through the distance matrix: P=E -1 Dis Where P is the state transfer matrix, E = diag (e1, e2, K, e m ) is a diagonal matrix, and the i-th diagonal element in E is m represents the number of generated spheres, Dis is the distance matrix based on spheres, Dis(i,j) is the distance matrix of spheres GB i GB j The distance between -1 represents the inverse matrix of the diagonal matrix E.
5. The auxiliary medical diagnosis method based on random walk of particles and balls according to claim 1, characterized in that: The step S4 is specifically as follows: S401, initialize time t=0, and set the probability distribution to m represents the number of spheres generated; S402, probability distribution s for time t (t) Perform Markov random walk with probability values; S 403, repeat step S 402 until s converges to a steady-state distribution s * , that is, s (t+1) -s (t) |≤10 -4 , otherwise the time t is increased by one and the process goes to step S402; S404, the steady-state distribution s * Each probability value represents the degree of abnormality of the corresponding ball: AD(GB k )=s * (k) Among them, AD(GB k ) represents the sphere GB k The degree of abnormality, s * (k) is the steady-state distribution s * The kth value of .
6. The auxiliary medical diagnosis method based on random walk of a particle sphere according to claim 1, characterized in that: The anomaly score of the sample in step S5 is defined as: AS(x i )=AD(GB k )×W(GB k ),x i ∈GB k Among them, AS(x i ) represents the sample x i The anomaly score, AD(GB k ) represents the sample x i Belong to the ball GB k The degree of abnormality, GB k The weight of |n k | indicates a sphere GB k The number of samples covered, |X| is the total number of samples in sample set X.