Rheumatism Activity Assessment System Based on Medical Information Processing
By introducing a method of determining weights to the nearest high-density point in the rheumatism activity assessment system, the problems of data scarcity and outlier interference are solved, and the evaluation accuracy and fine division ability are significantly improved.
Patent Information
- Application Number
- CN202510459352.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-04-14
AI Technical Summary
The existing rheumatism activity assessment system has problems of scarce data and outliers when dealing with severe rheumatism patients, resulting in poor assessment accuracy, especially when there are fewer samples of high-motion patients and large fluctuations, it is difficult to achieve fine division.
By introducing the distance to the nearest high-density point, noise interference is avoided; normalized labels are introduced to supplement a few samples of severe rheumatism patients; adaptive sampling is achieved based on density and sparseness indicators; weights are determined based on similarity calculations to ensure that the generation of synthetic samples is not only integrated with population-centered information, but also fully considering the clinical similarity between samples.
The accuracy of rheumatoid mobility assessment is significantly improved, ensuring a fine division of patients with severe rheumatism, and avoiding the generation of noise samples at the extreme edge of mobility.
Smart Images

Figure CN119993544B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing, and specifically refers to a rheumatoid activity assessment system based on medical information processing. Background Art
[0002] The goal of a rheumatoid activity assessment system is to accurately reflect the activity status of a patient's rheumatoid disease by objectively and quantitatively integrating multiple clinical indicators, so as to guide clinical treatment and monitor the progression of the disease. However, in general, for the scarce group of severe rheumatoid patients, the data of the rheumatoid activity assessment system is vulnerable to the interference of outliers and the scarcity of samples, making it difficult to support the state division, which in turn leads to poor accuracy of the rheumatoid activity assessment; when there are few samples and large fluctuations in high-activity patients, the general rheumatoid activity assessment system is easily affected by abnormal points or noise data and cannot achieve a fine division of severe rheumatoid patients. Summary of the Invention
[0003] In view of the above situation, to overcome the defects of the prior art, the present invention provides a rheumatoid activity assessment system based on medical information processing. For the problem that in general, for the scarce group of severe rheumatoid patients, the data of the rheumatoid activity assessment system is vulnerable to the interference of outliers and the scarcity of samples, making it difficult to support the state division, which in turn leads to poor accuracy of the rheumatoid activity assessment, this solution introduces the distance to the nearest high-density point to avoid noise interference; introduces a normalized label to supplement the minority class samples of severe rheumatoid patients; realizes adaptive sampling based on density and sparsity indicators; determines weights based on similarity calculation to ensure that when generating synthetic samples, the group center information is incorporated and the clinical similarity between samples is fully considered, avoiding the generation of noise samples at the extreme edges of the activity; thereby significantly improving the accuracy of subsequent rheumatoid activity assessment; for the problem that in general, when there are few samples and large fluctuations in high-activity patients, the rheumatoid activity assessment system is easily affected by abnormal points or noise data and cannot achieve a fine division of severe rheumatoid patients, this solution calculates the clinical proximity distance between each pair of patients to more accurately reflect the subtle differences in clinical status; updates the clustering center based on the clinical indicator coordinates and additional weights to ensure that the updated clustering center is more concentrated in the high-density area and truly reflects the central trend of rheumatoid activity; introduces a correction coefficient to adaptively adjust the correction strength using the local density gradient for clustering center correction to ensure that each activity level has a representative clustering center; thereby ensuring a fine division of subsequent rheumatoid activity assessment.
[0004] The technical solution adopted by the present invention is as follows: The rheumatoid activity assessment system based on medical information processing provided by the present invention includes a medical information collection module, a primary clustering module, a secondary clustering processing module, and a rheumatoid activity assessment module;
[0005] The medical information collection module collects the medical data of historical rheumatism patients; generates an initial medical data set through feature engineering processing;
[0006] The primary clustering module selects cluster centers by using the common feature similarity and the distance to high-density points, and optimizes the initial medical data set through density calculation and synthetic samples to obtain the final medical data set;
[0007] The secondary clustering processing module is based on the final medical data set, completes the secondary clustering processing by calculating the clinical neighborhood distance, constructing the clinical index coordinates, and iteratively updating and correcting the cluster centers;
[0008] The rheumatism activity assessment module performs cluster assignment on the medical data of real-time collected rheumatism patients based on the secondary clustering processing, and finally realizes the rheumatism activity assessment.
[0009] Further, in the medical information collection module, the medical data of the historical rheumatism patients includes clinical indicators, laboratory test data, and clinical status; the clinical status is used as a data label; the data label does not participate in the clustering operation; the clinical status includes remission status, low activity status, medium activity status, and high activity status, and the collected medical data of the historical rheumatism patients is subjected to feature engineering processing to obtain an initial medical data set.
[0010] Further, the primary clustering module performs primary clustering on the initial medical data set to further realize data optimization; specifically includes the following content:
[0011] Cluster center selection unit; Let GX(i,j) represent the set of common features of patient i and patient j; and define the similarity sim(·), which is expressed as:
[0012] ;
[0013] Among them, o is the feature index; and are the values of patient i and patient j on feature o respectively; define the distance to the nearest high-density point, which is expressed as:
[0014] ;
[0015] ;
[0016] Among them, is the distance from patient i to the nearest high-density cluster center; and are the densities of patient i and patient k respectively; and They are the distances between patient i and patient q, and between patient k and patient p respectively; M(·) is the neighborhood set of patients; the one with the largest decision value is selected as the clustering center, denoted as: ; where is the decision value;
[0017] Clustering density calculation unit; the density is denoted as:
[0018] ;
[0019] The sparsity is denoted as:
[0020] ;
[0021] Introduce the normalized label, and the sampling weight is denoted as:
[0022] ;
[0023] The number of synthetic samples is denoted as:
[0024] ;
[0025] where is the clustering density; is the clustering center, j is the clustering center index; K(·) is the Gaussian kernel function; is the value of the i-th sample in the j-th cluster; is all the samples in the j-th cluster; is the bandwidth parameter; is the sparsity of the j-th cluster; is the sampling weight of the j-th cluster; and are the sample generation ratios of the j-th cluster and the k-th cluster respectively; n is the total number of clusters; is the number of synthetic samples to be generated in the j-th cluster; Nm and Pm are the total number of final samples and the number of existing samples respectively; is the normalized label of the j-th clustering center;
[0026] Dataset optimization unit; generate synthetic samples between the clustering center and two randomly selected sample points within the cluster, and the weight is confirmed based on similarity; obtain the final medical dataset;
[0027] The synthetic sample is denoted as:
[0028] ;
[0029] ;
[0030] where is the synthetic sample; and are two samples randomly selected within the cluster; is the synthetic weight; is the overall weight.
[0031] Furthermore, the secondary clustering processing module performs secondary clustering on the optimized data of rheumatoid patients; specifically, it includes the following:
[0032] Clinical proximity calculation unit; initialize the cluster centers, randomly select H cluster centers, and calculate the clinical proximity between each pair of patients to measure the similarity between patients; expressed as:
[0033] ;
[0034] ;
[0035] where f(x) is the clinical proximity of patient x; C is the scale parameter; N is the total number of patient samples; is the reachable bandwidth; h is the bandwidth; is the Euclidean distance; is the bandwidth factor;
[0036] Clinical index coordinate construction unit; convert each patient into an index coordinate; select the remission state as the origin O; obtain the change in the rheumatoid activity of the patient, introduce the sigmoid mapping, and each index coordinate is expressed as:
[0037] ;
[0038] where, is the index coordinate of patient ; HJ is the improvement in rheumatoid activity; EH is the deterioration in rheumatoid activity; is the slope parameter; is the index mean;
[0039] Cluster center update unit; update the index coordinates of patients based on the centroid migration; introduce the additional weight to reduce the influence of outliers, and the cluster center update is expressed as:
[0040] ;
[0041] ;
[0042] where, is the updated cluster center; is the weight function;
[0043] ;
[0044] where, is the derivative of the Gaussian kernel function; is the cluster center to which patient i belongs;
[0045] Cluster center correction unit; During the transition of the rheumatoid activity level from low to high, a clinical bifurcation point will appear ; Correct the bifurcation point and the candidate centers around it; For the candidate cluster center clinical coordinates , according to its relationship with the bifurcation point and the sample points around it for correction, introducing a correction coefficient , adaptively correct the correction intensity based on the local density gradient, expressed as:
[0046] ;
[0047] ;
[0048] Among them, is the adjustment parameter; is at the local density gradient; and are the positions of the candidate cluster centers after and before correction respectively; is the set of all clinical branches at the bifurcation point; is the set branch; both p and q are branch indices; is at the bifurcation point on, the set of all clustering samples located on the branch ; is the position of the bifurcation point in the clinical index coordinate system; The candidate cluster center is the cluster center at the last iteration;
[0049] Cluster iteration unit; According to the corrected cluster center, reassign all patient samples, and assign the samples to the cluster center that is the closest in clinical proximity; Repeat the clustering process until the maximum number of iterations is reached or the clustering converges; Use the label corresponding to the cluster center as the cluster label.
[0050] Furthermore, the rheumatoid activity assessment module collects the medical data of rheumatoid patients in real time, performs clustering assignment based on the quadratic clustering processing module after feature engineering processing, and uses the corresponding cluster label after assignment as the rheumatoid activity assessment result.
[0051] The beneficial effects achieved by the present invention using the above solution are as follows:
[0052] (1)In view of the problems existing in the general rheumatoid activity assessment system, namely, for the scarce group of severe rheumatoid patients, their data is vulnerable to outliers and the sample is scarce, making it difficult to support the state division, and further leading to poor accuracy of rheumatoid activity assessment. This solution introduces the distance to the nearest high-density point to avoid noise interference; introduces normalized labels to supplement the minority samples of severe rheumatoid patients; realizes adaptive sampling based on density and sparsity indicators; determines weights based on similarity calculation to ensure that when generating synthetic samples, both the group center information is incorporated and the clinical similarity between samples is fully considered, avoiding the generation of noise samples at the extreme edges of the activity level; thus significantly improving the accuracy of subsequent rheumatoid activity assessment.
[0053] (2)In view of the problem that in the case of fewer samples and greater fluctuations of high-activity patients in the general rheumatoid activity assessment system, it is easily affected by outliers or noise data and cannot achieve a fine division of severe rheumatoid patients. This solution calculates the clinical proximity distance between each pair of patients to more accurately reflect the subtle differences in clinical status; updates the cluster center based on the clinical index coordinates and additional weights to ensure that the updated cluster center is more concentrated in the high-density area, truly reflecting the central tendency of rheumatoid activity; introduces a correction coefficient and adaptively adjusts the correction strength using the local density gradient to correct the cluster center, ensuring that each activity level has a representative cluster center; thus ensuring the fine division of subsequent rheumatoid activity assessment. Brief Description of the Drawings
[0054] Figure 1 It is a schematic flowchart of the rheumatoid activity assessment system based on medical information processing provided by the present invention;
[0055] Figure 2 It is a schematic flowchart of the secondary clustering processing module;
[0056] Figure 3 It is an effect diagram of k-means clustering;
[0057] Figure 4 It is an effect diagram of the clustering of this solution.
[0058] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention, but do not constitute a limitation to the present invention. Detailed Embodiments
[0059] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments; based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0060] In the description of the present invention, it should be understood that the terms "upper", "lower", "front", "rear", "left", "right", "top", "bottom", "inner", "outer", etc. indicating the orientation or positional relationship are based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the present invention.
[0061] Example 1, refer to Figure 1 , the rheumatism activity assessment system based on medical information processing provided by the present invention includes a medical information collection module, a primary clustering module, a secondary clustering processing module, and a rheumatism activity assessment module;
[0062] The medical information collection module collects the medical data of historical rheumatism patients; generates an initial medical data set through feature engineering processing; and sends the data to the primary clustering module;
[0063] The primary clustering module selects the clustering centers by using the common feature similarity and the distance to the high-density points, and optimizes the initial medical data set by density calculation and synthetic samples to obtain the final medical data set; and sends the data to the secondary clustering processing module;
[0064] The secondary clustering processing module is based on the final medical data set, completes the secondary clustering processing by calculating the clinical neighborhood distance, constructing the clinical index coordinates, and iteratively updating and correcting the clustering centers; and sends the data to the rheumatism activity assessment module;
[0065] The rheumatism activity assessment module performs clustering assignment on the medical data of real-time collected rheumatism patients based on the secondary clustering processing, and finally realizes the assessment of rheumatism activity.
[0066] Example 2, refer to Figure 1 , based on the above example, in the medical data collection module, the medical data of historical rheumatism patients includes clinical indicators, laboratory test data, and clinical status; the clinical status is used as the data label; the data label does not participate in the clustering operation; the clinical status includes remission status, low activity status, medium activity status, and high activity status; the remission status corresponds to normal or nearly normal; the clinical indicators include the number of swollen joints, the number of tender joints, the fever status, and the duration of morning stiffness; the laboratory test data includes white blood cell count, hemoglobin, platelet count, C-reactive protein, erythrocyte sedimentation rate, rheumatoid factor, and anti-cyclic citrullinated peptide antibody; the collected medical data of historical rheumatism patients is subjected to feature engineering processing to obtain the initial medical data set.
[0067] Example 3, refer to Figure 1, this embodiment is based on the above embodiment. The initial clustering module performs initial clustering on the initial medical data set to achieve data optimization. Specifically, it includes the following:
[0068] Clustering center selection unit; Let GX(i,j) represent the set of common features of patient i and patient j; and define the similarity sim(·), which is expressed as:
[0069] ;
[0070] where o is the feature index; and are the values of patient i and patient j on feature o respectively;
[0071] Define the distance to the nearest high-density point, which is expressed as:
[0072] ;
[0073] ;
[0074] where, is the distance from patient i to the nearest high-density clustering center; and are the densities of patient i and patient k respectively; and are the distances between patient i and patient q and between patient k and patient p respectively; M(·) is the neighborhood set of the patient; Select the one with the largest decision value as the clustering center, which is expressed as: ; where, is the decision value;
[0075] Clustering density calculation unit; Clusters with lower density represent rare patient categories of severe rheumatoid patients. Assign high sampling weights to sparse clusters to ensure the sample generation ratio of severe rheumatoid patients; Thus, more synthetic samples are generated to improve the deficiency of minority class samples; The density is expressed as:
[0076] ;
[0077] The sparsity is expressed as:
[0078] ;
[0079] Introduce a normalized label, assign a higher weight to the high-activity state, and the sampling weight is expressed as:
[0080] ;
[0081] The number of synthetic samples is expressed as:
[0082] ;
[0083] Among them, is the clustering density; is the clustering center, j is the clustering center index; K(·) is the Gaussian kernel function; is the value of the i-th sample in the j-th cluster; is all samples in the j-th cluster; is the bandwidth parameter; is the sparsity of the j-th cluster; is the sampling weight of the j-th cluster; and are the sample generation ratios of the j-th cluster and the k-th cluster respectively; n is the total number of clusters; is the number of synthetic samples to be generated for the j-th cluster; Nm and Pm are the total number of final samples and the number of existing samples respectively; is the normalized label of the j-th cluster center;
[0084] Dataset optimization unit; Generate synthetic samples between the clustering center and two randomly selected sample points within the cluster, and the weights are confirmed based on similarity to avoid generating noise samples in the extreme marginal areas of activity, ensuring that the generated samples can truly reflect the characteristics of rheumatic activity; Obtain the final medical dataset; The synthetic sample is expressed as: ;
[0085] ;
[0086] Among them, is the synthetic sample; and are two randomly selected samples within the cluster; is the synthetic weight; is the overall weight.
[0087] By performing the above operations, for the general rheumatic activity assessment system, there is a problem that for the scarce group of severe rheumatic patients, their data is vulnerable to outliers and sample scarcity, making it difficult to support the state division, which in turn leads to poor accuracy of rheumatic activity assessment. This solution introduces the distance to the nearest high-density point to avoid noise interference; introduces normalized labels to supplement the minority class samples of severe rheumatic patients; realizes adaptive sampling based on density and sparsity indicators; determines weights based on similarity calculation to ensure that when generating synthetic samples, both the group center information is incorporated and the clinical similarity between samples is fully considered, avoiding generating noise samples at the extreme margins of activity; thereby significantly improving the accuracy of subsequent rheumatic activity assessment.
[0088] Example 4, refer to Figure 1 and Figure 2, based on the above embodiment, the secondary clustering processing module performs secondary clustering on the optimized rheumatoid patient data, with the focus on correcting the clustering center deviation that may be caused by noise points, especially correcting the clustering center of rheumatoid patients with high activity; specifically, it includes the following content:
[0089] Clinical adjacent distance calculation unit; initialize the clustering center, randomly select H clustering centers, and calculate the clinical adjacent distance between each pair of patients to measure the similarity between patients; expressed as:
[0090] ;
[0091] ;
[0092] Among them, f(x) is the clinical adjacent distance of patient x; C is the scale parameter; N is the total number of patient samples; is the reachable bandwidth; h is the bandwidth; is the Euclidean distance; is the bandwidth factor;
[0093] Clinical index coordinate construction unit; by integrating key clinical indicators, convert each patient into an index coordinate to reflect the clinical status of the patient; select the remission state as the origin O; obtain the change in the rheumatoid activity of the patient, introduce the sigmoid mapping to reduce the influence of noise, and each index coordinate is expressed as:
[0094] ;
[0095] Among them, is patient 's index coordinate; HJ is the improvement of rheumatoid activity; EH is the deterioration of rheumatoid activity; is the slope parameter; is the index mean;
[0096] Clustering center update unit; update the index coordinates of patients based on the centroid migration to reflect the center of rheumatoid activity; introduce an additional weight , reduce the influence of outliers, and the clustering center update is expressed as:
[0097] ;
[0098] ;
[0099] Among them, is the updated clustering center; is the weight function;
[0100] ;
[0101] Among them, is the derivative of the Gaussian kernel function; is the cluster center to which patient i belongs;
[0102] Cluster center correction unit; During the transition from low to high rheumatoid activity, one or more clinical bifurcation points will appear , which is the point where the patient state distribution shows obvious branches or turns; It is necessary to correct these bifurcation points and the candidate centers around them so that there is only one representative center for each activity level in the end; For the candidate cluster center clinical coordinates , according to its relationship with the bifurcation point and the sample points around it, a correction is made, introducing a correction coefficient , and the correction intensity is adaptively corrected based on the local density gradient, expressed as:
[0103] ;
[0104] ;
[0105] Among them, is the adjustment parameter; is at local density gradient; and are the positions of the candidate cluster centers after and before correction respectively; is the set of all clinical branches at the bifurcation point; is the set branch; p and q are both branch indices; is at the bifurcation point , located on the branch set of all cluster samples; is the position of the bifurcation point in the clinical index coordinate system; The branch is the patient state development path. When the data is mapped to the coordinate system, the data points will form paths in different directions, reflecting different trends in the patient's disease progression; The candidate cluster center is the cluster center in the previous iteration;
[0106] Cluster iteration unit; According to the corrected cluster center, all patient samples are assigned again, and the samples are assigned to the cluster center that is the closest in clinical proximity; Repeat the clustering process until the maximum number of iterations is reached or the clustering converges; The label corresponding to the cluster center is used as the cluster label; If the clustering result is not good, parameter optimization is performed based on the particle swarm search algorithm.
[0107] By performing the above operations, in view of the problem that in the general rheumatoid activity assessment system, when the sample of patients with high activity is small and fluctuates greatly, it is easily affected by abnormal points or noise data and cannot achieve a fine classification of severe rheumatoid patients, this solution calculates the clinical proximity distance between each pair of patients to more accurately reflect the subtle differences in clinical status; updates the cluster centers based on the clinical index coordinates and additional weights to ensure that the updated cluster centers are more concentrated in the high-density areas and truly reflect the central tendency of rheumatoid activity; introduces a correction coefficient, adaptively adjusts the correction intensity using the local density gradient, and corrects the cluster centers to ensure that each activity level has a representative cluster center; and further ensures the fine classification of subsequent rheumatoid activity assessments.
[0108] Example 5, refer to Figure 1 , based on the above example, the rheumatoid activity assessment module collects the medical data of rheumatoid patients in real time. After being processed by feature engineering, it performs clustering assignment based on the secondary clustering processing module, and takes the corresponding cluster label after assignment as the rheumatoid activity assessment result.
[0109] Example 6, refer to Figure 3 and Figure 4 , Figure 3 is the effect diagram of k-means clustering; Figure 4 is the effect diagram of the clustering of this solution; the silhouette coefficient of the clustering result of this solution is 0.1105, and the silhouette coefficient of the kMeans clustering result is 0.068.
[0110] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device.
[0111] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention.
[0112] The above describes the present invention and its implementation manners. Such description is not restrictive. What is shown in the drawings is only one of the implementation manners of the present invention, and the actual structure is not limited thereto. In general, if those of ordinary skill in the art are inspired by it and, without departing from the gist of the present invention, design similar structural manners and embodiments to this technical solution without creative efforts, they shall fall within the protection scope of the present invention.
Claims
1. A rheumatic activity assessment system based on medical information processing, characterized in that: The system includes a medical information collection module, a primary clustering module, a secondary clustering processing module and a rheumatic activity assessment module; The medical information collection module collects historical medical data of rheumatism patients; The initial medical data set was generated after feature engineering; The initial clustering module selects cluster centers by using common feature similarity and distance to high-density points, and optimizes the initial medical data set by density calculation and synthetic sample generation to obtain the final medical data set; The secondary clustering processing module completes the secondary clustering processing based on the final medical data set by calculating the clinical proximity distance, constructing the clinical indicator coordinates, and iteratively updating and correcting the cluster center; The rheumatic activity assessment module performs clustering and allocation on the medical data of rheumatic patients collected in real time based on secondary clustering processing, and finally realizes the rheumatic activity assessment; In the medical information collection module, the historical rheumatic patient medical data includes clinical indicators, laboratory test data and clinical status; the clinical status is used as a data label; the data label does not participate in the clustering operation; the clinical status includes the remission state, low activity state, medium activity state and high activity state. The collected historical rheumatic patient medical data is subjected to feature engineering processing to obtain an initial medical data set; The initial clustering module performs initial clustering on the initial medical data set to achieve data optimization; specific Includes the following: Cluster center selection unit; let GX(i,j) represent the common feature set of patients i and j; and define the similarity sim(·), expressed as: ; Where o is the feature index; and are the values of patient i and patient j on feature o respectively; Define the distance to the nearest high-density point, expressed as: ; ; in, is the distance from patient i to the nearest high-density cluster center; and are the densities of patient i and patient k, respectively; and are the distances between patient i and patient q and between patient k and patient p respectively; M(·) is the patient neighborhood set; the cluster center is selected with the largest decision value, expressed as: ;in, is the decision value; Cluster density calculation unit; density is expressed as: ; Sparsity is expressed as: ; Introducing normalized labels, the sampling weights are expressed as: ; The number of synthetic samples is expressed as: ; in, is the cluster density; is the cluster center, j is the cluster center index; K(·) is the Gaussian kernel function; is the value of the i-th sample in the j-th cluster; are all samples in the jth cluster; is the bandwidth parameter; is the sparsity of the jth cluster; is the sampling weight of the jth cluster; and are the sample generation ratios of the jth cluster and the kth cluster respectively; n is the total number of clusters; is the number of synthetic samples that need to be generated for the jth cluster; Nm and Pm are the total number of final samples and the number of existing samples, respectively; is the normalized label of the jth cluster center; Dataset optimization unit.
2. The rheumatic activity assessment system based on medical information processing according to claim 1, characterized in that: The data set optimization unit generates a synthetic sample between the cluster center and two randomly selected sample points in the cluster, and the weight is determined based on the similarity; The final medical dataset is obtained; the synthetic sample is represented as: ; ; in, is a synthetic sample; and are two randomly selected samples within the cluster; is the composite weight; is the overall weight.
3. The rheumatic activity assessment system based on medical information processing according to claim 2, characterized in that: The secondary clustering processing module performs secondary clustering on the optimized rheumatism patient data; specifically includes the following contents: Clinical proximity distance calculation unit: initialize the cluster center, randomly select H cluster centers, calculate the clinical proximity distance between each pair of patients, and use it to measure the similarity between patients; expressed as: ; ; Where f(x) is the clinical proximity distance of patient x; C is the scale parameter; N is the total number of patient samples; is the achievable bandwidth; h is the bandwidth; is the Euclidean distance; is the bandwidth factor; Clinical indicator coordinate construction unit; convert each patient into an indicator coordinate; select the remission state as the origin O; obtain the change of patient rheumatic activity, introduce sigmoid mapping, and each indicator coordinate is expressed as: ; in, Is a patient The index coordinates; HJ is the improvement of rheumatic activity; EH is the deterioration of rheumatic activity; is the slope parameter; is the indicator mean; Cluster center update unit; update the patient's index coordinates based on center of gravity migration; introduce additional weights , reduce the influence of outliers, and the cluster center update is expressed as: ; ; in, is the updated cluster center; is the weight function; ; in, is the derivative of the Gaussian kernel function; is the cluster center to which patient i belongs; Cluster center correction unit; Clustering iteration unit; all patient samples are allocated again according to the corrected cluster centers, and the samples are allocated to the cluster centers closest to the clinical proximity distance; the clustering process is repeated until the maximum number of iterations is reached or the clustering converges; the label corresponding to the cluster center is used as the cluster label.
4. The rheumatic activity assessment system based on medical information processing according to claim 3, characterized in that: The cluster center correction unit; in the transition process from low to high rheumatic activity, a clinical bifurcation point will appear ; Correct the candidate centers around the bifurcation point; For the candidate cluster centers Clinical coordinates , according to its bifurcation point The relationship between the sample points around it is corrected by introducing the correction coefficient , based on the local density gradient adaptive correction correction strength, expressed as: ; ; in, is the adjustment parameter; is The local density gradient at ; and are the positions of candidate cluster centers after and before correction, respectively; is the set of all clinical branches at the bifurcation point; is a set branch; p and q are both branch indices; It is at the fork On the branch The set of all cluster samples; is the position of the bifurcation point in the clinical index coordinate system; the candidate cluster center is the cluster center at the last iteration.
5. The rheumatic activity assessment system based on medical information processing according to claim 4, characterized in that: The rheumatic activity assessment module collects medical data of rheumatic patients in real time, performs clustering allocation based on a secondary clustering processing module after feature engineering processing, and uses the corresponding cluster labels after allocation as rheumatic activity assessment results.
Citation Information
Patent Citations
Traditional Chinese medicine curative effect evaluation system based on artificial intelligence
CN119092056A