Rheumatism activity assessment system based on medical information processing

By introducing high-density point distance, normalized label, adaptive sampling and similarity weight calculation, combined with clinical proximity distance and cluster center correction, the problem of poor accuracy of the rheumatism activity assessment system in the evaluation of severe rheumatism patients is solved, and the fine division and evaluation accuracy of severe rheumatism patients is achieved is improved.

CN119993544AActive Publication Date: 2025-05-13QINGDAO UNIV
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202510459352.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-05-13
Estimated Expiration
2045-04-14

AI Technical Summary

Technical Problem

The existing rheumatism activity assessment system is susceptible to outliers and scarce samples when dealing with severe rheumatism patients, resulting in poor assessment accuracy. Especially when patients with high activity have fewer samples and large fluctuations, it is difficult to achieve fine division of patients with severe rheumatism.

Method used

By introducing the distance to the nearest high-density point, noise interference is avoided; normalized labels are introduced to supplement a few samples of severe rheumatism patients; adaptive sampling is achieved based on density and sparseness indicators; weights are determined based on similarity calculations to ensure that the generation of synthetic samples is not only integrated with population-centered information, but also fully considering the clinical similarity between samples. At the same time, by calculating clinical proximity distances, constructing clinical index coordinates, iterative updates and correcting cluster centers, ensuring that each level of activity has a representative cluster center.

Benefits of technology

It significantly improves the accuracy of rheumatoid activity assessment, ensures fine division of patients with severe rheumatism, reduces the impact of abnormal points or noise data, and improves the stability and reliability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119993544A_ABST
    Figure CN119993544A_ABST
Patent Text Reader

Abstract

The invention discloses a rheumatism activity assessment system based on medical information processing. The rheumatism activity assessment system comprises a medical information acquisition module, a primary clustering module, a secondary clustering processing module and a rheumatism activity assessment module. The invention belongs to the field of data processing, and particularly relates to a rheumatism activity assessment system based on medical information processing, which introduces a distance to a nearest high-density point to avoid noise interference; normalized labels are introduced, and minority class samples of the severe rheumatism patients are supplemented; adaptive sampling is realized based on density and sparsity indexes; by calculating the clinical proximity distance between each pair of patients, the clinical state subtle difference can be reflected more accurately; a correction coefficient is introduced, a local density gradient is utilized to adaptively adjust correction strength, clustering center correction is carried out, and it is ensured that each activity level has a representative clustering center; and thus, fine division of subsequent rheumatism activity assessment is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing, and in particular to a rheumatic activity assessment system based on medical information processing. Background Art

[0002] The goal of the rheumatic activity assessment system is to accurately reflect the activity status of a patient's rheumatic disease by integrating multiple clinical indicators in an objective and quantitative manner, thereby guiding clinical treatment and monitoring the progression of the disease. However, the general rheumatic activity assessment system has the problem that there is a scarce group of severe rheumatic patients, their data is easily interfered by outliers and samples are scarce, making it difficult to support the classification of states, which in turn leads to poor accuracy in rheumatic activity assessment; the general rheumatic activity assessment system has the problem that there are few samples of high-activity patients and the fluctuations are large, and it is easily affected by outliers or noise data, making it impossible to achieve a fine classification of severe rheumatic patients. Summary of the invention

[0003] In view of the above situation, in order to overcome the defects of the prior art, the present invention provides a rheumatic activity assessment system based on medical information processing. In view of the fact that general rheumatic activity assessment systems have a scarce group of severe rheumatic patients, whose data are easily interfered by outliers and have scarce samples, making it difficult to support state division, which in turn leads to poor accuracy of rheumatic activity assessment, this scheme introduces the distance to the nearest high-density point to avoid noise interference; introduces normalized labels to supplement minority samples of severe rheumatic patients; realizes adaptive sampling based on density and sparsity indicators; determines weights based on similarity calculations to ensure that the generation of synthetic samples not only incorporates the group center information, but also fully considers the clinical similarities between samples, thereby avoiding the generation of noise samples at the extreme edges of activity; and thus significantly improves Improve the accuracy of subsequent rheumatic activity assessment; in view of the fact that the general rheumatic activity assessment system has a small number of high-activity patient samples and large fluctuations, it is easily affected by abnormal points or noise data and cannot achieve a fine division of severe rheumatic patients. This solution calculates the clinical proximity distance between each pair of patients to more accurately reflect the subtle differences in clinical status; cluster centers are updated based on clinical indicator coordinates and additional weights to ensure that the updated cluster centers are more concentrated in high-density areas and truly reflect the central trend of rheumatic activity; a correction coefficient is introduced, and the correction strength is adaptively adjusted using the local density gradient to perform cluster center correction to ensure that each activity level has a representative cluster center; thereby ensuring the fine division of subsequent rheumatic activity assessment.

[0004] The technical solution adopted by the present invention is as follows: the rheumatic activity assessment system based on medical information processing provided by the present invention comprises a medical information acquisition module, a primary clustering module, a secondary clustering processing module and a rheumatic activity assessment module;

[0005] The medical information collection module collects historical rheumatism patient medical data; generates an initial medical data set through feature engineering processing;

[0006] The initial clustering module selects cluster centers by using common feature similarity and distance to high-density points, and optimizes the initial medical data set by density calculation and synthetic sample generation to obtain the final medical data set;

[0007] The secondary clustering processing module completes the secondary clustering processing based on the final medical data set by calculating the clinical proximity distance, constructing the clinical indicator coordinates, and iteratively updating and correcting the cluster center;

[0008] The rheumatic activity assessment module clusters and distributes the medical data of rheumatic patients collected in real time based on secondary clustering processing, and finally realizes the rheumatic activity assessment.

[0009] Furthermore, in the medical information collection module, the historical medical data of rheumatic patients include clinical indicators, laboratory test data and clinical status; the clinical status is used as a data label; the data label does not participate in clustering operations; the clinical status includes remission state, low activity state, medium activity state and high activity state. The collected historical medical data of rheumatic patients are subjected to feature engineering processing to obtain an initial medical data set.

[0010] Furthermore, the initial clustering module performs initial clustering on the initial medical data set to achieve data optimization; specifically, it includes the following contents:

[0011] Cluster center selection unit; let GX(i,j) represent the common feature set of patients i and j; and define the similarity sim(·), expressed as: ; Where o is the feature index; and are the values ​​of patient i and patient j on feature o respectively; the distance to the nearest high-density point is defined as: ; ; in, is the distance from patient i to the nearest high-density cluster center; and are the densities of patient i and patient k, respectively; and are the distances between patient i and patient q and between patient k and patient p respectively; M(·) is the patient neighborhood set; the cluster center is selected with the largest decision value, expressed as: ;in, is the decision value;

[0012] Cluster density calculation unit; density is expressed as: ; Sparsity is expressed as: ; Introducing normalized labels, the sampling weights are expressed as: ; The number of synthetic samples is expressed as: ; in, is the cluster density; is the cluster center, j is the cluster center index; K(·) is the Gaussian kernel function; is the value of the i-th sample in the j-th cluster; are all samples in the jth cluster; is the bandwidth parameter; is the sparsity of the jth cluster; is the sampling weight of the jth cluster; and are the sample generation ratios of the jth cluster and the kth cluster respectively; n is the total number of clusters; is the number of synthetic samples that need to be generated for the jth cluster; Nm and Pm are the total number of final samples and the number of existing samples, respectively; is the normalized label of the jth cluster center;

[0013] Dataset optimization unit; generating synthetic samples between the cluster center and two randomly selected sample points in the cluster, with weights confirmed based on similarity; obtaining the final medical dataset; The synthetic sample is represented as: ; ; in, is a synthetic sample; and are two randomly selected samples within the cluster; is the composite weight; is the overall weight.

[0014] Furthermore, the secondary clustering processing module performs secondary clustering on the optimized rheumatism patient data; specifically, it includes the following contents:

[0015] Clinical proximity distance calculation unit: initialize the cluster center, randomly select H cluster centers, calculate the clinical proximity distance between each pair of patients, and use it to measure the similarity between patients; expressed as: ; ; Where f(x) is the clinical proximity distance of patient x; C is the scale parameter; N is the total number of patient samples; is the achievable bandwidth; h is the bandwidth; is the Euclidean distance; is the bandwidth factor;

[0016] Clinical indicator coordinate construction unit; convert each patient into an indicator coordinate; select the remission state as the origin O; obtain the change of patient rheumatic activity, introduce sigmoid mapping, and each indicator coordinate is expressed as: ; in, Is a patient The index coordinates; HJ is the improvement of rheumatic activity; EH is the deterioration of rheumatic activity; is the slope parameter; is the indicator mean;

[0017] Cluster center update unit; update the patient's index coordinates based on center of gravity migration; introduce additional weights , reduce the influence of outliers, and the cluster center update is expressed as: ; ; in, is the updated cluster center; is the weight function; ; in, is the derivative of the Gaussian kernel function; is the cluster center to which patient i belongs;

[0018] Cluster center correction unit; clinical bifurcation point occurs during the transition from low to high rheumatic activity ; Correct the candidate centers around the bifurcation point; For the candidate cluster centers Clinical coordinates , according to its bifurcation point The relationship between the sample points around it is corrected by introducing the correction coefficient , based on the local density gradient adaptive correction correction strength, expressed as: ; ; in, is the adjustment parameter; is The local density gradient at ; and are the positions of candidate cluster centers after and before correction, respectively; is the set of all clinical branches at the bifurcation point; is a set branch; p and q are both branch indices; It is at the fork On the branch The set of all cluster samples; is the position of the bifurcation point in the clinical index coordinate system; the candidate cluster center is the cluster center at the last iteration;

[0019] Clustering iteration unit; all patient samples are allocated again according to the corrected cluster centers, and the samples are allocated to the cluster centers closest to the clinical proximity distance; the clustering process is repeated until the maximum number of iterations is reached or the clustering converges; the label corresponding to the cluster center is used as the cluster label.

[0020] Furthermore, the rheumatic activity assessment module collects medical data of rheumatic patients in real time, performs clustering allocation based on a secondary clustering processing module after feature engineering processing, and uses the corresponding cluster labels after allocation as rheumatic activity assessment results.

[0021] The beneficial effects achieved by the present invention using the above scheme are as follows:

[0022] (1) In view of the fact that the general rheumatic activity assessment system has a scarce group of severe rheumatic patients, whose data are easily disturbed by outliers and have scarce samples, making it difficult to support state classification, which in turn leads to poor accuracy in rheumatic activity assessment, this scheme introduces the distance to the nearest high-density point to avoid noise interference; introduces normalized labels to supplement minority samples of severe rheumatic patients; realizes adaptive sampling based on density and sparsity indicators; determines weights based on similarity calculation to ensure that the generation of synthetic samples not only incorporates the group center information, but also fully considers the clinical similarity between samples, avoiding the generation of noise samples at the extreme edge of activity; thereby significantly improving the accuracy of subsequent rheumatic activity assessment.

[0023] (2) In view of the problem that the general rheumatic activity assessment system has a small number of high-activity patient samples and large fluctuations, it is easily affected by abnormal points or noise data and cannot achieve a fine division of severe rheumatic patients. This scheme calculates the clinical proximity distance between each pair of patients to more accurately reflect the subtle differences in clinical status; updates the cluster centers based on the clinical indicator coordinates and additional weights to ensure that the updated cluster centers are more concentrated in high-density areas and truly reflect the central trend of rheumatic activity; introduces a correction coefficient and uses the local density gradient to adaptively adjust the correction strength to perform cluster center correction to ensure that each activity level has a representative cluster center; thereby ensuring the fine division of subsequent rheumatic activity assessment. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 A schematic diagram of the flow chart of the rheumatic activity assessment system based on medical information processing provided by the present invention;

[0025] Figure 2 It is a flowchart of the secondary clustering processing module;

[0026] Figure 3 This is the k-means clustering effect diagram;

[0027] Figure 4 This is the clustering effect diagram of this scheme.

[0028] The accompanying drawings are used to provide further understanding of the present invention and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention and do not constitute a limitation of the present invention. DETAILED DESCRIPTION

[0029] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments; based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0030] In the description of the present invention, it should be understood that terms such as “upper”, “lower”, “front”, “back”, “left”, “right”, “top”, “bottom”, “inside” and “outside” indicating directions or positional relationships are based on the directions or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific direction, be constructed and operated in a specific direction, and therefore should not be understood as limiting the present invention.

[0031] Example 1, see Figure 1The rheumatic activity assessment system based on medical information processing provided by the present invention includes a medical information acquisition module, a primary clustering module, a secondary clustering processing module and a rheumatic activity assessment module;

[0032] The medical information collection module collects historical rheumatism patient medical data; generates an initial medical data set through feature engineering processing; and sends the data to the initial clustering module;

[0033] The primary clustering module selects cluster centers by using common feature similarity and distance to high-density points, and optimizes the initial medical data set by density calculation and synthetic sample generation to obtain the final medical data set; and sends the data to the secondary clustering processing module;

[0034] The secondary clustering processing module completes the secondary clustering processing based on the final medical data set by calculating the clinical proximity distance, constructing the clinical indicator coordinates, and iteratively updating and correcting the cluster center; and sends the data to the rheumatic activity assessment module;

[0035] The rheumatic activity assessment module clusters and distributes the medical data of rheumatic patients collected in real time based on secondary clustering processing, and finally realizes the rheumatic activity assessment.

[0036] Example 2, see Figure 1 This embodiment is based on the above embodiment. In the medical data collection module, the historical medical data of rheumatism patients include clinical indicators, laboratory test data and clinical status; the clinical status is used as a data label; the data label does not participate in the clustering operation; the clinical status includes remission state, low activity state, medium activity state and high activity state; the remission state corresponds to normal or near normal; the clinical indicators include swollen joint count, tender joint count, fever state and morning stiffness duration; the laboratory test data include white blood cell count, hemoglobin, platelet count, C-reactive protein, erythrocyte sedimentation rate, rheumatoid factor and anti-cyclic citrullinated peptide antibody; the collected historical medical data of rheumatism patients are subjected to feature engineering processing to obtain an initial medical data set.

[0037] Example 3, see Figure 1 This embodiment is based on the above embodiment. The initial clustering module performs initial clustering on the initial medical data set to achieve data optimization. Specifically, it includes the following contents:

[0038] Cluster center selection unit; let GX(i,j) represent the common feature set of patients i and j; and define the similarity sim(·), expressed as: ; Where o is the feature index; and are the values ​​of patient i and patient j on feature o respectively; Define the distance to the nearest high-density point, expressed as: ; ; in, is the distance from patient i to the nearest high-density cluster center; and are the densities of patient i and patient k, respectively; and are the distances between patient i and patient q and between patient k and patient p respectively; M(·) is the patient neighborhood set; the cluster center is selected with the largest decision value, expressed as: ;in, is the decision value;

[0039] Cluster density calculation unit; clusters with lower density represent scarce patient categories of severe rheumatic patients, and high sampling weights are given to sparse clusters to ensure the sample generation ratio of severe rheumatic patients; thereby generating more synthetic samples to improve the shortage of minority class samples; the density is expressed as: ; Sparsity is expressed as: ; Normalized labels are introduced to give higher weights to high activity states, and the sampling weight is expressed as: ; The number of synthetic samples is expressed as: ; in, is the cluster density; is the cluster center, j is the cluster center index; K(·) is the Gaussian kernel function; is the value of the i-th sample in the j-th cluster; are all samples in the jth cluster; is the bandwidth parameter; is the sparsity of the jth cluster; is the sampling weight of the jth cluster; and are the sample generation ratios of the jth cluster and the kth cluster respectively; n is the total number of clusters; is the number of synthetic samples that need to be generated for the jth cluster; Nm and Pm are the total number of final samples and the number of existing samples, respectively; is the normalized label of the jth cluster center;

[0040] Dataset optimization unit; Generate synthetic samples between the cluster center and two randomly selected sample points in the cluster, and confirm the weight based on similarity to avoid generating noise samples in the edge area with extreme activity, so as to ensure that the generated samples can truly reflect the characteristics of rheumatic activity; Get the final medical data set; The synthetic sample is expressed as: ; ; in, is a synthetic sample; and are two randomly selected samples within the cluster; is the composite weight; is the overall weight.

[0041] By performing the above operations, in view of the fact that the general rheumatic activity assessment system has a scarce group of severe rheumatic patients, whose data are easily disturbed by outliers and have scarce samples, making it difficult to support state division, which in turn leads to poor accuracy of rheumatic activity assessment, this scheme introduces the distance to the nearest high-density point to avoid noise interference; introduces normalized labels to supplement minority samples of severe rheumatic patients; realizes adaptive sampling based on density and sparsity indicators; determines weights based on similarity calculation to ensure that the generation of synthetic samples not only incorporates the group center information, but also fully considers the clinical similarity between samples, avoiding the generation of noise samples at the extreme edge of activity; thereby significantly improving the accuracy of subsequent rheumatic activity assessment.

[0042] Example 4, see Figure 1 and Figure 2 This embodiment is based on the above embodiment. The secondary clustering processing module performs secondary clustering on the optimized rheumatic patient data, focusing on correcting the cluster center offset that may be caused by noise points, especially correcting the cluster center of highly active rheumatic patients; specifically, it includes the following contents:

[0043] Clinical proximity distance calculation unit: initialize the cluster center, randomly select H cluster centers, calculate the clinical proximity distance between each pair of patients, and use it to measure the similarity between patients; expressed as: ; ; Where f(x) is the clinical proximity distance of patient x; C is the scale parameter; N is the total number of patient samples; is the achievable bandwidth; h is the bandwidth; is the Euclidean distance; is the bandwidth factor;

[0044] Clinical indicator coordinate construction unit; by integrating key clinical indicators, each patient is converted into an indicator coordinate to reflect the patient's clinical status; the remission state is selected as the origin O; the change of patient rheumatic activity is obtained, and sigmoid mapping is introduced to reduce the impact of noise. Each indicator coordinate is expressed as: ; in, Is a patient The index coordinates; HJ is the improvement of rheumatic activity; EH is the deterioration of rheumatic activity; is the slope parameter; is the indicator mean;

[0045] Cluster center update unit; update the patient's index coordinates based on the center of gravity migration to reflect the center of rheumatic activity; introduce additional weights , reduce the influence of outliers, and the cluster center update is expressed as: ; ; in, is the updated cluster center; is the weight function; ; in, is the derivative of the Gaussian kernel function; is the cluster center to which patient i belongs;

[0046] Cluster center correction unit; one or more clinical bifurcation points occur during the transition from low to high rheumatic activity , which is the point where the distribution of patient status has obvious branches or turning points; these bifurcation points and the candidate centers around them need to be corrected so that there is only one representative center for each activity level; for the candidate cluster centers Clinical coordinates , according to its bifurcation point The relationship between the sample points around it is corrected by introducing the correction coefficient , based on the local density gradient adaptive correction correction strength, expressed as: ; ; in, is the adjustment parameter; is The local density gradient at ; and are the positions of candidate cluster centers after and before correction, respectively; is the set of all clinical branches at the bifurcation point; is a set branch; p and q are both branch indices; At the bifurcation point On the branch The set of all cluster samples; is the position of the bifurcation point in the clinical indicator coordinate system; the branch is the patient's status development path. When the data is mapped to the coordinate system, the data points will form paths along different directions, reflecting that the patient presents different trends in the progression of the disease; the candidate cluster center is the cluster center at the last iteration;

[0047] Clustering iteration unit; all patient samples are redistributed according to the corrected cluster centers, and the samples are assigned to the cluster centers closest to the clinical proximity distance; the clustering process is repeated until the maximum number of iterations is reached or the clustering converges; the labels corresponding to the cluster centers are used as cluster labels; if the clustering results are not good, parameter optimization is performed based on the particle swarm search algorithm.

[0048] By performing the above operations, in order to solve the problem that the general rheumatic activity assessment system has a small number of high-activity patient samples and large fluctuations, it is easily affected by abnormal points or noise data and cannot achieve a fine division of severe rheumatic patients. This solution calculates the clinical proximity distance between each pair of patients to more accurately reflect the subtle differences in clinical status; updates the cluster centers based on the clinical indicator coordinates and additional weights to ensure that the updated cluster centers are more concentrated in high-density areas and truly reflect the central trend of rheumatic activity; introduces a correction coefficient, uses the local density gradient to adaptively adjust the correction intensity, and performs cluster center correction to ensure that each activity level has a representative cluster center; thereby ensuring the fine division of subsequent rheumatic activity assessments.

[0049] Example 5, see Figure 1 This embodiment is based on the above embodiment. The rheumatic activity assessment module collects medical data of rheumatic patients in real time, performs clustering allocation based on the secondary clustering processing module after feature engineering processing, and uses the corresponding cluster label after allocation as the rheumatic activity assessment result.

[0050] Example 6, see Figure 3 and Figure 4 , Figure 3 This is the k-means clustering effect diagram; Figure 4 This is the clustering effect diagram of this scheme; the silhouette coefficient of the clustering result of this scheme is 0.1105, and the silhouette coefficient of the kMeans clustering result is 0.068.

[0051] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device.

[0052] While the embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that many changes, modifications, substitutions and variations can be made to the embodiments without departing from the principles and spirit of the invention.

[0053] The present invention and its embodiments are described above, and such description is not restrictive. The drawings show only one embodiment of the present invention, and the actual structure is not limited thereto. In short, if ordinary technicians in the field are inspired by it, without departing from the purpose of the invention, they can design a structure and embodiment similar to the technical solution without creativity, which should belong to the protection scope of the present invention.

Claims

1. A rheumatic activity assessment system based on medical information processing, characterized in that: The system includes a medical information collection module, a primary clustering module, a secondary clustering processing module and a rheumatic activity assessment module; The medical information collection module collects historical medical data of rheumatism patients; The initial medical data set was generated after feature engineering; The initial clustering module selects cluster centers by using common feature similarity and distance to high-density points, and optimizes the initial medical data set by density calculation and synthetic sample generation to obtain the final medical data set; The secondary clustering processing module completes the secondary clustering processing based on the final medical data set by calculating the clinical proximity distance, constructing the clinical indicator coordinates, and iteratively updating and correcting the cluster center; The rheumatic activity assessment module clusters and distributes the medical data of rheumatic patients collected in real time based on secondary clustering processing, and finally realizes the rheumatic activity assessment.

2. The rheumatic activity assessment system based on medical information processing according to claim 1, characterized in that: The initial clustering module performs initial clustering on the initial medical data set to achieve data optimization; specific Includes the following: Cluster center selection unit; let GX(i,j) represent the common feature set of patients i and j; and define the similarity sim(·), expressed as: ; Where o is the feature index; and are the values ​​of patient i and patient j on feature o respectively; Define the distance to the nearest high-density point, expressed as: ; ; in, is the distance from patient i to the nearest high-density cluster center; and are the densities of patient i and patient k, respectively; and are the distances between patient i and patient q and between patient k and patient p respectively; M(·) is the patient neighborhood set; the cluster center is selected with the largest decision value, expressed as: ;in, is the decision value; Cluster density calculation unit; density is expressed as: ; Sparsity is expressed as: ; Introducing normalized labels, the sampling weights are expressed as: ; The number of synthetic samples is expressed as: ; in, is the cluster density; is the cluster center, j is the cluster center index; K(·) is the Gaussian kernel function; is the value of the i-th sample in the j-th cluster; are all samples in the jth cluster; is the bandwidth parameter; is the sparsity of the jth cluster; is the sampling weight of the jth cluster; and are the sample generation ratios of the jth cluster and the kth cluster respectively; n is the total number of clusters; is the number of synthetic samples that need to be generated for the jth cluster; Nm and Pm are the total number of final samples and the number of existing samples, respectively; is the normalized label of the jth cluster center; Dataset optimization unit.

3. The rheumatic activity assessment system based on medical information processing according to claim 2, characterized in that: The data set optimization unit generates a synthetic sample between the cluster center and two randomly selected sample points in the cluster, and the weight is determined based on the similarity; The final medical dataset is obtained; the synthetic sample is represented as: ; ; in, is a synthetic sample; and are two randomly selected samples within the cluster; is the composite weight; is the overall weight.

4. The rheumatic activity assessment system based on medical information processing according to claim 3, characterized in that: The secondary clustering processing module performs secondary clustering on the optimized rheumatism patient data; specifically includes the following contents: Clinical proximity distance calculation unit: initialize the cluster center, randomly select H cluster centers, calculate the clinical proximity distance between each pair of patients, and use it to measure the similarity between patients; expressed as: ; ; Where f(x) is the clinical proximity distance of patient x; C is the scale parameter; N is the total number of patient samples; is the achievable bandwidth; h is the bandwidth; is the Euclidean distance; is the bandwidth factor; Clinical indicator coordinate construction unit; convert each patient into an indicator coordinate; select the remission state as the origin O; obtain the change of patient rheumatic activity, introduce sigmoid mapping, and each indicator coordinate is expressed as: ; in, Is a patient The index coordinates; HJ is the improvement of rheumatic activity; EH is the deterioration of rheumatic activity; is the slope parameter; is the indicator mean; Cluster center update unit; update the patient's index coordinates based on center of gravity migration; introduce additional weights , reduce the influence of outliers, and the cluster center update is expressed as: ; ; in, is the updated cluster center; is the weight function; ; in, is the derivative of the Gaussian kernel function; is the cluster center to which patient i belongs; Cluster center correction unit; Clustering iteration unit; all patient samples are allocated again according to the corrected cluster centers, and the samples are allocated to the cluster centers closest to the clinical proximity distance; the clustering process is repeated until the maximum number of iterations is reached or the clustering converges; the label corresponding to the cluster center is used as the cluster label.

5. The rheumatic activity assessment system based on medical information processing according to claim 4, characterized in that: The cluster center correction unit; in the transition process from low to high rheumatic activity, a clinical bifurcation point will appear ; Correct the candidate centers around the bifurcation point; For the candidate cluster centers Clinical coordinates , according to its bifurcation point The relationship between the sample points around it is corrected by introducing the correction coefficient , based on the local density gradient adaptive correction correction strength, expressed as: ; ; in, is the adjustment parameter; is The local density gradient at ; and are the positions of candidate cluster centers after and before correction, respectively; is the set of all clinical branches at the bifurcation point; is a set branch; p and q are both branch indices; It is at the fork On the branch The set of all cluster samples; is the position of the bifurcation point in the clinical index coordinate system; the candidate cluster center is the cluster center at the last iteration.

6. The rheumatic activity assessment system based on medical information processing according to claim 5, characterized in that: In the medical information collection module, the historical medical data of rheumatic patients include clinical indicators, laboratory test data and clinical status; the clinical status is used as a data label; the data label does not participate in clustering operations; the clinical status includes remission state, low activity state, medium activity state and high activity state. The collected historical medical data of rheumatic patients are subjected to feature engineering processing to obtain an initial medical data set.

7. The rheumatic activity assessment system based on medical information processing according to claim 6, characterized in that: The rheumatic activity assessment module collects medical data of rheumatic patients in real time, performs clustering allocation based on a secondary clustering processing module after feature engineering processing, and uses the corresponding cluster labels after allocation as rheumatic activity assessment results.

Citation Information

Patent Citations

  • Rheumatism immune disease data preprocessing method and system based on CL-FCM

    CN117315217A

  • Knowledge graph-driven medical large model diagnosis method

    CN118280562A

  • Multi-person attitude estimation method based on millimeter wave radar

    CN118865446A

  • Traditional Chinese medicine curative effect evaluation system based on artificial intelligence

    CN119092056A

  • Emergency treatment monitoring management system based on artificial intelligence

    CN119092143A