Patient evaluation classification method based on improved LDA topic model

Through improved LDA topic model and multi-step data processing methods, the difficulties of the prior art in extracting key information from patient evaluation are solved, and more accurate and reliable information extraction is achieved, supporting better public health management and decision-making.

CN120197615APending Publication Date: 2025-06-24CHENGZHEN MEDICAL TECHNOLOGY (WUXI) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311789883.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-22
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently extract key information from a large number of patient evaluations, especially when dealing with specific areas such as medical service evaluations, and faces issues of term processing, emotional tendency reflection and data bias overcome.

Method used

The improved LDA thesis model is adopted to extract and classify key information in patient evaluation through data preprocessing, multi-algorithm clustering, high confidence data selection and model training. Specific steps include data collection and preprocessing, multi-algorithm clustering, high confidence data selection, improved LDA topic model training and patient evaluation classification.

Benefits of technology

It achieves more efficient extraction of key information from patient evaluation, overcomes the limitations of the existing technology in handling medical service evaluation, improves the accuracy and credibility of information, and supports better public health management and decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120197615A_ABST
    Figure CN120197615A_ABST
Patent Text Reader

Abstract

The invention relates to a patient evaluation classification method based on an improved LDA topic model, and aims to efficiently and accurately classify patient written evaluations served by general practitioners. Firstly, patient evaluation is processed through a series of preprocessing steps including word segmentation, abbreviation conversion, stop word removal, number and special character removal and low-frequency and high-frequency vocabulary removal. Thirdly, analyzing the data by applying various unsupervised clustering algorithms to generate a plurality of clustering results, and calculating ambiguity so as to construct a high-confidence data set; and training an improved LDA topic model by using the data set, including calculating an inter-class scatter matrix and an intra-class scatter matrix, updating a projection matrix, re-clustering the data set by using the projection matrix, and determining 60 class center topics. And finally, determining a text topic category for a given user evaluation text. The innovation of the invention lies in that multiple clustering algorithms and high-confidence data selection are combined, the accuracy of topic model training is improved, and through 60 finely divided topic categories, the patient evaluation classification is more accurate, and the improvement of the medical service quality is facilitated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field:

[0001] The present invention relates to the analysis of medical user evaluations, and specifically to a patient evaluation classification method based on an improved Latent Dirichlet Allocation (LDA) topic model. Background Art:

[0002] With the development of Internet technology and the popularity of social media platforms, it has become increasingly common for patients to leave evaluations of medical services on online platforms. These evaluations contain rich information that can reflect the quality of medical services, patient satisfaction, and potential problems. However, due to the large number and diverse formats of these evaluations, traditional manual analysis methods are unable to efficiently process and extract the key information from them.

[0003] Currently, machine learning, especially topic models such as Latent Dirichlet Allocation (LDA), has been used for text analysis and information extraction. LDA can identify the latent topic structure from a large amount of text, helping to reveal the potential patterns in the text data. However, the standard LDA model faces specific challenges when dealing with specific domains, such as medical service evaluations. These challenges include how to effectively handle domain-specific terms, how to accurately reflect the sentiment tendency in the evaluations, and how to overcome the inherent biases in the data.

[0004] In addition, there are also limitations in applying the information extracted from text data to public health management and decision-making in the prior art. Although existing research has shown that using patient evaluations can improve the quality of medical services, there is still room for improvement in converting this information into practical operations and strategies in the existing methods. Especially in terms of how to handle the credibility and biases of large-scale anonymous online evaluations, the prior art has not provided an effective solution.

[0005] Therefore, an improved method is needed that can not only more effectively extract key information from a large number of patient evaluations, but also overcome the limitations of the prior art, and thus better support public health management and decision-making. Summary of the Invention:

[0006] In view of the deficiencies of the prior art, the present invention proposes a patient evaluation classification method based on an improved LDA topic model, which specifically includes the following steps to complete patient evaluation classification: S1: Data collection and preprocessing: Collect written evaluations of patients served by general practitioners and perform a series of preprocessing steps, including tokenization, lowercasing, removing stop words, removing numbers and special characters, and removing low-frequency and high-frequency words; S2: Multi-algorithm clustering: Use multiple unsupervised clustering algorithms to cluster the preprocessed data set, generate multiple different clustering results, and calculate a fuzziness for each data sample, which reflects the consistency between the various clustering results; S3: High-confidence data selection: According to the fuzziness calculated in S2, select the data samples with the lowest fuzziness from the preprocessed data set to form a high-confidence data set; S4: Improved LDA topic model training: Use the high-confidence data set to train the improved LDA topic model, including: S4-1: Calculate the between-class scatter matrix S b and the within-class scatter matrix S w ; S4-2: Update the projection matrix W by solving an optimization problem to maximize the function J(W), where the calculation formula of J(W) is: S4-3: Use the updated projection matrix W to recluster the high-confidence data set to update the clustering results, and determine the central topics of 60 categories accordingly; S5: Patient evaluation classification: For any given user evaluation text, use the improved LDA topic model trained in step S4 to perform the category judgment.

[0007] As a preferred technical solution of the present invention, the series of preprocessing steps in S1 are specifically:

[0008] S1-1: Tokenization: Use jieba to perform tokenization processing to split continuous written evaluations of patients into independent words;

[0009] S1-2: Lowercasing: Use the zhconv library to perform traditional-simplified conversion on independent words, and convert the existing traditional Chinese characters into simplified Chinese characters;

[0010] S1-3: Removing stop words: Use a Chinese stop word list to remove words including but not limited to "de", "le", "zai", and other common redundant words;

[0011] S1-4: Removing numbers and special characters: Remove all numbers and special characters, including: punctuation marks, letters, formulas, and special symbols;

[0012] S1-5: Remove low-frequency and high-frequency words: Among them, low-frequency words are those that appear less than 20 times in the entire dataset, and high-frequency words are those that appear 50,000 times in the entire dataset.

[0013] As a preferred technical solution of the present invention, the multi-algorithm clustering in S2 specifically includes K-means Clustering, Density-Based Spatial Clustering of Applications with Noise (DBSCAN), and Hierarchical Clustering. The fuzziness calculation formula is as follows:

[0014]

[0015] Among them, F j is the fuzziness of the j-th data, r is the number of clustering algorithms, specifically 3, K is the number of clusters, specifically 60, and H(i) is the calculation function of the i-th clustering algorithm.

[0016] As a preferred technical solution of the present invention, the specific steps for selecting high-confidence data in S3 are as follows:

[0017] S3-1: Sort all samples according to the fuzziness value, and the sample with the lowest fuzziness is ranked last;

[0018] S3-2: Set the fuzziness threshold to 80%, that is, select the first 80% of the sorted samples to form a high-confidence dataset;

[0019] S3-3: Use the selected high-confidence dataset for training or inference of the improved LDA topic model.

[0020] As a preferred technical solution of the present invention, the specific process of the model training in S4 is as follows:

[0021] T1: Initialize topic assignment: Randomly assign an initial topic to each word in the document set, and initialize the document-topic count matrix N d,t and the topic-word count matrix N t,w , and initialize the projection matrix W between topics;

[0022] T2: Calculate the between-class scatter matrix S b and the within-class scatter matrix S w , and its calculation formula is:

[0023]

[0024]

[0025] Among them, N tis the total number of words assigned to topic t, μ t is the average word vector of topic t, μ is the global average word vector, ν w is the vector representation of word w;

[0026] T3: Definition and maximization of the objective function: Specifically, the gradient ascent method is used to maximize J(W). In each iteration, W is updated according to the gradient of J(W). The specific update formula is:

[0027]

[0028] where W new is the updated W, η is the learning rate, is the gradient of J(W) with respect to the current W;

[0029] T4: Update the topic assignment based on the optimized W: Use the optimized W to update the topic assignment probability of each word; recount the matrix N d,t and the topic-word count matrix N t,w ;

[0030] T5: Iterate until convergence: Repeat steps T3 and T4 until the increase in J(W) is less than a preset threshold or the maximum number of iterations is reached;

[0031] T6: Output the final topic assignment: After the iteration ends, output the topic assignment based on the optimized W.

[0032] As a preferred technical solution of the present invention, the specific process of initializing the document-topic count matrix N d,t and the topic-word count matrix N t,w in T1 is as follows:

[0033] T1-1: Initialize all values in the document-topic count matrix N d,t and the topic-word count matrix N t,w to zero;

[0034] T1-2: Each word in each document is randomly assigned a topic;

[0035] T1-3: For each word assigned to topic t in document d, increment the corresponding count in N d,t by 1; for each word w assigned to topic t, increment the corresponding count in N t,w by 1.

[0036] As a preferred technical solution of the present invention, the specific 60 topic categories described in S4-3 specifically include: child treatment, practical medical practice, etc.

[0037] As a preferred technical solution of the present invention, by classifying the patient evaluation text based on the patient evaluation classification model, the following specific steps are included:

[0038] P1: Input the patient evaluation text Text input Perform the same data preprocessing operation as in step S1 to obtain the transformed vector as X;

[0039] P2: Use the parameters in the trained LDA model, including the document-topic count matrix N d,t 、the topic-term count matrix N t,w and the projection matrix W. For the input vector X, calculate its probability distribution under different topics, and its calculation formula is: p(t|X) = N d,t ×exp(x T WN t,w );

[0040] P3: According to the probability p(t|X) of each topic, select the topic with the highest probability

[0041] P4: According to the determined topic t max Map to the corresponding user evaluation category z, and output that the user evaluation category to which the evaluation text belongs is z.

[0042] Compared with the related prior art, compared with the prior art, the present application proposal has the following main technical advantages: The beneficial effects of the present invention are:

[0043] Comprehensive data preprocessing: The present invention, through comprehensive data preprocessing steps, including word segmentation, abbreviation conversion, stop word removal, digital and special character processing, and filtering of low-frequency and high-frequency words, more effectively prepares the text data, providing a cleaner and more accurate data basis for subsequent model training and analysis.

[0044] Combination of multiple clustering algorithms: The combination of multiple unsupervised clustering algorithms improves the accuracy and robustness of data classification. This method of comprehensively using multiple algorithms helps to more comprehensively reveal the hidden patterns in the data and improves the overall performance of the model.

[0045] Innovative model training method: Through the improved LDA topic model, combined with the calculation of the between-class scatter matrix and the within-class scatter matrix, and by maximizing a specific function to optimize the projection matrix, the present invention can more accurately capture and represent the topic structure in the text data.

[0046] Effective utilization of high-confidence data sets: Selecting high-confidence data samples according to the ambiguity provides a more reliable basis for model training. This step helps to improve the quality of model training and the accuracy of the final classification.

[0047] More accurate text classification: By comparing the processed feature vectors with the determined central points of the topic categories, the present invention can more accurately classify patient evaluations into relevant categories, which is crucial for understanding and improving the quality of medical services.

[0048] Strong adaptability: The method of the present invention is applicable to processing a large amount of text data, especially in fields such as medical service evaluation, and can effectively process and analyze large-scale user feedback, providing real-time and accurate insights for medical services.

[0049] The present invention provides a more effective and accurate patient evaluation classification and analysis tool by combining advanced text processing technologies and machine learning methods, which is of great significance for improving the quality of medical services and supporting public health decision-making. Description of the drawings:

[0050] Figure 1 is the flowchart of the method provided by the present invention;

[0051] Figure 2 is the two-dimensional diagram of 60 topics generated by the LDA model of the embodiment provided by the present invention. Detailed implementation manners:

[0052] The following further illustrates the present invention in conjunction with the drawings and embodiments. However, the present invention can be implemented in many different ways and should not be construed as being limited to the shown embodiments; on the contrary, these embodiments provide implementation manners that meet the applicable legal requirements for those skilled in the art.

[0053] Embodiment 1: This embodiment provides a specific implementation plan for a patient evaluation classification method based on an improved LDA topic model. As Figure 1 shown, the following is the classification process of the simulated patient evaluation text:

[0054] Data collection and preprocessing: Data collection: Select 1000 written evaluations of patients for general practitioner services. Word segmentation: Use jieba to perform word segmentation on each evaluation. Simplification: Use the zhconv library to convert traditional Chinese characters in the word segmentation results to simplified Chinese. Stop word removal: Use the Chinese stop word list to remove common irrelevant words such as "de", "le", etc. Removal of numbers and special characters: Remove all numbers, punctuation marks, letters, and special symbols. Removal of low-frequency and high-frequency words: Delete words that appear less than 20 times or more than 50,000 times in the dataset.

[0055] Multi-algorithm clustering: Apply DBSCAN, K-means, and hierarchical clustering algorithms to cluster the preprocessed data. Calculate the fuzziness of each sample to determine the consistency among the clustering algorithms.

[0056] High-confidence data selection: According to the ambiguity, 80% of the samples with the lowest ambiguity are selected from the preprocessed data as the high-confidence data set.

[0057] Improved LDA topic model training: Initialize the topic assignment, and establish the document-topic and topic-term count matrices. Calculate the between-class scatter matrix and the within-class scatter matrix. Optimize the projection matrix by the gradient ascent method to maximize the objective function. Use the optimized projection matrix to recluster the high-confidence data set to determine the central topics of 60 categories.

[0058] Evaluate text classification: Perform the same preprocessing on the newly collected patient evaluation text as in step 1. Use the trained projection matrix to transform the feature vectors. Compare the transformed vectors with the 60 central topics, and determine the text topic category based on the principle of the minimum distance.

[0059] Example 2: As Figure 2 shown, the two-dimensional graph of the 60 topics generated by the LDA model can also compare the similarity between topics. If the vocabulary selections represented by two topics are similar, they are considered similar; if there are very few common words in both topics, they are considered very different. Figure 2 Represents the relative similarity between the topics depicted on the two-dimensional plane. The node size indicates the topic popularity, and the node color indicates the topic cluster. The distance between topics is calculated by the cosine similarity score calculated for each topic. Relatively strong pairwise similarities are depicted by relatively small distances and thicker edges connecting the nodes. The result is a slightly elongated topic map, roughly divided into 2 groups, one on the left side of the graph and one on the right side.

[0060] Example 3: In this example, the present invention will introduce how to calculate the ambiguity through a set of specific data, and construct a high-confidence data set accordingly. The ambiguity here refers to the uncertainty that a data sample is assigned to the same cluster by different clustering algorithms. The high-confidence data set is composed of those data samples that remain consistent in different clustering results, which means that these samples are assigned to the same cluster or similar clusters in different clustering algorithms.

[0061] Suppose the data set is as follows: Text A, Text B, Text C, Text D, Text E.

[0062] Suppose the clustering results are as follows: K-means: Cluster 1 (A, B, C), Cluster 2 (D, E); DBSCAN: Cluster 1 (A, B), Cluster 2 (C, D, E); Hierarchical clustering: Cluster 1 (A, C), Cluster 2 (B, D, E).

[0063] For each text, the present invention will calculate the number of times it is consistently assigned to the same cluster in all clustering results. For text A: it belongs to cluster 1 in K-means, cluster 1 in DBSCAN, and cluster 1 in hierarchical clustering. Text A is assigned to cluster 1 in all clustering algorithms, so the fuzziness is 0. For text B: it belongs to cluster 1 in K-means, cluster 1 in DBSCAN, and cluster 2 in hierarchical clustering. Text B belongs to cluster 1 in two clustering algorithms and cluster 2 in one algorithm. Therefore, the fuzziness is calculated as 1 - (the maximum number of consistent classifications / the total number of clustering algorithms) = 1 - (2 / 3) = 1 / 3. For text C: it belongs to cluster 1 in K-means, cluster 2 in DBSCAN, and cluster 1 in hierarchical clustering. Text C belongs to cluster 1 in two clustering algorithms and cluster 2 in one algorithm, and the fuzziness is 1 / 3. For text D and text E: they belong to cluster 2 in K-means, cluster 2 in DBSCAN, and cluster 2 in hierarchical clustering. Text D and text E are assigned to cluster 2 in all clustering algorithms, so the fuzziness is 0.

[0064] Now, the present invention selects high-confidence data based on the calculated fuzziness. If the present invention sets the fuzziness threshold to 0, then text A, text D, and text E will be selected into the high-confidence data set because their classifications are consistent in all clustering algorithms. If the present invention sets the fuzziness threshold to be less than or equal to 1 / 3, then text A, text B, text C, text D, and text E will all be selected into the high-confidence data set because their fuzziness values do not exceed 1 / 3.

[0065] The calculated fuzziness results are: for text A: 0, for text B: 1 / 3, for text C: 1 / 3, for text D: 0, for text E: 0.

[0066] This high-confidence data set can be used to improve the training of the LDA topic model or for more refined patient evaluation analysis because the included samples show a high degree of consistency in different clustering algorithms.

[0067] Through the above embodiments, the specific application of the present invention from theory to practice can be seen, showing a complete workflow and method details. These steps together construct a systematic method for extracting, processing, and analyzing patient evaluation data and converting it into useful information for optimizing medical services and improving patient satisfaction. It can be clearly understood the importance of each operation step for the entire process and how these steps interact to support the final evaluation text classification.

[0068] The above embodiments merely illustrate several implementation manners of the present invention. The description thereof is relatively specific and detailed, but it should not be construed as a limitation to the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all fall within the protection scope of the present invention.

Claims

1. A patient evaluation classification method based on an improved LDA topic model, characterized in that Specifically, the following steps are included to complete the patient evaluation classification: S1: Data collection and preprocessing: Collect written evaluations of patients served by general practitioners and perform a series of preprocessing steps, including tokenization, lowercasing, removing stopwords, removing numbers and special characters, and removing low-frequency and high-frequency words; S2: Multi-algorithm clustering: Use multiple unsupervised clustering algorithms to cluster the preprocessed dataset, generate multiple different clustering results, and calculate a fuzziness for each data sample, which reflects the consistency among the clustering results; S3: High-confidence data selection: According to the fuzziness calculated in S2, select the data samples with the lowest fuzziness from the preprocessed dataset to form a high-confidence dataset; S4: Improved LDA topic model training: Use the high-confidence dataset to train the improved LDA topic model, including: S4-1: Calculate the between-class scatter matrix S b and the within-class scatter matrix S4-2: Update the projection matrix W by solving the optimization problem to maximize the function J(W), where the calculation formula of J(W) is: Among them, W represents the projection matrix, and W Τ represents the transpose of the projection matrix, and T r represents the trace of the matrix; S4-3: Use the updated projection matrix W to recluster the high-confidence dataset to update the clustering results, and determine the central topics of 60 categories accordingly; S5: Patient evaluation classification: For any given user evaluation text, use the improved LDA topic model trained in step S4 to perform the category judgment.

2. The patient evaluation classification method based on the improved LDA topic model according to claim 1, wherein The specific preprocessing steps described in S1 are: S1-1: Tokenization: Use jieba for tokenization to split continuous written evaluations of patients into independent words; S1-2: Lowercasing: Use the zhconv library to perform traditional-simplified conversion on independent words and convert the existing traditional Chinese characters into simplified Chinese characters; S1-3: Removing stopwords: Use a Chinese stopword list to remove words including but not limited to "de", "le", "zai", and other common redundant words; S1-4: Removing numbers and special characters: Remove all numbers and special characters, including punctuation, letters, formulas, and special symbols; S1-5: Removing low-frequency and high-frequency words: Among them, low-frequency words are those that appear less than 20 times in the entire dataset, and high-frequency words are those that appear 50,000 times in the entire dataset.

3. A patient evaluation classification method based on an improved LDA topic model according to claim 1, characterized in that The multi-algorithm clustering described in S2 specifically includes the K-means Clustering algorithm, the Density-Based Spatial Clustering of Applications with Noise (DBSCAN) algorithm, and the Hierarchical Clustering algorithm. The fuzziness calculation formula is: Among them, F j is the ambiguity of the j-th data, r is the number of clustering algorithms, specifically 3, K is the number of clusters, specifically 60, and H(i) is the calculation function of the i-th clustering algorithm.

4. A patient evaluation classification method based on an improved LDA topic model according to claim 1, characterized in that The specific steps for high-confidence data selection in S3 are: S3-1: Sort all samples according to the fuzziness value, and the sample with the lowest fuzziness is ranked last; S3-2: Set the fuzziness threshold to 80%, that is, select the first 80% of the sorted samples to form a high-confidence dataset; S3-3: Use the selected high-confidence dataset for the training or inference of the improved LDA topic model.

5. A patient evaluation classification method based on an improved LDA topic model according to claim 1, characterized in that, The specific process of the model training described in S4 is: T1: Initialize topic assignment: Randomly assign an initial topic to each term in the document set, and initialize the document-topic count matrix N d,t and the topic-term count matrix N t,w , and initialize the projection matrix W between topics; T2: Calculate the between-class scatter matrix $S$ b and the within-class scatter matrix $S$ w , and their calculation formulas are as follows: where N t is the total number of words assigned to topic t, μ t is the average word vector of topic t, μ is the global average word vector, ν w is the vector representation of word w; T3: Definition and maximization of the objective function: Specifically, the gradient ascent method is used to maximize J(W). In each iteration, W is updated according to the gradient of J(W), and the specific update formula is as follows: Among them, W new is the updated W, η is the learning rate, and is the gradient of J(W) with respect to the current W; T4: Update the topic assignment based on the optimized W: Use the optimized W to update the topic assignment probability for each term, and recount the matrix N d,t and the topic-term count matrix N t,w ; T5: Iterate until convergence: Repeat steps T3 and T4 until the increase in J(W) is less than a preset threshold or the maximum number of iterations is reached; T6: Output the final topic assignment: After the iteration ends, output the topic assignment based on the optimized W.

6. A method for classifying patient evaluations based on an improved LDA topic model according to claim 5, characterized in that The initialization document-topic count matrix N described in T1 d,t and the topic-term count matrix N t,w The specific process is as follows: T1-1: Initialize all values in the document-topic count matrix N d,t and the topic-term count matrix N t,w to zero; T1-2: Each term in each document is randomly assigned a topic; T1-3: For each term assigned to topic t in document d, increment the corresponding count in N d,t by 1; for each term w assigned to topic t, increment the corresponding count in N t,w by 1.

7. A patient evaluation classification method based on an improved LDA topic model according to claim 1, characterized in that The 60 topic categories described in S4-3 specifically include: child treatment, practical medical practices, etc.

8. A method for classifying patient evaluations based on an improved LDA topic model according to claim 1, characterized in that In step S5, the improved LDA topic model specifically completes patient evaluation classification through the following steps: P1: Input the patient evaluation text Text input Perform the same data preprocessing operation as in step S1 to obtain the transformed vector X; P2: Use the parameters in the trained LDA model, including the document-topic count matrix N d,t , the topic-term count matrix N t,w and the projection matrix W. For the input vector X, calculate its probability distribution under different topics, and its calculation formula is: p(t|X) = N d,t × exp(X T WN t,w ); P3: Select the topic \(t\) with the highest probability according to the probability \(p(t|X)\) for each topic max : P4: According to the determined topic t max Map it to the corresponding user evaluation category z, and output that the user evaluation category to which the evaluation text belongs is z.