Chronic pain feature recognition system based on m-n+ model
Patent Information
- Application Number
- CN202211615353.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-15
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2042-12-15
AI Technical Summary
[0004]本发明要解决的技术问题是:目前基于病历的慢性疼痛疾病特征提取方法有很多,但是这些方法存在病历数据离散化程度高、描述语言标准不统一、疾病特征的提取困难等问题
[0031]一)能从疼痛患者病例档案和多源多模态采集终端获取高维海量数据,对获取的数据进行管理、整合、分析和利用,通过数据挖掘和分析,探寻不同风险疼痛患者的差异化干预节点,制定个体化组合干预优化策略和持续改进的方案;
Smart Images

Figure CN115862844B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a chronic pain feature recognition technology based on the M-N+ model, which enables highly efficient and accurate identification of chronic pain features. Background Technology
[0002] The characteristics of chronic pain are the main basis for doctors to diagnose the type and level of pain. They are mainly contained in the chief complaint and various examination contents in the medical record. How to efficiently find the characteristics of chronic pain from massive electronic medical records to assist in diagnosis has always been a research hotspot in chronic pain data mining.
[0003] Chronic pain feature recognition refers to the process of identifying the most representative subset of features for a certain category from a dataset. Its principle mainly involves measuring the correlation between different features and the category, thereby selecting a subset of features with high relevance to the category from high-dimensional features. Generally, feature recognition methods fall into three categories: filtering, embedding, and wrapping. Filtering is independent of the learning algorithm, identifying feature subsets by filtering the dataset. Embedding performs feature recognition and learning simultaneously, selecting the optimal features during training. Wrapping incorporates the learning algorithm as part of feature selection. Filtering is the most commonly used feature recognition method; its main principle is to evaluate feature weights based on the inherent relationships within sample data, such as information gain and correlation coefficients. While these methods have played a role in feature extraction for analgesia, the high degree of discretization in chronic pain patient medical records and the lack of standardized descriptive language present challenges for feature extraction, reducing the accuracy of feature recognition. Summary of the Invention
[0004] The technical problem to be solved by this invention is that there are many methods for extracting chronic pain disease features based on medical records, but these methods have problems such as high degree of discretization of medical record data, inconsistent standardization of descriptive language, and difficulty in extracting disease features.
[0005] To address the aforementioned technical problems, the present invention provides a chronic pain feature recognition system based on the M-N+ model, characterized by comprising:
[0006] The data preprocessing module is used to preprocess and segment electronic medical record documents before storing them in text files;
[0007] M-N+ model: After running the M-N+ model on a text file, the medical record-disease distribution is obtained. Disease-characteristic distribution Two distribution matrices, through disease-feature distribution The feature distribution of diseases is obtained. The Bayesian method is used to count the optimal number of diseases K in the electronic medical record document. The calculated number of diseases K is then used as the input parameter of the M-N+ model. The formula for calculating the number of diseases K is shown below:
[0008]
[0009]
[0010] In the formula: P(w|s) represents the probability of the medical record-disease distribution; β represents the hyperparameter; Γ(Wβ) represents the conjugate binomial distribution of the pseudoword; Γ(β) represents the prior binomial distribution of the pseudoword; Represents the self-looping variable. Discrete function representing pseudowords; n i Γ(n) represents a pseudomorphic variable. i +Wβ) represents the discrete function of the true words; M represents the number of Gibbs samplings; P(w|K) represents the disease-feature distribution; s (i) p(w|s) represents the specific threshold range. (i) ) indicates a specific distribution.
[0011] Preferably, the data preprocessing module includes:
[0012] Data filtering unit: used to remove private and useless information from electronic medical record documents, retaining only information with a high density of disease characteristics;
[0013] Data discretization unit: Used to discretize continuous data in the data processed by the data filtering unit;
[0014] Word segmentation unit: Based on the medical lexicon, the data processed by the data discretization unit is segmented into words, and the segmented results are stored in a text file. In this process, the electronic medical record documents processed by the data discretization unit are manually annotated to obtain the complete medical terms appearing in the electronic medical record documents, and a medical terminology lexicon is built based on the obtained medical terms.
[0015] Preferably, in the M-N+ model, a joint probability formula is established among documents, diseases, and words, as shown in the following equation:
[0016]
[0017] In equation (1): P(θ, S, W|α, β) represents the joint probability among document, disease, and vocabulary; θ represents document attribute, S represents disease feature attribute, W represents vocabulary attribute, α and β are hyperparameters; P(θ|α) represents the deviation rate of document attribute; s n P(s) represents a discrete quantity representing a disease characteristic.n |θ) represents the goodness of fit between document attributes and disease features; w n P(w) represents the bias variable. n |s n ,β) represents the bias rate of feature fit;
[0018] Iterate through each word w in the medical record document d, calculate the marginal probability of word w, and obtain the probability of generating word w from medical record document d, as shown in formula (2):
[0019]
[0020] In the formula, P(w|α,β) represents the probability of feature fitting.
[0021] Preferably, the M-N+ model is used to train the medical record-disease distribution. Disease-characteristic distribution At that time, Gibbs sampling is used as shown in the following formula:
[0022]
[0023] In the formula: s i Indicates disease characteristic attributes, w i d represents the biased variable. i Let P(s) represent the case document, k represent the Gibbs model constant, and P(s) represent the case document. i =k|s i w i d i ) represents the Gibbs model; This represents the dynamic convergence matrix; K represents the number of diseases. W represents the static convergence matrix; W represents the lexical attributes.
[0024] Medical records - disease distribution Distribution of Disease Characteristics After m×n iterations, a stable case-disease distribution is finally obtained. Disease-characteristic distribution
[0025] Preferably, in each Gibbs sampling, the medical record-disease distribution The hidden Markov chains in the code are dynamically updated, and the update formula is shown below:
[0026]
[0027]
[0028] In the formula: θ m,s Represents the hidden Markov distribution of medical records and diseases; Represents the kernel variable of the unresolved problem in dimension s; α s Represents a positive hyperparameter sequence. Represents the disease-characteristic hidden Markov distribution; Represents the kernel variable for unresolved problems in n dimensions; β n represents the negative hyperparameter sequence; V represents the stochastic optimal control constant.
[0029] This invention forms a data loop through self-feedback and, based on a predetermined scenario, uses a waterfall flow to clean, extract, and standardize a high-dimensional set of chronic pain-related factors, ultimately achieving efficient and accurate identification of chronic pain features. This invention can simultaneously model the relationship between chronic pain patient medical records, chronic pain, and chronic pain patient features, deriving two distribution matrices: medical record-disease and disease-feature, thereby achieving the goal of disease feature identification. Experiments show that the disease feature identification accuracy of this invention is higher than that of the ID3 and C4.5 algorithms, achieving excellent results in chronic pain disease feature identification.
[0030] Compared with the prior art, the present invention has the following beneficial effects:
[0031] (i) It can acquire high-dimensional massive data from pain patient case files and multi-source multimodal acquisition terminals, manage, integrate, analyze and utilize the acquired data, explore differentiated intervention points for pain patients at different risks through data mining and analysis, and formulate individualized combined intervention optimization strategies and continuous improvement plans.
[0032] (ii) Collect massive amounts of data from multiple business systems at different levels, clean them, implement virtualized storage, and build a proprietary data DWH for the chronic pain disease management robot;
[0033] (iii) Using data mining techniques, algorithms are used to explore target information hidden in a large amount of multi-source data. Attached Figure Description
[0034] Figure 1 This illustrates the probability generation relationship;
[0035] Figure 2 The probability distribution of disease features generated by the M-N+ model is illustrated.
[0036] Figure 3 A comparison chart of the accuracy rates for disease feature identification. Detailed Implementation
[0037] The present invention will be further illustrated below with reference to specific embodiments. It should be understood that these embodiments are for illustrative purposes only and are not intended to limit the scope of the invention. Furthermore, it should be understood that after reading the teachings of this invention, those skilled in the art can make various alterations or modifications to the invention, and these equivalent forms also fall within the scope defined by the appended claims.
[0038] This invention is based on the M-N+ model. The M-N+ model is a 360-degree comprehensive case attribute representation generation model. This model identifies potential representation attributes in a case-related factor database by training two distributions: case document-representation attribute and representation attribute-vocabulary. The model does not require attribute labels during training, avoiding manual annotation. The core principle of the M-N+ model is that case documents generate a certain representation attribute with a specific probability, and the representation attribute, in turn, generates a certain key factor with a specific probability. Therefore, the metadata in case documents follows a multinomial distribution of topics, and topics follow a multinomial distribution of words, with the probability generation relationship as follows: Figure 1 As shown.
[0039] The M-N+ model treats case documents as a bag-of-words, without considering the logical or sequential relationships between words. The meanings of the parameters in the M-N+ model are shown in Table 1 below.
[0040]
[0041]
[0042] Table 1
[0043] Since w is an observable variable and topic z is a latent variable, the vocabulary w depends on topic z, and topic z depends on the case document-representation attribute distribution 0. Therefore, the probability that case document d generates representation vocabulary w is P(w|d)=P(z|d)P(w|z).
[0044] This invention focuses on the disease and uses the M-N+ model to model the dependencies between medical records, diseases, and features. Based on the probabilistic generation principle of generating diseases from medical records and features from diseases, it learns two distributions: chronic pain medical record archives – chronic pain and chronic pain – chronic pain features. This allows for efficient identification of the feature distribution of chronic pain diseases. Specifically, it includes the following:
[0045] Patient medical records document relevant information about a patient's diagnosis and treatment process, typically in textual form. Assuming a case involves several diseases, each with corresponding features, it can be deduced that there is a probabilistic dependency between the medical record, the disease, and the feature vocabulary. Here, the disease is a latent variable, and the vocabulary is an observable variable. Based on this, this paper applies the M-N+ model to the analysis of electronic medical records. By modeling the relationship between the medical record, the disease, and the feature vocabulary, it identifies the feature distribution of diseases within the medical record.
[0046] I. Define the medical record-disease distribution vector and the disease-feature distribution vector:
[0047] Suppose a medical record corpus D = {D1, D2, ..., D...} consists of M medical record texts. M}, D m Let m represent the m-th medical record in the medical record corpus D, where m = 1, 2, ..., M. Let the disease set S = {S1, S2, ..., S...} K}, S k Let represent the k-th disease in the disease set S, where k = 1, 2, ..., K, and K represents the total number of disease types in the disease set S. Let V = {V1, V2, ..., V...} be the vocabulary set consisting of all words in the medical record corpus D. N}, V n Let n represent the nth word in the vocabulary set V, where n = 1, 2, ..., N, and N represents the total number of words in the medical record corpus D.
[0048] Then there is a medical record-disease distribution vector. in, Indicates medical record D m The probability of generating the k-th disease. Indicates medical record D m The disease S is assigned to the kth disease. k The number of words, Table of Medical Records D m The total number of all words in the text.
[0049] Disease-Feature Distribution Vector S represents the kth disease k The probability of generating different feature words, where, S represents the kth disease k Generate the V of the nth word in the vocabulary set V. n probability, This indicates that the disease S is assigned to the k-th disease. k The nth word in the vocabulary set V n Quantity, Represents all diseases assigned to the k-th disease S. k The total number of words.
[0050] II. Learning the M-N+ Model
[0051] The M-N+ model first initializes the medical record-disease distribution. Disease-characteristic distribution Then, by traversing each medical record document and each word in the medical record document in the medical record corpus D, based on the activity word w in the disease-feature distribution... The change in probability is used to analyze the distribution of medical records and diseases. Disease-characteristic distribution The update process is described below:
[0052] For each medical record document in the medical record corpus D, from medical record-disease distribution Choose a disease S k This makes S k obey Distribution, among which, This represents a set of disease distribution vectors from multiple medical records. For each word in each medical record document, the disease-feature distribution is represented. Choose one word V n This makes V n obey distributed.
[0053] In the M-N+ model, in order to derive the probability of any word w in any medical record document d, it is first necessary to establish a joint probability formula among document, disease, and word, as shown in equation (1):
[0054]
[0055] In equation (1): P(θ, S, W|α, β) represents the joint probability among document, disease, and vocabulary; θ represents document attribute, S represents disease feature attribute, W represents vocabulary attribute, α and β are hyperparameters; P(θ|α) represents the deviation rate of document attribute; s n P(s) represents a discrete quantity representing a disease characteristic. n |θ) represents the goodness of fit between document attributes and disease features; w n P(w) represents the bias variable. n |s n β) represents the bias rate of feature fit;
[0056] Since equation (1) is a three-dimensional probability distribution, the marginal probability of word w is calculated to obtain the probability of generating word w from medical record document d, as shown in equation (2):
[0057]
[0058] In equation (2): P(w|α,β) represents the probability of feature fitting.
[0059] According to equation (2), iterate through each word w in the medical record document d.
[0060] Training medical records - disease distribution Disease-characteristic distribution The Gibbs sampling formula is shown in (3):
[0061]
[0062] In the formula: s i Indicates disease characteristic attributes, w i d represents the biased variable. i Let P(s) represent the case document, k represent the Gibbs model constant, and P(s) represent the case document. i =k|s i w i d i ) represents the Gibbs model; This represents the dynamic convergence matrix; K represents the number of diseases. W represents the static convergence matrix; W represents the lexical attributes.
[0063] In each sampling, the distribution of medical records and diseases The hidden Markov chains in the model are dynamically updated, and the update formulas are shown in equations (4) and (5):
[0064]
[0065]
[0066] In the formula: θ m,s Represents the hidden Markov distribution of medical records and diseases; Represents the kernel variable of the unresolved problem in dimension s; α s Represents a positive hyperparameter sequence. Represents the disease-characteristic hidden Markov distribution; Represents the kernel variable for unresolved problems in n dimensions; β n represents the negative hyperparameter sequence; V represents the stochastic optimal control constant.
[0067] The distribution of medical records and diseases is analyzed using formula (3). Disease-characteristic distribution After m×n iterations, a stable case-disease distribution is finally obtained. Disease-characteristic distribution
[0068] The experimental data for this embodiment was provided by the First Affiliated Hospital of Sun Yat-sen University. 63,252 electronic medical records of internal medicine inpatients from 2016 to 2021 were selected. The main content structure of the medical records includes basic patient information, chief complaint, present illness history, various physical examinations, diagnostic results, treatment methods and processes, etc.
[0069] Based on the above experimental data, the technical solution of the present invention further includes the following:
[0070] First, perform data preprocessing.
[0071] Since the characteristics of a disease are mainly distributed in the chief complaint, present medical history, examination results, and diagnosis, in order to eliminate the interference of irrelevant information, the medical record documents must first be processed to remove privacy and useless information, retaining only the patient's basic information, chief complaint, present medical history, examination results, diagnosis, and other content containing high density of disease characteristics, and then the data is discretized.
[0072] Step 1. Data Discretization
[0073] Some data in medical records are continuous, such as blood pressure, body temperature, and white blood cell count. Directly segmenting this data might yield useless results, so continuous data must be discretized. To improve the accuracy of disease feature recognition, discretization is performed manually. For example, if the medical record describes a patient's blood pressure as "diastolic blood pressure at 66," it is labeled as "diastolic blood pressure 65-70."
[0074] Step 2. Medical Glossary Construction
[0075] Unlike general text, the disease characteristics described in medical records are expressed using medical terminology. If a simple word segmentation method is used, the segmentation results not only fail to fully express the meaning of the features but also negatively impact the mining results. Therefore, ensuring the completeness of medical terminology is crucial for word segmentation of medical record text. For example, "no fever" and "blood pressure 65-70" are complete symptom descriptions, and we treat such phrases describing the patient's condition directly as "vocabulary words." To identify such descriptive phrases, we also need to manually annotate the discretized medical record documents and establish a medical terminology lexicon, which serves as the segmentation principle for our word segmentation tool.
[0076] Step 3. Medical record word segmentation
[0077] The processed electronic medical record documents were segmented and stop words were removed based on a medical terminology dictionary. The segmentation software used was ICTCLAS, a segmentation software developed by the Chinese Academy of Sciences. The segmentation results are stored in a text file named featureTxt, with one electronic medical record document per line.
[0078] Second, optimize the number of diseases.
[0079] Setting the number of diseases in the M-N+ model is a key aspect of this invention; setting it too high or too low will affect the accuracy of disease feature recognition. This invention first uses the currently popular Bayesian method to calculate the optimal number of diseases K in medical records, as shown in equations (6) and (7). Then, the calculated number of diseases K is used as the input parameter for the M-N+ model.
[0080]
[0081]
[0082] In the formula: P(w|s) represents the probability of the medical record-disease distribution; β represents the hyperparameter; Γ(Wβ) represents the conjugate binomial distribution of the pseudoword; Γ(β) represents the prior binomial distribution of the pseudoword; Represents the self-looping variable. Discrete function representing pseudowords; n i Γ(n) represents a pseudomorphic variable. i +Wβ) represents the discrete function of the true words; M represents the number of Gibbs samplings; P(w|K) represents the disease-feature distribution; s (i) p(w|s) represents the specific threshold range. (i) ) indicates a specific distribution.
[0083] The third point is the result of disease identification.
[0084] In the M-N+ model, the number of diseases K is obtained through optimization, with hyperparameters α and β set to 0.5 / K and 0.1, respectively. The disease feature threshold parameter disWord is set to 8, meaning each disease is represented by its 8 most probable feature words. After running the M-N+ model on the featureTxt dataset, the medical record-disease distribution is obtained. Disease-characteristic distribution Two distribution matrices, where the disease-feature distribution is used. The feature distribution of the diseases can be obtained. Since the identification results contain a large number of diseases, this embodiment selects the feature distributions of six diseases to illustrate the probability distribution of disease features generated by the M-N+ model, for easier explanation of the model. Details are as follows... Figure 2 As shown.
[0085] IV. Finally, evaluate the prediction accuracy.
[0086] To verify the accuracy of disease feature recognition in this invention, the data was divided into 10 equal parts using a 10-fold cross-validation method. The C4.5 algorithm and the ID3 algorithm were used as comparisons, and experiments were conducted on the experimental data to compare the disease feature recognition accuracy. Figure 3 As shown, Figure 3 The algorithm of this invention is labeled M-N+. (From...) Figure 3 The average disease feature recognition accuracy of the three algorithms, M-N+, C4.5 and ID3, is 81.72%, 79.74% and 77.26%, respectively. Therefore, the disease feature recognition accuracy of the M-N+ algorithm is better than that of the other two algorithms.
Claims
1. A chronic pain feature recognition system based on the M-N+ model, characterized in that, include: The data preprocessing module is used to preprocess and segment electronic medical record documents before storing them in text files; M-N+ model: After running the M-N+ model on a text file, the medical record-disease distribution is obtained. Disease-characteristic distribution Two distribution matrices, through disease-feature distribution The characteristic distribution of diseases is obtained, and the Bayesian method is used to count the optimal number of diseases in electronic medical record documents. The calculated number of diseases The number of diseases serves as an input parameter for the M-N+ model. The calculation formula is shown below: In the formula: Represents the probability distribution of medical records and diseases; Indicates hyperparameters, Indicates the conjugate binomial distribution of pseudowords; The prior binomial distribution of pseudowords; Represents the self-looping variable. Discrete functions representing pseudowords; Represents a pseudomorphic variable. Discrete functions representing true words; Indicates the number of Gibbs samples; Represents the disease-characteristic distribution; Indicates the specific threshold range, Indicates specific distribution; In the M-N+ model, a joint probability formula is established among documents, diseases, and words, as shown in the following equation: In formula (1): This represents the joint probability among documents, diseases, and words. Indicates document attributes, Indicates disease characteristics and attributes. Indicates lexical attributes, , For hyperparameters; Indicates the deviation rate of document attributes; Discrete quantities representing disease characteristics This indicates the goodness of fit between document attributes and disease characteristics; Indicates biased variables, Bias ratio, representing the feature fit; This indicates traversing medical record documents. Each word in Calculate vocabulary The marginal probability is used to obtain the medical record document. Generate vocabulary The probability of is shown in the following formula: In the formula, Indicates the probability of feature fitting; The M-N+ model is used to train medical records and disease distribution. Disease-characteristic distribution At that time, Gibbs sampling is used as shown in the following formula: In the formula: Indicates disease characteristics and attributes. Indicates biased variables, This refers to medical record documents. Represents the Gibbs model constants. Represents the Gibbs model; Represents the dynamic convergence matrix; Indicates the number of diseases; Represents the static convergence matrix; Indicates lexical attributes; Indicates the distribution of medical records and diseases Disease-characteristic distribution Perform m After n iterations, a stable case-disease distribution is finally obtained. Disease-characteristic distribution .
2. The chronic pain feature recognition system based on the M-N+ model as described in claim 1, characterized in that, The data preprocessing module includes: Data filtering unit: used to remove private and useless information from electronic medical record documents, retaining only information with a high density of disease characteristics; Data discretization unit: Used to discretize continuous data in the data processed by the data filtering unit; Word segmentation unit: Based on the medical lexicon, the data processed by the data discretization unit is segmented into words, and the segmented results are stored in a text file. In this process, the electronic medical record documents processed by the data discretization unit are manually annotated to obtain the complete medical terms appearing in the electronic medical record documents, and a medical terminology lexicon is built based on the obtained medical terms.
3. The chronic pain feature recognition system based on the M-N+ model as described in claim 1, characterized in that, In each Gibbs sampling, the distribution of medical records and diseases The hidden Markov chains in the code are dynamically updated, and the update formula is shown below: In the formula: Represents the hidden Markov distribution of medical records and diseases; Represents the kernel variable of the unresolved solution in dimension s; Represents a positive hyperparameter sequence. Representing disease-characteristic hidden Markov distribution; Represents the kernel variable of the unresolved problem in n dimensions; Represents a negative hyperparameter sequence; This represents the random optimal control constant.