Electronic medical record defect text classification method based on feature fusion

Through technical means such as feature fusion and weighted cross-entropy loss function, the problem of inability to capture deep semantics and category imbalance in the existing technology is solved, and efficient classification of electronic medical record defect texts and accurate identification of long-tail categories are achieved.

CN120179818APending Publication Date: 2025-06-20SOUTHWEAT UNIV OF SCI & TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510229047.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

Existing electronic medical record defective text classification methods cannot effectively capture the deep semantic information in the text and cannot cope with the problem of category imbalance, resulting in poor learning effects of low-frequency categories.

Method used

Using a method based on feature fusion, context-related word vectors are generated through preprocessing, time-dependent and local pattern features are extracted, mutual information is calculated for feature-weighted fusion, and model training is used using weighted cross entropy loss function and Adam optimizer to enhance the generalization ability of the model.

Benefits of technology

It significantly improves the model's recognition accuracy of long-tail categories, enhances text semantic understanding ability, improves classification performance and stability, and ensures high classification accuracy of low-frequency categories.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179818A_ABST
    Figure CN120179818A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of medical health informatics, and discloses an electronic medical record defect text classification method based on feature fusion, comprising the following steps: preprocessing electronic medical record data, including removing sensitive information, performing text word segmentation, and generating context-related word vectors by using a pre-trained BERT model; extracting features of the electronic medical record text, including extracting time sequence dependence features by using a Bi LSTM, extracting local mode features by using a CNN, and converting structured data into vectors by using a One-Hot encoding method; and calculating mutual information between each feature and the target category, weighting each feature according to a mutual information value, and fusing the weighted text features with the structured data features. By adopting the technical scheme of the weighted cross entropy loss function, the influence of each category is balanced, and the recognition precision of the model on the long-tail category is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical and health informatics, and particularly to an electronic medical record defect text classification method based on feature fusion. Background Art

[0002] An electronic medical record (EMR) is a digital system used by hospitals and medical institutions to record patients' health information and has become an important part of medical management. However, there may be some defects in the actual application of EMRs, such as input errors, inconsistencies, missing key information, or semantic ambiguities. These defects may affect the accuracy of medical decisions and the effectiveness of patient treatment. Therefore, identifying and fixing defects in EMRs has become a research focus in the medical field;

[0003] Most existing electronic medical record defect text classification methods rely on traditional feature extraction methods, such as the bag-of-words model (BOW) or TF-IDF, etc. Although these methods can extract certain text features, they do not fully consider the context relationship between words in the text. Especially for the identification of long-tail categories, existing methods usually over-bias towards frequently occurring categories, resulting in poor identification effects for minority categories. This simple way of processing text semantics cannot provide sufficient and accurate classification performance in the complex and changing medical field;

[0004] Classification algorithms such as Naive Bayes and Support Vector Machine are effective in some simple scenarios, but when the data becomes more complex, these methods often cannot capture the deep semantic information in the text. Especially in texts such as EMRs that contain medical entities and multiple semantic relationships, the performance of these models seems inadequate. Most existing solutions rely on manual features and rules, resulting in a significant decline in their performance when facing large-scale and diverse data. Therefore, the present invention provides an electronic medical record defect text classification method based on feature fusion to solve the deficiencies existing in the prior art. Summary of the Invention

[0005] Aiming at the deficiencies of the prior art, the present invention provides an electronic medical record defect text classification method based on feature fusion, which solves the problems in the prior art that it cannot capture the deep semantic information in the text, cannot effectively handle class imbalance, and cannot adjust weights according to the actual distribution of categories during the training process, which leads to poor learning effects for low-frequency categories.

[0006] To achieve the above objectives, the present invention is realized through the following technical solutions: An electronic medical record defect text classification method based on feature fusion, comprising the following steps:

[0007] Preprocess the defective data of electronic medical record texts, including removing sensitive information, performing text tokenization, and using a pre-trained BERT model to generate context-related word vectors;

[0008] Extract the features of the defective data of electronic medical record texts, including using Bi-LSTM to extract temporal dependence features, using CNN to extract local pattern features, and adopting the One-Hot encoding method to convert structured data into vectors;

[0009] Calculate the mutual information between each feature and the target category, weight each feature according to the mutual information value, and fuse the weighted defective data text features and structured data features to obtain a fused feature vector;

[0010] Train the model through a weighted cross-entropy loss function, where the weighted cross-entropy loss function dynamically calculates the class weights according to the frequency of the classes and gives larger weights to the long-tail classes;

[0011] Use the Adam optimizer for optimization during the training process, and combine Dropout and L2 regularization to enhance the generalization ability of the model;

[0012] Evaluate the model, and use accuracy, recall rate, and F1 value as performance evaluation indicators to ensure that the model has a high classification accuracy for low-frequency classes.

[0013] Preferably, the text tokenization includes using natural language processing tools to tokenize the text of the medical record defective data.

[0014] Preferably, the mutual information calculation formula is:

[0015] I(X i ; Y) = H(X i ) + H(Y) - H(X i , Y)

[0016] where X i is the i-th feature, Y is the target category, H(X i ) and H(Y) are the entropies of the feature and the category respectively, H(X i , Y) is the joint entropy of the feature and the category, and I(X i ; Y) represents the mutual information between the i-th feature X i and the target category Y.

[0017] Preferably, the fusion of the weighted defective data text features and structured data features further includes the following steps;

[0018] Fuse the defective data text features and structured data features by weighting according to their importance to ensure that the result after feature fusion maximally retains the key information;

[0019] After feature fusion, a feature selection method is used to further screen and compress irrelevant and redundant features, improving the model training efficiency;

[0020] An adaptive weight adjustment mechanism is introduced to dynamically adjust the weights of each feature according to the model feedback during the training process, optimizing the recognition ability for low-frequency categories;

[0021] The fused features optimized by feature weighting and selection are input into the classifier for the final classification of medical record defects.

[0022] Preferably, the formula for fusing the weighted feature fusion with the structured data features is:

[0023]

[0024] where, X i is the i-th feature, w i is the weighting coefficient of this feature, X fused is the fused feature vector, and M is the number of features.

[0025] Preferably, the calculation process of the weighted cross-entropy loss function includes the following steps:

[0026] Calculate the weight of each category, and the weight is dynamically adjusted according to the frequency of the category in the training data. The weight of the long-tail category is relatively large;

[0027] Use the weighted cross-entropy loss function to calculate the difference between the model output and the true label;

[0028] For each training sample, calculate the loss according to its category label. The loss of the long-tail category is relatively large;

[0029] Sum up the loss values of all samples, and adjust the model parameters through backpropagation to improve the recognition accuracy of the model for low-frequency categories.

[0030] Preferably, the calculation formula of the weighted cross-entropy loss function is:

[0031]

[0032] where, L represents the weighted cross-entropy loss value, w k is the weight of category k, y k is the category label, p k is the predicted probability of category k, and K is the number of categories.

[0033] Preferably, the update formula of the Adam optimizer is:

[0034]

[0035] Among them, θ t is a model parameter, and are respectively the first - order moment estimate and the second - order moment estimate of the gradient, η is the learning rate, ∈ is a constant to prevent division - by - zero errors, and θ t-1 represents the model parameter at time step t - 1.

[0036] Preferably, the evaluation of the model includes the following steps:

[0037] Calculate the proportion of samples correctly classified by the model in the total samples, reflecting the overall classification effect;

[0038] Calculate the proportion of positive class samples correctly predicted by the model in all positive class samples, evaluating the model's recognition ability for low - frequency categories;

[0039] Integrate the accuracy rate and the recall rate to evaluate the comprehensive performance of the model, especially in the case of class imbalance;

[0040] Adjust the hyperparameters in the training process to optimize the model performance, especially to improve the recognition accuracy of low - frequency categories.

[0041] The present invention also provides an electronic medical record defect text classification system based on feature fusion, including:

[0042] A pre - processing module: used to desensitize and segment the electronic medical record defect data, and generate word vectors through BERT;

[0043] A feature extraction module: used to extract defect data text features and structured data features, including BiLSTM layer, CNN layer and structured data processing;

[0044] A feature fusion module: used to perform mutual - information - optimized weighted fusion on the extracted defect data text features;

[0045] A classification module: used to train and optimize the model using a weighted cross - entropy loss function and calculate the classification result;

[0046] An evaluation module: used to evaluate the model performance and output accuracy rate, recall rate and F1 - value metrics.

[0047] The present invention provides an electronic medical record defect text classification method based on feature fusion.

[0048] It has the following beneficial effects:

[0049] 1. The technical solution of the present invention adopts a weighted cross-entropy loss function, achieving the balance of the influence of various categories and improving the recognition accuracy of the model for long-tail categories. Compared with the solution using the standard cross-entropy loss function in the prior art, the problem of class imbalance is not considered, which easily leads to the model being overly biased towards frequent categories and ignoring the effective recognition of minority categories. Through the weighting mechanism, the present invention significantly improves the performance of the model on long-tail categories.

[0050] 2. The technical solution of the present invention adopts the BERT model to generate context-related word vectors, achieving in-depth semantic understanding and feature extraction of the text of electronic medical record defect data. Compared with the solution based on a simple bag-of-words model in the prior art, it cannot fully capture the semantic connections between contexts, resulting in information loss or misunderstanding. The BERT model provides context-sensitive feature representations, significantly enhancing the semantic expression ability of the text.

[0051] 3. The present invention adopts a feature extraction method combining BiLSTM and CNN, achieving efficient extraction of temporal dependencies and local pattern features in the text. Compared with single feature extraction methods in the prior art, traditional methods often ignore the multi-level information of complex texts. The present invention significantly improves the accuracy and comprehensiveness of feature extraction by simultaneously using BiLSTM to capture temporal information and CNN to extract local patterns.

[0052] 4. The technical solution of the present invention adopts feature weighted fusion and mutual information optimization, ensuring the priority of highly relevant features in model training. Compared with the solution using simple weighting or average fusion in the prior art, which lacks in-depth analysis of feature relevance, the present invention effectively ensures the dominant position of key features in the model and improves the classification accuracy and stability by calculating mutual information and weighted feature fusion. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 is the flowchart of the steps of the present invention;

[0054] Figure 2 is the flowchart for establishing the medical record defect suggestion atlas of the present invention;

[0055] Figure 3 is the medical record quality control flowchart of the present invention;

[0056] Figure 4 is the system structure diagram of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0057] Next, in combination with the accompanying drawings of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0058] Please refer to the attached Figure 1 - attached Figure 3 , the embodiment of the present invention provides a method for classifying defective text of electronic medical records based on feature fusion. By removing sensitive information and performing text tokenization on the defective data of electronic medical records, and using the BERT model to generate context-related word vectors, effective preprocessing of the data and extraction of semantic representations are completed. Combining the BiLSTM and CNN models to extract the text features of the defective data and perform One-Hot encoding to convert the structured data realizes multi-level feature extraction and improves the expression ability of the features. By adopting a weighted cross-entropy loss function and optimizer adjustment, the problem of class imbalance is solved, and the recognition accuracy of the model for long-tail categories is enhanced. Finally, by optimizing the model through the feedback of comprehensive evaluation indicators, its accuracy, generalization ability and stability are improved, ensuring high efficiency and robustness in practical applications. Specifically, it includes the following steps:

[0059] S1. Preprocess the defective data of electronic medical records, including removing sensitive information and performing text tokenization preprocessing;

[0060] S2. Extract the features of the text of the defective data of electronic medical records;

[0061] S3. Calculate the mutual information between each feature and the target category, and fuse the text features and structured data features of the defective data;

[0062] S4. Train the model through a weighted cross-entropy loss function;

[0063] S5. Use the Adam optimizer for optimization during the training process;

[0064] S6. Evaluate the model.

[0065] For step S1, in this embodiment, the preprocessing of the defective data of electronic medical records mainly includes two aspects: removing sensitive information and extracting text features. Among them, removing sensitive information is the key to ensuring data compliance, and extracting text features provides suitable input for subsequent feature fusion and model training.

[0066] First, to ensure data privacy and compliance, sensitive information in electronic medical record data must be removed. Specifically, the content to be removed includes, but is not limited to, personal privacy information such as patient names, ID numbers, contact information, etc. These information must be excluded during the processing to comply with privacy protection laws and regulations. For other parts of the medical record that may involve privacy, the system should identify them through rules or models and perform corresponding data cleaning and masking processing.

[0067] Generally, the method of removing these sensitive information can be to automatically identify and delete them through regular expressions or deep learning techniques. In addition, in some scenarios where privacy information needs to be retained, the data may be encrypted or anonymized to further ensure data security.

[0068] After removing the sensitive information, the electronic medical record text will enter the text tokenization (NER) stage. Text tokenization is a basic task in natural language processing. It decomposes long texts into individual words or phrases for subsequent semantic analysis and feature extraction. In electronic medical records, text tokenization not only needs to consider common vocabulary but also needs to specially process professional terms in the medical field, such as disease names, symptom descriptions, and treatment methods, etc.

[0069] As an option, text tokenization can use existing natural language processing tools (such as NLTK or spaCy). These tools can already handle Chinese text tokenization well. Especially in the medical field, there are many specialized medical dictionaries and tools that can be used to enhance the accuracy of tokenization. In some embodiments, part-of-speech tagging can also be performed on the tokenization results to capture important grammatical information in the medical record, such as nouns, verbs, etc.

[0070] After completing text tokenization, the next step is to convert the text data into numerical features that the model can understand. The BERT model is used to convert the text into context-sensitive word vectors. Different from traditional bag-of-words models and Word2Vec, BERT can capture the context relationship of words through a bidirectional Transformer architecture, thus generating a more accurate word vector representation.

[0071] In this stage, a pre-trained BERT model is used in the present invention to generate word vectors for the defective text of electronic medical records. In some embodiments, the BERT model can be further fine-tuned according to the medical record data to adapt to specific task requirements. The word vectors generated by BERT can better retain the semantic information of each word, and the vector can change dynamically according to the context, thereby enhancing the semantic understanding ability of the model.

[0072] Specifically, BERT captures the relationships between words by converting each word in the medical record into a vector representation and adjusts the word meanings based on the context. For example, in the sentence "The patient feels nauseous and has a headache", BERT can embed "nauseous" and "headache" into the corresponding vector spaces according to the context, thus providing more accurate input features for subsequent model training.

[0073] For step S2, in this embodiment, BiLSTM (Bidirectional Long Short-Term Memory Network) is used to extract the temporal dependence features in the text. Through its bidirectional structure, BiLSTM can capture the forward and backward context information in the text simultaneously. For the chronological information (such as the change of the condition, the description of symptoms, etc.) in the defective text of the electronic medical record, BiLSTM can effectively learn and capture these temporal relationships.

[0074] Specifically, the output of BiLSTM is the context representation of each input word, and the calculation process is as follows:

[0075] h t = BiLSTM(x t ) = LSTM forward (x t ) + LSTM backward (x t )

[0076] where x t represents the input word vector at the t-th moment, h t is the output of the BiLSTM network at this moment, and LSTM forward and LSTM backward represent the forward and backward LSTM units of the BiLSTM network respectively.

[0077] CNN (Convolutional Neural Network) is used to extract the local pattern features in the text, especially suitable for identifying specific phrases or patterns in the text, such as the description of the condition or symptoms. In the electronic medical record, the expressions of diseases and symptoms often present specific phrase structures. Through convolutional operations, CNN can automatically identify these phrase-level features. The calculation process of CNN is:

[0078] f i = CNN(x t ) = ReLU(W·x t + b)

[0079] where x t represents the local features of the input text, W and b are the convolutional kernel and the bias term respectively, and f i is the convolutional result, and ReLU represents the activation function.

[0080] In this embodiment, structured data (such as a patient's diagnosis information, medications, treatment plans, etc.) is converted into vector form through One-Hot encoding. One-Hot encoding maps each discrete value (e.g., different disease types) to a high-dimensional binary vector. For example, if a patient's diagnosis result has three possible values, "hypertension", "diabetes", and "coronary heart disease", One-Hot encoding represents these values as:

[0081] HighBloodPressure = [1, 0, 0], Diabetes = [0, 1, 0],

[0082] CoronaryHeartDisease = [0, 0, 1]

[0083] Among them, HighBloodPressure = [1, 0, 0] indicates that this sample belongs to the "hypertension" category, Diabetes = [0, 1, 0] indicates that this sample belongs to the "diabetes" category, and CoronaryHeartDisease = [0, 0, 1] indicates that this sample belongs to the "coronary heart disease" category. In this way, structured data is transformed into numerical vectors and can be combined with text features for further processing.

[0084] For step S3, feature fusion adopts a weighted fusion method to combine features from text and structured data. The basic idea of this weighted fusion is to assign different weights to each feature according to the importance of the feature, ensuring that the model can focus on the most representative features. To quantify the relevance of each feature, we use mutual information (MI) technology to evaluate the dependence relationship between each feature and the target category and calculate the weights accordingly.

[0085] Generally, mutual information is used to quantify the mutual dependence between two random variables. In the feature fusion process, we calculate the mutual information between each feature and the target category and assign different weights to the features according to the mutual information values. The higher the mutual information value of a feature, the stronger its correlation with the target category. Therefore, the weight of this feature in the fusion process will be greater, making its contribution to the model classification result higher.

[0086] Specifically, the mutual information calculation formula is as follows:

[0087] I(X i ; Y) = H(X i ) + H(Y) - H(X i , Y)

[0088] Among them, X i is the i-th feature, Y is the target category, H(Xi ) and H(Y) are the entropies of features and classes respectively, and H(X i , Y) is the joint entropy of features and classes. By calculating the mutual information, the correlation between features and target classes can be quantified, and the weight of each feature in the fusion process can be determined.

[0089] After obtaining the weights of each feature, the text features and structured data features are weighted and fused, and the calculation formula is:

[0090]

[0091] where X i is the i-th feature, w i is the weight of this feature, X fused is the fused feature vector, and M is the total number of features. After weighted fusion, the combination of text features and structured data features can provide more comprehensive medical record information to help the model classify better.

[0092] As an option, to improve the effect of feature fusion, the present invention further introduces a feature selection and adaptive weight adjustment mechanism. The purpose of feature selection is to remove redundant or irrelevant features to ensure that the model focuses on the most valuable features. After feature fusion, some feature selection algorithms (such as L1 regularization or information gain method) can be used to screen the fused features to remove those features that contribute less to the classification result, thereby improving the training efficiency and accuracy of the model

[0093] Through the adaptive weight adjustment mechanism, the model can dynamically adjust the weight of each feature according to the feedback during the training process. During the training process, the change of the weight can make the model pay more attention to the long-tail classes and optimize the recognition effect of low-frequency classes. This mechanism ensures that the model can still maintain a high classification accuracy when facing the long-tail distribution problem.

[0094] The final feature vector after weighted fusion and feature selection is input into the classifier for further training and prediction. The classifier uses these fused features to perform the defect classification task on the electronic medical record data. Compared with the traditional method, the fused feature vector contains richer information, so the classifier can more accurately identify complex defect types.

[0095] For step S4, the purpose of the weighted cross-entropy loss function is to dynamically adjust the weights of each class according to the class distribution, especially to assign higher weights to the long-tail classes (i.e., those classes with fewer samples) to ensure their importance in model training. This method can effectively alleviate the dominant influence of large-class data on the model when dealing with class imbalance, and then improve the classification ability of low-frequency classes.

[0096] In general, the cross-entropy loss function is the standard loss function used in classification tasks and is optimized by calculating the difference between the predicted results and the true labels. In datasets with imbalanced classes, the traditional cross-entropy loss function may cause the model to overfit to frequently occurring classes while ignoring the long-tail classes with fewer samples. To address this issue, the present invention adopts a weighted cross-entropy loss function, which adjusts the influence of each class by assigning weights to each class, thereby achieving balance between classes.

[0097] Specifically, the calculation formula of the weighted cross-entropy loss function is as follows:

[0098]

[0099] where L represents the weighted cross-entropy loss, K is the total number of classes, w k is the weight of class k, y k is the label value of class k in the true label, and p k is the predicted probability of the model for class k. The core idea of this formula is to dynamically adjust the weight w k of each class according to the frequency of the class or other measurement criteria to ensure that the long-tail classes with fewer samples are not ignored.

[0100] As an option, in practical applications, the weight w k of the class can be calculated by inverse frequency. Specifically, the occurrence frequency of the class in the training data can be calculated, and then the weight of each class can be obtained. The weight calculation formula is:

[0101]

[0102] where f k represents the occurrence frequency of class k, and w k is the weight of class k. Classes with lower occurrence frequencies will receive larger weights, thus prompting the model to pay more attention to these low-frequency classes.

[0103] Specifically, during the training process, according to the calculated class weights, the model adjusts the model parameters according to the weighted cross-entropy loss function every time the parameters are updated. This weighting mechanism ensures the importance of the long-tail classes in training and prevents them from being overwhelmed by the number of samples of the mainstream classes. For example, in the task of classifying electronic medical records, some rare disease classes may have only a very small number of training samples. By assigning larger weights to these classes, the model can pay more attention to these classes, thereby improving its recognition ability.

[0104] In a possible implementation, if the classes in the training data are highly imbalanced, inverse frequency weighting can be used to more significantly increase the influence of low-frequency classes. For example, if the frequency of a certain class is 10%, its weight might be set to 10 times, thereby enhancing the weight of this class in model training and ensuring that it can have a greater impact on the training results.

[0105] As another option, during model training, the loss value calculated using the weighted cross-entropy loss function will update the model's parameters through the backpropagation algorithm. Specifically, the gradient of the model is calculated through the derivative of the loss function, and the backpropagation process will adjust the model parameters of the classes with larger weights according to the root weights of each class. In this way, the classes with larger weights will occupy a more important position in model training.

[0106] The derivative of the weighted cross-entropy loss function is calculated as follows:

[0107]

[0108] where, represents the gradient of the loss function L with respect to the model parameter θ, θ is the parameter of the model, is the gradient of the model prediction probability, y k is the label value of class k in the true label, p k is the prediction probability of the model for class k. Through backpropagation, the model will adjust the gradient of each parameter according to the weights of the classes, ensuring that the model fully considers each class during training, especially the performance of long-tail classes.

[0109] By using the weighted cross-entropy loss function, the present invention can effectively address the deficiencies of traditional classification methods when faced with class imbalance. Especially in the recognition of long-tail classes, it can greatly improve the accuracy of the model. The weighted cross-entropy loss function not only helps the model better understand the minority classes in the data but also dynamically balances the contributions of different classes to the final result during training, avoiding the overdominant role of high-frequency classes in model training.

[0110] In some embodiments, the weighted cross-entropy loss function can be combined with other optimization algorithms (such as the Adam optimizer) to further improve the training efficiency and model performance. Through this method, the present invention can maintain high-efficiency and high-accuracy classification performance when processing large-scale electronic medical record data, especially in the recognition of minority classes, significantly enhancing the application value of the model.

[0111] For step S5, the complete process of model training and optimization is introduced. By means of backpropagation of model parameters, selection of optimizers, and use of regularization techniques, the classification performance and generalization ability of the model are further improved.

[0112] Generally, in the training of deep learning models, the core of the training process is to calculate the gradient of the loss function through the backpropagation algorithm and update the model's parameters based on the gradient information. By continuously adjusting the model's parameters through the optimization algorithm, the model can gradually approach the optimal solution and finally achieve the classification task. In the present invention, the Adam optimizer is adopted in the optimization process, combined with techniques such as Dropout and L2 regularization to prevent overfitting and enhance the generalization ability of the model.

[0113] In this embodiment, the Adam optimizer is used for model training. The Adam optimizer is one of the most commonly used optimization algorithms in the field of deep learning at present. It combines the advantages of the momentum method and the adaptive learning rate and can achieve better optimization effects with fewer iterations.

[0114] The core idea of the Adam optimizer is to adaptively adjust the learning rate of each parameter by calculating the first-order moment estimate (i.e., the mean of the gradient) and the second-order moment estimate (i.e., the variance of the gradient) of the gradient of each parameter. Specifically, the update formula of the Adam optimizer is:

[0115]

[0116] where θ t represents the model parameters, and are the first-order moment estimate and the second-order moment estimate of the gradient respectively, η is the learning rate, and ∈ is a constant to prevent division by zero errors. Through this update formula, the Adam optimizer can adaptively adjust the learning rate of each parameter according to the first and second moments of the gradient during the training process, avoiding the trouble of manually adjusting the learning rate.

[0117] In some embodiments, the momentum factor and decay coefficient of the Adam optimizer can be adjusted according to the actual situation to further improve the training efficiency and effect.

[0118] As an option, in order to enhance the generalization ability of the model and avoid overfitting of the model to the training data, Dropout and L2 regularization techniques are also introduced in this embodiment.

[0119] Specifically, Dropout is an effective regularization method. By randomly "discarding" some neurons during the training process, it forces the model not to overly rely on certain features, thereby enhancing its adaptability to new data. In each round of training, Dropout will randomly ignore the outputs of some neurons, which can force the network to learn more generalizable features.

[0120] L2 regularization controls the complexity of the model by adding a penalty term to the weights of the model. The purpose of L2 regularization is to prevent the weights from being too large, thereby preventing the model from overfitting to the training data. The calculation formula for the L2 regularization term is:

[0121]

[0122] where λ is the regularization coefficient, θ i is the i-th parameter in the model, and N is the total number of parameters. By adding the L2 regularization term, the model will tend to select smaller weights during training, thereby avoiding over-reliance on certain features and improving the generalization ability of the model.

[0123] Feedback and adjustment during training are very important. During the training process, the performance of the model is usually evaluated based on the loss values and accuracies of the training set and the validation set. By monitoring these metrics, the model can gradually adjust hyperparameters such as the learning rate and the Dropout rate.

[0124] Specifically, after each training epoch, the model adjusts the learning rate according to the performance of the validation set. For example, if the accuracy on the validation set does not increase significantly within a certain number of epochs, the learning rate can be decreased to enable the model to optimize more precisely. Usually, learning rate decay methods can be used to gradually reduce the learning rate so that the model can converge more stably in the later stage of training.

[0125] Generally, during the training process, the dataset is first divided into a training set, a validation set, and a test set. In each round of training, the model evaluates the difference between the prediction results and the true labels by calculating the weighted cross-entropy loss function and updates the model parameters through the backpropagation algorithm. After each update, the validation set is used to evaluate the generalization performance of the model and determine whether early stopping or hyperparameter adjustment is needed.

[0126] In some embodiments, the Batch Normalization method is used during the training process, which further accelerates the convergence of the model and prevents the occurrence of gradient vanishing or explosion during training. Batch Normalization standardizes the input of each layer, keeping the input of each layer within an appropriate range during the network training process, thereby improving the stability and efficiency of training.

[0127] For step S6, a model evaluation process is introduced, and a series of evaluation metrics are used to comprehensively measure the performance of the model. By evaluating the performance on the training set and the validation set, it is possible to understand whether there is an overfitting phenomenon in the model, as well as the accuracy and stability of the model when dealing with complex and unseen samples. To comprehensively evaluate the model, commonly used evaluation metrics such as accuracy, recall, and F1-score are used to measure the classification ability of the model, especially in the identification of long-tail categories.

[0128] Generally, precision, recall, and F1-score are three commonly used metrics for evaluating the performance of classification models, and they can comprehensively reflect the classification effect of the model. For the task of electronic medical record classification, these metrics are particularly important because the model not only requires high accuracy but also needs to have good recognition ability for low-frequency categories in the case of class imbalance.

[0129] Specifically, accuracy is the most intuitive evaluation metric for classification effect, which measures the proportion of samples correctly classified by the model in the total number of samples. The formula for calculating accuracy is:

[0130]

[0131] where TP is the number of true positive samples, TN is the number of true negative samples, FP is the number of false positive samples, and FN is the number of false negative samples. Accuracy can reflect the overall classification performance of the model, but in the case of data class imbalance, accuracy cannot comprehensively measure the ability of the model.

[0132] As an option, when there is class imbalance in the data, recall becomes a more important evaluation metric. Recall reflects the ability of the model to identify positive (or minority) class samples, which is the ability of the model to identify all positive class samples. The formula for calculating recall is:

[0133]

[0134] where the definitions of TP and FN are the same as those described above. A higher recall indicates that the model can effectively identify long-tail categories and thus perform well in the classification of minority categories.

[0135] As a metric that combines accuracy and recall, the F1-score is particularly valuable in tasks with class imbalance. The F1-score is the harmonic mean of accuracy and recall, which can balance these two metrics to a certain extent and avoid the influence of extreme values of a certain metric on the overall performance. The formula for calculating the F1-score is:

[0136]

[0137] where Precision is the precision rate, which indicates how many of the samples predicted as positive classes are truly positive class samples. The formula for calculating it is:

[0138]

[0139] The F1 value combines the performance of the model in terms of accuracy and recall, and is suitable for comprehensive evaluation when the categories are unbalanced.

[0140] Generally, after each round of training, the model is evaluated on the validation set. The validation set is a dataset that has not been seen during model training and is used to test the performance of the model in actual applications. The model will evaluate its ability to classify various medical record defects by calculating indicators such as accuracy, recall, and F1 value.

[0141] Specifically, at the end of each training cycle, the model will make predictions on the validation set, compare them with the true labels, and calculate the loss of the prediction results. Through the loss function and evaluation indicators, the model can judge its performance on low-frequency categories. If the accuracy in the validation set is high, but the recall rate is low, it means that the model may perform poorly on a few categories. At this time, further measures can be taken to adjust the training strategy, such as adjusting the category weights, using more samples, or re-adjusting the feature extraction method.

[0142] As an option, the evaluation process can also use cross-validation technology for a more comprehensive model evaluation. Cross-validation divides the dataset into multiple subsets, takes turns using one of the subsets as the validation set, and uses the remaining subsets for training. In this way, the stability and accuracy of the model under different data partitions can be more comprehensively evaluated.

[0143] The ability to classify long-tail categories is an important evaluation goal. Electronic medical record data often suffers from class imbalance, where some categories have fewer samples and other categories have more samples. To evaluate the model's performance on long-tail categories, we can use a weighted evaluation method to give higher weights to minority categories to reflect their importance to the classification results. In this way, the model's classification effect on long-tail categories can be measured more accurately.

[0144] Specifically, in the long-tail category, F1 value and recall rate are indicators that need special attention. A high recall rate means that the model can identify more minority class samples, while a high F1 value means that the model is better at balancing precision and recall rate. Therefore, in the evaluation of long-tail categories, F1 value is used as a key performance indicator.

[0145] As an option, after each evaluation cycle, the model can also automatically adjust the hyperparameters of the training process according to the evaluation results. For example, if the recall rate of a certain category is low, the recognition ability of the model for this category can be improved by increasing the sample weight or the sample size of this category. In addition, during the training process, the learning rate can also be dynamically adjusted according to the changes in the evaluation metrics to ensure the stable convergence of the model during the optimization process.

[0146] Based on the evaluation, the performance of the model is further improved through various means, especially when dealing with long-tail categories, to ensure that the model has stronger generalization ability. Generally, the optimization process of the model not only depends on the results of training and evaluation, but also requires the adjustment of hyperparameters, the application of regularization methods, and the improvement of training strategies to ensure that the model can achieve the optimal effect in actual applications.

[0147] In this embodiment, hyperparameter adjustment is an important part of model optimization. In deep learning, the selection of hyperparameters has an important impact on the training speed and final effect of the model. Common hyperparameters include learning rate, Batch size, number of training epochs, Dropout rate, etc. The settings of these hyperparameters have a significant impact on the convergence speed, stability, and generalization ability of the model.

[0148] Specifically, in this embodiment, first, we optimize the hyperparameters through the GridSearch or RandomSearch method. These two methods can traverse multiple hyperparameter combinations to find the optimal hyperparameter configuration. For example, the selection of the learning rate is crucial for the convergence speed of the Adam optimizer. An overly large learning rate may cause the model to skip the optimal solution, while an overly small learning rate may cause the training to be too slow or even stagnate. Therefore, selecting a suitable learning rate range can effectively accelerate the convergence process of the model and improve the final performance.

[0149] As an option, another commonly used adjustment method is learning rate decay, that is, gradually reducing the learning rate during the training process. Learning rate decay helps to avoid model instability caused by an overly large learning rate in the later stage of training and can help the model finely adjust the parameters when approaching the optimal solution, thereby improving the performance of the model.

[0150] Regularization techniques are used to prevent overfitting and ensure that the model can perform well on unseen data. Regularization methods such as L2 regularization (weight decay) and Dropout are very effective means.

[0151] Specifically, L2 regularization restricts the magnitude of the model parameters by adding a penalty term of the sum of the squares of the weights to the loss function, thereby preventing the model from being too complex and overfitting the training data. The formula for L2 regularization is:

[0152]

[0153] where λ is the regularization coefficient, θ i represents the i-th parameter in the model, and N is the total number of parameters. By adding the L2 regularization term, the training process of the model will be penalized, avoiding excessive parameter values and reducing the risk of overfitting.

[0154] As another option, the Dropout technique is also an important means to prevent overfitting in the present invention. Dropout randomly discards some neurons during the training process, forcing the model not to overly rely on certain features, thereby enhancing its adaptability to new samples. In this way, Dropout effectively avoids overfitting of the model to the training data and improves the generalization ability of the model.

[0155] After each round of training and optimization, the performance evaluation results of the model will be used to feedback and adjust the strategy of the next round of training. According to the evaluation results, hyperparameters such as the learning rate, regularization term, and Dropout rate of the model may be dynamically adjusted. For example, when the recall rate of the model on certain categories is low, the sample weights of these categories may be increased or higher-intensity training may be performed to further improve the model's recognition ability for these categories.

[0156] Specifically, when the accuracy and recall rate of the model on the validation set reach a balance point, the training can be terminated early by the Early Stopping method to prevent the model from overfitting on the training set. The Early Stopping method will automatically stop the training when the performance on the validation set no longer improves, thereby saving computing resources and ensuring the generalization ability of the model.

[0157] Please refer to the appendix Figure 4 , the present invention also provides an electronic medical record defect text classification system based on feature fusion. This system combines the advantages of deep learning models and traditional machine learning algorithms. Through the extraction and weighted fusion of multi-modal features, it ensures the accurate recognition of long-tail categories and the efficient classification of complex medical record data, realizes high-precision and low-latency automatic classification of electronic medical record defects, and effectively improves the generalization ability of the model and the ability to process complex data, including the following modules:

[0158] The preprocessing module is responsible for the preliminary processing of the original electronic medical record data, including removing sensitive information, text tokenization, and word vector generation. This module performs the following tasks:

[0159] Removing Sensitive Information: Use natural language processing (NLP) techniques to automatically detect and remove parts of the medical record that involve patient privacy, such as names, ID numbers, phone numbers, etc., to ensure compliance with privacy protection laws and regulations. Text Tokenization: Apply existing Chinese tokenization tools (such as NLTK, spaCy, etc.) to tokenize the medical record. BERT Model to Generate Word Vectors: Generate context-aware word vectors for the medical record text through the BERT model, enabling the model to handle semantic ambiguities and understand the multi-level semantics of medical entities such as disease names and symptoms.

[0160] The feature extraction module aims to extract effective features from the preprocessed data, including the extraction of text features and structured data features. This module performs the following tasks:

[0161] BiLSTM to Extract Temporal Dependence Features: Use bidirectional long short-term memory networks (BiLSTM) to extract features from text data to capture the temporal relationships present in the text, especially time series information such as disease progression and symptom changes in the medical record. CNN to Extract Local Pattern Features: Use convolutional neural networks (CNN) to extract local patterns from the text to help the model identify fixed patterns or phrases, such as symptom descriptions and treatment methods for specific diseases. One-Hot Encoding to Convert Structured Data: Convert structured data (such as patient age, gender, past medical history, etc.) into numerical vectors through the One-Hot encoding method for easy fusion with text features.

[0162] The feature fusion module is used to fuse the features extracted from text and structured data to optimize the feature representation. This module performs the following tasks:

[0163] Weighted Feature Fusion: According to the results of mutual information calculation, perform weighted fusion on the extracted text features and structured data features. Mutual information calculation helps evaluate the correlation of each feature with the target class and weights the features according to the weights to ensure that important features have a higher influence after fusion. Feature Selection and Redundancy Removal: Use feature selection techniques (such as L1 regularization or information gain method) to screen the weighted features, remove redundant and irrelevant features, and improve the training efficiency and effect of the model. Adaptive Weight Adjustment Mechanism: Dynamically adjust the weights of each feature according to the feedback of the model during the training process, especially for the optimization of low-frequency categories, to enhance the recognition ability of long-tail categories.

[0164] The classification module is the core part of the system and is trained and classified using the weighted cross-entropy loss function and optimization algorithm. This module performs the following tasks:

[0165] Training with Weighted Cross-Entropy Loss Function: The model is trained using the weighted cross-entropy loss function, which dynamically adjusts the class weights according to the frequency of classes in the training data, enabling the model to pay more attention to the long-tail classes. Model Optimization: Combining the Adam optimizer, multiple backpropagation is performed to optimize the model parameters. During training, the learning rate is dynamically adjusted, and Dropout and L2 regularization are combined to enhance the generalization ability of the model. Optimization of Long-Tail Classes: For the long-tail classes in the training data, the system adaptively adjusts the weights to ensure that the low-frequency classes receive sufficient training resources, thereby improving the recognition accuracy of the minority classes.

[0166] The evaluation module is used to evaluate the performance of the trained model and provide feedback and optimization for the model through various metrics. This module performs the following tasks:

[0167] Accuracy, Recall, and F1-Score Evaluation: Evaluate metrics such as the accuracy, recall, and F1-score of the model to comprehensively measure the performance of the model, especially in the case of class imbalance. Evaluate the performance of the model in identifying various medical record defects through these metrics. Long-Tail Class Evaluation: Specifically evaluate the recognition effect of long-tail classes through a weighted evaluation method to ensure that the classification accuracy of these classes is not lower than that of common classes. Performance Feedback and Optimization: According to the evaluation results, dynamically adjust the hyperparameters (such as the learning rate, class weights, etc.) during the training process, and apply an early stopping strategy during training to prevent overfitting.

[0168] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for classifying defect text of electronic medical records based on feature fusion, characterized in that: The following steps are involved: Preprocess the electronic medical record text defect data, including removing sensitive information, performing text segmentation, and using the pre-trained BERT model to generate context-related word vectors; Extract features of electronic medical record text defect data, including using BiLSTM to extract temporal dependency features, using CNN to extract local pattern features, and using One-Hot encoding to convert structured data into vectors; Calculate the mutual information between each feature and the target category, and weight each feature according to the mutual information value. Fuse the weighted defect data text features with the structured data features to obtain a fused feature vector. The model is trained using a weighted cross entropy loss function that dynamically calculates category weights based on the frequency of the categories, giving larger weights to long-tail categories; The Adam optimizer is used for optimization during the training process, and Dropout and L2 regularization are combined to enhance the generalization ability of the model; The model was evaluated using precision, recall, and F1 value as performance evaluation indicators to ensure that the model had high classification accuracy for low-frequency categories.

2. According to the feature fusion-based electronic medical record defect text classification method of claim 1, it is characterized in that: The text segmentation includes using a natural language processing tool to segment the medical record defect data text.

3. The electronic medical record defect text classification method based on feature fusion according to claim 1 is characterized in that: The mutual information calculation formula is: I(X i ;Y)=H(X i )+H(Y)-H(X i ,Y) Among them, X i is the i-th feature, Y is the target category, H(X i ) and H(Y) are the entropy of features and categories respectively, H(X i ,Y) is the joint entropy of features and categories, I(X i ; Y) represents the i-th feature X i The mutual information between the target category Y.

4. The electronic medical record defect text classification method based on feature fusion according to claim 1 is characterized in that: The step of fusing the weighted defect data text features with the structured data features further includes the following steps: The defect data text features and structured data features are weighted and fused according to their importance to ensure that the result of feature fusion retains the key information to the maximum extent; After feature fusion, feature selection methods are used to further filter and compress irrelevant and redundant features to improve model training efficiency; Introducing an adaptive weight adjustment mechanism to dynamically adjust the weight of each feature based on model feedback during training to optimize the recognition of low-frequency categories; The fused features after feature weighting and selection optimization are input into the classifier for the final classification of medical record defects.

5. The electronic medical record defect text classification method based on feature fusion according to claim 4 is characterized in that: The formula for fusing the weighted features with the structured data features is: Among them, X i is the i-th feature, w i is the weight coefficient of the feature, X fused is the fused feature vector, and M is the number of features.

6. The electronic medical record defect text classification method based on feature fusion according to claim 1 is characterized in that: The calculation process of the weighted cross entropy loss function includes the following steps: Calculate the weight of each category. The weight is dynamically adjusted according to the frequency of the category in the training data. The weight of the long-tail category is relatively large. The weighted cross entropy loss function is used to calculate the difference between the model output and the true label; For each training sample, the loss is calculated according to its category label, and the loss of long-tail categories is larger; The loss values ​​of all samples are weighted and summed, and the model parameters are adjusted through back propagation to improve the model's recognition accuracy for low-frequency categories.

7. The electronic medical record defect text classification method based on feature fusion according to claim 6 is characterized in that: The calculation formula of the weighted cross entropy loss function is: Among them, w k is the weight of category k, y k is the category label, p k is the predicted probability of category k, and K is the number of categories.

8. The electronic medical record defect text classification method based on feature fusion according to claim 1 is characterized in that: The update formula of the Adam optimizer is: Among them, θ t are model parameters, and are the first-order moment estimate and the second-order moment estimate of the gradient, η is the learning rate, ∈ is a constant to prevent division by zero errors, and θ t-1 represents the model parameters at time step t-1.

9. The electronic medical record defect text classification method based on feature fusion according to claim 1 is characterized in that: The model evaluation comprises the following steps: The proportion of samples correctly classified by the calculation model to the total samples reflects the overall classification effect; Calculate the proportion of positive samples correctly predicted by the model to all positive samples to evaluate the model's ability to recognize low-frequency categories; Combine precision and recall to evaluate the overall performance of the model, especially in the case of class imbalance; Adjust hyperparameters during training to optimize model performance, especially to improve recognition accuracy of low-frequency categories.

10. A system for classifying defect text of electronic medical records based on feature fusion, applied to a method for classifying defect text of electronic medical records based on feature fusion as claimed in any one of claims 1 to 9, characterized in that: include: Preprocessing module: used to desensitize and segment electronic medical record defect data, and generate word vectors through BERT; Feature extraction module: used to extract defect data text features and structured data features, including BiLSTM layer, CNN layer and structured data processing; Feature fusion module: used to perform mutual information optimization weighted fusion on the extracted defect data text features; Classification module: used to train and optimize the model using the weighted cross entropy loss function and calculate the classification results; Evaluation module: used to evaluate model performance and output accuracy, recall and F1 value indicators.

Citation Information

Cited By

  • Deep learning prediction method for rock freeze-thaw damage based on text embedding

    CN120470950A