Full-life-cycle resident electronic health record data processing method and system
By preprocessing and formatting the received residents' electronic health record data, and using NLP technology and machine learning models for label classification and semantic correlation, the problems of inefficiency and inadequacy in the existing technology are solved, and efficient and accurate health record data processing and label classification are achieved.
Patent Information
- Application Number
- CN202510410135.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art is inefficient in generating personal health marks, prone to human errors, and is unable to adapt to changing health data and disease classification standards, resulting in insufficient flexibility and adaptability of the system.
By preprocessing and formatting the received residents' electronic health profile data, health features are extracted to form health profile feature vectors, text data is processed using NLP technology and integrated with structured data, label classification and semantic association are used using machine learning models and graph convolution networks, label relationship vectors are generated, and the model is optimized through iterative training.
It improves the efficiency and accuracy of health file data processing, enhances the adaptability and flexibility of the system, and can adjust the label classification standards synchronized with medical research progress in real time to ensure the uniqueness and continuous update of health file.
Smart Images

Figure CN119993362A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical information processing, and in particular to a method and system for processing electronic health record data of residents throughout their life cycle. Background Art
[0002] In health record management, personal health identification services, especially those for key groups, are very important. Personal health identification is a native objective label generated in accordance with unified technical standards and business rules in various relevant business application systems (such as chronic disease management, maternal and child health care, etc.), which can be automatically pushed to the homepage of residents' electronic health records. It has the characteristics of dynamic update, non-arbitrary change and traceability. Personal health identification can be used to accurately identify key groups and provide proactive health management. However, the existing native objective labels generated according to unified technical standards and business rules are not only inefficient, but also prone to human errors, especially when processing large amounts of residents' health information.
[0003] Although some existing classification systems can usually be based on simple keyword matching and static rules, they are still unable to adapt to the ever-changing health data and disease classification standards, resulting in insufficient flexibility and adaptability of the system. For example, the diagnostic criteria for a disease may change with the progress of medical research, and the traditional label classification method cannot be quickly adjusted to adapt to these changes, resulting in a disconnect between health data and actual needs.
[0004] In addition, electronic health records are not limited to traditional disease diagnosis and treatment records, but also include more health information, such as living habits, environmental factors, exercise status, etc. Among these rich information, achieving the uniqueness of health record creation has become a difficult problem in the medical and health information system. Summary of the invention
[0005] In view of the deficiencies in the prior art, the present invention provides a method and system for processing electronic health record data of residents throughout their life cycle, which solves the problems mentioned in the background technology.
[0006] To achieve the above objectives, the present invention is implemented through the following technical solutions: a method for processing electronic health record data of residents throughout their life cycle, comprising the following steps:
[0007] S1, preprocessing and formatting the received resident electronic health record data, and then extracting the resident health characteristics to form a health record feature vector F1;
[0008] S2, processing the received health record feature vector F1 using NLP technology, including word segmentation, part-of-speech tagging and named entity recognition, and integrating it with the health record feature vector F1 to reorganize it into a health reorganized feature vector F2;
[0009] S3, training the pre-built machine learning model by acquiring historical residents' electronic health record data, and using the trained machine learning model to classify the health reorganization feature vector F2 to generate health record classification results Cha with several labels;
[0010] S4, by using the graph convolutional network to perform semantic association and relationship extraction on the labels in the health record classification result Cha, obtain the potential association between different labels, and form a label relationship vector Rha;
[0011] S5. The process of iteratively training the machine learning model and graph convolutional network method through the new resident electronic health record data received at a fixed period, and updating the label relationship vector Rha.
[0012] Preferably, said S1 includes S11 and S12;
[0013] S11. The received resident electronic health record data is a heterogeneous data set from multiple sources, including structured data and unstructured data. The resident electronic health record data is preprocessed, including data cleaning preprocessing, data denoising preprocessing, data supplementation missing value preprocessing and data formatting preprocessing. In the data cleaning preprocessing process, duplicate resident electronic health record data will be eliminated, and the data formatting preprocessing will unify the resident electronic health record data from multiple sources into standard structured data to form the initial health record data set F0;
[0014] The structured data includes diagnosis results, physical examination data and hospitalization records, and the unstructured data includes medical records, medical reports and doctors' notes.
[0015] Preferably, S12, performing a feature extraction phase on the preprocessed initial health record data set F0, by extracting health features from the formatted initial health record data set F0, wherein the extracted health features include extracting numerical field information from the structured data; extracting keywords and topics from the unstructured data, and marking numerical features for the keywords and topics, wherein the numerical features include quantity and length;
[0016] Among them, the numerical field information includes age, weight, blood sugar and blood pressure;
[0017] By integrating the extracted health features, the health record feature vector F1 is obtained.
[0018] Preferably, S2 includes S21 and S22;
[0019] S21, processing the text data in the received health record feature vector F1, performing word segmentation, part-of-speech tagging and named entity tagging on the text data by using NLP technology, and marking it as a text feature vector Ftext;
[0020] The text feature vector Ftext is obtained by the following NLP processing formula:
[0021] Ftext={Tokens(Q),POSTags(Q),Entities(Q)};
[0022] In the formula, Q represents text data, Tokens(Q) represents the set of word segmentation results for text data, POSTags(Q) represents the set of part-of-speech tagging results for text data, and Entities(Q) represents the set of named entity results for text data;
[0023] The word segmentation is to divide the text data into words and phrases, and then divide the long text in the text data into short text;
[0024] The part-of-speech tagging is used to mark the parts of speech of the divided words and phrases, and the parts of speech include nouns, verbs and adjectives, which are used to reflect the structure and meaning of the text data;
[0025] The named entities are used to identify entity names in text data, and the entity names include disease names, drug names, and medical institutions.
[0026] Preferably, S22, integrating the acquired text feature vector Ftext with the health record feature vector F1 to reorganize into a health reorganized feature vector F2;
[0027] The healthy reorganization feature vector F2 is obtained by the following integration formula:
[0028]
[0029] In the formula, It represents adjacent concatenation operations, specifically, the text feature vector Ftext and the health record feature vector F1 are merged into a new vector.
[0030] Preferably, said S3 includes S31;
[0031] S31, training the machine learning model by using the stored historical resident electronic health record data, and after the training is completed, inputting the health reorganization feature vector F2 and the label set Cinit into the machine learning model for classification, and generating health record classification results Cha with several labels for the health reorganization feature vector F2;
[0032] Among them, the label set Cinit is composed of the stored historical residents' electronic health record data and health labels preset by medical experts, including high blood sugar and heart disease;
[0033] The health record classification result Cha is obtained by the following training method:
[0034] Cha=M(F2,Cinit);
[0035] Where M represents the machine learning model.
[0036] Preferably, the S4 includes S41;
[0037] S41, converting the health record classification result Cha into a graph structure, wherein each node in the graph structure represents a label in the health record classification result Cha, and constructing edges between nodes according to the similarity and co-occurrence probability between the labels;
[0038] The edges between labels are then weighted and information is transferred through the graph convolutional network. Each label node exchanges information with adjacent label nodes through the connections in the graph, thereby realizing the semantic association between labels and generating the enhanced label representation Hv(l+1) after semantic enhancement.
[0039] The enhanced label indicates that Hv(l+1) is obtained by the following calculation formula:
[0040]
[0041] Wherein, Hv(l+1) represents the enhanced label representation of node v in the l+1 layer, specifically represents the feature vector of node v in the graph convolutional network. Each layer of convolution operation will update the representation of the node by transferring and aggregating the information of neighboring nodes. σ represents the nonlinear activation function, including ReLU and Sigmoid activation function. u∈N(v) represents that node u is in the neighbor set of node v. N(v) represents the neighbor set of node v, that is, the neighboring nodes connected to node v. A in Avu represents the adjacency matrix, specifically represents the element Avu of the adjacency matrix, representing the connection relationship between node v and node u. Among them, Avu=1 represents that node v and node u are connected by an edge, Avu=0 represents that node v and node u are not connected by an edge, dv and du represent the degrees of node v and node u respectively, that is, the number of neighbors of node v and node u. W(l) represents the weight matrix of the lth layer. hu(l) represents the representation of node u in the lth layer. It is the output of the graph convolution operation of the previous layer and represents the feature vector of node u.
[0042] Preferably, the S4 includes S42, calculating the similarity R(i, j) between the tags based on the acquired enhanced tag representation Hv(l+1), reflecting the similarity between the tag i and the tag j in the space vector, and forming a tag relationship vector Rha by aggregating the similarity R(i, j) between each tag;
[0043] The similarity R(i, j) is obtained by the following calculation formula:
[0044]
[0045] Where H(i) and H(j) represent the enhanced label representations of label i and label j, respectively.
[0046] Preferably, the S5 includes S51;
[0047] S51. Compare the actual label relationship provided by the doctor with the new resident electronic health record data received at a fixed period to verify whether the label relationship vector Rha is consistent with the actual medical situation. When the label relationship vector Rha is verified to be inconsistent with the actual medical situation, iteratively train the machine learning model and the graph convolutional network method, and update the acquisition process of the label relationship vector Rha until the fixed period verifies that the label relationship vector Rha is consistent with the actual medical situation, and stop iteratively training the machine learning model and the graph convolutional network method.
[0048] Among them, the iterative training of machine learning models and graph convolutional network methods includes adjusting the weight matrix W and the number and depth of graph convolutional layers.
[0049] A full-life cycle resident electronic health record data processing system, comprising an archive data processing module, an archive data identification module, an archive data training and analysis module, an archive label generation module and an iterative optimization module;
[0050] The archive data processing module preprocesses and formats the received resident electronic health archive data, and then extracts the resident health characteristics to form a health archive feature vector F1;
[0051] The archive data recognition module processes the received health archive feature vector F1 using NLP technology, including word segmentation, part-of-speech tagging and named entity recognition, and integrates it with the health archive feature vector F1 to reorganize it into a health reorganized feature vector F2;
[0052] The file data training and analysis module trains the pre-built machine learning model by acquiring historical residents' electronic health file data, and uses the trained machine learning model to classify the health reorganization feature vector F2 to generate health file classification results Cha with several labels;
[0053] The file label generation module uses a graph convolutional network to perform semantic association and relationship extraction on the labels in the health file classification result Cha, obtains the potential association between different labels, and forms a label relationship vector Rha;
[0054] The iterative optimization module iteratively trains the machine learning model and graph convolutional network method through new resident electronic health record data received at a fixed period, and updates the acquisition process of the label relationship vector Rha.
[0055] The present invention provides a method and system for processing electronic health record data of residents throughout their life cycle, which has the following beneficial effects:
[0056] (1) By converting unstructured data into structured data and integrating it with the health record feature vector F1 into the health reorganization feature vector F2, the problem that the static rules based on keyword matching in the existing system cannot adapt to the ever-changing health data and disease classification standards is effectively compensated. In addition, the use of machine learning models to classify the health reorganization feature vector F2 and automatically assign labels to health records not only improves the accuracy of classification, but also overcomes the limitations of low efficiency of traditional methods. Through the semantic association and relationship extraction of labels through graph convolutional networks, the system can identify and mine the potential associations between labels and generate label relationship vectors Rha, making the system more adaptable and flexible, and able to adjust the label classification standards in real time in line with the progress of medical research. Finally, by regularly receiving new residents' electronic health record data and optimizing the machine learning model and graph convolutional network method through iterative training, the accuracy and real-time performance of the label relationship vector Rha are continuously improved, solving the fragmentation of traditional health records. After data sorting, it becomes a globally unique health record. The continuously added health records can still achieve distributed uniqueness.
[0057] (2) By forming the health record feature vector F1, the content of the entire health record can more clearly and accurately reflect the individual's health status. Furthermore, NLP technology is used to perform word segmentation, part-of-speech tagging and named entity recognition on text data. The extracted text feature vector Ftext can deeply mine key health information in the text, such as disease names, drug names and medical institutions, so that unstructured data can be fully utilized. Finally, the text feature vector Ftext is integrated with the health record feature vector F1 to generate the health reorganization feature vector F2. This fusion data representation method provides richer and more accurate data support for subsequent machine learning and graph convolutional network models, which helps to better reveal the potential connections between different health record data.
[0058] (3) Through machine learning training of the health reorganization feature vector F2 and the label set Cinit, the health record classification results Cha with multiple labels are generated, which effectively solves the challenge of single label classification and inability to handle complex health conditions in traditional health record systems. In particular, through the multi-label classification method, each health record can be assigned multiple labels according to the multiple health conditions it involves, thereby providing more comprehensive and refined health information, which can be adjusted in time to cope with changes in medical knowledge and clinical practice. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] Figure 1 A schematic diagram of the steps of a method for processing electronic health records data of residents throughout their life cycle according to the present invention;
[0060] Figure 2 This is a schematic diagram of a block diagram of a resident electronic health record data processing system for the entire life cycle of the present invention. DETAILED DESCRIPTION
[0061] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0062] Example 1
[0063] The present invention provides a method for processing electronic health records data of residents throughout their life cycle. Figure 1 , including the following steps:
[0064] S1, preprocessing and formatting the received resident electronic health record data, and then extracting the resident health characteristics to form a health record feature vector F1;
[0065] S2, processing the received health record feature vector F1 using NLP technology, including word segmentation, part-of-speech tagging and named entity recognition, and integrating it with the health record feature vector F1 to reorganize it into a health reorganized feature vector F2;
[0066] S3, training the pre-built machine learning model by acquiring historical residents' electronic health record data, and using the trained machine learning model to classify the health reorganization feature vector F2 to generate health record classification results Cha with several labels;
[0067] S4, by using the graph convolutional network to perform semantic association and relationship extraction on the labels in the health record classification result Cha, obtain the potential association between different labels, and form a label relationship vector Rha;
[0068] S5. The process of iteratively training the machine learning model and graph convolutional network method through the new resident electronic health record data received at a fixed period, and updating the label relationship vector Rha.
[0069] In this embodiment, by preprocessing and formatting the received resident electronic health record data and extracting the health record feature vector F1 from it, not only the efficiency of data processing is improved, but also the standardization and accuracy of a large amount of health data are ensured, thereby avoiding errors caused by manual input or fixed rules. Secondly, the NLP technology is used to process the health record feature vector F1 for word segmentation, part-of-speech tagging and named entity recognition, converting unstructured data into structured data, and integrating it with the health record feature vector F1 into a health reorganization feature vector F2, which effectively makes up for the problem that the static rules based on keyword matching in the existing system cannot adapt to the changing health data and disease classification standards. In addition, the machine learning model is used to classify the health reorganization feature vector F2 and automatically assign labels to health records, which not only improves the accuracy of classification, but also overcomes the limitations of low efficiency of traditional methods. Through the graph convolutional network for semantic association and relationship extraction of labels, the system can identify and mine the potential associations between labels and generate label relationship vectors Rha, so that the system has stronger adaptability and flexibility, and can adjust the label classification standards in real time in synchronization with the progress of medical research. Finally, by regularly receiving new residents' electronic health record data and optimizing the machine learning model and graph convolutional network method through iterative training, the accuracy and real-time performance of the label relationship vector Rha are continuously improved. Overall, this method not only effectively improves the automation and accuracy of data processing and classification, but also enhances the system's ability to respond to emerging health information, solves the problems of low efficiency, poor adaptability, and high error rate in traditional systems, and lays a solid foundation for the further development of future medical and health information systems.
[0070] Example 2
[0071] This embodiment is explained in Example 1, please refer to Figure 1 , specifically: the S1 includes S11 and S12;
[0072] S11. The received resident electronic health record data is a heterogeneous data set from multiple sources, including structured data and unstructured data. The resident electronic health record data is preprocessed, including data cleaning preprocessing, data denoising preprocessing, data supplementation missing value preprocessing and data formatting preprocessing. In the data cleaning preprocessing process, duplicate resident electronic health record data will be eliminated, and the data formatting preprocessing will unify the resident electronic health record data from multiple sources into standard structured data to form the initial health record data set F0;
[0073] The structured data includes diagnosis results, physical examination data and hospitalization records, and the unstructured data includes medical records, medical reports and doctors' notes.
[0074] S12, performing a feature extraction phase on the preprocessed initial health record data set F0, by extracting health features from the formatted initial health record data set F0, wherein the extracted health features include extracting numerical field information from the structured data; extracting keywords and topics from the unstructured data, and marking numerical features for the keywords and topics, wherein the numerical features include quantity and length;
[0075] Among them, the numerical field information includes age, weight, blood sugar and blood pressure;
[0076] By integrating the extracted health features, the health record feature vector F1 is obtained.
[0077] The S2 includes S21 and S22;
[0078] S21, processing the text data in the received health record feature vector F1, performing word segmentation, part-of-speech tagging and named entity tagging on the text data by using NLP technology, and marking it as a text feature vector Ftext;
[0079] The text feature vector Ftext is obtained by the following NLP processing formula:
[0080] Ftext={Tokens(Q),POSTags(Q),Entities(Q)};
[0081] In the formula, Q represents text data, Tokens(Q) represents the set of word segmentation results for text data, POSTags(Q) represents the set of part-of-speech tagging results for text data, and Entities(Q) represents the set of named entity results for text data;
[0082] The word segmentation is to divide the text data into words and phrases, and then divide the long text in the text data into short text;
[0083] The part-of-speech tagging is used to mark the parts of speech of the divided words and phrases, and the parts of speech include nouns, verbs and adjectives, which are used to reflect the structure and meaning of the text data;
[0084] The named entities are used to identify entity names in text data, and the entity names include disease names, drug names, and medical institutions.
[0085] S22, integrating the acquired text feature vector Ftext with the health record feature vector F1 to reorganize into a health reorganized feature vector F2;
[0086] The healthy reorganization feature vector F2 is obtained by the following integration formula:
[0087]
[0088] In the formula, It represents adjacent concatenation operations, specifically, the text feature vector Ftext and the health record feature vector F1 are merged into a new vector.
[0089] In this embodiment, by performing data cleaning, denoising, supplementing missing values and formatting preprocessing on heterogeneous data sets from different sources, the structured data and unstructured data are unified into a standard health record data set F0, which effectively solves the problem of inconsistent data from different medical systems and platforms, thereby improving the efficiency and consistency of data processing. Then, by extracting numerical field information from structured data, extracting keywords and topics from unstructured data, and marking features such as quantity and length, a health record feature vector F1 is formed, so that the content of the entire health record can more clearly and accurately reflect the health status of the individual. Further, NLP technology is used to perform word segmentation, part-of-speech tagging and named entity recognition on text data, and the extracted text feature vector Ftext can deeply mine key health information in the text, such as disease names, drug names and medical institutions, so that unstructured data can be fully utilized. Finally, the text feature vector Ftext is integrated with the health record feature vector F1 to generate a health reorganization feature vector F2. This fusion data representation method provides more abundant and accurate data support for subsequent machine learning and graph convolutional network models, which helps to better reveal the potential connection between different health record data. Overall, this method not only improves the accuracy of data analysis through comprehensive and multi-dimensional processing of health data, but also enhances the flexibility and adaptability of the system in processing complex health information.
[0090] Example 3
[0091] This embodiment is explained in Example 2. Please refer to Figure 1 , specifically: the S3 includes S31;
[0092] S31, training the machine learning model by using the stored historical resident electronic health record data, and after the training is completed, inputting the health reorganization feature vector F2 and the label set Cinit into the machine learning model for classification, and generating health record classification results Cha with several labels for the health reorganization feature vector F2;
[0093] Among them, several labels are generated because the health reorganization feature vector F2 of a health record may involve multiple diseases or health states, so a multi-label classification method is needed, and the model will assign multiple labels to each health record;
[0094] Among them, the label set Cinit is composed of the stored historical residents' electronic health record data and health labels preset by medical experts, including high blood sugar and heart disease;
[0095] The training includes a data input step, a parameter initialization step, a loss function step, an optimization step, a classification output step and a multi-label classification step;
[0096] Among them, the data input step: the stored historical resident electronic health record data, health reorganization feature vector F2 and label set Cinit are used as input data to the machine learning model;
[0097] Initialization parameter step: Initialize the weight vector and bias term;
[0098] Loss function step: After training, the machine learning model evaluates the error between its predicted health record classification results and the true label by calculating the loss function. The loss function usually uses the cross entropy loss function.
[0099] Optimization step: After training is completed, the weight vector and bias term are continuously adjusted to minimize the loss function through optimization methods such as backpropagation and gradient descent. At each training iteration, the machine learning model calculates the gradient of the loss and uses an optimization algorithm, such as the stochastic gradient descent (SGD) optimization algorithm;
[0100] Classification output step: Once the training is completed, the machine learning model will perform a multi-label classification step on the input health reorganization feature vector F2 according to the learned weight vector and bias term, and output the health record classification result Cha;
[0101] Multi-label classification step: Since each health reorganization feature vector F2 may involve multiple health problems, a multi-label classification method is used. The prediction result of each label is converted into a probability value through the Sigmoid function, and finally the set threshold (usually 0.5) is used to determine whether each label is related to the health record;
[0102] The health record classification result Cha is obtained by the following training method:
[0103] Cha=M(F2,Cinit);
[0104] In the formula, M represents the machine learning model, including support vector set SVM model, random forest model and neural network model;
[0105] Among them, M(F2, Cinit) specifically outputs an activation value through the machine learning model, which represents the preliminary prediction value of each label. Then, through the activation value, the Sigmoid function is used to convert the activation value into a probability value, obtain the predicted probability value of the label, and then integrate all the labels in the healthy reorganization feature vector F2;
[0106] The conversion method is as follows:
[0107] Z(j)=m(j) T *F2+b(j);
[0108] Wherein, Z(j) represents the activation value corresponding to the j-th label, specifically the initial prediction value of the machine learning model for the j-th label, indicating whether the label is related to the health record; m(j) represents the weight vector of the j-th label, specifically indicating that the machine learning model will learn the weight of each label, and this weight also determines the strength of the relationship between the label and the input feature; b(j) represents the bias term of the j-th label, specifically a constant term in the machine learning model; T represents the transpose operation.
[0109] By inputting the health reorganization feature vector F2 and the label set Cinit into the machine learning model for training, the final generated health record classification result Cha assigns multiple labels to each health record. The advantage of this multi-label classification method is that it can simultaneously identify multiple health states or diseases involved in a health record, thereby achieving more accurate health status analysis. This method not only improves the flexibility and adaptability of the model, but also reflects the complex relationship between multiple health problems.
[0110] The S4 includes S41;
[0111] S41, converting the health record classification result Cha into a graph structure, wherein each node in the graph structure represents a label in the health record classification result Cha, and constructing edges between nodes according to the similarity and co-occurrence probability between the labels;
[0112] The edges between labels are then weighted and information is transferred through the graph convolutional network. Each label node exchanges information with adjacent label nodes through the connections in the graph, thereby realizing the semantic association between labels and generating the enhanced label representation Hv(l+1) after semantic enhancement.
[0113] The enhanced label indicates that Hv(l+1) is obtained by the following calculation formula:
[0114]
[0115] Where Hv(l+1) represents the enhanced label representation of node v at the l+1 layer, specifically the feature vector of node v in the graph convolutional network. Each layer of convolution operation updates the representation of the node by transferring and aggregating the information of neighboring nodes. Hv(l) is the representation of node v at the lth layer, that is, the feature representation of node v at the previous layer. By repeatedly applying convolution operations, the representation of the node is enhanced and updated layer by layer. σ represents the nonlinear activation function, including ReLU and Sigmoid activation functions. The purpose is to introduce nonlinearity so that the network can learn more complex features. It acts on the aggregated neighbor information to update the representation of node v. u∈N(v) means that node u is in the neighbor set of node v. N(v) represents the neighbor set of node v, that is, node v is connected to The A in Avu represents the adjacency matrix, specifically the element Avu of the adjacency matrix, which represents the connection relationship between node v and node u, where Avu=1 means that node v and node u are connected by an edge, Avu=0 means that node v and node u are not connected by an edge, dv and du represent the degrees of node v and node u respectively, that is, the number of neighbors of node v and node u, W(l) represents the weight matrix of the lth layer. Each layer of graph convolutional network has a corresponding weight matrix used to adjust and learn the feature information of each node. hu(l) represents the representation of node u in the lth layer. It is the output of the graph convolution operation of the previous layer, representing the feature vector of node u. Through this feature vector, node v can obtain information from node u, thereby updating its own representation.
[0116] The S4 includes S42, calculating the similarity R(i, j) between the tags based on the acquired enhanced tag representation Hv(l+1), reflecting the similarity between the tag i and the tag j in the space vector, and forming a tag relationship vector Rha by aggregating the similarity R(i, j) between each tag;
[0117] The similarity R(i, j) is obtained by the following calculation formula:
[0118]
[0119] Where H(i) and H(j) represent the enhanced label representations of label i and label j, respectively.
[0120] By converting the health record classification result Cha into a graph structure and processing it using a graph convolutional network, the potential semantic associations between labels can be effectively captured. Each label is represented by a node in the graph, and the edges between nodes are constructed based on label similarity and the probability of co-occurrence. The graph convolutional network updates the feature representation of label nodes through multi-layer information transmission and aggregation, and gradually enhances the association information between labels, thereby generating an enhanced label representation Hv(l+1) after semantic enhancement. This graph structure construction and information transmission mechanism not only strengthens the relationship between labels, but also learns complex label features through nonlinear activation functions, improving the performance and adaptability of the model. In addition, by calculating the similarity R(i,j) between labels and generating a label relationship vector Rha, the model can quantify the similarity between labels, further improving the accuracy of label classification and the intelligent level of health record management. Through multi-layer information dissemination, the connection between complex health states is strengthened, providing stronger support for personalized health prediction and multi-dimensional health risk assessment.
[0121] The S5 includes S51;
[0122] S51. Compare the actual label relationship provided by the doctor with the new resident electronic health record data received at a fixed period to verify whether the label relationship vector Rha is consistent with the actual medical situation. When the label relationship vector Rha is verified to be inconsistent with the actual medical situation, iteratively train the machine learning model and the graph convolutional network method, and update the acquisition process of the label relationship vector Rha until the fixed period verifies that the label relationship vector Rha is consistent with the actual medical situation, and stop iteratively training the machine learning model and the graph convolutional network method.
[0123] Among them, iterative training of machine learning models and graph convolutional network methods includes adjusting the weight matrix W and the number and depth of graph convolutional layers;
[0124] Among them, the weight matrix W is adjusted continuously through the back-propagation algorithm and the gradient descent method. By minimizing the loss function, the model automatically updates the weight matrix to make the relationship between features and labels more accurate. Specifically, after each training, the gradient of the loss function is calculated, and the value of the weight matrix is adjusted according to the learning rate to reduce the prediction error;
[0125] Adjusting the number and depth of graph convolution layers is usually chosen during the design phase of the network. By increasing the number of graph convolution layers, the network can capture more complex semantic associations between labels, especially when the label relationships are more complex. Increasing the number of layers may bring more information propagation and feature learning, but it may also lead to overfitting. Therefore, it is necessary to adjust the appropriate depth based on the performance of the validation set after training. In addition, when the depth of the graph convolution layer increases, regularization techniques are usually used to avoid overfitting and improve the generalization ability of the model. After training, these parameters are continuously adjusted through methods such as cross-validation, and the optimal configuration is selected by evaluating the performance of the model on the validation set, thereby obtaining the best model that can capture complex label relationships and avoid overfitting.
[0126] In this embodiment, by performing machine learning training on the health reorganization feature vector F2 and the label set Cinit, a health record classification result Cha with multiple labels is generated, which effectively solves the challenge of single label classification and inability to handle complex health states in the traditional health record system. In particular, through the multi-label classification method, each health record can be assigned multiple labels according to the multiple health conditions involved, thereby providing more comprehensive and refined health information. Next, the label set Cinit is converted into a graph structure, and by constructing edges between labels and performing weighted processing, the graph convolution network can transfer information between label nodes, further semantically enhancing the label representation Hv(l+1). This process ensures that the potential semantic connections between labels can be fully mined and improves the semantic consistency between labels. Further, the similarity R(i, j) between labels is calculated, and the label relationship vector Rha is formed according to the label similarity aggregation, which provides effective semantic support for further analysis and association of health records. Finally, by regularly receiving new resident electronic health record data and comparing it with the actual label relationship provided by the doctor, the accuracy of the label relationship vector Rha can be verified, and the model can be optimized through the feedback mechanism to ensure the applicability and accuracy of the label relationship vector Rha in actual medical situations. Through continuous iterative training and optimization, including adjusting the weight matrix W and the number and depth of graph convolutional layers, this method can continuously improve the accuracy and stability of the model, and ultimately achieve accurate and efficient health data management and analysis. Overall, this method not only improves the intelligent processing capabilities of health record labels, but also enhances the adaptability and flexibility of the model, and can be adjusted in time to cope with changes in medical knowledge and clinical practice.
[0127] Example 4
[0128] A data processing system for electronic health records of residents throughout their life cycle, please refer to Figure 2,Specifically: including archive data processing module, archive data identification module, archive data training and analysis module, archive label generation module and iterative optimization module;
[0129] The archive data processing module preprocesses and formats the received resident electronic health archive data, and then extracts the resident health characteristics to form a health archive feature vector F1;
[0130] The archive data recognition module processes the received health archive feature vector F1 using NLP technology, including word segmentation, part-of-speech tagging and named entity recognition, and integrates it with the health archive feature vector F1 to reorganize it into a health reorganized feature vector F2;
[0131] The file data training and analysis module trains the pre-built machine learning model by acquiring historical residents' electronic health file data, and uses the trained machine learning model to classify the health reorganization feature vector F2 to generate health file classification results Cha with several labels;
[0132] The file label generation module uses a graph convolutional network to perform semantic association and relationship extraction on the labels in the health file classification result Cha, obtains the potential association between different labels, and forms a label relationship vector Rha;
[0133] The iterative optimization module iteratively trains the machine learning model and graph convolutional network method through new resident electronic health record data received at a fixed period, and updates the acquisition process of the label relationship vector Rha.
[0134] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for processing electronic health records data of residents throughout their life cycle, characterized by: The following steps are involved: S1, preprocessing and formatting the received resident electronic health record data, and then extracting the resident health characteristics to form a health record feature vector F1; S2, processing the received health record feature vector F1 using NLP technology, including word segmentation, part-of-speech tagging and named entity recognition, and integrating it with the health record feature vector F1 to reorganize it into a health reorganized feature vector F2; S3, training the pre-built machine learning model by acquiring historical residents' electronic health record data, and using the trained machine learning model to classify the health reorganization feature vector F2 to generate health record classification results Cha with several labels; S4, by using the graph convolutional network to perform semantic association and relationship extraction on the labels in the health record classification result Cha, obtain the potential association between different labels, and form a label relationship vector Rha; S5. The process of acquiring the label relationship vector Rha is updated by iteratively training the machine learning model and graph convolutional network method through the new resident electronic health record data received at a fixed period.
2. A method for processing electronic health records of residents throughout their life cycle according to claim 1, characterized in that: Said S1 includes S11 and S12; S11. The received resident electronic health record data is a heterogeneous data set from multiple sources, including structured data and unstructured data. The resident electronic health record data is preprocessed, including data cleaning preprocessing, data denoising preprocessing, data supplementation missing value preprocessing and data formatting preprocessing. In the data cleaning preprocessing process, duplicate resident electronic health record data will be eliminated, and the data formatting preprocessing will unify the resident electronic health record data from multiple sources into standard structured data to form the initial health record data set F0; The structured data includes diagnosis results, physical examination data and hospitalization records, and the unstructured data includes medical records, medical reports and doctors' notes.
3. A method for processing electronic health records of residents throughout their life cycle according to claim 2, characterized in that: S12, performing a feature extraction phase on the preprocessed initial health record data set F0, by extracting health features from the formatted initial health record data set F0, wherein the extracting health features includes extracting numerical field information from the structured data; Extracting keywords and topics from the unstructured data, and marking numerical features for the keywords and topics, where the numerical features include quantity and length; Among them, the numerical field information includes age, weight, blood sugar and blood pressure; By integrating the extracted health features, the health record feature vector F1 is obtained.
4. A method for processing electronic health records of residents throughout their life cycle according to claim 3, characterized in that: The S2 includes S21 and S22; S21, processing the text data in the received health record feature vector F1, performing word segmentation, part-of-speech tagging and named entity tagging on the text data by using NLP technology, and marking it as a text feature vector Ftext; The text feature vector Ftext is obtained by the following NLP processing formula: Ftext={Tokens(Q),POSTags(Q),Entities(Q)}; In the formula, Q represents text data, Tokens(Q) represents the set of word segmentation results for text data, POSTags(Q) represents the set of part-of-speech tagging results for text data, and Entities(Q) represents the set of named entity results for text data; The word segmentation is to divide the text data into words and phrases, and then divide the long text in the text data into short text; The part-of-speech tagging is used to mark the parts of speech of the divided words and phrases, and the parts of speech include nouns, verbs and adjectives, which are used to reflect the structure and meaning of the text data; The named entities are used to identify entity names in text data, and the entity names include disease names, drug names, and medical institutions.
5. A method for processing electronic health records of residents throughout their life cycle according to claim 4, characterized in that: S22, integrating the acquired text feature vector Ftext with the health record feature vector F1 to reorganize into a health reorganized feature vector F2; The healthy reorganization feature vector F2 is obtained by the following integration formula: In the formula, It represents adjacent concatenation operations, specifically, the text feature vector Ftext and the health record feature vector F1 are merged into a new vector.
6. A method for processing electronic health records of residents throughout their life cycle according to claim 1, characterized in that: The S3 includes S31; S31, training the machine learning model by using the stored historical resident electronic health record data, and after the training is completed, inputting the health reorganization feature vector F2 and the label set Cinit into the machine learning model for classification, and generating health record classification results Cha with several labels for the health reorganization feature vector F2; Among them, the label set Cinit is composed of the stored historical residents' electronic health record data and health labels preset by medical experts, including high blood sugar and heart disease; The health record classification result Cha is obtained by the following training method: Cha=M(F2,Cinit); Where M represents the machine learning model.
7. A method for processing electronic health records of residents throughout their life cycle according to claim 6, characterized in that: The S4 includes S41; S41, converting the health record classification result Cha into a graph structure, wherein each node in the graph structure represents a label in the health record classification result Cha, and constructing edges between nodes according to the similarity and co-occurrence probability between the labels; The edges between labels are then weighted and information is transferred through the graph convolutional network. Each label node exchanges information with adjacent label nodes through the connections in the graph, thereby realizing the semantic association between labels and generating a semantically enhanced enhanced label representation Hv(l+1). The enhanced label representation Hv(l+1) is obtained by the following calculation formula: Wherein, Hv(l+1) represents the enhanced label representation of node v in the l+1 layer, specifically represents the feature vector of node v in the graph convolutional network. Each layer of convolution operation will update the representation of the node by transferring and aggregating the information of neighboring nodes. σ represents the nonlinear activation function, including ReLU and Sigmoid activation function. u∈N(v) represents that node u is in the neighbor set of node v. N(v) represents the neighbor set of node v, that is, the neighboring nodes connected to node v. A in Avu represents the adjacency matrix, specifically represents the element Avu of the adjacency matrix, representing the connection relationship between node v and node u. Among them, Avu=1 represents that node v and node u are connected by an edge, Avu=0 represents that node v and node u are not connected by an edge, dv and du represent the degrees of node v and node u respectively, that is, the number of neighbors of node v and node u. W(l) represents the weight matrix of the lth layer. hu(l) represents the representation of node u in the lth layer. It is the output of the graph convolution operation of the previous layer and represents the feature vector of node u.
8. A method for processing electronic health records of residents throughout their life cycle according to claim 7, characterized in that: The S4 includes S42, calculating the similarity R(i, j) between tags based on the acquired enhanced tag representation Hv(l+1), reflecting the similarity between tag i and tag j in the space vector, and forming a tag relationship vector Rha by aggregating the similarity R(i, j) between each tag; The similarity R(i, j) is obtained by the following calculation formula: Where H(i) and H(j) represent the enhanced label representations of label i and label j, respectively.
9. A method for processing electronic health records of residents throughout their life cycle according to claim 1, characterized in that: The S5 includes S51; S51. Compare the actual label relationship provided by the doctor with the new resident electronic health record data received at a fixed period to verify whether the label relationship vector Rha is consistent with the actual medical situation. When the label relationship vector Rha is verified to be inconsistent with the actual medical situation, iteratively train the machine learning model and the graph convolutional network method, and update the acquisition process of the label relationship vector Rha until the fixed period verifies that the label relationship vector Rha is consistent with the actual medical situation, and stop iteratively training the machine learning model and the graph convolutional network method. Among them, the iterative training of machine learning models and graph convolutional network methods includes adjusting the weight matrix W and the number and depth of graph convolutional layers.
10. A system for processing electronic health records data of residents throughout their life cycle, applied to a method for processing electronic health records data of residents throughout their life cycle as claimed in any one of claims 1 to 9, characterized in that: It includes an archive data processing module, an archive data identification module, an archive data training and analysis module, an archive label generation module and an iterative optimization module; The archive data processing module preprocesses and formats the received resident electronic health archive data, and then extracts the resident health characteristics to form a health archive feature vector F1; The archive data recognition module processes the received health archive feature vector F1 using NLP technology, including word segmentation, part-of-speech tagging and named entity recognition, and integrates it with the health archive feature vector F1 to reorganize it into a health reorganized feature vector F2; The file data training and analysis module trains the pre-built machine learning model by acquiring historical residents' electronic health file data, and uses the trained machine learning model to classify the health reorganization feature vector F2 to generate health file classification results Cha with several labels; The file label generation module uses a graph convolutional network to perform semantic association and relationship extraction on the labels in the health file classification result Cha, obtains the potential association between different labels, and forms a label relationship vector Rha; The iterative optimization module iteratively trains the machine learning model and graph convolutional network method through new resident electronic health record data received at a fixed period, and updates the acquisition process of the label relationship vector Rha.