An electronic medical record management method and system based on AI computing

By introducing AI computing technology into the electronic medical record system, designing general templates and using patient data classification models, the problems of inefficiency and high error rates of traditional electronic medical record systems are solved, and the automated generation and quality control of high-quality electronic medical records are realized.

CN119724463BActive Publication Date: 2025-05-13WUHAN TONGBU YUANFANG INFORMATION TECH DEV CO
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510226413.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-05-13
Estimated Expiration
2045-02-27

AI Technical Summary

Technical Problem

Traditional electronic medical record systems are difficult to automatically generate high-quality electronic medical records, are inefficient and prone to error introduction, especially when dealing with complex clinical judgments and unstructured descriptive content.

Method used

Using an electronic medical record management method based on AI computing, we automatically identify and classify patient data by designing a general template, collecting and preprocessing patient data from multiple terminals, using patient data classification models and named entity recognition models, generating high-quality electronic medical records, and performing abnormal detection and quality control through the medical record content graph network.

Benefits of technology

It realizes automatic generation of high-quality electronic medical records, reduces the work burden of medical staff, reduces the risk of human error, ensures the high quality and reliability of electronic medical records, and improves the accuracy and completeness of medical records.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119724463B_ABST
    Figure CN119724463B_ABST
Patent Text Reader

Abstract

The present invention provides an electronic medical record management method and system based on AI calculation, the method comprising the following steps: designing a general template for electronic medical records; preprocessing patient data; dividing the preprocessed patient data into patient numerical data and patient text data; marking key content items in all general content items in the general template; performing abnormal data detection on patient numerical data and patient text data based on a target general template; extracting patient data entities and content item entities from a target electronic medical record using a named entity recognition model; constructing a medical record content graph network of a target electronic medical record by combining patient data entities and content item entities; performing abnormal content detection on a target electronic medical record based on the medical record content graph network, and if the target electronic medical record passes the abnormal content detection, saving the target electronic medical record. The present invention has the effect of intelligently generating high-quality electronic medical records.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of medical data processing, and specifically relates to an electronic medical record management method and system based on AI computing. Background Art

[0002] With the in-depth application of information technology in the medical field, electronic medical record systems have become the core infrastructure of modern medical institutions. These systems are not only used to record and manage patients' medical information, but also play a key role in supporting clinical decision-making, promoting medical collaboration, and improving medical quality. However, with the explosive growth of medical data and the increasing complexity of data types, traditional electronic medical record systems mainly consider data entry and storage, which leads to the fact that the generation process of medical records still relies heavily on manual operations, which is not only inefficient but also prone to errors. On the other hand, since medical information contains a large amount of professional terms, complex clinical judgments, and unstructured descriptive content, the automated generation of high-quality electronic medical records has become a huge challenge. Summary of the invention

[0003] The present invention provides an electronic medical record management method and system based on AI computing to solve the problem of difficulty in automatically generating high-quality electronic medical records.

[0004] In a first aspect, the present invention provides an electronic medical record management method based on AI computing, the method comprising the following steps:

[0005] Designing a universal template for electronic medical records, wherein the universal template includes a plurality of universal content items;

[0006] Collect complete patient data through multiple terminals and pre-process the patient data;

[0007] Inputting the preprocessed patient data into a patient data classification model, completing identification and classification of the patient data through the patient data classification model, and dividing the patient data into patient numerical data and patient text data according to the identification and classification results, wherein the patient text data includes patient key text data and patient non-key text data;

[0008] Marking key content items in all the universal content items in the universal template according to the recognition and classification results to obtain a target universal template;

[0009] Performing abnormal data detection on the patient numerical data and the patient text data based on the target universal template, and if both the patient numerical data and the patient text data pass the abnormal data detection, generating a target electronic medical record based on the target universal template and in combination with the patient numerical data and the patient text data;

[0010] Extracting patient data entities and content item entities from the target electronic medical record using a named entity recognition model;

[0011] Constructing a medical record content graph network of the target electronic medical record by combining the patient data entity and the content item entity;

[0012] The target electronic medical record is detected for abnormal content based on the medical record content graph network, and if the target electronic medical record passes the abnormal content detection, the target electronic medical record is saved.

[0013] Optionally, the collecting complete patient data through multiple terminals and preprocessing the patient data includes the following steps:

[0014] Collect complete patient data through multiple terminals, including outpatient doctor terminal, resident doctor terminal, nurse terminal, hospital inspection terminal, hospital imaging terminal and patient client terminal;

[0015] cleaning duplicate data in the patient data;

[0016] Standardizing the patient data based on a preset medical term ontology library and using natural language processing technology;

[0017] The standardized patient data is subjected to data structuring processing.

[0018] Optionally, the step of inputting the preprocessed patient data into a patient data classification model, completing identification and classification of the patient data through the patient data classification model, and dividing the patient data into patient numerical data and patient text data according to the identification and classification results comprises the following steps:

[0019] Inputting the preprocessed patient data into a patient data classification model, identifying numerical data and text data in the patient data through a logistic regression module in the patient data classification model to obtain a first recognition and classification result, and dividing the patient data into patient numerical data and patient text data according to the first recognition and classification result;

[0020] Inputting the patient text data into a department classification module in the patient data classification model, outputting a second recognition and classification result through the department classification module, and determining the core department to which the patient text data belongs according to the second recognition and classification result, wherein the department classification module is constructed based on a Transformer architecture;

[0021] According to the second identification and classification result output by the department classification module, and using the data annotation module in the patient data classification model, the key patient text data corresponding to the core department is annotated in the patient text data.

[0022] Optionally, performing abnormal data detection on the patient numerical data and the patient text data based on the target general template, if both the patient numerical data and the patient text data pass the abnormal data detection, generating a target electronic medical record based on the target general template and in combination with the patient numerical data and the patient text data comprises the following steps:

[0023] Performing abnormal value detection on the patient numerical data according to a preset numerical threshold list, and if there is any one or more patient numerical data exceeding the preset numerical threshold, determining that the patient numerical data fails the abnormal data detection;

[0024] If there is no patient numerical data exceeding the preset numerical threshold, determining that the patient numerical data passes the abnormal data detection;

[0025] Matching all of the patient key text data to the corresponding key content items in the target universal template;

[0026] If any one or more of the key content items fail to successfully match the patient key text data, it is determined that the patient text data fails the abnormal data detection;

[0027] If all of the key content items match at least one of the patient key text data, a target electronic medical record is generated based on the target universal template and in combination with the patient numerical data and the patient text data.

[0028] Optionally, the extracting the patient data entity and the content item entity from the target electronic medical record by using a named entity recognition model comprises the following steps:

[0029] Extracting a content item entity and a first patient data entity from the target electronic medical record respectively by using a pattern matching rule method in combination with a predefined medical record dictionary and a medical dictionary;

[0030] After deleting the patient data corresponding to the first patient data entity in the target electronic medical record, if there is still target patient data to be entity extracted in the target electronic medical record, extracting the target patient data into a second patient data entity through a pre-trained entity extraction model, wherein the entity extraction model is constructed based on a conditional random field model or a support vector machine;

[0031] The first patient data entity and the second patient data entity are integrated into a patient data entity.

[0032] Optionally, the step of combining the patient data entity and the content item entity to construct a medical record content graph network of the target electronic medical record comprises the following steps:

[0033] An entity relationship recognition model for identifying entity association relationships is added to the entity extraction model to construct an entity comprehensive recognition model, and a first loss function of the entity extraction model and a second loss function of the entity relationship recognition model are integrated into a comprehensive loss function of the entity comprehensive recognition model, wherein the entity relationship recognition model is constructed based on a support vector machine;

[0034] Generate an entity relationship training set based on historical patient data pre-stored in a hospital database, and use the entity relationship training set to train the entity comprehensive recognition model until the comprehensive loss function converges to a minimum value;

[0035] Performing entity normalization processing on the patient data entity;

[0036] Inputting the patient data entity after entity normalization processing into the trained entity comprehensive recognition model, and extracting the data entity association relationship in the patient data entity through the entity comprehensive recognition model, wherein the data entity association relationship includes disease-symptom relationship, drug-indication relationship, treatment-effect relationship and disease-risk factor relationship;

[0037] Extracting the medical record content association relationship between the patient data entity and the content item entity according to the matching relationship between the patient data and all content items in the target electronic medical record;

[0038] The patient data entity and the content item entity are used as entity graph nodes, and the data entity association relationship and the medical record content association relationship are used as entity graph node edges between the entity graph nodes to construct a medical record content graph network of the target electronic medical record.

[0039] Optionally, performing abnormal content detection on the target electronic medical record based on the medical record content graph network, and if the target electronic medical record passes the abnormal content detection, saving the target electronic medical record comprises the following steps:

[0040] Using a multi-layer perceptron to identify whether there is a local information contradiction in the medical record content graph network, wherein the local information contradiction includes a node-level information contradiction and an edge-level information contradiction;

[0041] Using a graph pooling technique to obtain multiple medical record content subgraph representations of the medical record content graph network, and identifying whether there is a global information contradiction in the medical record content graph network based on all the medical record content subgraph representations;

[0042] If the medical record content graph network does not have the local information contradiction and the global information contradiction, extracting a key subgraph network from the medical record content graph network according to the key content items and the patient key text data;

[0043] Extracting high-dimensional topological features from the key subgraph network based on a feature mapping method;

[0044] Inputting the high-dimensional topological features into a pre-trained deep abnormal content recognition model, and judging whether the target electronic medical record has abnormal content according to the high-dimensional feature recognition results output by the deep abnormal content recognition model, wherein the deep abnormal content recognition model is constructed based on a graph neural network model;

[0045] If the target electronic medical record does not have the abnormal content, it is determined that the target electronic medical record passes the abnormal content detection, and the target electronic medical record is saved.

[0046] Optionally, the extracting high-dimensional topological features from the key subgraph network based on the feature mapping method comprises the following steps:

[0047] For any entity graph node in the key subgraph network, use the entity graph node as a first entity graph node, and use any entity graph node of a different type in the key subgraph network except the first entity graph node as a second entity graph node;

[0048] Mapping node features of the first entity graph node and the second entity graph node to a unified high-dimensional feature space to obtain high-dimensional node features;

[0049] In combination with the high-dimensional node feature and the type of association relationship between the first entity graph node and the second entity graph node, constructing a node heterogeneous attention between the first entity graph node and the second entity graph node in the high-dimensional feature space;

[0050] In combination with the association relationship type and the node types of the first entity graph node and the second entity graph node, constructing an information conduction function of the first entity graph node in the high-dimensional feature space;

[0051] Obtaining neighbor aggregation information of the first entity graph node based on the node heterogeneous attention and the information conduction function;

[0052] After mapping the neighbor aggregation information of all the entity graph nodes to the original feature space, high-dimensional topological features of the key subgraph network are obtained by aggregation.

[0053] In a second aspect, the present invention also provides an electronic medical record management system based on AI computing, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the electronic medical record management method based on AI computing as described in the first aspect is implemented.

[0054] In a third aspect, the present invention further provides a computer-readable storage medium, in which a computer program is stored. When the computer program is loaded and executed by a processor, the electronic medical record management method based on AI calculation described in the first aspect is adopted.

[0055] The beneficial effects of the present invention are:

[0056] By utilizing the patient data classification model and the target universal template, the present invention can automatically identify and classify patient data, and generate targeted electronic medical records according to actual needs, which not only reduces the workload of medical staff, but also reduces the risk of human error. Secondly, the present invention establishes a multi-level quality control mechanism, including abnormal detection of patient data and content detection of generated medical records, which ensures the high quality and reliability of electronic medical records. In particular, by constructing a medical record content graph network, the present invention can perform deeper semantic analysis and consistency checks, which greatly improves the accuracy and completeness of medical records. The intelligent processing method of the present invention makes the format of electronic medical records more standardized and standardized, which not only facilitates subsequent data analysis and utilization, but also promotes information sharing and collaboration among different medical institutions. On the other hand, the present invention is also very flexible and can generate personalized electronic medical records according to the characteristics and needs of different patients, which helps to provide more accurate medical services. By introducing advanced artificial intelligence technologies, such as named entity recognition models and patient data classification models, the intelligence level of the electronic medical record system is greatly improved, laying the foundation for the future development of medical informationization. In summary, the present invention not only solves the problems of traditional electronic medical record systems in terms of efficiency, accuracy and consistency, but also provides strong support for improving the overall quality of medical services, supporting precision medical decision-making and promoting medical big data analysis. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] Figure 1 This is a flowchart of an electronic medical record management method based on AI computing in one of the embodiments of the present application.

[0058] Figure 2 This is a schematic diagram of the structure of a patient data classification model in one embodiment of the present application. DETAILED DESCRIPTION

[0059] The following will be combined with the drawings in the embodiments of the present application to clearly describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments in the present application belong to the scope of protection of this application.

[0060] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described here, and the objects distinguished by "first", "second", etc. are generally of one type, and the number of objects is not limited. For example, the first object can be one or more. In addition, "and / or" in the specification and claims represents at least one of the connected objects, and the character " / " generally indicates that the objects associated with each other are in an "or" relationship.

[0061] Figure 1 FIG. 1 is a flowchart of an electronic medical record management method based on AI computing in one embodiment. It should be understood that although Figure 1 The steps in the flowchart are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, Figure 1 At least part of the steps in the above method may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least part of the sub-steps or stages of other steps. Figure 1 As shown, the electronic medical record management method based on AI calculation disclosed in the present invention specifically includes the following steps:

[0062] S101. Design a common template for electronic medical records.

[0063] Among them, the general template contains multiple general content items, which cover all aspects of medical records, such as basic patient information, chief complaint, current medical history, past medical history, physical examination, auxiliary examination, diagnosis, treatment plan, etc. The design of these content items needs to strictly follow the standards and specifications of the medical industry, while also taking into account the special needs of different departments and diseases to ensure the universality and adaptability of the template. The design process of the general template is a complex and systematic work involving multiple stages. First, a comprehensive needs analysis is required to deeply understand the clinical workflow and ensure that the needs of all parties are fully understood. Secondly, based on the results of the needs analysis, a preliminary template is designed. At this stage, the organizational structure and data format of the content items need to be considered to ensure that the template can meet clinical needs and facilitate data entry and extraction. During the design process, special attention should be paid to the flexibility and scalability of the template to adapt to the needs of different platforms or departments. An effective method is to adopt a modular design, setting common content items as mandatory items and special content items as optional items. This can reserve custom space for different departments or special cases while ensuring the integrity of basic information. For example, for chronic disease management, an additional long-term follow-up module can be set up; for the emergency department, an injury assessment module can be added.

[0064] In addition, the ease of use of the template is also an important consideration. Intuitive interface design can be adopted, such as using drop-down menus, radio buttons and other controls to reduce manual input and improve efficiency. At the same time, an intelligent prompt function can be provided to automatically recommend relevant options or remind important information that may be missed based on the information already filled in. For example, when certain symptoms are entered, the system can prompt relevant inspection suggestions. In terms of technical implementation, you can consider using structured data formats such as XML or JSON to define templates, which can facilitate template modification and version control. At the same time, you can use template engine technology, such as Velocity or FreeMarker, to achieve dynamic rendering and data filling of templates.

[0065] S102. Collect complete patient data through multiple terminals and pre-process the patient data.

[0066] Among them, multiple terminals include multiple data sources such as outpatient doctor end, resident doctor end, nurse end, hospital inspection end, hospital imaging end and patient client end, aiming to comprehensively collect various types of medical information of patients and form a complete patient data set. In the process of data collection, it is necessary to focus on the integrity, accuracy and timeliness of the data. To ensure data integrity, a mandatory item reminder function can be set in the system to avoid the omission of important information. For example, for key items such as the patient's basic information, chief complaint, current medical history, etc., a mandatory filling mechanism can be set, and this information can only be submitted after it is fully filled in. At the same time, an intelligent prompt function can be set to automatically recommend relevant information that may need to be supplemented based on the information that has been filled in.

[0067] After data collection is completed, the preprocessing process is a key step to ensure data quality, including data cleaning, format conversion, and preliminary verification. Data cleaning aims to remove noise data, handle missing values ​​and outliers. For missing values, different processing strategies can be adopted according to the specific situation, such as mean filling, median filling, or machine learning-based prediction filling. For example, for continuous data such as blood pressure values, the average value of the patient's historical records can be used for filling; for categorical data such as gender, the mode can be used for filling. For more complex abnormal patterns, machine learning algorithms such as Isolation Forest or density-based methods such as DBSCAN can be used. These methods can identify abnormal points in high-dimensional space and are particularly suitable for anomaly detection of multivariate data. For example, when analyzing various physiological indicators of patients, a single indicator may be within the normal range, but the combination of multiple indicators may show abnormal patterns.

[0068] Format conversion is another important part of preprocessing, which aims to unify data from different sources into a standard format for subsequent processing. This may involve the unification of date formats, conversion of measurement units, unification of coding standards, etc. For example, different systems may use different disease coding standards (such as ICD-9, ICD-10), and a mapping relationship needs to be established for conversion. When performing format conversion, special attention should be paid to maintaining the semantic consistency of the data to avoid information loss or errors caused by conversion.

[0069] In the preprocessing process, data desensitization is an important link that cannot be ignored, aiming to protect patient privacy. Various desensitization techniques can be used, such as data masking (such as replacing names with codes), data replacement (such as randomly exchanging certain fields of different records), data perturbation (such as adding random noise to numerical data), etc. The choice of desensitization method requires a balance between privacy protection and data availability.

[0070] S103. Input the preprocessed patient data into the patient data classification model, complete the recognition and classification of the patient data through the patient data classification model, and divide the patient data into patient numerical data and patient text data according to the recognition and classification results.

[0071] Among them, in terms of model selection, a variety of machine learning algorithms can be used, such as support vector machines (SVM), random forests, or deep learning models such as long short-term memory networks (LSTM). Choosing a suitable model requires considering factors such as the characteristics of the data, the complexity of the classification task, and the interpretability of the model. For example, for highly structured data, such as laboratory test results, traditional machine learning algorithms such as random forests can be used; while for unstructured text data, such as medical records, deep learning models such as LSTM or BERT (Bidirectional Encoder Representations from Transformers) may be more suitable. For numerical data, normalization or standardization may be required to eliminate scale differences between different features. For categorical data, encoding processing is required, such as One-Hot encoding or label encoding. For text data, features can be extracted using technologies such as Bag of Words or Word Embedding.

[0072] The model training process requires a large amount of labeled medical data. In practical applications, there may be a problem of insufficient labeled data. To solve this problem, transfer learning or semi-supervised learning methods can be used. For example, the model can be pre-trained on a large-scale general medical dataset and then fine-tuned on a small-scale dataset for a specific task. In addition, data enhancement techniques can be used, such as synonym replacement for text data or adding small noise to numerical data to expand the training set. For multi-classification problems, these indicators can be calculated using macro-average or micro-average methods. In addition, the ROC curve and AUC value can be used to evaluate the overall performance of the model.

[0073] In practical applications, the model first extracts the features of the input data and then classifies it based on the learned patterns. The classification results divide the patient data into patient numerical data and patient text data, where the patient text data is further divided into key text data and non-key text data. Key text data usually includes important information such as the diagnosis results and treatment plans of the corresponding department, while non-key text data may be some descriptive supplementary information of the department.

[0074] S104. Mark the key content items in all the general content items in the general template according to the recognition and classification results to obtain the target general template.

[0075] Among them, this process involves customizing the generic template to highlight the most relevant information for a specific patient and a specific department. The labeling process can adopt a strategy that combines rule-based methods and machine learning methods. The rule-based method labels certain content items as key items by default, such as chief complaint, diagnosis, etc., based on preset medical knowledge and clinical experience. The machine learning method can learn the association between data from different departments and the importance of content items based on historical data. For example, a decision tree algorithm or a random forest algorithm can be used to train a classification model with patient characteristics as input and the importance of content items as output. The performance of the model can be evaluated by cross-validation, and indicators such as accuracy, precision, and recall can be calculated. In practical applications, each content item can be assigned an importance score, and those with scores exceeding a threshold are marked as key content items. The setting of the threshold can be dynamically adjusted according to specific needs. For example, the ROC curve can be used to find the optimal threshold, that is, to maximize the difference between the true positive rate (TPR) and the false positive rate (FPR). The labeling process also needs to consider the particularities of different departments and diseases, and department-specific rules or models can be set. During the generation of the target generic template, key content items may be visually emphasized, such as by using different colors or bold fonts.

[0076] S105. Perform abnormal data detection on the patient numerical data and the patient text data based on the target general template. If both the patient numerical data and the patient text data pass the abnormal data detection, generate a target electronic medical record based on the target general template and in combination with the patient numerical data and the patient text data.

[0077] Among them, abnormal data detection uses a variety of technical methods, including statistical methods, machine learning methods, and domain-specific rules. For numerical data, the Z-score method or the interquartile range (IQR) method can be used to detect outliers. For text data, natural language processing techniques, such as word frequency statistics and topic models, can be used to identify abnormal word combinations or semantic structures. In addition, deep learning models, such as autoencoders, can be applied to learn the normal patterns of data and mark data that deviate from these patterns as abnormal. In practical applications, abnormal data detection needs to take into account the particularity of the medical field, such as some rare cases may produce data that seems abnormal but is actually correct. Therefore, it is possible to combine expert knowledge bases to establish domain-specific rule systems to improve the accuracy of detection. If an abnormality is detected, the system will generate a warning message, requiring the operator to review or provide additional explanations. Only when all data pass the anomaly detection will the next step of generating the target electronic medical record be entered. The process of generating the target electronic medical record is to combine the patient data with the target general template to form a complete and standardized electronic medical record document. This process involves formatting, structuring, and visualization of data. For example, template engine technology can be used to dynamically fill patient data into a predefined template structure. At the same time, natural language generation (NLG) technology can also be applied to convert certain numerical data into easy-to-understand descriptive text.

[0078] S106. Use a named entity recognition model to extract patient data entities and content item entities from the target electronic medical record.

[0079] Among them, the construction and application of named entity recognition (NER) models involve multiple technical aspects, including model selection, data preprocessing, feature engineering, model training and evaluation, and post-processing. Named entity recognition models can adopt a variety of technical methods, mainly including rule-based methods, machine learning methods, and deep learning methods.

[0080] Rule-based methods use predefined dictionaries and pattern matching rules to identify entities. For example, you can build a dictionary containing common disease names, symptoms, drug names, etc., and then use regular expressions or other string matching techniques to find these entities in the text. The advantages of rule-based methods are that they are intuitive, highly interpretable, and suitable for highly structured content. However, they are less flexible and have difficulty handling variants and newly emerging entities.

[0081] Machine learning methods include Conditional Random Fields (CRF) and Support Vector Machines (SVM). These methods can learn contextual features and improve recognition accuracy.

[0082] Deep learning methods have performed well in NER tasks. Commonly used models include bidirectional long short-term memory networks (Bi-LSTM), convolutional neural networks (CNN), and models that combine attention mechanisms such as BERT (Bidirectional Encoder Representations from Transformers). Among them, the Bi-LSTM-CRF model is a widely used architecture that combines the feature extraction capabilities of Bi-LSTM and the sequence modeling capabilities of CRF.

[0083] In practical applications, an ensemble learning strategy can be used to combine the advantages of multiple models. For example, a rule-based approach can be used to first identify clear entities, and then a deep learning model can be used to handle complex situations. This approach can balance accuracy and efficiency while maintaining a certain degree of interpretability.

[0084] Model training requires a large amount of annotated medical text data. The following strategies can be used:

[0085] Transfer learning: Pre-train the model on a large-scale dataset in a general field, such as using a medical literature database such as PubMed, and then fine-tune it on a small-scale dataset for a specific task.

[0086] Semi-supervised learning: learning with a small amount of labeled data and a large amount of unlabeled data. For example, you can use the self-training method to first train the initial model with a small amount of labeled data, then use the model to predict the unlabeled data, add the high-confidence prediction results to the training set, and continuously iterate this process.

[0087] Active learning: The most informative samples are selected through algorithms for manual labeling, achieving the greatest performance improvement with the lowest labeling cost.

[0088] Data augmentation: Perform synonym replacement, back translation and other operations on existing annotated data to expand the training set.

[0089] In the process of entity recognition, the particularity of medical terms, such as abbreviations, synonyms, etc., also needs to be considered. A medical terminology knowledge base can be built to assist in entity recognition and normalization. For example, "MI" can be recognized and normalized as "Myocardial Infarction".

[0090] Post-processing is an important step to improve the quality of recognition results. It can include the following operations:

[0091] Entity Boundary Adjustment: Correct possible boundary errors, such as removing punctuation marks that were mistakenly included.

[0092] Entity type consistency check: Ensure that the type tag of the same entity is consistent throughout the document.

[0093] Entity linking: Linking identified entities to standard medical terminologies such as SNOMED CT or ICD-10.

[0094] Overlapping entity handling: Resolve issues with nested or overlapping entities, such as "chronic kidney disease" containing "kidney disease".

[0095] In actual deployment, the efficiency and scalability of the model also need to be considered. Model compression techniques, such as knowledge distillation or quantization, can be used to reduce model size and inference time. At the same time, a distributed computing framework can be used to process large-scale data.

[0096] S107. Construct a medical record content graph network of the target electronic medical record by combining the patient data entity and the content item entity.

[0097] Among them, the process of building a medical knowledge graph can be divided into several key stages: entity normalization, relationship extraction, knowledge fusion, knowledge representation and storage, and verification and maintenance of the knowledge graph. The goal of entity normalization is to map the extracted entities to standardized medical terms or concepts. Relationship extraction aims to identify the semantic relationships between entities, which is essential for building a meaningful knowledge graph. Knowledge fusion is the process of integrating knowledge extracted from different sources into a consistent knowledge graph. For storage, you can consider using specialized graph databases such as Neo4j, JanusGraph, etc., which are optimized for the storage and query of graph structured data. After the construction is completed, the knowledge graph needs to be verified and continuously maintained.

[0098] S108. Perform abnormal content detection on the target electronic medical record based on the medical record content graph network. If the target electronic medical record passes the abnormal content detection, save the target electronic medical record.

[0099] In one embodiment, collecting complete patient data through multiple terminals and preprocessing the patient data includes the following steps:

[0100] Collect complete patient data through multiple terminals, including outpatient doctor terminal, resident doctor terminal, nurse terminal, hospital laboratory terminal, hospital imaging terminal and patient client terminal;

[0101] Cleaning duplicate data from patient data;

[0102] Standardize patient data based on a preset medical terminology ontology library and use natural language processing technology;

[0103] The standardized patient data is structured.

[0104] In this embodiment, at the outpatient doctor end, the doctor can record the patient's initial diagnosis information, symptoms and preliminary diagnosis through the electronic medical record system. The resident doctor end records more detailed medical records and treatment responses during hospitalization. The nurse end is used to record changes in patient status and nursing measures observed during the nursing process. The records of these three doctors and nurses together form a complete medical record of the patient during hospitalization. The hospital inspection end is responsible for uploading various medical test results, including blood tests, urine analysis, etc. The hospital imaging end is responsible for integrating and storing medical imaging data such as CT and MRI. These inspection and imaging data provide doctors with reliable diagnostic basis. Finally, the patient client allows patients to upload self-measured data, such as blood sugar monitoring, heart rate monitoring, and even feedback from health questionnaires, so as to fully understand the patient's symptoms and living habits. Implementing this multi-terminal data collection step can effectively avoid the limitations of single-source data, improve the accuracy and richness of data, and ensure information exchange and sharing between medical personnel and patients.

[0105] Cleaning duplicate data from patient data is an important step to ensure data accuracy and consistency. During the data collection process, duplicate data may appear due to different input sources or manual input errors. To address this problem, data deduplication algorithms are often used. First, it is necessary to determine the duplicate data criteria, which include a series of features such as the text content of the data, date, patient ID, etc. For example, possible duplicate records can be identified by comparing the combination of name, date of birth, and hospital ID. After initial identification, deduplication algorithms such as fuzzy matching algorithms are used to further detect data records that are slightly different in spelling but are essentially similar. Fuzzy matching can calculate the similarity between strings, such as the Levenshtein distance, which determines the similarity between two strings by calculating the minimum number of editing steps to transform one string into another. If the distance between two strings is less than a certain threshold, they can be considered duplicates.

[0106] Standardizing patient data based on a preset medical term ontology library and using natural language processing technology is an important step to ensure data uniformity and standardization. Natural language processing (NLP) technology is at the core of this step, converting unstructured text data into a recognizable standard form through semantic analysis and entity recognition. The ontology library is one of the key infrastructures in this process, which contains the terms and relationships used in the medical field. At the beginning of processing, the system matches the free text in the patient data with the ontology library to identify standardized vocabulary. For example, "heart attack" and "myocardial infarction" in patient records are uniformly identified, and these terms may be synonyms. Semantic analysis can ensure that the complex requirements of the text are accurately understood and processed through part-of-speech tagging, syntactic analysis, and entity recognition.

[0107] Structuring standardized patient data is an important step in effectively using data for in-depth analysis. The core of structuring is to convert standardized text information into a format that the database can easily understand and process. This step usually involves using a relational database or a non-relational database for data storage, and this choice mainly depends on the scenario and analysis requirements of the data. First, it is necessary to design the data entry format. For example, electronic medical records can be divided into multiple structural fields, such as basic patient information, diagnosis results, treatment plans, and drug use. In the database, these fields are represented by tables and associated through foreign key relationships. In data entry, the ETL (Extract, Transform, Load) process can be used to implement it, that is, extracting data from the data source, cleaning and transforming it, and finally loading it into the target database. Data structuring makes it faster to call and count information.

[0108] In one embodiment, the pre-processed patient data is input into a patient data classification model, the patient data is identified and classified by the patient data classification model, and the patient data is divided into patient numerical data and patient text data according to the identification and classification results, including the following steps:

[0109] Inputting the preprocessed patient data into a patient data classification model, identifying numerical data and text data in the patient data through a logistic regression module in the patient data classification model, obtaining a first recognition and classification result, and dividing the patient data into patient numerical data and patient text data according to the first recognition and classification result;

[0110] Input the patient text data into the department classification module in the patient data classification model, output the second recognition and classification result through the department classification module, and determine the core department to which the patient text data belongs according to the second recognition and classification result. The department classification module is built based on the Transformer architecture;

[0111] According to the second identification and classification result output by the department classification module, the key patient text data corresponding to the core department is annotated in the patient text data using the data annotation module in the patient data classification model.

[0112] In this embodiment, refer to Figure 2 , input the preprocessed patient data into the patient data classification model, and identify the numerical data and text data in the patient data through the logistic regression module in the patient data classification model. Specifically, logistic regression is a commonly used binary classification algorithm, which is used here to distinguish between numerical data and text data. In specific implementation, it is first necessary to extract features from the input data. For example, regular expressions can be used to match numeric patterns, or to detect whether text-specific punctuation marks are contained. For each data item, the extracted features may include the proportion of numbers, the proportion of letters, the proportion of special characters, etc. These features form a vector as the input of the logistic regression model. The core of the logistic regression model is a linear function plus a sigmoid activation function, and its mathematical expression is:

[0113]

[0114] Where X is the input feature vector, θ is the model parameter, and Y=1 indicates the probability that the data is of numerical type. During the model training process, the parameter θ is optimized using methods such as maximum likelihood estimation or gradient descent. In the prediction stage, if P(Y=1|X) is greater than the preset threshold (usually 0.5), it is judged as numerical data, otherwise it is text data. The advantage of this method is that it has fast calculation speed and is easy to implement and explain. Through this step, patient data is effectively divided into patient numerical data and patient text data, laying the foundation for subsequent refined processing. Numerical data may include various test indicators, vital signs, etc., while text data may contain symptom descriptions, diagnostic opinions, etc.

[0115] The patient text data is input into the department classification module in the patient data classification model, and the second recognition classification result is output through the department classification module. The department classification module is built on the Transformer architecture, which is a powerful deep learning model that is particularly suitable for processing sequence data such as text. The core of Transformer is the self-attention mechanism (Self-Attention), which allows the model to take into account the context of the entire sequence when processing each element in the sequence. In specific implementation, the text data must first be converted into a numerical representation that the model can understand, usually using word embedding (WordEmbedding) technology. For example, pre-trained medical field word vectors such as BioWord2Vec can be used. The encoder part of the Transformer model contains multiple identical layers, each layer has two sub-layers: a multi-head self-attention mechanism and a feedforward neural network. The calculation formula of the multi-head self-attention mechanism is:

[0116]

[0117] Where Q, K, and V represent query, key, and value respectively. is the dimension of the key. This allows the model to capture long-distance dependencies in the text. In the department classification task, the output layer of the model is usually a softmax classifier that maps high-dimensional features to probability distributions of various departments. The cross-entropy loss function is used during model training, and the parameters are updated through backpropagation and optimizers (such as Adam). In the prediction stage, the patient text data is input, and the model outputs the probability of each department, and the one with the highest probability is selected as the classification result. This method can accurately classify patient text data into corresponding core departments, such as internal medicine, surgery, obstetrics and gynecology, etc. The implementation of this step greatly improves the efficiency of organizing patient data, so that subsequent diagnosis and treatment and research can be more accurately targeted at the characteristics of specific departments.

[0118] The data annotation module usually adopts sequence annotation methods, such as conditional random fields (CRF) or a combination of bidirectional long short-term memory networks (Bi-LSTM) and CRF. Taking Bi-LSTM-CRF as an example, the model can effectively capture contextual information and consider the dependencies between labels. First, the text is converted into a sequence of word vectors and then input into the Bi-LSTM layer. The forward and backward hidden states of the Bi-LSTM are connected to form the feature representation of each word. These features are then input into the CRF layer, which considers the transition probability of adjacent labels and outputs the optimal label sequence. The objective function of CRF is usually a log-likelihood function in the form of: logP(y|x)=Σ(φ(y_i,x,i))-logZ(x), where φ is the feature function and Z(x) is the normalization factor. During the training process, the model learns specific terms, symptom descriptions, and diagnostic patterns for different departments. For example, for cardiology, the model may pay special attention to symptom descriptions such as "chest pain" and "palpitations"; for orthopedics, it may pay more attention to words such as "joint pain" and "fractures". In actual applications, the model will assign a label to each word, such as B-symptom (symptom start), I-symptom (symptom inside), O (other), etc.

[0119] In one embodiment, abnormal data detection is performed on patient numerical data and patient text data based on a target general template. If both the patient numerical data and the patient text data pass the abnormal data detection, generating a target electronic medical record based on the target general template and in combination with the patient numerical data and the patient text data includes the following steps:

[0120] Perform abnormal value detection on the patient numerical data according to a preset numerical threshold list, and if there is any one or more patient numerical data exceeding the preset numerical threshold, it is determined that the patient numerical data fails the abnormal data detection;

[0121] If there is no patient numerical data exceeding the preset numerical threshold, it is determined that the patient numerical data passes the abnormal data detection;

[0122] Match all patient key text data to the corresponding key content items in the target universal template;

[0123] If any one or more key content items fail to successfully match the patient's key text data, it is determined that the patient's text data fails the abnormal data detection;

[0124] If all key content items match at least one patient key text data, a target electronic medical record is generated based on the target universal template and in combination with the patient numerical data and the patient text data.

[0125] In this embodiment, the preset numerical threshold list is usually formulated by medical experts based on a large amount of clinical data and medical research results, and includes the normal range of various physiological indicators. For example, for the body temperature of an adult, the normal range may be set to 36.1°C to 37.2°C; the normal range of blood pressure may be set to systolic pressure 90-140mmHg and diastolic pressure 60-90mmHg. During the detection process, each numerical data is compared with its corresponding threshold range. For more complex indicators, multiple conditions may need to be considered, such as blood sugar levels may need to set different thresholds according to different states such as fasting and postprandial. In practical applications, hash tables or dictionary structures can be used to store these thresholds for fasting and comparison. If any value is found to exceed its preset threshold during the detection process, the system will mark the entire patient numerical data set as failing the abnormal data detection.

[0126] If there is no patient numerical data that exceeds the preset numerical threshold, the patient numerical data is judged to have passed the abnormal data detection. At this stage, the system will conduct a comprehensive review of all patient numerical data to ensure that each indicator falls within the preset normal range. This process can be achieved by traversing all numerical data items, and each item needs to be compared with its corresponding threshold range. For example, for a series of blood test results, including red blood cell count, white blood cell count, hemoglobin, platelet count, etc., each item needs to be compared with its specific normal range. If all data items pass the inspection and do not exceed their respective threshold ranges, the entire patient numerical data set is marked as having passed the abnormal data detection. This comprehensive and rigorous inspection mechanism ensures the reliability and consistency of the data entering the subsequent analysis and processing stages.

[0127] All patient key text data are matched to the corresponding key content items in the target general template. The matching process involves natural language processing (NLP) technology, mainly including steps such as text classification, named entity recognition (NER), and relationship extraction. First, the patient key text data is classified into large categories of the template through text classification algorithms such as support vector machines (SVM) or deep learning models such as BERT. Then, the key entities in the text, such as disease names, symptoms, drugs, etc., are identified using NER technology. NER can use models such as conditional random fields (CRF) or bidirectional long short-term memory networks (Bi-LSTM). For example, for the sentence "The patient has intermittent headaches in the past three days, accompanied by mild nausea", the NER model may identify "headache" and "nausea" as symptom entities. Next, the relationship between entities, such as the duration and severity of symptoms, is determined through relationship extraction technology. This can be achieved using dependency syntactic analysis or neural network models based on attention mechanisms. Finally, the extracted information is mapped to the corresponding fields in the template. This process may involve fuzzy matching algorithms such as cosine similarity or edit distance to handle differences in synonyms or expressions. For example, the formula for calculating the cosine similarity of two strings is: similarity=(A·B) / (||A||||B||), where A and B are the word vector representations of the two strings. In this way, the symptom of "headache" can be accurately matched to the "complaint" or "symptom" field in the template.

[0128] If any one or more key content items fail to match the patient's key text data, the patient's text data is judged to have failed the abnormal data detection. In this process, the system checks each key content item in the target general template to ensure that they can find corresponding information in the patient's key text data. This check can be implemented by setting a flag array, each element of which corresponds to a key content item in the template and is initially set to false. When a content item successfully matches the text data, the corresponding flag is set to true. The check process can be expressed in pseudo code as: foreachkey_itemintemplate:ifnotmatched(key_item,patient_text_data):returnfalse;returntrue. The matched function may involve complex text matching algorithms, such as fuzzy matching or semantic similarity calculation. For example, for the key content item "past history", if no relevant information is found in the patient's text data, even negative information such as "no special past history" will be considered as a failed match. This strict checking mechanism ensures that each electronic medical record contains all the necessary information. If any key content item is found to be unmatched, the system will mark the entire patient text data as failing the abnormal data detection. The significance of this approach is that it can promptly detect missing data problems that may be caused by incomplete information collection, omissions in doctor records, or system processing errors.

[0129] If all key content items match at least one patient key text data, the target electronic medical record is generated based on the target universal template and combined with the patient numerical data and patient text data. At this stage, the system first confirms that all key content items have been successfully matched, which can be achieved by checking the flag array set in the previous step to ensure that all elements in the array are true. Next, the system begins to fill the patient's numerical data and text data into the target universal template. This process is usually completed using a template engine, such as Apache Velocity or FreeMarker. The template engine allows data to be dynamically inserted into a predefined document structure. For example, for the "${patient.name}" placeholder in the template, the system will replace it with the actual patient's name. For numerical data, such as blood pressure, body temperature, etc., the corresponding fields can be filled directly. For text data, the system needs to perform more complex processing, which may involve natural language generation (NLG) technology. For example, integrate discrete symptom descriptions into coherent paragraphs, or synthesize multiple test results into a concise summary. In this process, the system may also need to normalize and standardize the data, such as unifying all units of measurement or converting drug names to standard generic names. In addition, the system may need to select different sub-templates according to different departments or diseases to meet the needs of specific types of medical records. Finally, the system will generate a complete, high-quality, and structured electronic medical record document.

[0130] In one embodiment, extracting patient data entities and content item entities from a target electronic medical record using a named entity recognition model includes the following steps:

[0131] A pattern matching rule method is used in combination with a predefined medical record dictionary and a medical dictionary to extract a content item entity and a first patient data entity from the target electronic medical record, respectively;

[0132] After deleting the patient data corresponding to the first patient data entity in the target electronic medical record, if there is still target patient data to be extracted in the target electronic medical record, extracting the target patient data into a second patient data entity through a pre-trained entity extraction model, where the entity extraction model is constructed based on a conditional random field model or a support vector machine;

[0133] The first patient data entity and the second patient data entity are integrated into a patient data entity.

[0134] In this embodiment, the pattern matching rule method is a text processing technology based on predefined patterns and rules, which can identify and extract text fragments that conform to specific patterns. In this process, it is first necessary to build a medical record dictionary containing common medical record terms, formats and structures, as well as a medical dictionary covering a wide range of medical terms, disease names, symptom descriptions, etc. These dictionaries are usually stored in the form of a tree structure or hash table for fast search and matching. Pattern matching rules can be defined using regular expressions. For example, the rule for extracting the patient's age may be "\d{1,3} years old", and the rule for extracting blood pressure may be "\d{2,3} / \d{2,3}mmHg". For the extraction of content item entities, the system scans the entire electronic medical record to find text paragraphs that match the predefined pattern, such as fixed-format titles such as "Chief Complaint:", "Present Medical History:". For the extraction of the first patient data entity, the system combines the medical dictionary to identify and extract specific medical terms and values. For example, in the sentence "patient's blood pressure is 140 / 90mmHg", the system will recognize "blood pressure" as a medical term and "140 / 90mmHg" as the corresponding numerical value. This process can be implemented by a finite state automaton (FSA) or a decision tree algorithm to handle complex matching rules.

[0135] After deleting the patient data corresponding to the first patient data entity in the target electronic medical record, if there is still target patient data to be extracted in the target electronic medical record, the target patient data is extracted as the second patient data entity through the pre-trained entity extraction model. This step uses more complex machine learning techniques to process data that is difficult to capture with simple rules. The entity extraction model is built based on the conditional random field (CRF) model or the support vector machine (SVM), both of which are powerful tools for processing sequence labeling tasks. The CRF model is particularly suitable for processing sequence data, which takes into account the dependencies between labels. The core idea of ​​CRF is to construct a conditional probability model p(y|x), where x is the input sequence (i.e., text) and y is the corresponding label sequence. The objective function of CRF is usually the log-likelihood function, in the form of: L(θ)=Σlogp(y|x;θ), where θ is the model parameter. During the training process, the gradient ascent method or the quasi-Newton method is used to maximize this objective function. SVM is a binary classification model that can be extended to multi-classification problems through a one-to-many strategy. The core of SVM is to find an optimal hyperplane to separate data points of different categories. Its objective function can be expressed as: min(1 / 2)||w||^2+CΣξi, where w is the weight vector, C is the penalty parameter, and ξi is the slack variable. Both models require a large amount of labeled data for training, usually including various medical terms, symptom descriptions, test results, etc. In practical applications, the system first preprocesses the remaining text, including word segmentation, part-of-speech tagging, etc. Then, a series of features are generated for each word, such as the word itself, part of speech, whether it contains numbers, whether it is in the medical dictionary, etc. These features are used as input to the model. The model predicts a label for each word, such as B-symptom (symptom start), I-symptom (symptom inside), O (other), etc. In this way, the system can identify those irregular expressions or newly emerging medical terms, greatly improving the coverage and accuracy of entity extraction. This machine learning-based method is more flexible than the pure rule-based method and can adapt to a variety of text styles and expressions, especially when processing unstructured or semi-structured medical texts.

[0136] The integration process of integrating the first patient data entity and the second patient data entity into a patient data entity involves multiple complex subtasks, including entity alignment, conflict resolution, data standardization, and data structuring. First, entity alignment is required, that is, identifying and merging the same or similar entities from different extraction methods (rule-based and machine learning-based). This can be achieved through string matching algorithms or semantic similarity calculations. If the similarity exceeds a preset threshold, it is considered to be the same entity. Second, possible data conflicts need to be resolved. When the information of the same entity extracted by the two methods is inconsistent, the system needs to formulate a strategy to decide which value to adopt. For example, a weighted average mechanism can be set to give a higher weight to a more reliable source. Then, data standardization is performed to convert data in different formats, units, or expressions into a unified standard form. This may involve unit conversion (such as converting imperial units to metric units), terminology standardization (such as converting common names to standard medical terms), etc. Finally, the integrated data is organized into a structured format, such as JSON or XML, for subsequent processing and storage.

[0137] In one embodiment, combining the patient data entity and the content item entity to construct a medical record content graph network of the target electronic medical record includes the following steps:

[0138] An entity relationship recognition model for identifying entity association relationships is added to the entity extraction model to construct an entity comprehensive recognition model, and the first loss function of the entity extraction model and the second loss function of the entity relationship recognition model are integrated into a comprehensive loss function of the entity comprehensive recognition model. The entity relationship recognition model is constructed based on a support vector machine.

[0139] Generate an entity relationship training set based on historical patient data pre-stored in the hospital database, and use the entity relationship training set to train the entity comprehensive recognition model until the comprehensive loss function converges to a minimum value;

[0140] Perform entity normalization on patient data entities;

[0141] The normalized patient data entities are input into the trained entity comprehensive recognition model, and the data entity association relationships in the patient data entities are extracted through the entity comprehensive recognition model. The data entity association relationships include disease-symptom relationships, drug-indication relationships, treatment-effect relationships, and disease-risk factor relationships.

[0142] Extracting the medical record content association relationship between the patient data entity and the content item entity based on the matching relationship between the patient data and all content items in the target electronic medical record;

[0143] Taking patient data entities and content item entities as entity graph nodes, and data entity association relationships and medical record content association relationships as entity graph node edges between entity graph nodes, a medical record content graph network of the target electronic medical record is constructed.

[0144] In this embodiment, an entity relationship recognition model for identifying entity association relationships is added on the basis of the entity extraction model to construct an entity comprehensive recognition model. The core of this step is to integrate the two tasks of entity recognition and relationship recognition into a unified framework. The entity extraction model is usually based on sequence labeling methods, such as conditional random fields (CRF) or bidirectional long short-term memory networks (Bi-LSTM), and its loss function can be expressed as cross entropy loss: L1=-Σy_ilog(p(y_i|x)), where y_i is the true label and p(y_i|x) is the probability predicted by the model. The entity relationship recognition model is built based on support vector machines (SVM), and its goal is to find an optimal hyperplane in a high-dimensional feature space to separate relationships of different categories. The loss function of SVM can be expressed as hinge loss: L2=Σmax(0,1-y_i(w·x_i+b)), where w is the weight vector and b is the bias term. In order to integrate these two models, a comprehensive loss function needs to be designed: L=αL1+βL2+γR(θ), where α and β are weight coefficients to balance the importance of the two tasks, R(θ) is a regularization term used to prevent overfitting, and γ is a regularization coefficient. The training process of this comprehensive model involves optimizing both entity extraction and relationship recognition tasks at the same time, which can be achieved through the framework of multi-task learning. In practical applications, gradient descent or its variants (such as Adam optimizer) can be used to minimize the comprehensive loss function. The advantage of this integration method is that it can exploit the interdependence between entity recognition and relationship recognition tasks to improve the overall performance. For example, knowing that there is a "treatment-effect" relationship between two entities can help more accurately identify the types of these two entities.

[0145] Next, historical patient data is extracted from the hospital database, which usually contains a large number of electronic medical records, examination reports, and treatment records. Next, entities and relationships are annotated using automatic annotation tools combined with manual verification. For example, predefined medical dictionaries and rules can be used to preliminarily annotate entities, and then dependency syntax analysis or distance rules can be used to preliminarily identify the relationships between entities. The generated training set is usually in a specific format, such as JSON or XML, and each sample contains original text, entity annotations, and relationship annotations. The training process uses a small batch stochastic gradient descent method, and each batch contains a certain number of samples (such as 32 or 64). In each training round (epoch), the model performs a complete traversal of all training samples. For each batch, the value and gradient of the comprehensive loss function are calculated, and then the model parameters are updated. The calculation formula of the comprehensive loss function is: L=αL1+βL2+γR(θ), where L1 is the loss of entity recognition, L2 is the loss of relationship recognition, and R(θ) is the regularization term. The parameter update formula is: θ=θ-η▽L(θ), where η is the learning rate, which can be dynamically adjusted using an adaptive learning rate algorithm such as Adam. This training method can make full use of the hospital's rich historical data, allowing the model to learn the entity and relationship patterns in real medical scenarios, and improve the accuracy and robustness of the model in practical applications.

[0146] Entity normalization of patient data entities is a key step to ensure data quality and consistency. This process involves multiple subtasks, including entity disambiguation, synonym processing, unit unification, and format standardization. First, entity disambiguation aims to solve the problem that the same expression may refer to different entities. For example, "surgery" may refer to a specific surgical procedure or the name of a department. The disambiguation process can use contextual information and knowledge bases to determine the most likely reference by calculating word vector similarity or using probabilistic graphical models. Synonym processing is to unify entities with different expressions but the same meaning into a standard form. This can be achieved through a pre-built medical synonym dictionary, or semantic similarity can be calculated using word embedding technology. For example, "myocardial infarction", "heart attack", "MI", etc. are standardized as "myocardial infarction". Unit unification is to convert numerical values ​​of different units into a unified standard unit.

[0147] The normalized patient data entities are input into the trained entity comprehensive recognition model, and the data entity association relationships in the patient data entities are extracted through the entity comprehensive recognition model. The core of this step is to use the trained model to identify and extract various types of entity relationships, including disease-symptom relationships, drug-indication relationships, treatment-effect relationships, and disease-risk factor relationships. The entity comprehensive recognition model is usually based on deep learning architectures, such as long short-term memory networks (LSTM) or transformer models, combined with attention mechanisms to capture long-distance dependencies between entities. The input of the model is the normalized patient data entities, which have been converted into word vectors or character-level embeddings. For each pair of entities that may have a relationship, the model generates a relationship representation vector. This vector is converted into a probability distribution of relationship categories through a fully connected layer and a softmax function. For example, for the entity pair "headache" and "migraine", the model may output {"symptom-disease": 0.8,"irrelevant": 0.2}, indicating that there is an 80% probability that a symptom-disease relationship exists. In actual application, a probability threshold (such as 0.5) can be set to determine whether a relationship is established. For complex and long texts, the model may need to use sliding window technology to ensure that the relationship between distant entities is captured. In addition, the model may also use a multi-head attention mechanism to focus on multiple related entities at the same time to improve the accuracy of relationship recognition. For example, when identifying the drug-indication relationship between "aspirin" and "heart disease", the model may also pay attention to the keyword "prevention". The identified relationship is usually expressed in the form of a triple, such as (aspirin, used for prevention, heart disease).

[0148] According to the matching relationship between the patient data and all content items in the target electronic medical record, the medical record content association relationship between the patient data entity and the content item entity is extracted. This process first requires a clear definition of the content item structure of the electronic medical record, which usually includes basic patient information, chief complaint, current medical history, past medical history, physical examination, auxiliary examination, diagnosis, treatment plan, etc. Each content item can be regarded as an entity, and the patient data entity related to it forms an association relationship. The matching process can adopt a strategy combining multiple technical methods. First, a rule-based method can be used to preliminarily match content items and patient data through predefined templates and keywords. For example, the "chief complaint" section usually contains symptom descriptions, and regular expressions can be used to match symptom-related words. Secondly, a text similarity algorithm, such as TF-IDF (term frequency-inverse document frequency) combined with cosine similarity, can be used to calculate the similarity between patient data and the description of each content item. In addition, more complex semantic matching algorithms, such as word embedding-based methods or deep learning models (such as BERT), can be used to capture deeper semantic associations. In practical applications, it may be necessary to consider multiple matching indicators comprehensively and use weighted average or voting mechanisms to make the final decision. The matching process also needs to consider the hierarchical relationship and logical order between content items. For example, the disease mentioned in "Past History" may affect the formulation of "Diagnosis" and "Treatment Plan". Therefore, this temporal and causal relationship needs to be considered when establishing associations.

[0149] The medical record content graph network of the target electronic medical record is constructed by taking patient data entities and content item entities as entity graph nodes, and taking data entity associations and medical record content associations as entity graph node edges. This graph network is essentially a knowledge graph, where each node represents an entity (such as a patient data entity or a content item entity), and each edge represents the relationship between entities. The construction process first requires defining the structure of the graph. An attribute graph model is usually used, in which attributes can be attached to each node and edge. For example, a patient data entity node may have attributes such as "type" (such as symptoms, test results), "timestamp", and a relationship edge may have attributes such as "relationship type" and "confidence". The graph construction process can be divided into several steps:

[0150] Node creation: Create a unique node for each patient data entity and content item entity. This requires an efficient entity recognition and deduplication mechanism, and a hash table can be used to quickly check whether an entity already exists.

[0151] Edge connection: Create edges based on the previously identified data entity associations and medical record content associations. This step needs to consider the directionality of the relationship, for example, the "symptom-disease" relationship is directed, while the "complication" relationship may be undirected.

[0152] Attribute assignment: Add corresponding attributes to nodes and edges. This may involve data type conversion and normalization to ensure the consistency of attribute values.

[0153] Graph optimization: Optimize the constructed graph, including removing redundant edges, merging similar nodes, etc. Graph algorithms such as transitive closure can be used to infer and add implicit relationships.

[0154] In the implementation process, a graph database (such as Neo4j) or a specialized graph processing library (such as NetworkX) can be used to store and manipulate the graph structure. After the graph network is constructed, various analysis and query operations can be performed. For example, the shortest path algorithm can be used to find the potential association between two symptoms, or the community detection algorithm can be used to identify closely related symptom groups. Graph visualization is also an important aspect, and algorithms such as force-directed layout can be used to generate intuitive graphical representations. The advantage of this graph network representation is that it can intuitively display the associations between complex medical information and support efficient information retrieval and reasoning.

[0155] In one implementation, abnormal content detection is performed on a target electronic medical record based on a medical record content graph network. If the target electronic medical record passes the abnormal content detection, saving the target electronic medical record includes the following steps:

[0156] Use multi-layer perceptron to identify whether there are local information contradictions in the medical record content graph network. Local information contradictions include node-level information contradictions and edge-level information contradictions.

[0157] Use graph pooling technology to obtain multiple medical record content subgraph representations of the medical record content graph network, and identify whether there is a global information contradiction in the medical record content graph network based on all medical record content subgraph representations;

[0158] If there is no local information contradiction and global information contradiction in the medical record content graph network, a key subgraph network is extracted from the medical record content graph network according to key content items and key text data of patients;

[0159] Extract high-dimensional topological features from key subgraph networks based on feature mapping methods;

[0160] The high-dimensional topological features are input into a pre-trained deep abnormal content recognition model, and the target electronic medical record is judged whether there is abnormal content according to the high-dimensional feature recognition results output by the deep abnormal content recognition model. The deep abnormal content recognition model is built based on a graph neural network model.

[0161] If the target electronic medical record does not contain abnormal content, it is determined that the target electronic medical record passes the abnormal content detection, and the target electronic medical record is saved.

[0162] In this embodiment, a multilayer perceptron (MLP) is a feedforward neural network consisting of an input layer, one or more hidden layers, and an output layer. In this embodiment, MLP is used to detect information contradictions at the node level and edge level. For contradictions at the node level, the input may include features such as the attribute vector of the node, the type and number of edges directly connected to it, etc. For example, a node represents "high blood pressure", but its attributes show that blood pressure is normal, which may be a contradiction. For contradictions at the edge level, the input may include features of the two end nodes of the edge, the type and attributes of the edge, etc. For example, an edge represents a "treatment" relationship, connecting "diabetes" and "increase sugar intake", which is obviously a contradiction. The structure of MLP may be as follows: the number of neurons in the input layer is equal to the dimension of the feature vector, there may be multiple hidden layers, and the number of neurons in each layer may be about 2 / 3 of the input dimension, and the output layer has two neurons, which respectively represent the probability of the existence and non-existence of contradictions. The activation function can select ReLU, and the output layer uses the softmax function. In this way, local information contradictions in medical records can be effectively identified, providing an important basis for subsequent data cleaning and quality control.

[0163] Graph pooling is a technique for compressing a large graph into a small graph while retaining the key structural information of the original graph. In this application, multiple graph pooling methods can be used, such as topological pooling (TopKPooling) or attention pooling (SAGPooling). Taking SAGPooling as an example, it uses a graph attention network (GAT) to learn the importance scores of nodes, and then selects the nodes with the highest scores to form a subgraph. The mathematical expression is as follows:

[0164]

[0165] Where X is the node feature matrix, is a learnable parameter, σ is a nonlinear activation function, k is the number of nodes retained, and N is the number of nodes in the original graph. By applying this pooling operation multiple times, multiple subgraph representations of different scales can be obtained. Each subgraph representation captures certain aspects of the original image and can be regarded as a "summary" of the medical record content.

[0166] To identify global information contradictions, these subgraph representations can be input into a graph-level classifier, such as a graph isomorphism network (GIN). Each layer of GIN can be expressed as: h^(k)=MLP^(k)(h^(k-1)+Σ(h^(k-1)_neighbor)). where h^(k) is the node representation of the kth layer. The output of the last layer can be obtained by global average pooling to obtain a graph-level representation, and then passed through a fully connected layer and a softmax function to obtain the final classification result. The advantage of this approach is that it can capture long-distance dependencies and complex patterns across the entire medical record, thereby identifying global contradictions that may be overlooked by local analysis. For example, it may find inconsistencies between medical history descriptions and final diagnoses, or mismatches between treatment plans and overall disease severity.

[0167] If there are no local information contradictions and global information contradictions in the medical record content graph network, a key subgraph network is extracted from the medical record content graph network based on the key content items and the patient's key text data. The extraction process can use a graph traversal algorithm, such as depth-first search (DFS) or breadth-first search (BFS), starting from the key node, and exploring the nodes and edges directly related to it. For example, starting from the "chief complaint" node, it may be connected to the related "symptom" node, then to the "diagnosis" node, and finally to the "treatment" node. During the traversal process, heuristic rules can be used to decide whether to include a node or edge in the key subgraph. These rules can include: the type of node, the degree of the node, the weight of the edge, and the attribute value of the node or edge. An importance function I(v) can also be defined to evaluate the importance of node v:

[0168]

[0169] Among them, w1, w2, w3, w4 are weight coefficients, which can be optimized by machine learning methods such as logistic regression.

[0170] Extracting high-dimensional topological features from key subgraph networks based on feature mapping methods is a complex process of converting graph structural information into vector representations. The purpose of this step is to capture the structural characteristics of the graph so that subsequent machine learning algorithms can process this information. There are many feature mapping methods, including the extraction of node-level features, edge-level features, and graph-level features. A commonly used feature mapping method is the GraphKernel. The graph kernel function K(G1, G2) calculates the similarity of two graphs, or uses graph embedding techniques such as node2vec or graph2vec to map the entire graph to a fixed-dimensional vector space. These methods are usually based on deep learning and can automatically learn graph representations. The result of feature mapping is a high-dimensional vector whose dimensions may reach hundreds or thousands, and each dimension represents a specific topological feature of the graph. This high-dimensional representation can capture the complex structural information of the graph and provide rich input for subsequent anomaly detection.

[0171] The high-dimensional topological features are input into a pre-trained deep abnormal content recognition model, which is built on a graph neural network (GNN) and can effectively process graph structured data. The core idea of ​​the GNN model is to learn the representation of nodes through a message passing mechanism. In this embodiment, architectures such as graph convolutional networks (GCN) or graph attention networks (GAT) can be used. Taking GCN as an example, its basic operation can be expressed as: H^(l+1)=σ(D^(-1 / 2)AD^(-1 / 2)H^(l)W^(l)). Where A is the adjacency matrix, D is the degree matrix, H^(l) is the node feature of the lth layer, W^(l) is a learnable weight matrix, and σ is a nonlinear activation function.

[0172] The model may contain multiple such layers, each of which aggregates neighbor information and updates node representations. The output of the last layer can be seen as a representation of the entire graph, which is used for anomaly detection tasks. Anomaly detection can be achieved through classification methods, treating the problem as a binary classification task and directly outputting the probability of whether the medical record is abnormal or not. The last layer of the model can be a fully connected layer followed by a softmax function. The model uses a large number of labeled normal and abnormal medical record samples for learning during the training phase, and uses the cross entropy loss function and back propagation algorithm to optimize the parameters. During the inference phase, the high-dimensional topological features of the target electronic medical record are input into the model to obtain the anomaly probability. The advantage of this method is that it can capture complex anomaly patterns and is not limited to simple rule matching. For example, it may identify inconsistencies between symptoms, diagnoses, and treatments, or discover rare but potentially dangerous drug interactions. Through this deep learning approach, the accuracy and efficiency of abnormal content detection can be greatly improved.

[0173] If the target electronic medical record does not contain abnormal content, the target electronic medical record is determined to have passed the abnormal content detection and the target electronic medical record is saved. Specifically, the system will assign a unique identification code to this medical record, which may be in the form of a UUID (universally unique identifier). This identification code will be used for all subsequent related operations to ensure the uniqueness and traceability of the medical record. The preservation process of medical records involves multiple aspects. The first is the encrypted storage of data, which may use algorithms such as the Advanced Encryption Standard (AES) to ensure data security. The second is data backup, which may adopt a multi-copy storage strategy to store data in multiple physical locations at the same time to prevent data loss. In addition, data version control needs to be considered to record the history of each modification in order to track the change process of medical records.

[0174] In one embodiment, extracting high-dimensional topological features from a key subgraph network based on a feature mapping method includes the following steps:

[0175] For any entity graph node in the key subgraph network, the entity graph node is used as a first entity graph node, and any entity graph node of a different type other than the first entity graph node in the key subgraph network is used as a second entity graph node;

[0176] Mapping node features of the first entity graph node and the second entity graph node to a unified high-dimensional feature space to obtain high-dimensional node features;

[0177] Combining the high-dimensional node features and the type of association relationship between the first entity graph node and the second entity graph node, constructing node heterogeneous attention between the first entity graph node and the second entity graph node in the high-dimensional feature space;

[0178] In combination with the association relationship type and the node type of the first entity graph node and the second entity graph node, an information conduction function of the first entity graph node is constructed in the high-dimensional feature space;

[0179] The neighbor aggregation information of the first entity graph node is obtained based on the node heterogeneous attention and information conduction function;

[0180] After mapping the neighbor aggregation information of all entity graph nodes to the original feature space, the high-dimensional topological features of the key subgraph network are aggregated.

[0181] In this embodiment, for any entity graph node in the key subgraph network, the entity graph node is used as the first entity graph node, and any different type of entity graph node in the key subgraph network except the first entity graph node is used as the second entity graph node. Specifically, a nested loop can be used to traverse all possible node pairs, or a more efficient sampling strategy can be used, such as random walk or neighborhood sampling. This process can be expressed as a function f(G,v)→{(v,u)|u∈N(v),type(u)≠type(v)}, where G is the key subgraph network, v is the first entity graph node, N(v) is the neighbor set of v, and type(·) returns the node type. This pairing method allows the analysis of interactions between different types of entities in subsequent steps, such as the relationship between symptoms and diagnosis, the relationship between diagnosis and treatment, etc. Next, the node features of the first entity graph node and the second entity graph node are mapped to a unified high-dimensional feature space to obtain high-dimensional node features. This step is the key to achieving heterogeneous graph information fusion. In the context of medical electronic medical records, different types of nodes may have original features of different dimensions and scales. For example, a symptom node may contain features such as severity and duration, while a diagnosis node may contain features such as disease code and diagnosis probability. In order to make these heterogeneous features comparable and fusible, they need to be mapped into a common high-dimensional space.

[0182] Specifically, for each node v, its original feature vector x_v is mapped to a high-dimensional space through a transformation function φ: h_v=φ(x_v)=σ(W·x_v+b). Where W is the weight matrix, b is the bias vector, and σ is a nonlinear activation function (such as ReLU, tanh, etc.). This transformation can be multi-layered, and each layer can be expressed as: h_v^(l)=σ(W^(l)·h_v^(l-1)+b^(l)). Where l represents the number of layers. The output h_v^(L) of the last layer is the high-dimensional feature representation of node v. In order to process the features of different types of nodes, specific transformation functions can be designed for each type of node. For example:

[0183] h_symptom=φ_symptom(x_symptom)

[0184] h_diagnosis=φ_diagnosis(x_diagnosis)

[0185] This ensures that different types of node features are appropriately mapped to the same high-dimensional space. The advantage of this mapping method is that it can unify medical information of different types and scales into a common representation space, so that subsequent attention mechanisms and information transfer can be carried out on a consistent basis. This is crucial for capturing complex medical entity relationships.

[0186] Next, combining the high-dimensional node features and the type of association relationship between the first entity graph node and the second entity graph node, the node heterogeneous attention between the first entity graph node and the second entity graph node is constructed in the high-dimensional feature space. This step aims to capture the complex interactions between different types of medical entities. In the context of medical electronic medical records, the types of association relationships may include "cause", "symptom-diagnosis", "diagnosis-treatment", etc. The node heterogeneous attention mechanism allows the model to dynamically adjust the attention allocation according to different relationship types, thereby more accurately capturing the correlation between medical entities. The specific implementation can be based on the multi-head attention mechanism. For the first entity graph node i and the second entity graph node j, and the relationship type r between them, the heterogeneous attention can be expressed as: α_ij^r=softmax(e_ij^r)=exp(e_ij^r) / Σ_kexp(e_ik^r). Where e_ij^r is the attention score, which can be calculated as follows: e_ij^r=LeakyReLU(a_r^T[W_rh_i||W_rh_j]). Here, W_r is a relation-specific transformation matrix, a_r is a relation-specific attention vector, || represents vector concatenation, and LeakyReLU is the activation function. The multi-head attention mechanism can further improve the expressiveness of the model:

[0187] α_ij^r,k=softmax(e_ij^r,k)

[0188] h_i'=||_k=1^Kσ(Σ_jα_ij^r,kW_r^kh_j)

[0189] Where K is the number of attention heads and σ is the non-linear activation function.

[0190] In practical applications, specific attention calculation methods can be designed for different types of relationships. For example, for the "symptom-diagnosis" relationship, more attention may be paid to the severity and duration of the symptoms; while for the "diagnosis-treatment" relationship, more attention may be paid to the certainty of the diagnosis and the applicability of the treatment. The advantage of this heterogeneous attention mechanism is that it can flexibly handle complex relationships between medical entities. For example, when analyzing the relationship between the "headache" symptom and the "migraine" diagnosis, the model can dynamically adjust the attention weights based on the characteristics of the symptoms (such as frequency, intensity) and the characteristics of the diagnosis (such as probability, related test results). This enables the model to better understand and represent complex medical relationship networks. By constructing this node heterogeneous attention, the model can more accurately capture the mutual influence between different types of medical entities, providing a strong foundation for subsequent information aggregation and decision support.

[0191] Next, the information transmission function of the first entity graph node is constructed in the high-dimensional feature space by combining the type of association relationship and the node types of the first entity graph node and the second entity graph node. In the context of heterogeneous graphs, the information transmission function needs to be able to flexibly handle different types of nodes and multiple relationship characteristics between them. This is particularly important for medical electronic medical record analysis, because different types of nodes (such as symptoms, diagnoses) and different types of relationships (such as "diagnosis-treatment") require information to flow in different ways. The message passing framework in graph neural networks (GNNs) can be used to build the information transmission function. For each pair of nodes i and j, assuming that there is a relationship r between them, the information transmission function can be expressed as a message passing process from node j to node i. Specifically, it can be implemented through the following formalized process: m_i^r=Σ_j∈N(i)α_ij^rh_j^r. Among them, m_i^r represents the information aggregated by node i from all related nodes j, α_ij^r is the node heterogeneous attention weight calculated previously, and h_j^r is the feature representation of node j after being transformed by relationship r.

[0192] In order to make the message passing mechanism of different relationship types more flexible, a specific transfer function can be defined for each relationship r. This is usually achieved by combining linear transformation and nonlinear activation: h_i^(t+1)=σ(W_i^rm_i^r+b_i^r). Here, W_i^r and b_i^r are relationship-specific weights and biases, and σ is an activation function (such as ReLU). This transmission mechanism ensures that the propagation of information between nodes can reasonably reflect the combined effects of node type and relationship characteristics. In addition, in order to enhance the expressiveness of the model, multi-layer information aggregation with skip connection can be used: h_i^(t+1)=σ(W_i^rm_i^r+b_i^r)+h_i^(t). This structure helps to alleviate the gradient vanishing problem that may occur during the multi-layer transmission of information.

[0193] The design of this information conduction function enables the model to efficiently combine node heterogeneous attention in high-dimensional feature space and map the complex interactive relationship between different types of nodes into a clear information flow. For example, the symptom "fever" and the diagnosis "flu" may be combined with the temperature and duration of fever, the test results of flu and case data. Through appropriate conduction functions, the model can effectively aggregate these heterogeneous information in high-dimensional space. The advantage of this approach is that it can achieve effective information exchange between multi-relationship and multi-type nodes, improving the model's ability to express complex networks. The successful construction of the information conduction function lays the foundation for subsequent neighbor aggregation and full-graph feature extraction, and is the core of information flow and relationship modeling. Through this mechanism, efficient information exchange can be carried out between different entities in the key subgraph network to ensure the accurate extraction of high-dimensional topological features.

[0194] In a heterogeneous graph, different types of nodes and relationships may convey different types of information, so accurately aggregating this information is crucial to capturing the overall characteristics of the node. For the target node i, the aggregated information from all its neighboring nodes is calculated. Using the previously defined heterogeneous attention mechanism and information conduction function, the most relevant information is selected from the neighboring nodes and aggregated. The specific mathematical expression is: h_i^'=Σ_j∈N(i)α_ij^rh_j^r. Among them, h_i^' is the feature representation of node i after aggregation, α_ij^r represents the attention weight transferred from node j to node i, and h_j^r is the neighbor node feature processed by the specific information conduction function of relationship type r. By weighted summation, it is ensured that the information of different neighbors is effectively aggregated according to their association strength.

[0195] In order to avoid information loss during layer transmission, residual connections can be introduced to combine aggregate information with original node features: h_i^(t+1)=g(h_i^(t),h_i^')=σ(W_i[h_i^(t)||h_i^']+b_i). Where g is the aggregation function, which can be a simple addition or a more complex combination, such as using a convolution operation, W_i and b_i are the weights and biases of the linear transformation, and σ is a nonlinear activation function. In this way, it can be ensured that the aggregate information of each layer not only depends on the characteristics of the neighboring nodes, but also takes into account the historical characteristics of the target node itself, so as to better maintain the integrity of the information. The advantage of this neighbor aggregation method is that it can dynamically adapt to the heterogeneity of medical data. For example, information from different clinical sources (such as laboratory results, doctor observations) can be weighted and aggregated through a heterogeneous attention mechanism, while the information transmission function ensures that these heterogeneous information can be filtered and integrated in a consistent manner, ultimately forming a comprehensive representation of the target node.

[0196] After mapping the neighbor aggregation information of all entity graph nodes to the original feature space, the high-dimensional topological features of the key subgraph network are aggregated. Summarizing the node-level information obtained in the previous steps helps to understand the characteristics of the entire key subgraph network at a more macro level. This process first involves mapping the high-dimensional aggregate features of the nodes back to the original feature space or other specified low-dimensional space for subsequent comprehensive analysis and visualization. This mapping can be achieved using linear dimensionality reduction techniques, such as principal component analysis (PCA), or nonlinear ones such as t-SNE. This transformation retains the main features of the topological information obtained in the high-dimensional space, while reducing the dimensionality of the feature vector and improving computational efficiency. Next, the mapped features need to be aggregated according to application requirements, which can usually be achieved through summation, averaging, maximization, etc. The aggregated global features comprehensively reflect the structure and attribute information of each node and its neighborhood.

[0197] The present invention also discloses an electronic medical record management system based on AI computing, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the electronic medical record management method based on AI computing as described in any one of the above embodiments is implemented.

[0198] Among them, the processor can adopt a central processing unit (CPU). Of course, according to actual usage, other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. can also be adopted. The general-purpose processor can adopt a microprocessor or any conventional processor, etc., and this application does not impose any restrictions on this.

[0199] Among them, the memory can be an internal storage unit of a computer device, such as a hard disk or memory of a computer device, or an external storage device of a computer device, such as a plug-in hard disk, a smart memory card (SMC), a secure digital card (SD) or a flash memory card (FC) equipped on the computer device, and the memory can also be a combination of an internal storage unit and an external storage device of a computer device. The memory is used to store computer programs and other programs and data required by the computer device. The memory can also be used to temporarily store data that has been output or is to be output, and this application does not impose any restrictions on this.

[0200] An embodiment of the present application further discloses a computer-readable storage medium, and the computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the electronic medical record management method based on AI calculation in the above embodiment is adopted.

[0201] Among them, the computer program can be stored in a machine-readable medium, the computer program includes computer program code, the computer program code can be in the form of source code, object code, executable file or certain middleware, etc. The machine-readable medium includes any entity or device that can carry the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal and software distribution medium, etc. It should be noted that the machine-readable medium includes but is not limited to the above-mentioned components.

[0202] Among them, through this computer-readable storage medium, the comprehensive fault detection method for transmission lines in the above embodiment is stored in a computer-readable storage medium, and is loaded and executed on a processor to facilitate the storage and application of the above method.

[0203] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.

[0204] A person skilled in the art should understand that the discussion of any of the above embodiments is merely illustrative and is not intended to imply that the scope of protection of the present application is limited to these examples. In line with the concept of the present application, the technical features in the above embodiments or different embodiments may be combined, the steps may be implemented in any order, and there are many other variations of different aspects of one or more embodiments of the present application as above, which are not provided in detail for the sake of simplicity.

[0205] One or more embodiments of the present application are intended to cover all such substitutions, modifications and variations that fall within the broad scope of the present application. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of one or more embodiments of the present application should be included in the protection scope of the present application.

Claims

1. An electronic medical record management method based on AI computing, characterized in that: The steps include: Designing a universal template for electronic medical records, wherein the universal template includes a plurality of universal content items; Collect complete patient data through multiple terminals and pre-process the patient data; Inputting the preprocessed patient data into a patient data classification model, completing identification and classification of the patient data through the patient data classification model, and dividing the patient data into patient numerical data and patient text data according to the identification and classification results, wherein the patient text data includes patient key text data and patient non-key text data; Marking key content items in all the universal content items in the universal template according to the recognition and classification results to obtain a target universal template; Performing abnormal data detection on the patient numerical data and the patient text data based on the target universal template, and if both the patient numerical data and the patient text data pass the abnormal data detection, generating a target electronic medical record based on the target universal template and in combination with the patient numerical data and the patient text data; Extracting patient data entities and content item entities from the target electronic medical record using a named entity recognition model; An entity relationship recognition model for identifying entity association relationships is added to the entity extraction model to construct an entity comprehensive recognition model, and a first loss function of the entity extraction model and a second loss function of the entity relationship recognition model are integrated into a comprehensive loss function of the entity comprehensive recognition model, wherein the entity relationship recognition model is constructed based on a support vector machine; Generate an entity relationship training set based on historical patient data pre-stored in a hospital database, and use the entity relationship training set to train the entity comprehensive recognition model until the comprehensive loss function converges to a minimum value; Performing entity normalization processing on the patient data entity; Inputting the patient data entity after entity normalization processing into the trained entity comprehensive recognition model, and extracting the data entity association relationship in the patient data entity through the entity comprehensive recognition model, wherein the data entity association relationship includes disease-symptom relationship, drug-indication relationship, treatment-effect relationship and disease-risk factor relationship; Extracting the medical record content association relationship between the patient data entity and the content item entity according to the matching relationship between the patient data and all content items in the target electronic medical record; The patient data entity and the content item entity are used as entity graph nodes, and the data entity association relationship and the medical record content association relationship are used as entity graph node edges between the entity graph nodes to construct a medical record content graph network of the target electronic medical record; Using a multi-layer perceptron to identify whether there is a local information contradiction in the medical record content graph network, wherein the local information contradiction includes a node-level information contradiction and an edge-level information contradiction; Using a graph pooling technique to obtain multiple medical record content subgraph representations of the medical record content graph network, and identifying whether there is a global information contradiction in the medical record content graph network based on all the medical record content subgraph representations; If the medical record content graph network does not have the local information contradiction and the global information contradiction, extracting a key subgraph network from the medical record content graph network according to the key content items and the patient key text data; Extracting high-dimensional topological features from the key subgraph network based on a feature mapping method; Inputting the high-dimensional topological features into a pre-trained deep abnormal content recognition model, and judging whether the target electronic medical record has abnormal content according to the high-dimensional feature recognition results output by the deep abnormal content recognition model, wherein the deep abnormal content recognition model is constructed based on a graph neural network model; If the target electronic medical record does not have the abnormal content, it is determined that the target electronic medical record passes the abnormal content detection, and the target electronic medical record is saved.

2. The electronic medical record management method based on AI calculation according to claim 1, characterized in that: The method of collecting complete patient data through multiple terminals and preprocessing the patient data comprises the following steps: Collect complete patient data through multiple terminals, including outpatient doctor terminal, resident doctor terminal, nurse terminal, hospital inspection terminal, hospital imaging terminal and patient client terminal; cleaning duplicate data in the patient data; Standardizing the patient data based on a preset medical term ontology library and using natural language processing technology; The standardized patient data is subjected to data structuring processing.

3. The electronic medical record management method based on AI calculation according to claim 2, characterized in that: The step of inputting the preprocessed patient data into a patient data classification model, completing identification and classification of the patient data through the patient data classification model, and dividing the patient data into patient numerical data and patient text data according to the identification and classification results comprises the following steps: Inputting the preprocessed patient data into a patient data classification model, identifying numerical data and text data in the patient data through a logistic regression module in the patient data classification model to obtain a first recognition and classification result, and dividing the patient data into patient numerical data and patient text data according to the first recognition and classification result; Inputting the patient text data into a department classification module in the patient data classification model, outputting a second recognition and classification result through the department classification module, and determining the core department to which the patient text data belongs according to the second recognition and classification result, wherein the department classification module is constructed based on a Transformer architecture; According to the second identification and classification result output by the department classification module, and using the data annotation module in the patient data classification model, the key patient text data corresponding to the core department is annotated in the patient text data.

4. The electronic medical record management method based on AI calculation according to claim 3 is characterized in that: The step of performing abnormal data detection on the patient numerical data and the patient text data based on the target universal template, and if both the patient numerical data and the patient text data pass the abnormal data detection, generating a target electronic medical record based on the target universal template and in combination with the patient numerical data and the patient text data comprises the following steps: Performing abnormal value detection on the patient numerical data according to a preset numerical threshold list, and if there is any one or more patient numerical data exceeding the preset numerical threshold, determining that the patient numerical data fails the abnormal data detection; If there is no patient numerical data exceeding the preset numerical threshold, determining that the patient numerical data passes the abnormal data detection; Matching all of the patient key text data to the corresponding key content items in the target universal template; If any one or more of the key content items fail to successfully match the patient key text data, it is determined that the patient text data fails the abnormal data detection; If all of the key content items match at least one of the patient key text data, a target electronic medical record is generated based on the target universal template and in combination with the patient numerical data and the patient text data.

5. The electronic medical record management method based on AI calculation according to claim 1, characterized in that: The step of extracting patient data entities and content item entities from the target electronic medical record using a named entity recognition model comprises the following steps: Extracting a content item entity and a first patient data entity from the target electronic medical record respectively by using a pattern matching rule method in combination with a predefined medical record dictionary and a medical dictionary; After deleting the patient data corresponding to the first patient data entity in the target electronic medical record, if there is still target patient data to be entity extracted in the target electronic medical record, extracting the target patient data into a second patient data entity through a pre-trained entity extraction model, wherein the entity extraction model is constructed based on a conditional random field model or a support vector machine; The first patient data entity and the second patient data entity are integrated into a patient data entity.

6. The electronic medical record management method based on AI calculation according to claim 1, characterized in that: The extracting of high-dimensional topological features from the key subgraph network based on the feature mapping method comprises the following steps: For any entity graph node in the key subgraph network, use the entity graph node as a first entity graph node, and use any entity graph node of a different type in the key subgraph network except the first entity graph node as a second entity graph node; Mapping node features of the first entity graph node and the second entity graph node to a unified high-dimensional feature space to obtain high-dimensional node features; In combination with the high-dimensional node feature and the type of association relationship between the first entity graph node and the second entity graph node, constructing a node heterogeneous attention between the first entity graph node and the second entity graph node in the high-dimensional feature space; In combination with the association relationship type and the node types of the first entity graph node and the second entity graph node, constructing an information conduction function of the first entity graph node in the high-dimensional feature space; Obtaining neighbor aggregation information of the first entity graph node based on the node heterogeneous attention and the information conduction function; After mapping the neighbor aggregation information of all the entity graph nodes to the original feature space, high-dimensional topological features of the key subgraph network are obtained by aggregation.

7. An electronic medical record management system based on AI computing, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the electronic medical record management method based on AI calculation as described in any one of claims 1 to 6 is implemented.

8. A computer-readable storage medium having a computer program stored therein, characterized in that: When the computer program is loaded and executed by the processor, the electronic medical record management method based on AI calculation described in any one of claims 1 to 6 is adopted.

Citation Information

Patent Citations

  • Medical information processing method and device

    CN118098471A

  • Medical record document review method and system based on electronic medical record management system

    CN118468886A