AI-based medical record intelligent analysis and pre-filling system

By using an AI-based intelligent medical record parsing and pre-filling system, deep learning and medical knowledge graphs are employed for deep semantic parsing and multi-objective optimization. This solves the problem of insufficient uncertainty parsing capability in existing intelligent medical record parsing methods, thereby improving the accuracy and security of medical record data.

CN120823938AActive Publication Date: 2025-10-21XIAN GEOMETRY DIGITAL INFORMATION TECH CO LTD

Patent Information

Application Number
CN202511320110.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-16
Publication Date
2025-10-21
Estimated Expiration
2045-09-16

AI Technical Summary

Technical Problem

Existing intelligent medical record parsing methods based on statistical machine learning paradigms lack the ability to accurately analyze uncertainties and negative statements in clinical texts, leading to misjudgments and affecting the accuracy and security of electronic medical record data.

Method used

An AI-based intelligent medical record parsing and pre-filling system is adopted, including a data acquisition and preprocessing module, an uncertainty feature parsing module, a contextual influence analysis module, and a multi-objective decision optimization module. Through deep learning and medical knowledge graphs, deep semantic parsing, contextual analysis, and multi-objective optimization are performed to generate structured medical data with deterministic labels.

Benefits of technology

It enables the deconstruction and digital representation of the degree of certainty in clinical texts, ensuring the accuracy and logical consistency of structured data, avoiding the risk of misjudgment, and improving the accuracy of medical record data processing and the reliability of clinical applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120823938A_ABST
    Figure CN120823938A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of medical information, and particularly discloses an AI-based medical record intelligent analysis and pre-filling system, which quantifies the certainty degree of clinical expression by constructing a medical entity uncertainty characteristic spectrum, and analyzes the semantic influence of quantified modifiers in combination with the context influence, so that the accuracy of medical record analysis and pre-filling is improved. A multi-objective optimization algorithm is adopted to cooperatively balance clinical safety, data integrity and filling efficiency objectives; and finally, performing ontology alignment and logic verification based on the medical knowledge graph, and implementing a differential pre-filling strategy according to a determinacy level. According to the method, the uncertainty semantics in the medical text can be accurately analyzed, the accuracy and clinical reliability of medical record data processing are improved, and meanwhile, self-adaptive optimization of intelligent pre-filling is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of medical information technology, and in particular to an AI-based intelligent medical record parsing and pre-filling system. Background Art

[0002] With the advancement of medical informatization, electronic medical record systems have been widely used in various medical institutions, generating massive amounts of medical text data. Automatically extracting key information from this unstructured medical text and implementing intelligent pre-reporting have become key technical challenges in improving medical data quality and clinical work efficiency.

[0003] The existing technology has the following deficiencies: Existing intelligent medical record parsing methods based on the statistical machine learning paradigm lack the ability to accurately identify and parse the deep semantics of uncertain and negative expressions commonly found in clinical texts during the core information extraction process. This problem makes it easy for the system to confuse the certainty of the diagnosis during entity recognition and relationship extraction, resulting in misjudgments at the clinical semantic level. These misjudgments are further amplified during subsequent structured processing and pre-reporting, generating electronic medical record data that deviates from clinical authenticity, thus posing a serious safety hazard and hindering the critical transition from laboratory validation to reliable clinical application. Summary of the Invention

[0004] The purpose of the present invention is to provide an AI-based intelligent medical record analysis and pre-filling system to solve the problems in the above background.

[0005] The purpose of the present invention can be achieved through the following technical solutions: An AI-based medical record intelligent parsing and pre-filling system, including: A data acquisition and preprocessing module receives medical record text data through a medical text data interface, and performs preprocessing and medical terminology standardization on the medical record text data; An uncertainty feature parsing module, which performs deep semantic analysis on standardized medical record text data, constructs a medical entity uncertainty feature spectrum, and calculates a semantic certainty offset value; A context influence analysis module, which performs a spatiotemporal joint analysis of the modifying context in the standardized medical record text data, constructs a clinical expression behavior profile, and calculates the context modification influence value; A multi-objective decision optimization module, which fuses the semantic deterministic offset value and the contextual modification influence value into a medical semantic feature vector, and inputs the medical semantic feature vector into a pre-trained medical record parsing model for multi-objective decision optimization; An intelligent pre-filling execution module performs ontology alignment and logical verification based on the optimized medical semantic feature vector and the medical knowledge graph, generates structured medical data with deterministic tags, and drives differentiated pre-filling strategies based on the level of deterministic tags, specifically including: automatically filling in high-certainty data, generating pending review tags for low-certainty data, and providing candidate filling content.

[0006] As a further solution of the present invention: the pre-processing of medical record text data and the standardization of medical terminology specifically include: Receive heterogeneous medical record data, perform sentence breakpoint detection and medical narrative paragraph segmentation on unstructured text; A deep scanner is used to identify medical entity boundaries and pre-label categories of segmented text units; Launch a multi-level medical term normalization pipeline, match the basic medical terminology database, and perform term semantic disambiguation; The semantic faults generated during the standardization process are reconstructed in context based on the knowledge graph to generate standardized text output with unified medical coding.

[0007] As a further solution of the present invention: performing deep semantic analysis on standardized medical record text data to construct a medical entity uncertainty feature spectrum specifically includes: A context-aware deep convolutional network is used to extract multi-scale semantic features from standardized text and generate a basic semantic feature matrix. The context segments containing suspicious words are weighted to form uncertainty-enhanced feature vectors; Establish a medical entity uncertainty association map and map the uncertainty-enhanced feature vectors with the disease-symptom probability relationship in the medical knowledge base; The multi-dimensional uncertainty feature spectrum of medical entities is generated by integrating the multi-scale semantic feature matrix, uncertainty enhanced feature vector and probability relationship mapping results.

[0008] As a further solution of the present invention: the calculation process of the semantic deterministic offset value is: Based on the uncertainty feature spectrum of medical entities, feature distribution dispersion index, context consistency index and knowledge base matching confidence index are extracted; Calculate the degree of dispersion of the eigenvalue distribution in the dimensions of certainty and uncertainty; The characteristic value is the specific value of each dimension of the uncertainty characteristic spectrum of the medical entity; Evaluate the degree of consistency between entity representation and the semantics of the overall medical record context; The distribution dispersion, the consistency of context semantics and the knowledge base matching are weighted and fused to generate a quantized semantic deterministic offset value.

[0009] As a further solution of the present invention: the spatiotemporal joint analysis of the modifying context in the standardized medical record text data to construct a clinical expression behavior profile specifically includes: A spatiotemporal convolutional neural network is used to perform multi-dimensional scanning of medical record texts, capturing the distribution patterns and temporal evolution characteristics of modifying context in text sequences. Statistical analysis of the spatiotemporal distribution of negative words, degree adverbs, and uncertainty expressions around specific medical entities; Match the captured distribution patterns with the preset clinical expression templates to identify the physician's expression habit types; Integrate spatiotemporal distribution characteristics and expression habit types to generate a multidimensional clinical expression behavior portrait that includes the intensity of the modification context, distribution breadth and evolution pattern.

[0010] As a further solution of the present invention: the calculation process of the context modification influence value is: Based on the clinical expression behavior profile, we extract the modification context intensity index, distribution breadth index and temporal stability index; Calculate the influence of negative words and uncertainty expressions on the semantics of medical entities, which is recorded as the modifying context intensity; Evaluate the scope and distance of influence of modifying context in the text, recorded as distribution breadth; The intensity, distribution breadth and temporal stability of the modifying context are weighted and aggregated to generate a quantitative context modification influence value.

[0011] As a further solution of the present invention: the process of constructing the medical semantic feature vector is: Standardize the semantic certainty offset value and contextual modification influence value to eliminate dimensional differences; Automatically adjust the fusion weight coefficient of the deterministic offset value and the contextual modification influence value according to the type characteristics of the current medical entity; Establish a nonlinear mapping relationship between the deterministic offset value and the contextual modification influence value to generate a fused feature representation with context-awareness; The weighted certainty offset value and contextual modification influence value are concatenated with the fusion feature representation to generate a medical semantic feature vector containing semantic certainty and contextual influence.

[0012] As a further solution of the present invention: the medical semantic feature vector is input into the pre-trained medical record parsing model to perform multi-objective decision optimization, specifically including: Synchronously optimize clinical safety goals, data integrity goals, and reporting efficiency goals through a multi-objective optimization function; Use the Pareto optimal solution search algorithm to find the best balance point among multiple optimization objectives and generate a set of candidate decision solutions; Use clinical knowledge constraint validators to verify the medical rationality of candidate solutions and eliminate decision-making solutions that do not meet clinical standards; Based on the needs of real-time application scenarios, select the best medical data analysis solution from verified candidate solutions.

[0013] As a further solution of the present invention: the ontology alignment and logic verification based on the optimized medical semantic feature vector and the medical knowledge graph specifically includes: Perform semantic similarity matching between medical semantic feature vectors and concept nodes in the medical knowledge graph; Conduct clinical validation of alignment results, including symptom-disease correlation verification and treatment plan rationality check; Ensure that the generated structured medical data remains consistent with the overall semantic environment of the medical record; Based on the alignment and verification results, each medical entity is labeled with a corresponding certainty level mark.

[0014] As a further solution of the present invention: the differentiated pre-filling strategy driven by the certainty marker level specifically includes: Directly fill in electronic medical record fields for highly certain data; Add visual review marks to low-certainty data and lock the corresponding fields; Generate candidate filling options that meet clinical standards based on medical knowledge graphs and similar case analysis; Adaptively adjust the priority and filling strategy of pre-filled fields based on real-time medical scenario requirements.

[0015] Beneficial effects of the present invention: (1) This invention deconstructs and digitally represents the degree of certainty in clinical texts by constructing a characteristic spectrum of medical entity uncertainty and quantifying the influence of contextual modification. It then relies on medical knowledge graphs to perform ontology alignment and logic verification to ensure that structured data conforms to terminology standards and maintains clinical logic consistency. Finally, it drives a differentiated pre-filling strategy based on the level of certainty tagging. While automatically processing high-certainty data to improve efficiency, it initiates review of low-certainty data and provides knowledge-driven candidate options, thereby avoiding the risk of misjudgment at the root and improving the accuracy of medical record data processing and the reliability of clinical applications.

[0016] (2) The present invention incorporates the three-dimensional objectives of clinical safety, data integrity and reporting efficiency into a unified mathematical optimization framework by establishing a multi-objective optimization function. The Pareto optimal solution search algorithm is used to find non-dominated solution sets in multiple mutually constrained objective spaces to generate candidate decision plans with different trade-off characteristics. For high-risk clinical scenarios, the system automatically increases the weight coefficients of clinical safety and data integrity; for high-efficiency demand scenarios, the weight configuration of the reporting efficiency target is appropriately optimized. This adaptive optimization ensures that the system always provides the solution that best matches the current clinical scenario at the output end, which not only ensures the accuracy and security of data in critical medical scenarios, but also achieves operational efficiency optimization in conventional scenarios, ultimately forming a set of accurate decision support systems that can intelligently adapt to the needs of different medical environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The present invention will be further described below with reference to the accompanying drawings.

[0018] Figure 1 It is a flow chart of the system of the present invention. DETAILED DESCRIPTION

[0019] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.

[0020] See also Figure 1 As shown, the present invention is an AI-based medical record intelligent parsing and pre-filling system, comprising: A data acquisition and preprocessing module receives medical record text data through a medical text data interface, and performs preprocessing and medical terminology standardization on the medical record text data; An uncertainty feature parsing module, which performs deep semantic analysis on standardized medical record text data, constructs a medical entity uncertainty feature spectrum, and calculates a semantic certainty offset value; A context influence analysis module, which performs a spatiotemporal joint analysis of the modifying context in the standardized medical record text data, constructs a clinical expression behavior profile, and calculates the context modification influence value; A multi-objective decision optimization module, which fuses the semantic deterministic offset value and the contextual modification influence value into a medical semantic feature vector, and inputs the medical semantic feature vector into a pre-trained medical record parsing model for multi-objective decision optimization; An intelligent pre-filling execution module performs ontology alignment and logical verification based on the optimized medical semantic feature vector and the medical knowledge graph, generates structured medical data with deterministic tags, and drives differentiated pre-filling strategies based on the level of deterministic tags, specifically including: automatically filling in high-certainty data, generating pending review tags for low-certainty data, and providing candidate filling content.

[0021] The data acquisition and preprocessing module serves as the data entry point and quality assurance foundation for the entire system. This module connects to medical data sources such as the hospital information system (HIS), laboratory information system (LIS), and radiology information system (RIS) through a medical text data interface, receiving medical record text data from these systems. This data includes, but is not limited to, outpatient medical records, inpatient medical records, laboratory test reports, discharge summaries, and other different types of medical documents. Due to the diverse sources of this data, its formats and standards vary and may include structured, semi-structured, and unstructured data. The module first parses and recognizes this heterogeneous data, extracting the text content that requires processing. For non-text data, such as scanned or image-based medical records, the module uses optical character recognition technology to convert it into processable medical record text data. During data reception, the module records metadata such as the data source, type, and receipt time, and assigns a unique identifier to each data object for subsequent tracking and processing. The module also provides data buffering and capacity adjustment to cope with potential data peaks from medical data sources and ensure stable system operation.

[0022] The received heterogeneous medical record data is first processed for sentence breakpoint detection and medical narrative paragraph segmentation. Sentence breakpoint detection is based on sentence boundary identification techniques from natural language processing, optimized for the specific characteristics of medical text. The system employs a deep learning-based sequence annotation model, incorporating common punctuation patterns, abbreviation patterns, and grammatical structural features in medical text to accurately identify sentence boundaries within the text. The model not only considers obvious markers such as periods and question marks, but also recognizes the discursive function of common punctuation marks such as ellipsis and semicolons. Based on sentence segmentation, the system further performs medical narrative paragraph segmentation, which is based on the structural characteristics and semantic coherence of the medical document. The system identifies natural paragraph boundaries within the medical narrative by analyzing thematic consistency, temporal features, and variations in medical concept density. Each paragraph typically revolves around a specific clinical theme or event, such as the chief complaint, history of present illness, or physical examination findings. This segmentation provides more structured and ordered text units for subsequent deep semantic analysis, improving the accuracy and efficiency of subsequent processing.

[0023] Deep Scanner is used to identify medical entity boundaries and pre-label categories in segmented text units. Deep Scanner is a specialized text processing component based on a deep neural network architecture. Its core is a bidirectional long short-term memory network combined with a conditional random field model, trained on a large amount of medical text. This model identifies the start and end locations of medical entities by analyzing each word in the text and its context. The system uses a system of entity categories encompassing key medical entity types, including diseases and diagnoses, symptoms and signs, examinations and tests, medications and treatments, body parts, and medical procedures. Deep Scanner not only considers lexical features during processing but also incorporates character-level embedding representations to better handle undocumented terms and abbreviations commonly found in medical text. The model utilizes an attention mechanism to enhance focus on key information, enabling it to identify composite entities and nested entity structures spanning multiple words. The pre-labeling process generates a type probability distribution and boundary confidence score for each entity. These intermediate results provide important reference information for subsequent standardization.

[0024] A multi-level medical terminology normalization pipeline was launched, consisting of three main processing layers: term identification, concept mapping, and semantic disambiguation. At the term identification layer, the system uses a term identifier based on a maximum matching algorithm and a modified Trie data structure to identify possible medical term segments from text. The system is equipped with a base medical terminology database containing over one million medical term entries. This database integrates multiple standard medical terminology systems, including the International Classification of Diseases, the Systematic Nomenclature of Medicine, and the Standards for Clinical Terminology. At the concept mapping layer, identified term segments are preliminarily matched against standard concepts in the terminology database to generate a set of candidate concepts. Each candidate concept is assigned a match score and a semantic similarity assessment. At the semantic disambiguation layer, the system leverages the contextual information of medical entities to select the most appropriate standard term from multiple candidate concepts. The disambiguation process considers linguistic features such as clinical indicators, negation words, and degree modifiers in the context, as well as co-occurrence patterns of medical entities and clinical practice patterns. Through these three layers of processing, the system maps medical terms with diverse expressions in text to standardized medical concepts.

[0025] The system uses a knowledge graph to reconstruct the semantic gaps created during the standardization process. The medical terminology standardization process can disrupt certain semantic connections and contextual coherence in the original text, resulting in semantic gaps. To address this issue, the system uses context reconstruction technology based on the medical knowledge graph. The medical knowledge graph contains a large number of semantic relationships and clinical association rules between medical concepts, such as the relationship between diseases and symptoms, the relationship between drugs and indications, and the connection between examination items and diagnostic objectives. The system analyzes the position and relationships of standardized medical entities in the knowledge graph and reconstructs the semantic connections between entities. This process involves identifying implicit clinical logical relationships in the text, completing semantic information obscured by the standardization process, and verifying the rationality of the standardization results within the clinical context. Context reconstruction not only repairs semantic gaps but also enriches the semantic representation of the text, providing more complete and accurate input for subsequent in-depth semantic analysis. The reconstructed text maintains the coherence and integrity of the clinical narrative while also using standardized terminology.

[0026] Generating standardized text output with uniform medical codes is the final step of the entire preprocessing process. After the aforementioned processing steps, the system produces structured standardized text output in which each medical entity is assigned a uniform medical code. These codes are derived from internationally standardized medical coding systems, such as ICD-10 for disease diagnosis, LOINC for laboratory test items, and RxNorm for drug ingredients. The standardized text is represented using a hierarchical structure, preserving the narrative order and paragraph structure of the original text while adding a standardized semantic annotation layer. The output data is formatted in JSON-LD, adhering to the Semantic Web standards for healthcare, ensuring machine readability and interoperability. Each medical entity includes rich annotations, including its original text representation, standardized concept identifiers, coding system information, confidence scores, and location information. This standardized output provides high-quality structured input data for downstream medical entity uncertainty analysis and contextual influence analysis, and is a critical foundation for the effective operation of the entire system.

[0027] The uncertainty feature parsing module performs deep semantic parsing on preprocessed, standardized medical record text data using a context-aware deep convolutional neural network architecture. The network comprises eight convolutional layers, each using three different kernel sizes (3×3, 5×5, and 7×7) to extract multi-granular semantic features in parallel. Convolutions are performed on a sequence of 256-dimensional word vectors with a stride of 1 and uniform padding to maintain sequence length. Each convolutional layer is followed by a normalization layer and an exponential linear unit activation function. The output feature maps are then concatenated through a maximum pooling layer for dimensionality reduction to form a multi-scale feature representation. The resulting basic semantic feature matrix has a dimension of 512×768, where 512 represents the sequence length and 768 represents the feature dimension.

[0028] A 327-entry medical uncertainty dictionary was constructed to address suspicious expressions appearing in texts, covering common clinical uncertainty expressions such as "suspected," "possible," "not ruled out," "to be investigated," and "considered." A bidirectional long short-term memory network combined with a conditional random field sequence labeling model was used to identify uncertainty terms and their modification scope. When an uncertainty term was detected, the grammatical dependency distance between the term and the surrounding medical entities was calculated, and an attention weight ranging from 0.1 to 0.9 was assigned based on the dependency type (e.g., negation, degree, or speculation). A gated attention mechanism was used to dynamically adjust the representation strength of the feature vector, forming an uncertainty-enhanced feature vector with a dimension of 256.

[0029] The construction of the uncertainty association graph for medical entities is based on probabilistic relational data from a medical knowledge base. Approximately 4.5 million disease-symptom association pairs were extracted from the Unified Medical Language System, with each record containing co-occurrence frequency, relative risk ratio, and diagnostic specificity indicators. A graph convolutional network was used to map medical entities into the knowledge space, and their similarity with standard concept nodes was calculated. An embedding representation of the entity node was generated using a random walk algorithm with an embedding dimension of 128. The uncertainty-enhanced feature vector was then subjected to dot-product similarity calculations with the knowledge base embedding. A softmax function was then used to generate a probability distribution representing the strength of association between the entity and different disease concepts.

[0030] The feature spectrum synthesis process uses a three-way feature fusion architecture. The multi-scale semantic feature matrix is ​​reduced to 256 dimensions through global average pooling, and the uncertainty-enhanced feature vector is adjusted to 256 dimensions through a fully connected layer. The probability relationship mapping result is directly used as a 256-dimensional input. The three feature sources are first added element-by-element, and then feature cross-calculations are performed through a fully connected layer with 128 hidden units. Residual connections are used to retain the original feature information, and finally, layer normalization is performed to generate a 512-dimensional medical entity uncertainty feature spectrum. This feature spectrum contains 32 feature groups, each with 16 eigenvalues, describing different types of uncertainty patterns.

[0031] Three quantitative metrics are extracted from the uncertainty feature spectrum of medical entities. The feature distribution dispersion metric is obtained by calculating the variance within each feature group and the covariance matrix between groups, with the determinant value representing the overall degree of dispersion. The contextual consistency metric uses cosine similarity to calculate the match between entity features and global context features, while considering similarity at both the local window (the five words before and after) and global document levels. The knowledge base match confidence metric is based on the pre-trained BERT matching model, which inputs an entity description and a knowledge base concept definition and outputs a match confidence score between 0 and 1.

[0032] The degree of distribution dispersion is quantified using a multidimensional analysis method. Principal component analysis is used to extract the first two principal components in the 512-dimensional feature space as the certainty-uncertainty dimension. The projected distribution of the eigenvalues ​​along these two dimensions is calculated, and the probability distribution function is fitted using kernel density estimation. The degree of dispersion is ultimately quantified as the entropy of the distribution function. Skewness (a measure of distribution asymmetry) and kurtosis (a measure of distribution sharpness) are also calculated as auxiliary indicators. Higher entropy values ​​indicate lower certainty and a more dispersed distribution.

[0033] Contextual consistency was assessed using a multi-level comparative analysis. The Sentence-BERT model was used to generate semantic vector representations of medical records with a dimension of 1024. Euclidean distance and cosine similarity were calculated between the medical entity vector and the document vector to measure absolute distance and directional consistency, respectively. The semantic compatibility between the entity and its direct modifiers (adjectives and adverbs) was also analyzed, with a pre-trained compatibility model outputting a compatibility score between 0 and 1. For time-series text, an LSTM model was used to analyze the temporal evolution of entity representations to detect any contradictory representations.

[0034] The semantic certainty offset value is generated using dynamic weighted fusion. The three basic indicators are first normalized using the min-max method to eliminate dimensional differences. The weight coefficients are learned through a gradient boosting decision tree model, with the basic weights being: distribution dispersion 0.4, context fit 0.3, and knowledge base fit 0.3. Dynamic adjustments are made based on entity type: diagnostic entities increase the knowledge base weight by 0.1, symptom entities increase the context weight by 0.1, and treatment entities increase the distribution dispersion weight by 0.1. The weighted summation result is converted to a standard score between 0 and 1 using a sigmoid function, with four decimal places retained as the final output value. Scores exceeding 0.7 are considered high-certainty entities, while scores below 0.3 are considered low-certainty entities. Intermediate values ​​need to be determined based on specific clinical scenarios.

[0035] The contextual influence analysis module uses a three-layer spatiotemporal convolutional neural network architecture to perform a spatiotemporal joint analysis of the modifying context in standardized medical record text data. The first layer uses 128 3×3 two-dimensional convolution kernels to scan the spatial dimension of the text sequence and capture the modifying relationships between adjacent words. The second layer uses 64 5×5 convolution kernels to expand the receptive field to identify long-distance dependencies. The third layer uses a time domain convolution layer, using 32 7×1 convolution kernels to slide along the text sequence to analyze the temporal evolution pattern of the modifying context. The ReLU activation function is used after each convolution layer, and max pooling is used for downsampling. The network output contains 64 feature maps, each of which corresponds to a specific modifying context distribution pattern.

[0036] To analyze the distribution characteristics of negative words, adverbs of degree, and uncertainty expressions, we first established a modifier dictionary containing 423 entries, including 87 negative words, 156 adverbs of degree, and 180 uncertainty expressions. A sliding window-based statistical method was used, with 10 words forward and 10 words backward, centered on the medical entity, as analysis windows. Within each window, the frequency, positional distribution, and density variation of each type of modifier were calculated. Kernel density estimation was used to generate a probability distribution map of the modifiers in text space, and its spectral characteristics were analyzed using discrete Fourier transforms to capture periodic patterns in the distribution.

[0037] When matching captured distribution patterns with clinical expression templates, a template matching algorithm based on an attention mechanism is used. The preset template library contains 12 categories of clinical expression patterns, each containing 20 to 50 typical expression examples. A bidirectional attention mechanism is used to calculate the relevance weights between the current text and each template example, and a weighted summation is used to obtain a matching score. The matching process considers three dimensions: lexical overlap, grammatical structure similarity, and semantic relevance. Ultimately, the template type with the highest matching score is selected as the classification result of the doctor's expression habits, and a matching confidence score is output.

[0038] Clinical expression behavior profiles are generated using feature fusion technology. Spatiotemporal distribution features (64 dimensions), expression habit type (12 dimensions), and matching confidence (1 dimension) are concatenated to generate an initial 77-dimensional feature vector. Principal component analysis reduces the dimensionality to 32 dimensions, retaining 95% of the variance. Three specialized metrics are added: modifier context strength (calculated based on modifier density), distribution breadth (calculated based on the standard deviation of the modifier distribution), and evolution pattern (based on temporal feature extraction), ultimately resulting in a 35-dimensional multidimensional clinical expression behavior profile. Each dimension undergoes min-max normalization to ensure values ​​range between 0 and 1.

[0039] Three core indicators were extracted based on the clinical description behavior profile. The modifier context strength index was calculated by calculating the weighted sum of the 1st to 10th dimension features in the behavior profile, with the weight coefficient determined through ridge regression training. The distribution breadth index was calculated by combining the standard deviation and entropy of the 11th to 20th dimension features, reflecting the degree of dispersion of the modifier context. The temporal stability index was calculated by analyzing the coefficient of variation and autocorrelation coefficient of the 21st to 30th dimension features, measuring the stability of the modification pattern over time.

[0040] The impact strength of negative words and uncertainty expressions is calculated using a quantification method based on attention weights. A pre-trained language model is used to obtain a context-aware embedding representation for each modifier. The cosine similarity between the embedding and the medical entity embedding is then calculated. This similarity is converted using a sigmoid function and used as the baseline impact strength. The grammatical distance between the modifier and the entity is also considered, with the impact strength attenuated by 0.1 for each additional lexeme in the distance. The final impact strength is the product of the baseline strength and the distance attenuation factor, ranging from 0 to 1.

[0041] The influence and distance of modifying context are assessed using a graph propagation-based algorithm. The text is constructed as a lexical graph, with nodes representing words and edges representing grammatical dependencies. Starting with the modifying word, a random walk algorithm is used to calculate the probability of influence propagation. The influence range is defined as the proportion of words with a propagation probability exceeding 0.2 to the total number of words. The distance is calculated by taking the harmonic mean of the shortest grammatical path lengths from the modifying word to each affected word.

[0042] The context-modified influence value is generated using a three-layer neural network with weighted aggregation. The input layer receives three indicators: influence intensity (1D), influence range (1D), and influence distance (1D). The hidden layer contains 8 neurons and uses the tanh activation function. The output layer uses the softmax function to generate the final influence value, which ranges from 0 to 1. The network weights are obtained by training with labeled samples using the Adam optimizer, with a learning rate set to 0.001 and a batch size of 32. The final output influence value is retained to four decimal places to ensure that the calculation accuracy meets clinical application requirements.

[0043] In the multi-objective decision optimization module, the module is responsible for integrating and optimizing the semantic features generated by the preceding modules. This module first performs normalization preprocessing on the semantic certainty offset and contextual modification influence values ​​to eliminate the effects of dimensionality differences. This normalization utilizes the Z-score normalization method. Specifically, the mean μ and standard deviation σ of each feature in the training dataset are calculated. The original feature value x is then converted to a standard distribution of (x-μ) / σ. The mean of the semantic certainty offset value typically ranges from 0.45 to 0.55, with a standard deviation between 0.15 and 0.25. The mean of the contextual modification influence value typically ranges from 60.35 to 0.45, with a standard deviation of approximately 0.18 to 0.28. This normalization ensures that the two feature values ​​are within the same numerical range, laying the foundation for subsequent weighted fusion. During the normalization process, the system monitors changes in the feature value distribution in real time. If a distribution shift exceeds a threshold of 0.1, it automatically triggers parameter recalibration to ensure the stability of the normalization effect.

[0044] Based on the characteristics of the current medical entity type, the system uses a dynamic weighted fusion mechanism to automatically adjust the fusion weights of the semantic certainty offset and contextual influence. This dynamic weighted fusion mechanism is based on a predefined medical entity type weight mapping table. This table includes eight major entity types, including disease diagnosis, clinical symptoms, examinations and tests, and medications, each of which is further subdivided into 36 subcategories. For each entity type, the system is configured with a different combination of weights. For example, for disease diagnosis entities, the weight of the semantic certainty offset is set to 0.6, and the weight of the contextual influence is set to 0.4; while for examination and test entities, the weights are 0.3 and 0.7. The weights are determined by analyzing the statistical characteristics of a large amount of annotated data and manually calibrated based on clinical expert experience. In actual operation, the system automatically determines the type of the entity being processed through the entity type recognition module and then retrieves the corresponding weight from the weight mapping table. Furthermore, the system incorporates context-adaptive adjustment. When unique clinical scenarios are detected (such as emergency department medical records), the weights are automatically fine-tuned, typically within a range of plus or minus 0.15.

[0045] Establishing a nonlinear mapping between certainty offsets and contextual influence values ​​is achieved through a specially designed cross-modal attention fusion layer. This fusion layer consists of two main components: a feature interaction layer and an attention generation layer. The feature interaction layer first concatenates the two normalized feature values ​​and then passes them through a fully connected network layer with 128 hidden units, which uses the ReLU activation function to capture the complex nonlinear relationships between features. The attention generation layer calculates the contribution weight of each feature to the final representation and uses a softmax function to generate an attention distribution. This process generates a context-aware fused feature representation with 64 dimensions. This representation not only contains information about the original features but also encodes the interaction patterns between features, better reflecting the comprehensive certainty state of the medical entity. The fusion layer's parameters are obtained through end-to-end training on labeled data, with the training objective of minimizing the difference between the certainty levels of the fused features and those of the manually annotated features.

[0046] The weighted deterministic offset values ​​and contextual influence values ​​are concatenated with the fused feature representation to generate the final medical semantic feature vector. This concatenation utilizes a channel-wise concatenation approach, combining the weighted original features (2-dimensional) with the fused feature representation (64-dimensional) into a 66-dimensional composite feature vector. This feature vector retains the interpretability of the original features while incorporating the deep feature representations derived from deep learning. To further enhance feature representation, the system also performs dimensionality reduction on the concatenated feature vector, compressing it to 32 dimensions via a fully connected layer with 32 output units. The resulting medical semantic feature vector incorporates comprehensive information on semantic certainty and contextual influence, providing a rich feature foundation for subsequent multi-objective decision optimization. Each dimension of the feature vector has a clear clinical semantic meaning. For example, the first 16 dimensions primarily reflect the deterministic characteristics of the entity itself, while the last 16 dimensions encode the influence patterns of contextual modifications.

[0047] A multi-objective optimization function is used to simultaneously optimize clinical safety, data integrity, and reporting efficiency. Clinical safety is quantified by calculating the degree of compliance of decision outcomes with clinical guidelines, using a rule-based matching approach with built-in clinical safety rules. Data integrity is measured by evaluating the completeness and information content of populated fields, using a weighted combination of information entropy and field fill rate as an indicator. Reporting efficiency is assessed by calculating the time complexity of the decision-making process and system resource consumption. The weights of the three objectives are dynamically adjusted based on the application scenario: in outpatient settings, the weight of the reporting efficiency objective is set to 0.5, the clinical safety objective to 0.3, and the data integrity objective to 0.2. In inpatient ward settings, the weight of the clinical safety objective is increased to 0.6, the data integrity objective to 0.3, and the reporting efficiency objective to 0.1. The multi-objective optimization function uses a weighted summation method to linearly combine the three objective functions into a single comprehensive optimization objective.

[0048] A Pareto optimal solution search algorithm is employed to find the optimal balance between multiple optimization objectives. The system uses an improved non-dominated sorting genetic algorithm (NSGA-II) as the search framework. The algorithm first initializes a population of 100 candidate solutions, each representing a possible set of decision parameter configurations. New solutions are then generated through genetic operations such as crossover and mutation, with a crossover probability of 0.9 and a mutation probability of 0.1. During each evolutionary generation, the algorithm performs non-dominated sorting and crowding calculations on the solutions to select the optimal solution on the Pareto front. The search continues for 50 generations, ultimately outputting a set of candidate decision solutions containing 20 to 30 non-dominated solutions. Each candidate solution exhibits different trade-offs across the three optimization objectives, with some prioritizing clinical safety and others prioritizing reporting efficiency.

[0049] Candidate solutions are verified for medical plausibility using a clinical knowledge constraint validator. Built on a clinical knowledge graph and medical logic rules, the validator incorporates three levels of validation: syntactic validation checks terminology usage and formatting; semantic validation checks the rationality of relationships between medical concepts; and pragmatic validation checks applicability to clinical scenarios. During the validation process, the system generates a compliance score for each solution, and solutions with scores below a threshold of 0.8 are eliminated.

[0050] Based on the needs of real-time application scenarios, the optimal medical data parsing solution is selected from verified candidate solutions. The selection process utilizes a multi-attribute decision-making approach, specifically the TOPSIS (Top of Score) algorithm. The algorithm first constructs a decision matrix, with rows representing candidate solutions and columns representing evaluation metrics (including the achievement of three optimization objectives and compliance scores). The algorithm then determines the weights for each metric, dynamically adjusting them based on the real-time application scenario: during the morning outpatient peak, the weighting of reporting efficiency is increased; when processing critically ill patient records, the weighting of clinical safety is increased. The algorithm then calculates the distance between each solution and both the positive and negative ideal solutions, ultimately selecting the solution with the highest relative proximity as the optimal solution. The entire decision-making process takes less than 200 milliseconds, ensuring the system's real-time responsiveness. The selected optimal solution is output as structured medical data parsing instructions, which are then delivered to the subsequent intelligent pre-filling execution module.

[0051] The intelligent pre-filling execution module, the final output link of the system, is responsible for converting the results of the previous processing into usable clinical data. This module performs ontology alignment based on the optimized medical semantic feature vectors and the medical knowledge graph. This ontology alignment utilizes a semantic similarity calculation model based on an attention mechanism. This model first maps the medical semantic feature vectors into the same vector space as the medical knowledge graph. This mapping process is implemented using a fully connected neural network with 64 hidden units, which uses the tanh activation function to convert the 32-dimensional medical semantic feature vectors into a 128-dimensional graph space vector. Similarity is calculated using the cosine similarity algorithm, specifically calculating the cosine of the angle between the two vectors. The similarity threshold is set to 0.85. When the similarity reaches this threshold, the system considers the medical entity to be successfully matched with the knowledge graph concept node. For each medical entity, the system retrieves the top three candidate concept nodes with the highest similarity and records their matching scores. These candidate concepts and their matching scores will provide important reference for subsequent clinical validation.

[0052] Verifying the clinical plausibility of alignment results is a critical step in ensuring data quality. The verification process is divided into two main parts: symptom-disease association verification and treatment plan plausibility check. Symptom-disease association verification is based on probabilistic relationships in the medical knowledge graph. The system calculates the strength of the clinical association between symptoms and diseases, with the association strength threshold set at 0.6. For matching pairs with association strengths below this threshold, the system will issue a warning. Treatment plan plausibility check involves verification of multiple aspects, including drug-disease indications, drug-drug interactions, and consistency between treatment and test results. During the verification process, the system generates a plausibility score for each alignment result, ranging from 0 to 1. Alignment results with a score below 0.7 are marked for manual review. Results with a score between 0.7 and 0.9 are accepted with a note, and results with a score above 0.9 are adopted directly.

[0053] Ensuring that the generated structured medical data remains consistent with the overall semantic environment of the medical record is achieved through contextual consistency detection. Contextual consistency detection is based on a recurrent neural network architecture and can capture long-term dependencies in medical record text. The system maintains a medical record-level state vector that encodes the semantic environment and clinical context information of the entire medical record. For each medical entity, the system calculates its semantic consistency score with the medical record state vector, using a similarity measurement method weighted by an attention mechanism. The consistency score threshold is set to 0.75, and entities below this threshold are marked as potentially having contextual inconsistencies. In addition, the system uses a temporal consistency checker to verify the temporal logic rationality of medical events, such as ensuring that the diagnosis time is no later than the treatment time and the examination time is after the onset of symptoms. These checks ensure that the structured data generated is not only locally correct, but also reasonable and coherent in the global context.

[0054] Based on the alignment and verification results, the system labels each medical entity with a corresponding certainty level tag. Certainty levels are divided into three levels: high certainty (confidence greater than 0.9), medium certainty (confidence between 0.7 and 0.9), and low certainty (confidence less than 0.7). The labeling process is based on a weighted comprehensive score of multiple factors, including semantic similarity score (weight 0.4), clinical plausibility score (weight 0.3), contextual consistency score (weight 0.2), and data quality indicators (weight 0.1). The final certainty tag for each medical entity includes not only the level classification, but also a detailed confidence score and a description of the decision basis. These tags provide a decision basis for subsequent differentiated pre-filling strategies, enabling the system to take appropriate treatment methods based on different levels of certainty.

[0055] When performing auto-population operations on high-certainty data, the system uses direct mapping to populate EMR fields. The system maintains a field mapping rule library containing multiple mapping rules that define the correspondence between medical entity types and EMR form fields. During the populating process, the system populates both standard terminology codes and original text descriptions to ensure both machine and human readability of the data. For each auto-population operation, the system records the populating time, data source, and certainty score for subsequent audit and traceability.

[0056] For low-certainty data, the system adds a visual review indicator and locks the corresponding field. The review indicator uses a color-coded system: yellow indicates that attention is required (medium certainty), and red indicates that immediate action is required (low certainty). This locking mechanism prevents accidental modification of these fields, and they can only be unlocked for editing after authorized manual review. The system also provides detailed review guidance, including an analysis of the causes of uncertainty, recommended treatment suggestions, and relevant clinical reference information. These designs ensure data security while providing clinicians with comprehensive decision-making support.

[0057] Based on the medical knowledge graph and similar case analysis, the system generates candidate filling options that meet clinical standards. The candidate generation process adopts a method that combines collaborative filtering and knowledge reasoning. The system first retrieves the standard concepts most relevant to the current medical entity in the medical knowledge graph and generates knowledge-based candidate options. It then searches for similar cases in the historical case database, extracts the treatment methods of the same entities in these cases, and generates instance-based candidate options. Finally, the candidate options are comprehensively scored and ranked through a sorting algorithm. The scoring takes into account factors such as clinical applicability, frequency of use, and compliance with the latest guidelines. The system provides 3 to 5 candidate filling options for each low-certainty data, each of which comes with a detailed description and confidence score.

[0058] Based on the real-time medical scenario requirements, the system adaptively adjusts the priority and filling strategy of pre-filled fields. The system monitors multiple factors such as workload, user roles, and clinical scenarios in real time. During peak outpatient hours, the system will prioritize the automatic filling of high-certainty data to improve reporting efficiency; in intensive care scenarios, the system will increase the strictness of clinical safety checks and lower the certainty threshold for automatic filling. The adjustment strategy is based on preset rules and machine learning models, and can dynamically optimize system behavior according to actual conditions. The system re-evaluates scenario parameters every 5 minutes to ensure that the filling strategy is always optimally matched to current clinical needs. This adaptive design enables the system to maintain optimal performance in different clinical environments.

[0059] The working principle of the present invention is as follows: multi-source heterogeneous medical text data is received and standardized through the data acquisition and preprocessing module, and entity recognition and multi-level medical terminology normalization are performed using a deep scanner; the uncertainty feature analysis module constructs the uncertainty feature spectrum of medical entities and calculates the semantic certainty offset value, uses the context-aware deep convolutional network to extract multi-scale features, and combines with the medical knowledge base to perform probabilistic relationship mapping; the context influence analysis module uses a spatiotemporal convolutional network to analyze the modifying context, constructs a clinical expression behavior portrait and calculates the context modification influence value; the multi-objective decision optimization module fuses the above-mentioned feature values ​​into a medical semantic feature vector, and uses a multi-objective optimization function and a Pareto optimal solution search algorithm to collaboratively optimize clinical safety, data integrity and reporting efficiency; finally, the intelligent pre-filling execution module performs ontology alignment and logical verification based on the medical knowledge graph, implements differentiated pre-filling strategies according to the certainty mark level, automatically fills in high-certainty data, generates review marks for low-certainty data and provides candidate options, thereby achieving precise and adaptive intelligent parsing and pre-filling of medical records.

[0060] The above is a detailed description of an embodiment of the present invention. However, the content described is only a preferred embodiment of the present invention and should not be considered to limit the scope of the present invention. All equivalent changes and improvements made within the scope of the present invention should still fall within the scope of the patent coverage of the present invention.

Claims

1. An AI-based medical record intelligent analysis and pre-filling system, characterized by: include: A data acquisition and preprocessing module receives medical record text data through a medical text data interface, and performs preprocessing and medical terminology standardization on the medical record text data; An uncertainty feature parsing module, which performs deep semantic analysis on standardized medical record text data, constructs a medical entity uncertainty feature spectrum, and calculates a semantic certainty offset value; A context influence analysis module, which performs a spatiotemporal joint analysis of the modifying context in the standardized medical record text data, constructs a clinical expression behavior profile, and calculates the context modification influence value; A multi-objective decision optimization module, which fuses the semantic deterministic offset value and the contextual modification influence value into a medical semantic feature vector, and inputs the medical semantic feature vector into a pre-trained medical record parsing model for multi-objective decision optimization; An intelligent pre-filling execution module performs ontology alignment and logical verification based on the optimized medical semantic feature vector and the medical knowledge graph, generates structured medical data with deterministic tags, and drives differentiated pre-filling strategies based on the level of deterministic tags, specifically including: automatically filling in high-certainty data, generating pending review tags for low-certainty data, and providing candidate filling content.

2. The AI-based medical record intelligent analysis and pre-filling system according to claim 1 is characterized in that: The pre-processing of medical record text data and the standardization of medical terminology specifically include: Receive heterogeneous medical record data, perform sentence breakpoint detection and medical narrative paragraph segmentation on unstructured text; A deep scanner is used to identify medical entity boundaries and pre-label categories of segmented text units; Launch a multi-level medical term normalization pipeline, match the basic medical terminology database, and perform term semantic disambiguation; The semantic faults generated during the standardization process are reconstructed in context based on the knowledge graph to generate standardized text output with unified medical coding.

3. The AI-based medical record intelligent analysis and pre-filling system according to claim 1 is characterized in that: The process of performing deep semantic analysis on standardized medical record text data and constructing a medical entity uncertainty feature spectrum specifically includes: A context-aware deep convolutional network is used to extract multi-scale semantic features from standardized text and generate a basic semantic feature matrix. The context segments containing suspicious words are weighted to form uncertainty-enhanced feature vectors; Establish a medical entity uncertainty association map and map the uncertainty-enhanced feature vectors with the disease-symptom probability relationship in the medical knowledge base; The multi-dimensional uncertainty feature spectrum of medical entities is generated by integrating the multi-scale semantic feature matrix, uncertainty enhanced feature vector and probability relationship mapping results.

4. The AI-based medical record intelligent analysis and pre-filling system according to claim 1 is characterized in that: The calculation process of the semantic deterministic offset value is: Based on the uncertainty feature spectrum of medical entities, feature distribution dispersion index, context consistency index and knowledge base matching confidence index are extracted; Calculate the degree of dispersion of the eigenvalue distribution in the dimensions of certainty and uncertainty; The characteristic value is the specific value of each dimension of the uncertainty characteristic spectrum of the medical entity; Evaluate the degree of consistency between entity representation and the semantics of the overall medical record context; The distribution dispersion, the consistency of context semantics and the knowledge base matching are weighted and fused to generate a quantized semantic deterministic offset value.

5. The AI-based medical record intelligent analysis and pre-filling system according to claim 1 is characterized in that: The spatiotemporal joint analysis of the modifying context in the standardized medical record text data to construct a clinical expression behavior portrait specifically includes: A spatiotemporal convolutional neural network is used to perform multi-dimensional scanning of medical record texts, capturing the distribution patterns and temporal evolution characteristics of modifying context in text sequences. Statistical analysis of the spatiotemporal distribution of negative words, degree adverbs, and uncertainty expressions around specific medical entities; Match the captured distribution patterns with the preset clinical expression templates to identify the physician's expression habit types; Integrate spatiotemporal distribution characteristics and expression habit types to generate a multidimensional clinical expression behavior portrait that includes the intensity of the modification context, distribution breadth and evolution pattern.

6. The AI-based medical record intelligent analysis and pre-filling system according to claim 1 is characterized in that: The calculation process of the context modification influence value is as follows: Based on the clinical expression behavior profile, we extract the modification context intensity index, distribution breadth index and temporal stability index; Calculate the influence of negative words and uncertainty expressions on the semantics of medical entities, which is recorded as the modifying context intensity; Evaluate the scope and distance of influence of modifying context in the text, recorded as distribution breadth; The intensity, distribution breadth and temporal stability of the modifying context are weighted and aggregated to generate a quantitative context modification influence value.

7. The AI-based medical record intelligent analysis and pre-filling system according to claim 1 is characterized in that: The construction process of the medical semantic feature vector is as follows: Standardize the semantic certainty offset value and contextual modification influence value to eliminate dimensional differences; Automatically adjust the fusion weight coefficient of the deterministic offset value and the contextual modification influence value according to the type characteristics of the current medical entity; Establish a nonlinear mapping relationship between the deterministic offset value and the contextual modification influence value to generate a fused feature representation with context-awareness; The weighted certainty offset value and contextual modification influence value are concatenated with the fusion feature representation to generate a medical semantic feature vector containing semantic certainty and contextual influence.

8. The AI-based medical record intelligent analysis and pre-filling system according to claim 1 is characterized in that: The medical semantic feature vector is input into the pre-trained medical record parsing model for multi-objective decision optimization, specifically including: Synchronously optimize clinical safety goals, data integrity goals, and reporting efficiency goals through a multi-objective optimization function; Use the Pareto optimal solution search algorithm to find the best balance point among multiple optimization objectives and generate a set of candidate decision solutions; Use clinical knowledge constraint validators to verify the medical rationality of candidate solutions and eliminate decision-making solutions that do not meet clinical standards; Based on the needs of real-time application scenarios, select the best medical data analysis solution from verified candidate solutions.

9. The AI-based medical record intelligent analysis and pre-filling system according to claim 1 is characterized in that: The ontology alignment and logic verification based on the optimized medical semantic feature vector and the medical knowledge graph specifically includes: Perform semantic similarity matching between medical semantic feature vectors and concept nodes in the medical knowledge graph; Conduct clinical validation of alignment results, including symptom-disease correlation verification and treatment plan rationality check; Ensure that the generated structured medical data remains consistent with the overall semantic environment of the medical record; Based on the alignment and verification results, each medical entity is labeled with a corresponding certainty level mark.

10. The AI-based medical record intelligent analysis and pre-filling system according to claim 1 is characterized in that: The differentiated pre-filling strategy driven by the certainty mark level specifically includes: Directly fill in electronic medical record fields for highly certain data; Add visual review marks to low-certainty data and lock the corresponding fields; Generate candidate filling options that meet clinical standards based on medical knowledge graphs and similar case analysis; Adaptively adjust the priority and filling strategy of pre-filled fields based on real-time medical scenario requirements.

Citation Information

Patent Citations

  • Medical record structured analysis method based on medical field entities

    CN110032648A

  • Personal electronic medical record-based uncertainty knowledge graph automatic construction method

    CN120108753A

  • Intelligent nursing record generation method and device

    CN120542403A

  • Deep learning large model-based medical record information extraction and analysis method and system

    CN120579548A

  • Methods, apparatus and products for semantic processing of text

    EP2639749A1

Cited By

  • Medical record text processing method and device, equipment and medium

    CN121096687A

  • General medical record generation method and system and storage medium

    CN121122548A