An AI-based medical record intelligent analysis and pre-reporting system
By constructing an AI-based intelligent medical record parsing system, the problem of insufficient parsing of uncertain and negative statements in existing medical record parsing methods has been solved, realizing efficient, accurate and secure intelligent pre-filling of medical record data, and adapting to the needs of various medical environments.
Patent Information
- Application Number
- CN202511320110.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-16
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2045-09-16
AI Technical Summary
Existing intelligent medical record parsing methods based on statistical machine learning paradigms lack the ability to accurately analyze uncertainties and negative statements in clinical texts, leading to misjudgments and data inaccuracies, which affect the quality and security of electronic medical records.
An AI-based intelligent medical record parsing and pre-filling system is adopted, including a data acquisition and preprocessing module, an uncertainty feature parsing module, a contextual influence analysis module, a multi-objective decision optimization module, and an intelligent pre-filling execution module. Through deep learning and knowledge graph technology, it constructs a spectrum of uncertainty features of medical entities and a profile of clinical expression behavior, performs multi-objective decision optimization and ontology alignment, and generates structured medical data with deterministic labels.
It enables the deconstruction and digital representation of the certainty of clinical texts, ensuring the accuracy and logical consistency of structured data, improving the accuracy of medical record data processing and the reliability of clinical applications, adapting to the needs of different medical environments, and reducing the risk of misjudgment.
Smart Images

Figure CN120823938B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical information technology, specifically to an AI-based intelligent medical record parsing and pre-filling system. Background Technology
[0002] With the deepening of medical informatization, electronic medical record systems have been widely used in various medical institutions, generating massive amounts of medical text data. How to automatically extract key information from these unstructured medical texts and achieve intelligent pre-filling has become a key technical challenge for improving the quality of medical data and the efficiency of clinical work.
[0003] The existing technology has the following shortcomings:
[0004] Existing intelligent medical record parsing methods based on statistical machine learning paradigms lack the ability to accurately identify and analyze the deep semantics of uncertain and negative expressions commonly found in clinical texts during the core information extraction process. This problem causes the system to easily confuse the definitive state of diagnosis during entity recognition and relation extraction, resulting in misjudgments at the clinical semantic level. These misjudgments are further amplified in subsequent structured processing and pre-filling stages, generating electronic medical record data that deviates from clinical reality, thus creating serious security risks and hindering the crucial leap from laboratory validation to reliable clinical application. Summary of the Invention
[0005] The purpose of this invention is to provide an AI-based intelligent medical record parsing and pre-filling system to solve the problems mentioned above.
[0006] The objective of this invention can be achieved through the following technical solutions:
[0007] An AI-based intelligent medical record parsing and pre-filling system includes:
[0008] The data acquisition and preprocessing module receives medical record text data through a medical text data interface and performs preprocessing and medical terminology standardization on the medical record text data.
[0009] The uncertainty feature parsing module performs deep semantic parsing on standardized medical record text data, constructs a medical entity uncertainty feature spectrum, and calculates semantic determinism offset values.
[0010] The contextual influence analysis module performs spatiotemporal joint analysis on the modifier context in standardized medical record text data, constructs a clinical expression behavior profile, and calculates the contextual modifier influence value.
[0011] A multi-objective decision optimization module is provided, which integrates semantic deterministic offset value and contextual modification influence value into a medical semantic feature vector, and inputs the medical semantic feature vector into a pre-trained medical record parsing model for multi-objective decision optimization.
[0012] The intelligent pre-filling execution module performs ontology alignment and logical verification based on the optimized medical semantic feature vector and medical knowledge graph, generates structured medical data with deterministic labels, and drives differentiated pre-filling strategies according to the deterministic label level. Specifically, it performs automatic filling for high deterministic data, generates pending review labels for low deterministic data, and provides candidate filling content.
[0013] As a further aspect of the present invention: the preprocessing and medical terminology standardization of the medical record text data specifically includes:
[0014] Receive heterogeneous medical record data and perform sentence breakpoint detection and medical narrative paragraph segmentation on unstructured text;
[0015] A depth scanner was used to identify medical entity boundaries and pre-label categories of segmented text units.
[0016] Initiate a multi-level medical terminology normalization pipeline, match it with the basic medical terminology database, and perform semantic disambiguation of the terms;
[0017] The semantic gaps generated during the standardization process are reconstructed using a knowledge graph-based context reconstruction method to generate standardized text output with unified medical coding.
[0018] As a further aspect of the present invention: the deep semantic analysis of standardized medical record text data to construct a medical entity uncertainty feature spectrum specifically includes:
[0019] A context-aware deep convolutional network is used to extract multi-scale semantic features from standardized text, generating a basic semantic feature matrix.
[0020] Weight enhancement is applied to contextual segments containing potentially suggestive words to form uncertainty-enhanced feature vectors;
[0021] Establish an uncertainty association graph for medical entities and map the uncertainty enhancement feature vectors to the disease-symptom probability relationships in the medical knowledge base;
[0022] By integrating multi-scale semantic feature matrices, uncertainty-enhanced feature vectors, and probabilistic relationship mapping results, a multi-dimensional uncertainty feature spectrum for medical entities is generated.
[0023] As a further aspect of the present invention: the calculation process of the semantic deterministic offset value is as follows:
[0024] Based on the uncertainty feature spectrum of medical entities, feature distribution dispersion index, context consistency index and knowledge base matching confidence index are extracted;
[0025] Calculate the degree of dispersion of eigenvalues across deterministic and uncertain dimensions;
[0026] The feature values are specific values for each dimension of the uncertainty feature spectrum of medical entities;
[0027] Assess the degree of consistency between the entity representation and the overall semantic context of the medical record;
[0028] We perform weighted fusion of the degree of distribution dispersion, the degree of semantic consistency with the context, and the degree of knowledge base matching to generate a quantified semantic deterministic offset value.
[0029] As a further aspect of the present invention: the spatiotemporal joint analysis of the descriptive context in standardized medical record text data to construct a clinical expression behavior profile specifically includes:
[0030] Spatiotemporal convolutional neural networks are used to perform multi-dimensional scanning of medical record texts to capture the distribution patterns and temporal evolution features of modifying contexts in the text sequence;
[0031] The spatiotemporal distribution characteristics of negative words, degree adverbs, and expressions of uncertainty around a specific medical entity;
[0032] The captured distribution patterns are matched with preset clinical expression templates to identify the doctor's expression habits.
[0033] By integrating spatiotemporal distribution characteristics and expression habit types, a multidimensional clinical expression behavior profile is generated, which includes the intensity of modifying context, the breadth of distribution, and the evolution pattern.
[0034] As a further aspect of the present invention: the calculation process for the influence value of contextual modification is as follows:
[0035] Based on clinical expression behavior profiles, we extracted indicators of modifier context intensity, distribution breadth, and temporal stability.
[0036] The strength of the influence of negative words and uncertain expressions on the semantics of medical entities is calculated and denoted as the modifier context strength.
[0037] The scope and distance of influence of modifying context in the text are assessed and denoted as the distribution breadth.
[0038] We perform weighted aggregation of the intensity, distribution breadth, and temporal stability of modifying contexts to generate a quantified value of the influence of contextual modifications.
[0039] As a further aspect of the present invention: the construction process of the medical semantic feature vector is as follows:
[0040] Standardize the semantic deterministic offset value and the contextual modification influence value to eliminate dimensional differences;
[0041] The fusion weighting coefficient of deterministic offset value and contextual modification influence value is automatically adjusted based on the type characteristics of the current medical entity.
[0042] Establish a nonlinear mapping relationship between deterministic offset values and contextual modification influence values to generate a fusion feature representation with context awareness capabilities;
[0043] The weighted deterministic offset value and contextual modification influence value are concatenated with the fused feature representation to generate a medical semantic feature vector that includes semantic determinism and contextual influence.
[0044] As a further aspect of the present invention: the step of inputting medical semantic feature vectors into a pre-trained medical record parsing model for multi-objective decision optimization specifically includes:
[0045] The clinical safety objective, data integrity objective, and reporting efficiency objective are simultaneously optimized through a multi-objective optimization function.
[0046] The Pareto optimal solution search algorithm is used to find the best balance point among multiple optimization objectives and generate a set of candidate decision schemes.
[0047] A clinical knowledge-constrained validator is used to verify the medical rationality of candidate solutions and eliminate decision-making solutions that do not conform to clinical norms.
[0048] Based on the needs of real-time application scenarios, the optimal medical data parsing solution is selected from the validated candidate solutions.
[0049] As a further aspect of the present invention: the ontology alignment and logical verification based on the optimized medical semantic feature vector and medical knowledge graph specifically includes:
[0050] Perform semantic similarity matching between medical semantic feature vectors and concept nodes in the medical knowledge graph;
[0051] The alignment results were clinically validated, including symptom-disease correlation verification and treatment plan rationality check.
[0052] Ensure that the generated structured medical data remains consistent with the overall semantic environment of the medical records;
[0053] Based on the alignment and verification results, each medical entity is labeled with a corresponding deterministic level tag.
[0054] As a further aspect of the present invention: the differential pre-filling strategy driven by deterministic labeling level specifically includes:
[0055] For highly deterministic data, the electronic medical record fields can be filled directly.
[0056] Add visual audit markers to low-deterministic data and lock the corresponding fields;
[0057] Based on medical knowledge graphs and similar case analysis, candidate fill options that conform to clinical standards are generated;
[0058] Based on the needs of real-time medical scenarios, the priority and filling strategy of pre-filled fields are adaptively adjusted.
[0059] The beneficial effects of this invention are:
[0060] (1) This invention deconstructs and digitally represents the degree of certainty in clinical texts by constructing a spectrum of uncertain features of medical entities and quantifying the influence of contextual modification; then, it relies on medical knowledge graphs to perform ontology alignment and logical verification to ensure that structured data conforms to terminology standards and maintains consistency with clinical logic; finally, it drives a differentiated pre-filling strategy based on the level of certainty labeling, which improves efficiency by automatically processing high-certainty data, while initiating review of low-certainty data and providing knowledge-driven candidate options, thereby avoiding the risk of misjudgment at the source and improving the accuracy of medical record data processing and the reliability of clinical applications.
[0061] (2) This invention establishes a multi-objective optimization function, incorporating the three dimensions of clinical safety, data integrity, and reporting efficiency into a unified mathematical optimization framework. The Pareto optimal solution search algorithm is employed to find non-dominated solution sets in multiple mutually constraining objective spaces, generating candidate decision schemes with different trade-off characteristics. For high-risk clinical scenarios, the system automatically increases the weight coefficients of clinical safety and data integrity; for scenarios with high efficiency requirements, the weight configuration of the reporting efficiency objective is appropriately optimized. This adaptive optimization ensures that the system always provides the solution most suitable for the current clinical scenario at the output end, guaranteeing both data accuracy and security in critical medical scenarios and operational efficiency optimization in routine scenarios, ultimately forming a precise decision support system that can intelligently adapt to the needs of different medical environments. Attached Figure Description
[0062] The invention will now be further described with reference to the accompanying drawings.
[0063] Figure 1 This is a flowchart of the system of the present invention. Detailed Implementation
[0064] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0065] Please see Figure 1 As shown, this invention is an AI-based intelligent medical record parsing and pre-filling system, comprising:
[0066] The data acquisition and preprocessing module receives medical record text data through a medical text data interface and performs preprocessing and medical terminology standardization on the medical record text data.
[0067] The uncertainty feature parsing module performs deep semantic parsing on standardized medical record text data, constructs a medical entity uncertainty feature spectrum, and calculates semantic determinism offset values.
[0068] The contextual influence analysis module performs spatiotemporal joint analysis on the modifier context in standardized medical record text data, constructs a clinical expression behavior profile, and calculates the contextual modifier influence value.
[0069] A multi-objective decision optimization module is provided, which integrates semantic deterministic offset value and contextual modification influence value into a medical semantic feature vector, and inputs the medical semantic feature vector into a pre-trained medical record parsing model for multi-objective decision optimization.
[0070] The intelligent pre-filling execution module performs ontology alignment and logical verification based on the optimized medical semantic feature vector and medical knowledge graph, generates structured medical data with deterministic labels, and drives differentiated pre-filling strategies according to the deterministic label level. Specifically, it performs automatic filling for high deterministic data, generates pending review labels for low deterministic data, and provides candidate filling content.
[0071] In the data acquisition and preprocessing module, the data entry point and quality assurance foundation of the entire system are crucial. This module connects to medical data sources such as Hospital Information Systems (HIS), Laboratory Information Systems (LIS), and Radiological Information Systems (RIS) via a medical text data interface, receiving medical record text data from these systems. This data includes, but is not limited to, different types of medical documents such as outpatient medical records, inpatient medical records, laboratory test reports, and discharge summaries. Due to the diverse sources of this data, their formats and standards are not uniform, potentially including structured, semi-structured, and unstructured data. The module first parses and identifies this heterogeneous data, extracting the text content that needs processing. For non-text format data, such as scanned or image-formatted medical records, the module uses optical character recognition technology to convert them into processable medical record text data. During data reception, the module records metadata such as the data source, type, and reception time, and assigns a unique identifier to each data object for subsequent tracking and processing. This module also features data buffering and capacity adjustment to handle potential data peaks from medical data sources, ensuring stable system operation.
[0072] The received heterogeneous medical record data first undergoes sentence breakpoint detection and medical narrative paragraph segmentation. Sentence breakpoint detection is based on sentence boundary recognition technology from natural language processing, optimized for the specific characteristics of medical texts. The system employs a deep learning-based sequence labeling model, combining common punctuation usage patterns, abbreviation patterns, and grammatical structure features found in medical texts to accurately identify sentence boundaries. The model considers not only obvious markers like periods and question marks but also the discourse function of special punctuation marks such as ellipses and semicolons commonly found in medical texts. Building upon sentence segmentation, the system further performs medical narrative paragraph segmentation, based on the structured features and semantic coherence of medical documents. The system identifies natural paragraph boundaries in the medical narrative by analyzing the text's thematic consistency, temporal characteristics, and changes in the density of medical concepts. Each paragraph typically revolves around a specific clinical theme or event, such as the chief complaint, present medical history, or physical examination findings. This segmentation provides more structured and ordered text units for subsequent deep semantic analysis, helping to improve the accuracy and efficiency of subsequent processing.
[0073] A depth scanner is used to identify medical entity boundaries and pre-label categories in segmented text units. The depth scanner is a specialized text processing component based on a deep neural network architecture. Its core is a bidirectional long short-term memory network combined with a conditional random field model trained on a large amount of medical text. This model identifies the start and end positions of medical entities by analyzing each word in the text and its context. The entity category system used by the system includes major medical entity types such as diseases and diagnoses, symptoms and signs, examinations and tests, drugs and treatments, body parts, and medical procedures. During processing, the depth scanner considers not only word-level features but also incorporates character-level embedding representations to better handle common out-of-vocabulary words and abbreviations in medical text. The model strengthens its focus on key information through an attention mechanism, enabling it to identify complex entities and nested entity structures spanning multiple words. The pre-labeling process generates a type probability distribution and boundary confidence score for each entity; these intermediate results provide important reference information for subsequent standardization processing.
[0074] A multi-level medical terminology normalization pipeline is initiated, comprising three main processing levels: terminology recognition, concept mapping, and semantic disambiguation. At the terminology recognition level, a terminology recognizer based on a maximum matching algorithm and an improved Trie data structure identifies potential medical terminology fragments from the text. The system is equipped with a basic medical terminology database containing over one million medical terminology entries, integrating multiple standard medical terminology systems, including international standards such as the International Classification of Diseases (ICD), Systematic Nomenclature in Medicine (SNOMA), and Standard Clinical Terminology. At the concept mapping level, the identified terminology fragments are initially matched with standard concepts in the terminology database, generating a set of candidate concepts. Each candidate concept carries a matching score and a semantic similarity assessment value. At the semantic disambiguation level, the system utilizes the contextual information of medical entities to select the most appropriate standard term from multiple candidate concepts. The disambiguation process considers linguistic features such as clinical indicators, negations, and degree modifiers in the context, as well as co-occurrence patterns and clinical practice rules of medical entities. Through these three levels of processing, the system maps medical terms with diverse expressions in the text to standardized medical concepts.
[0075] This system employs knowledge graph-based context reconstruction to address semantic gaps arising during the standardization process. The standardization of medical terminology can disrupt certain semantic connections and contextual coherence in the original text, resulting in semantic gaps. To address this issue, the system utilizes context reconstruction technology based on a medical knowledge graph. Medical knowledge graphs contain numerous semantic relationships and clinical association rules between medical concepts, such as the association between diseases and symptoms, drugs and indications, and examination items and diagnostic purposes. The system reconstructs the semantic connections between entities by analyzing the position and relationships of standardized medical entities within the knowledge graph. This process includes identifying implicit clinical logical relationships in the text, completing semantic information blurred by the standardization process, and verifying the rationality of the standardization results in the clinical context. Context reconstruction not only repairs semantic gaps but also enriches the semantic representation of the text, providing a more complete and accurate input for subsequent deep semantic analysis. The reconstructed text maintains the coherence and integrity of the clinical narrative while retaining standardized terminology.
[0076] Generating standardized text output with unified medical codes is the final step in the entire preprocessing process. After the above processing steps, the system produces structured standardized text output, in which each medical entity is assigned a unified medical code. These codes come from internationally standardized medical coding systems, such as ICD-10 for disease diagnosis, LOINC for laboratory testing items, and RxNorm for drug components. The standardized text is represented in a hierarchical structure, preserving the narrative order and paragraph structure of the original text while adding a standardized semantic annotation layer. The output data uses JSON-LD format, following the Semantic Web standard in the healthcare field, ensuring machine readability and interoperability of the data. Each medical entity includes its original text expression, standardized concept identifier, coding system information, confidence score, and location information, among other rich annotations. This standardized output provides high-quality structured input data for downstream medical entity uncertainty analysis and contextual influence analysis, and is a crucial foundation for the effective operation of the entire system.
[0077] In the uncertainty feature parsing module, the module is responsible for deep semantic parsing of preprocessed standardized medical record text data, employing a context-aware deep convolutional neural network architecture. This network contains eight convolutional layers, each using three different kernel sizes (3×3, 5×5, 7×7) to extract multi-granularity semantic features in parallel. Convolutional operations are performed on 256-dimensional word vector sequences with a stride of 1, using the same padding method to maintain sequence length. Each convolutional layer is followed by a normalization layer and an exponential linear unit activation function. The output feature maps are then dimensionality-reduced through a max-pooling layer and concatenated to form a multi-scale feature representation. The final generated basic semantic feature matrix has a dimension of 512×768, where 512 represents the sequence length and 768 represents the feature dimension.
[0078] To address potentially problematic expressions in text, a medical uncertainty dictionary with 327 entries was constructed, covering commonly used clinical uncertainty expressions such as "suspected," "possible," "not ruled out," "to be investigated," and "considered." A sequence labeling model combining a bidirectional long short-term memory network and a conditional random field was employed to identify uncertain words and their modification ranges. When an uncertain word was detected, the grammatical dependency distance between the word and surrounding medical entities was calculated, and attention weights ranging from 0.1 to 0.9 were assigned based on the dependency relationship type (e.g., negation modification, degree modification, speculative modification). A gated attention mechanism was used to dynamically adjust the representation strength of the feature vector, forming an uncertainty-enhanced feature vector with a dimension of 256.
[0079] The construction of the uncertainty association graph for medical entities is based on probabilistic relationship data from a medical knowledge base. Approximately 4.5 million disease-symptom association pairs were extracted from the Unified Medical Language System (UDLS), with each record including co-occurrence frequency, relative risk ratio, and diagnostic specificity indicators. A graph convolutional network was used to map medical entities to the knowledge space, calculating their similarity to standard concept nodes. An embedding representation of the entity nodes was generated using a random walk algorithm, with an embedding dimension of 128. The uncertainty-enhanced feature vector was then subjected to dot product similarity calculation with the knowledge base embeddings, and a probability distribution was generated using a softmax function to represent the association strength between the entity and different disease concepts.
[0080] The feature spectrum synthesis process employs a three-way feature fusion architecture. The multi-scale semantic feature matrix is reduced to 256 dimensions through global average pooling, and the uncertainty enhancement feature vector is adjusted to 256 dimensions through a fully connected layer. The probability relation mapping result is directly used as the 256-dimensional input. The three feature sources are first added element-wise, and then feature cross-computation is performed through a fully connected layer with 128 hidden units. Residual connections are used to preserve the original feature information, and finally, layer normalization generates a 512-dimensional medical entity uncertainty feature spectrum. This feature spectrum contains 32 feature groups, each with 16 feature values, describing different types of uncertainty patterns.
[0081] Three quantitative indicators are extracted from the uncertainty feature spectrum of medical entities. The feature distribution dispersion indicator is obtained by calculating the variance within each feature group and the covariance matrix between groups, using the determinant value to represent the overall dispersion. The context consistency indicator uses cosine similarity to calculate the matching degree between entity features and global context features, considering both local window (five words before and after) and global document similarity levels. The knowledge base matching confidence indicator is based on a pre-trained BERT matching model, taking entity descriptions and knowledge base concept definitions as input, and outputting a matching confidence score between 0 and 1.
[0082] The dispersion of the distribution is quantified using a multi-dimensional analysis method. In a 512-dimensional feature space, the first two principal components are extracted through principal component analysis as the deterministic-uncertainty dimension. The projected distribution of eigenvalues on these two dimensions is calculated, and the probability distribution function is fitted using kernel density estimation. The dispersion is ultimately quantified as the entropy value of the distribution function, while skewness (measuring distribution asymmetry) and kurtosis (measuring distribution sharpness) are calculated as auxiliary indicators. A higher entropy value indicates a lower degree of determinism and a more dispersed distribution.
[0083] Contextual fit assessment employs multi-level comparative analysis. A Sentence-BERT model is used to generate semantic vector representations of medical record documents, with a dimension of 1024. Euclidean distance and cosine similarity are calculated between medical entity vectors and document vectors, measuring absolute distance and directional consistency, respectively. Simultaneously, semantic compatibility between entities and their direct modifiers (adjectives and adverbs) is analyzed, and a compatibility score between 0 and 1 is output using a pre-trained compatibility model. For time-series texts, an LSTM model is used to analyze the changing patterns of entity representations over time, detecting any contradictory statements.
[0084] The semantic deterministic offset values are generated using dynamic weighted fusion. The three basic indicators are first standardized using a min-max process to eliminate dimensional differences. The weight coefficients are learned through a gradient boosting decision tree model, with the basic weights being: distribution dispersion 0.4, context fit 0.3, and knowledge base matching 0.3. These weights are dynamically adjusted based on entity type: diagnostic entities receive an additional 0.1 knowledge base weight, symptom entities receive an additional 0.1 context weight, and treatment entities receive an additional 0.1 distribution dispersion weight. The weighted sum is converted into a standard score between 0 and 1 using a sigmoid function, retaining four decimal places as the final output value. Scores above 0.7 are considered high-deterministic entities, and scores below 0.3 are considered low-deterministic entities; intermediate values require further judgment based on the specific clinical scenario.
[0085] The contextual influence analysis module employs a three-layer spatiotemporal convolutional neural network architecture for spatiotemporal joint analysis of modifying contexts in standardized medical record text data. The first layer uses 128 3×3 two-dimensional convolutional kernels to scan the spatial dimension of the text sequence, capturing the modifying relationships between adjacent words. The second layer uses 64 5×5 convolutional kernels to expand the receptive field and identify long-distance dependencies. The third layer is a temporal convolutional layer using 32 7×1 convolutional kernels that slide along the text sequence to analyze the temporal evolution patterns of modifying contexts. Each convolutional layer is followed by a ReLU activation function and max pooling for downsampling. The network output contains 64 feature maps, each corresponding to a specific modifying context distribution pattern.
[0086] To statistically analyze the distribution characteristics of negative words, degree adverbs, and expressions of uncertainty, a modifier dictionary with 423 entries was first established, including 87 negative words, 156 degree adverbs, and 180 expressions of uncertainty. A sliding window-based statistical method was employed, selecting 10 words before and after the medical entity as the analysis window. The frequency, positional distribution, and density changes of each type of modifier within each window were calculated. Kernel density estimation was used to generate a probability distribution map of the modifiers in the text space, and their spectral characteristics were analyzed using discrete Fourier transform to capture the periodic patterns of the distribution.
[0087] When matching the captured distribution patterns with clinical expression templates, an attention-based template matching algorithm is used. The pre-set template library contains 12 categories of clinical expression patterns, with 20-50 typical expression examples in each category. A bidirectional attention mechanism is used to calculate the relevance weights between the current text and each template example, and the matching score is obtained by weighted summation. The matching process considers three dimensions: lexical overlap, grammatical structure similarity, and semantic relevance. Finally, the template type with the highest matching score is selected as the classification result of the doctor's expression habits, and a matching confidence score is output.
[0088] The generation of clinical expression behavior profiles employs feature fusion technology. Spatiotemporal distribution features (64 dimensions), expression habit types (12 dimensions), and matching confidence (1 dimension) are concatenated to obtain an initial feature vector of 77 dimensions. Principal component analysis is then used to reduce the dimensionality to 32 dimensions, retaining 95% of the variance information. Based on this, three specific indicators are added: modifier context strength (calculated based on modifier density), distribution breadth (calculated based on modifier distribution standard deviation), and evolution pattern (based on temporal feature extraction), ultimately forming a 35-dimensional multidimensional clinical expression behavior profile. Each dimension undergoes min-max standardization to ensure that the values range from 0 to 1.
[0089] Three core indicators were extracted based on clinical verbal behavior profiling. The intensity of the modifier context was calculated by weighting the features from dimensions 1-10 of the behavioral profile, with the weight coefficients determined through ridge regression training. The breadth of distribution was calculated by combining the standard deviation and entropy values of features from dimensions 11-20, reflecting the dispersion of the modifier context. The temporal stability indicator was calculated by analyzing the coefficients of variation and autocorrelation coefficients of features from dimensions 21-30, measuring the stability of the modifier pattern over time.
[0090] The influence strength of negative words and expressions of uncertainty is calculated using an attention-weighted quantification method. First, a context-aware embedding representation of each modifier is obtained through a pre-trained language model. Then, its cosine similarity to the medical entity embedding is calculated. The similarity value, after being transformed by a sigmoid function, serves as the base influence strength. Simultaneously, the grammatical distance between the modifier and the entity is considered; for each additional word position in the distance, the influence strength decreases by 0.1. The final influence strength is the product of the base strength and the distance decay factor, ranging from 0 to 1.
[0091] The impact scope and effective distance of modifier contexts are assessed using a graph propagation-based algorithm. The text is constructed as a lexical graph, where nodes represent words and edges represent grammatical dependencies. Starting with modifiers, the propagation probability is calculated using a random walk algorithm. The impact scope is defined as the proportion of words with a propagation probability exceeding 0.2 out of the total vocabulary. The effective distance is obtained by calculating the harmonic mean of the shortest grammatical path lengths from modifiers to each affected word.
[0092] The generation of contextual modification influence values employs a weighted aggregation using a three-layer neural network. The input layer receives three metrics: influence strength (1D), influence range (1D), and influence distance (1D). The hidden layer contains eight neurons using the tanh activation function. The output layer generates the final influence value, ranging from 0 to 1, using the softmax function. The network weights are obtained through training with labeled samples using the Adam optimizer with a learning rate of 0.001 and a batch size of 32. The final output influence value is rounded to four decimal places to ensure computational precision meets clinical application requirements.
[0093] In the multi-objective decision optimization module, the module is responsible for fusing and optimizing the semantic features generated by the preceding modules. This module first performs standardization preprocessing on the semantic deterministic offset values and contextual modification influence values to eliminate the impact of dimensional differences. The standardization process uses Z-score normalization, specifically: first, the mean μ and standard deviation σ of each feature in the training dataset are calculated; then, the original feature value x is converted to a standard distribution of (x-μ) / σ. For the semantic deterministic offset values, the mean is typically between 0.45 and 0.55, and the standard deviation is between 0.15 and 0.25; the mean of the contextual modification influence values is generally between 0.35 and 0.45, and the standard deviation is approximately 0.18 to 0.28. This standardization ensures that the two feature values are within the same numerical range, laying the foundation for subsequent weighted fusion. During the standardization process, the system monitors the distribution changes of the feature values in real time. When a distribution offset exceeding a threshold of 0.1 is detected, parameter recalibration is automatically triggered to ensure the stability of the standardization effect.
[0094] Based on the characteristics of the current medical entity type, the system employs a dynamic weighted fusion mechanism to automatically adjust the fusion weight coefficients of deterministic offset values and contextual modification influence values. This dynamic weighted fusion mechanism is based on a predefined medical entity type weight mapping table, which includes eight major entity types such as disease diagnosis, clinical symptoms, examinations and tests, and drug treatments, each further subdivided into 36 subcategories. For each entity type, the system configures different combinations of weight coefficients. For example, for disease diagnosis entities, the weight coefficient for semantic deterministic offset values is set to 0.6, and the weight coefficient for contextual modification influence values is 0.4; while for examination and test entities, the weight allocation is 0.3 and 0.7. The weight coefficients are determined by analyzing the statistical characteristics of a large amount of labeled data, combined with manual calibration based on clinical expert experience. In actual operation, the system automatically determines the type of the currently processed entity through the entity type recognition module, and then retrieves the corresponding weight coefficients from the weight mapping table. Furthermore, the system introduces context-adaptive adjustment; when a special clinical scenario is detected (such as emergency department medical records), the weight coefficients are automatically fine-tuned, with the adjustment range generally controlled within ±0.15.
[0095] Establishing a nonlinear mapping between deterministic offset values and contextual modification influence values is achieved through a specially designed cross-modal attention fusion engine. This fusion engine comprises two main components: a feature interaction layer and an attention generation layer. The feature interaction layer first concatenates two standardized feature values, then passes them through a fully connected network layer with 128 hidden units, employing the ReLU activation function to capture the complex nonlinear relationships between features. The attention generation layer calculates the contribution weight of each feature to the final representation, using a softmax function to generate the attention distribution. During this process, the system generates a context-aware fused feature representation with a 64-dimensional dimension. This representation not only contains information from the original features but also encodes the patterns of interaction between features, better reflecting the overall deterministic state of the medical entity. The parameters of the fusion engine are obtained through end-to-end training on labeled data, with the training objective being to minimize the difference between the deterministic levels of the fused features and those of manually labeled data.
[0096] The weighted deterministic offset value and contextual influence value are concatenated with the fused feature representation to generate the final medical semantic feature vector. The concatenation process uses a channel connection method, combining the weighted original features (2-dimensional) with the fused feature representation (64-dimensional) into a 66-dimensional comprehensive feature vector. This feature vector retains the interpretability of the original features while incorporating deep feature representations obtained through deep learning. To further enhance feature representation capabilities, the system also performs dimensionality reduction on the concatenated feature vector, compressing it to 32 dimensions through a fully connected layer with 32 output units. The final generated medical semantic feature vector contains comprehensive information on semantic determinism and contextual influence, providing a rich feature foundation for subsequent multi-objective decision optimization. Each dimension of the feature vector has a clear clinical semantic meaning; for example, the first 16 dimensions mainly reflect the deterministic features of the entity itself, while the latter 16 dimensions encode more of the influence patterns of contextual modification.
[0097] The system simultaneously optimizes clinical safety, data integrity, and data entry efficiency objectives using a multi-objective optimization function. Clinical safety is quantified by calculating the conformity of decision-making results with clinical guidelines, specifically through rule-based matching, with built-in clinical safety rules. Data integrity is measured by evaluating the completeness and information content of filled fields, using a weighted combination of information entropy and field completion rate as the indicator. Data entry efficiency is evaluated by calculating the time complexity of the decision-making process and system resource consumption. The weights of the three objectives are dynamically adjusted according to the application scenario: in outpatient settings, the weights are set to 0.5 for data entry efficiency, 0.3 for clinical safety, and 0.2 for data integrity; while in inpatient settings, the weights are increased to 0.6 for clinical safety, 0.3 for data integrity, and 0.1 for data entry efficiency. The multi-objective optimization function uses a weighted summation method to linearly combine the three objective functions into a single comprehensive optimization objective.
[0098] A Pareto optimality search algorithm is employed to find the optimal balance among multiple optimization objectives. The system uses an improved non-dominated sorting genetic algorithm (NSGA-II) as its search framework. The algorithm first initializes a population containing 100 candidate solutions, each representing a set of possible decision parameter configurations. Then, new solutions are generated through genetic operations such as crossover and mutation, with a crossover probability of 0.9 and a mutation probability of 0.1. During each generation of evolution, the algorithm performs non-dominated sorting and crowding calculations on the solutions, selecting the optimal solution on the Pareto front. The search process continues for 50 generations, ultimately outputting a set of candidate decision schemes containing 20–30 non-dominated solutions. Each candidate scheme exhibits different trade-offs among the three optimization objectives; some schemes prioritize clinical safety, while others focus more on reporting efficiency.
[0099] A clinical knowledge-constrained validator is used to verify the medical rationality of candidate solutions. The validator is built based on a clinical knowledge graph and medical logic rules, and includes three levels of verification: syntactic verification checks the usage and formatting accuracy; semantic verification checks the rationality of relationships between medical concepts; and pragmatic verification checks the applicability to the clinical scenario. During the verification process, the system generates a compliance score for each solution, and solutions with scores below the threshold of 0.8 are directly eliminated.
[0100] Based on the needs of real-time application scenarios, the optimal medical data parsing solution is selected from validated candidate solutions. The selection process employs a multi-attribute decision-making method, specifically using the TOPSIS (Top-Approximation-Ideal-Solution Ranking) algorithm. The algorithm first constructs a decision matrix, where rows represent candidate solutions and columns represent evaluation indicators (including the achievement of three optimization objectives, compliance scores, etc.). Then, the weights of each indicator are determined, dynamically adjusted according to the real-time application scenario: during morning outpatient peak hours, the weight of the data entry efficiency indicator increases; when processing critically ill patient records, the weight of the clinical safety indicator increases. Next, the distance between each solution and the positive and negative ideal solutions is calculated, and the solution with the highest relative proximity is selected as the optimal solution. The entire decision-making process is controlled within 200 milliseconds to ensure the system's real-time responsiveness. The selected optimal solution outputs structured medical data parsing instructions, which are then delivered to the subsequent intelligent pre-entry execution module.
[0101] In the intelligent pre-filling execution module, which is the final output stage of the system, the pre-processing results are transformed into usable clinical data. This module performs ontology alignment between the optimized medical semantic feature vectors and the medical knowledge graph. The ontology alignment process employs an attention-based semantic similarity calculation model, which first maps the medical semantic feature vectors to the same vector space as the medical knowledge graph. This mapping process is implemented through a fully connected neural network with 64 hidden units, using the tanh activation function to convert the 32-dimensional medical semantic feature vectors into 128-dimensional graph space vectors. Similarity calculation uses a cosine similarity algorithm, specifically calculating the cosine of the angle between two vectors, with a similarity threshold set at 0.85. When the similarity reaches this threshold, the system considers the medical entity to have successfully matched with a concept node in the knowledge graph. For each medical entity, the system retrieves the top three candidate concept nodes with the highest similarity and records their respective matching scores. These candidate concepts and their matching scores provide important reference for subsequent clinical rationality verification.
[0102] Clinically validating the alignment results is a crucial step in ensuring data quality. The validation process consists of two main parts: symptom-disease association verification and treatment rationality checks. Symptom-disease association verification is based on probabilistic relationships in a medical knowledge graph. The system calculates the clinical association strength between symptoms and diseases, with a threshold of 0.6. For matches with an association strength below this threshold, the system issues a warning. Treatment rationality checks involve verification of multiple aspects, including drug-disease indications, drug-drug interactions, and consistency between treatment and examination results. During validation, the system generates a rationality score for each alignment result, ranging from 0 to 1. Alignment results with scores below 0.7 are marked as requiring manual review; scores between 0.7 and 0.9 are accepted with added comments; and scores above 0.9 are directly adopted.
[0103] Ensuring the consistency of the generated structured medical data with the overall semantic environment of the medical record is achieved through contextual consistency detection. Contextual consistency detection, based on a recurrent neural network architecture, captures long-term dependencies within the medical record text. The system maintains a medical record-level state vector, which encodes the semantic environment and clinical context information of the entire medical record. For each medical entity, the system calculates its semantic consistency score with the medical record state vector, using an attention-weighted similarity metric. A consistency score threshold of 0.75 is set; entities below this threshold are flagged as potentially having contextual inconsistencies. Furthermore, the system uses a temporal consistency checker to verify the temporal logic of medical events, such as ensuring that diagnosis is no later than treatment and examination occurs after symptom onset. These checks ensure that the final generated structured data is not only locally correct but also logical and coherent within the global context.
[0104] Based on the alignment and verification results, the system assigns a corresponding certainty level label to each medical entity. The certainty level is divided into three levels: high certainty (confidence greater than 0.9), medium certainty (confidence between 0.7 and 0.9), and low certainty (confidence less than 0.7). The labeling process is based on a weighted comprehensive score of multiple factors, including semantic similarity score (weight 0.4), clinical rationality score (weight 0.3), contextual consistency score (weight 0.2), and data quality indicators (weight 0.1). The final certainty label obtained for each medical entity not only includes the level classification but also a detailed confidence score and explanation of the decision-making basis. These labels provide a decision-making foundation for subsequent differentiated pre-filling strategies, enabling the system to adopt appropriate processing methods based on different levels of certainty.
[0105] When performing autofill operations on highly deterministic data, the system uses direct mapping to populate electronic medical record fields. The system maintains a field mapping rule base containing multiple mapping rules that define the correspondence between medical entity types and electronic medical record form fields. During the autofill process, the system simultaneously populates standard terminology codes and raw text descriptions to ensure both machine and human readability. For each autofill operation, the system records the filling time, data source, and determinism score for subsequent auditing and traceability.
[0106] For low-certainty data, the system adds visual review indicators and locks the corresponding fields. The review indicators use a color-coded system: yellow indicates data requiring attention (medium certainty), and red indicates data requiring immediate action (low certainty). The locking mechanism prevents these fields from being accidentally modified; editing is only possible after authorized human review. The system also provides detailed review guidelines, including uncertainty cause analysis, recommended treatment suggestions, and relevant clinical reference information. These designs ensure data security while providing clinicians with ample decision support.
[0107] Based on a medical knowledge graph and similar case analysis, the system generates candidate fill options that conform to clinical norms. The candidate generation process employs a combination of collaborative filtering and knowledge reasoning. First, the system retrieves the most relevant standard concepts to the current medical entity from the medical knowledge graph, generating knowledge-based candidate options. Then, it searches for similar cases in the historical case database, extracting how these cases handled the same entities, generating instance-based candidate options. Finally, a ranking algorithm comprehensively scores and ranks the candidate options, considering factors such as clinical applicability, frequency of use, and compliance with the latest guidelines. The system provides 3 to 5 candidate fill options for each low-determinism data point, each accompanied by a detailed description and confidence score.
[0108] Based on the needs of real-time medical scenarios, the system adaptively adjusts the priority and filling strategy of pre-filled fields. The system monitors various factors in real time, including workload, user roles, and clinical scenarios. During peak outpatient hours, the system prioritizes the automatic filling of high-determinism data to improve filling efficiency; in intensive care settings, the system increases the rigor of clinical safety checks and lowers the determinism threshold for automatic filling. The adjustment strategy is based on preset rules and machine learning models, enabling dynamic optimization of system behavior according to actual conditions. The system re-evaluates scenario parameters every 5 minutes to ensure that the filling strategy always maintains an optimal match with current clinical needs. This adaptive design allows the system to maintain optimal performance in different clinical environments.
[0109] The working principle of this invention is as follows: A data acquisition and preprocessing module receives and standardizes multi-source heterogeneous medical text data, employing a depth scanner for entity recognition and multi-level medical terminology normalization. An uncertainty feature analysis module constructs a medical entity uncertainty feature spectrum and calculates semantic deterministic offset values. A context-aware deep convolutional network is used to extract multi-scale features, which are then mapped probabilistically using a medical knowledge base. A contextual influence analysis module analyzes modifying contexts using a spatiotemporal convolutional network, constructing a clinical expression behavior profile and calculating contextual modification influence values. A multi-objective decision optimization module fuses the above feature values into a medical semantic feature vector, employing a multi-objective optimization function and a Pareto optimal solution search algorithm for synergistic optimization of clinical safety, data integrity, and filling efficiency. Finally, an intelligent pre-filling execution module performs ontology alignment and logical verification based on a medical knowledge graph, implementing differentiated pre-filling strategies according to the deterministic label level. High-deterministic data is automatically filled, while low-deterministic data generates review labels and provides candidate options, thereby achieving precise and adaptive intelligent medical record parsing and pre-filling.
[0110] The foregoing has provided a detailed description of one embodiment of the present invention, but this description is merely a preferred embodiment and should not be construed as limiting the scope of the invention. All equivalent variations and modifications made within the scope of the claims of this invention should still fall within the patent coverage of this invention.
Claims
1. An AI-based medical record intelligent analysis and pre-reporting system, characterized in that, The method comprises the following steps: A data acquisition and preprocessing module receives medical record text data through a medical text data interface, and pre-processes and standardizes medical terms of the medical record text data; An uncertainty feature analysis module performs deep semantic analysis on the standardized medical record text data, constructs a medical entity uncertainty feature spectrum, and calculates a semantic certainty offset value; The calculation process of the semantic certainty offset value is as follows: Based on the medical entity uncertainty feature spectrum, the feature distribution dispersion index, the context consistency index, and the knowledge base matching confidence index are extracted; The distribution dispersion degree of the feature value in the certainty and uncertainty dimensions is calculated; The feature value is the specific numerical value of each dimension of the medical entity uncertainty feature spectrum; The degree of agreement of the entity expression with the overall medical record context semantics is evaluated; The distribution dispersion degree, the degree of agreement of the context semantics, and the knowledge base matching degree are weighted and fused to generate a quantitative semantic certainty offset value; A context influence analysis module performs spatio-temporal joint analysis on the modifying context in the standardized medical record text data, constructs a clinical expression behavior portrait, and calculates a context modification influence value; The calculation process of the context modification influence value is as follows: Based on the clinical expression behavior portrait, the modifying context intensity index, the distribution breadth index, and the temporal stability index are extracted; The influence intensity of negative words and uncertain expressions on medical entity semantics is calculated, which is denoted as the modifying context intensity; The influence range and action distance of the modifying context in the text are evaluated, which is denoted as the distribution breadth; The modifying context intensity, the distribution breadth, and the temporal stability are weighted and aggregated to generate a quantitative context modification influence value; A multi-objective decision optimization module fuses the semantic certainty offset value and the context modification influence value into a medical semantic feature vector, and inputs the medical semantic feature vector into a pre-trained medical record analysis model for multi-objective decision optimization; An intelligent pre-filling execution module performs ontology alignment and logical verification based on the optimized medical semantic feature vector and the medical knowledge graph, generates structured medical data with certainty labels, and drives differentiated pre-filling strategies according to the certainty label level, specifically including: performing automatic filling on high-certainty data, generating a to-be-reviewed label for low-certainty data, and providing candidate filling content.
2. The AI-based medical record intelligent analysis and pre-reporting system according to claim 1, characterized in that, The pre-processing and medical term standardization of the medical record text data specifically includes: Receiving heterogeneous medical record data, detecting sentence breakpoints and medical narrative paragraph segmentation for unstructured text; Using a deep scanner to identify medical entity boundaries and perform class pre-labeling on the segmented text units; Starting a multi-level medical term normalization pipeline, matching a basic medical term library, and performing term semantic disambiguation; Performing context reconstruction based on a knowledge graph on the semantic discontinuity generated in the standardization process to generate standardized text output with unified medical coding. 3.The AI-based medical record intelligent analysis and pre-reporting system according to claim 1, characterized in that, The deep semantic analysis of the standardized medical record text data and the construction of the medical entity uncertainty feature spectrum specifically include: The standardized text is subjected to multi-scale semantic feature extraction by a context-aware deep convolutional network to generate a basic semantic feature matrix; The context fragments containing suspected nature words are subjected to weight enhancement to form an uncertainty-enhanced feature vector; An uncertainty correlation graph of medical entities is established to map the uncertainty-enhanced feature vector with the disease-symptom probability relationship in the medical knowledge base; The multi-scale semantic feature matrix, the uncertainty-enhanced feature vector and the probability relationship mapping result are integrated to generate a multi-dimensional uncertainty feature spectrum of medical entities.
4. The AI-based medical record intelligent analysis and pre-reporting system according to claim 1, characterized in that, The spatiotemporal joint analysis of the modifying context in the standardized medical record text data is performed to construct a clinical expression behavior portrait, specifically including: The medical record text is subjected to multi-dimensional scanning by a spatiotemporal convolutional neural network to capture the distribution pattern and time evolution characteristics of the modifying context in the text sequence; The spatiotemporal distribution characteristics of negative words, degree adverbs and uncertain expressions around specific medical entities are counted; The captured distribution pattern is matched with a preset clinical expression template to identify the type of expression habit of the doctor; The spatiotemporal distribution characteristics and expression habit type are integrated to generate a multi-dimensional clinical expression behavior portrait containing the intensity, distribution breadth and evolution pattern of the modifying context.
5. The AI-based medical record intelligent analysis and pre-reporting system according to claim 1, characterized in that, The construction process of the medical semantic feature vector is as follows: The semantic certainty offset value and context modification influence value are standardized to eliminate dimensional differences; The fusion weight coefficients of the certainty offset value and the context modification influence value are automatically adjusted according to the type characteristics of the current medical entity; A nonlinear mapping relationship between the certainty offset value and the context modification influence value is established to generate a fusion feature representation with context awareness; The weighted certainty offset value and context modification influence value are spliced with the fusion feature representation to generate a medical semantic feature vector containing semantic certainty and context influence.
6. The AI-based medical record intelligent analysis and pre-reporting system according to claim 1, characterized in that, The medical semantic feature vector is input into a pre-trained medical record analysis model for multi-objective decision optimization, specifically including: The clinical safety target, data integrity target and filling efficiency target are simultaneously optimized by a multi-objective optimization function; The best balance point is found among multiple optimization targets by using a Pareto optimal solution search algorithm to generate a candidate decision scheme set; The candidate schemes are subjected to medical rationality verification by a clinical knowledge constraint verifier to eliminate decision schemes that do not conform to clinical standards; Based on the demand side of real-time application scenarios, the optimal medical data analysis scheme is selected from the verified candidate schemes.
7. The AI-based medical record intelligent analysis and pre-reporting system according to claim 1, characterized in that, The ontology alignment and logical verification between the optimized medical semantic feature vector and the medical knowledge graph are performed, specifically including: The medical semantic feature vector is subjected to semantic similarity matching with the concept nodes in the medical knowledge graph; The alignment result is subjected to clinical rationality verification, including symptom-disease correlation verification and treatment scheme rationality check; The generated structured medical data is ensured to be coherent with the overall semantic environment of the medical record; According to the alignment and verification results, each medical entity is labeled with a corresponding certainty level mark. 8.The AI-based medical record intelligent analysis and pre-reporting system according to claim 1, characterized in that, The differential pre-filling strategy is driven according to the certainty mark level, specifically including: The high-certainty data is directly filled into the electronic medical record fields; Add visual audit marks to low-certainty data and lock corresponding fields; Based on medical knowledge graph and similar case analysis, generate candidate filling options that conform to clinical standards; Adaptively adjust the priority and filling strategy of pre-report fields according to real-time medical scene needs.
Citation Information
Patent Citations
Medical record structured analysis method based on medical field entities
CN110032648A
Deep learning large model-based medical record information extraction and analysis method and system
CN120579548A