Real-time quality control and coding method for medical record home page data based on multi-dimensional verification
By combining multi-dimensional tensor models and reinforcement learning with knowledge graphs, real-time quality control and coding of medical record homepage data are achieved, solving the problems of update lag and resource dependence in existing technologies, improving the real-time and accuracy of the coding system, reducing maintenance costs, and improving the stability and efficiency of the system.
Patent Information
- Application Number
- CN202510747386.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-09-12
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing medical record homepage data coding system lacks real-time performance, lags behind in knowledge updates, and has insufficient graph expansion capabilities, resulting in inaccurate coding suggestions and an inability to meet the timely quality control requirements in the clinical environment. It also relies excessively on manual intervention and computing resources, resulting in low efficiency.
A multi-dimensional tensor model and reinforcement learning method are used, combined with knowledge graphs to perform real-time quality control and encoding of medical record homepage data. The potential characteristics of medical record data are analyzed through tensor decomposition and factor matrix, and the verification rules and edge weights are dynamically adjusted to achieve incremental updates and automatic expansion of the graph, avoiding resource occupation and stability issues caused by overall reconstruction.
It achieves dynamic adaptation of the medical knowledge graph and optimization of relationship accuracy, reduces system resource usage, improves the real-time and accuracy of coding suggestions, reduces maintenance costs, solves the problem of uncontrolled expansion of the graph in high-frequency update scenarios, and improves the stability and efficiency of the coding system.
Smart Images

Figure CN120636662A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical information technology, and in particular to a multi-dimensional verification method for real-time quality control and coding of medical record homepage data. Background Art
[0002] As hospital information systems continue to improve, medical record front-page data is widely used in a variety of scenarios, including medical management, medical insurance payment review, and clinical research. The diagnosis and surgical procedure codes on medical record front pages, as key data reflecting the patient's diagnosis and treatment process, have a direct impact on the quality of downstream medical data use. Consequently, medical institutions and information system vendors at all levels generally use coding review systems or semi-automated quality control modules to assist with coding tasks.
[0003] Currently, common coding assistance methods primarily rely on rule templates, empirical logic, or static mapping libraries to recommend and verify diagnostic and surgical codes. While these solutions offer some standardization capabilities, they often suffer from issues such as delayed updates and poor recognition of new disease and surgical procedures. This is especially true given the increasing diversity of clinical terminology and the volume of data, which is causing these existing methods to exhibit significant shortcomings in scalability and adaptability.
[0004] In other real-world deployments, to improve data quality, some systems have introduced knowledge graphs or entity association networks as a foundational reasoning framework, attempting to enhance coding matching accuracy using structured knowledge. However, these graphs are often statically constructed, built once and used over time, and lack the ability to update structure and relationships based on real-time medical record data. As a result, in situations where new diseases or surgical procedures are constantly emerging, the graph content deviates from actual business scenarios, rendering coding suggestions ineffective and ultimately impacting the overall usability of the system.
[0005] Furthermore, existing coding systems based on graphs or multidimensional rules often require manual intervention to maintain knowledge when processing newly added medical record data. The node expansion and edge weighting processes lack automated mechanisms, resulting in a high degree of system dependency and continuously increasing maintenance costs. This not only limits efficiency but also makes it easy for errors to spread due to human error, further weakening the system's coding assistance capabilities.
[0006] To avoid knowledge expansion and computational redundancy, some systems employ a periodic full-scale reconstruction of the atlas to refresh knowledge content. However, this strategy relies heavily on computing resources, has a long update cycle, and lacks real-time performance, failing to meet the timely quality control requirements of clinical settings. If this update strategy lags behind, the entire coding system will become stale and unable to make accurate recommendations.
[0007] Furthermore, current coding quality control processes are generally insensitive to "contextual co-occurrence." This means that the probability of co-occurrence between diagnoses and procedures based on real case data is overlooked during the coding recommendation process, resulting in recommendations lacking practical relevance. This is particularly true for cases involving complex diagnoses and multiple surgical procedures, where recommendation systems often experience quality issues such as misclassification and omission. Summary of the Invention
[0008] In response to the shortcomings of the existing technology, the present invention provides a real-time quality control and coding method for medical record homepage data with multi-dimensional verification, which solves the problems in the existing technology of lack of real-time coding suggestions, delayed knowledge updating and insufficient graph expansion capabilities.
[0009] To achieve the above objectives, the present invention is implemented through the following technical solutions: a multi-dimensional verification method for real-time quality control and coding of medical record homepage data, comprising the following steps:
[0010] S1. Obtain medical record front page data from a hospital information system or an electronic medical record system, wherein the medical record front page data includes structured data, unstructured data, and time series data;
[0011] S2. Classify the medical record homepage data according to the disease classification standard, divide the medical record homepage data into different disease categories, and construct a multi-dimensional tensor model related to each disease for subsequent data processing and feature extraction;
[0012] S3. Applying a tensor decomposition method to the multi-dimensional tensor model constructed based on the medical record homepage data to decompose it into a core tensor and multiple factor matrices. The core tensor is used to represent the potential features of the medical record homepage data, and the factor matrix is used to describe the correlation between the dimensions in the medical record homepage data.
[0013] S4. Based on the core tensor and factor matrix obtained in step S3, the reconstruction error of the medical record homepage data in the tensor space is calculated, and the error is compared with a preset threshold. If the reconstruction error is greater than the threshold, the medical record homepage data is determined to be abnormal data, triggering subsequent abnormal data processing and verification optimization;
[0014] S5. Using reinforcement learning to optimize the verification rules for abnormal data, dynamically adjust the weight distribution of each data feature in the verification rules according to the characteristics of the abnormal data in step S4, thereby achieving dynamic optimization of the data verification strategy for the medical record homepage;
[0015] S6. Based on the verification rules optimized in step S5, a knowledge graph is constructed. The knowledge graph includes diseases, surgeries, complications, and their interrelationships. Based on the knowledge graph, coding reasoning is performed on the medical record homepage data. The reasoning model generates coding suggestions related to the disease type based on the analysis of the medical record homepage data, and applies them to the coding optimization of the medical record homepage data.
[0016] S7. Based on the coding suggestions generated in step S6, the knowledge graph is updated in combination with the real-time changes of the medical record homepage data, and the medical entities and their relationships in the knowledge graph are incrementally updated to ensure that the knowledge graph can be consistent with the latest medical record homepage data.
[0017] Preferably, the dimensions of the multidimensional tensor model in step S2 include at least four dimensions of age, primary diagnosis, surgical method, cost, and length of hospital stay, and the rank of the tensor decomposition is dynamically selected for different disease categories. The specific rules are:
[0018] If the disease type is tumor, the rank is ≥5; if it is other diseases, the rank is =3.
[0019] Preferably, the tensor decomposition method in step S3 is Tucker decomposition, which is expressed as:
[0020]
[0021] Among them, the core tensor The dimension is consistent with the dynamically selected rank value, and the factor matrix A (i) Solve it iteratively using the alternating least squares (ALS) method. is the original multi-dimensional tensor, × N The product of tensors over their Nth mode.
[0022] Preferably, the threshold value setting method in step S4 is: based on the distribution of the reconstruction error of the historical medical record data, the dynamic threshold value θ is calculated using kernel density estimation. th , so that the abnormal data determination meets the significance level γ<0.05.
[0023] Preferably, the reinforcement learning method in step S5 includes:
[0024] Design a dual-objective reward function:
[0025]
[0026] Among them, R(s t ,a t ) is in state s t Next, perform action a t Instant rewards received; acdlovd The number of resolved conflicts in the current state; nconflict_total is the total number of conflicts detected in the current state; is the conflict resolution rate; n steps The number of steps required for the doctor to correct the conflict according to the system's suggestions; n max The maximum number of operation steps allowed by the system; is the operational efficiency indicator; α and β are weight coefficients, which are dynamically adjusted through the Nash equilibrium game model to achieve a balance between the conflict resolution rate and operational efficiency.
[0027] Preferably, the method for constructing the knowledge graph in step S6 includes:
[0028] Compute node embedding vectors for diseases, surgeries, and complications through graph neural networks;
[0029] Assign weights to related edges based on the graph attention mechanism.
[0030] Preferably, the knowledge graph update in step S7 is an incremental update, specifically including:
[0031] According to the co-occurrence frequency of diseases and surgeries in the new medical record data, adjust the weight w of the edge in the knowledge graph ij , and its update formula is:
[0032] in, is the edge weight between node i and node j after update; is the edge weight between node i and node j before updating; λ is the decay factor; n co-occur (i, j) is the number of times node i and node j co-occur in the newly added data; n new_cases is the total number of cases with newly added data; is the co-occurrence frequency of node i and node j in the newly added data;
[0033] When new disease or surgery nodes are added, the knowledge graph is dynamically expanded through the graph embedding algorithm.
[0034] Preferably, the specific method for dynamically adjusting the weight of the verification rule in step S5 is:
[0035] The weight assignment of verification rules is modeled as a Markov decision process, where the state space includes the current number of conflicts and the eigenvector, and the action space is the weight adjustment strategy;
[0036] The optimal strategy is solved through the Q-learning algorithm.
[0037] Preferably, the coding suggestions related to the disease types generated in step S6 should be accompanied by interpretability feedback, specifically including:
[0038] Locating abnormal dimensions based on tensor reconstruction error;
[0039] The correction path is generated through the knowledge graph, and the feedback is in the form of: "A conflict between [Dimension X] and [Dimension Y] is detected, and it is recommended to check [Field Z]."
[0040] The present invention provides a multi-dimensional verification method for real-time quality control and encoding of medical record homepage data. It has the following beneficial effects:
[0041] 1. The present invention adopts an incremental edge weight adjustment mechanism driven by co-occurrence frequency, combined with an adjustable attenuation factor parameter λ, to effectively realize the dynamic adaptation of the medical knowledge graph and the optimization of relationship accuracy, and achieves the technical effect of continuously synchronizing the strength of entity relationships without reconstructing the overall graph structure. Compared with the update method of the existing technology that requires periodic overall reconstruction of the graph, this solution avoids the large-scale occupation of system resources and structural instability, and solves the technical bottleneck of difficult to ensure graph consistency during the update process.
[0042] 2. The present invention introduces a graph embedding model for generating structural representations of new nodes, and automatically establishes connection relationships based on semantic similarity, thereby realizing plug-and-play graph expansion of medical entities. The effect is that new nodes can be quickly integrated into the graph without being "isolated" and can participate in coding reasoning tasks. Unlike traditional graph update mechanisms based on manual rules or static mapping, this method does not rely on human intervention, and solves the problems of low efficiency in the node expansion process and delayed response of upstream and downstream tasks.
[0043] 3. The present invention maintains the stability and reasoning efficiency of the knowledge graph after dynamic updates through embedding vector normalization processing and graph structure pruning strategy. This ensures that even in the face of high-frequency data updates, the system will not have relationship redundancy or graph structure loss of control. Compared with the existing methods where the accumulated edge weights are unbalanced or the growth of nodes leads to a decrease in query performance, this solves the core difficulty of "uncontrolled expansion" of medical graphs in high-frequency update scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 Flowchart of the present invention. DETAILED DESCRIPTION
[0045] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the present specification. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0046] Please see the attached Figure 1The embodiment of the present invention provides a multi-dimensional verification method for real-time quality control and coding of medical record homepage data, comprising the following steps:
[0047] S1. Obtain medical record front page data from a hospital information system or an electronic medical record system, wherein the medical record front page data includes structured data, unstructured data, and time series data;
[0048] This step, as the starting point of the method of the present invention, mainly realizes the extraction, classification, formatting and preprocessing of the medical record homepage data in the hospital information system (HIS) or electronic medical record system (EMR). Its core is to build a structured data set in a unified format, support the operation of multiple modules such as subsequent tensor construction, conflict detection, knowledge graph update and intelligent coding recommendation, and form a closed-loop data-driven mechanism. This step combines data interface calls, information extraction, time series data alignment and dynamic missing completion strategies to ensure the comprehensiveness and high availability of data.
[0049] In this embodiment, the medical record homepage data interacts with the hospital business system through a standardized interface, and the acquisition methods include three channels: database query, interface subscription and log analysis.
[0050] In one possible implementation, the data source module is configured as follows:
[0051] The structured data source is connected to the HIS database through the JDBC protocol and executes incremental extraction statements such as SELECT * FROM homepage_table WHERE discharge_date >= CURDATE() - 1;
[0052] The unstructured data source uses HTTP API to call the medical record information interface and uses the medical NLP module to extract entities;
[0053] Time series data monitors the hospital's real-time data bus through the MQ message queue, and parses "ADT" events in the HL7v2 format for time-tagged synchronization.
[0054] The data categories are broken down as follows:
[0055] Structured data fields include, but are not limited to:
[0056] Basic information: age, gender, hospital number, admission and discharge time, and length of hospital stay;
[0057] Medical coding: main diagnosis ICD-10, major surgery ICD-9-CM-3;
[0058] Economic fields: total cost, drug cost, surgery cost, nursing cost;
[0059] Coding assistance: surgical incision level, anesthesia method, diagnosis and treatment category.
[0060] Unstructured text fields include:
[0061] Chief complaint (e.g., "abdominal pain for 3 days with fever");
[0062] Free-format clinical texts such as current medical history, diagnosis and treatment records, and surgical records;
[0063] The physician's remarks field contains implicit factors such as subjective inferences and excluded diagnoses.
[0064] Time series data fields include:
[0065] Laboratory data (such as blood routine, liver function);
[0066] Vital signs (such as temperature, pulse, blood pressure);
[0067] Medication data (e.g., intravenous fluid records, antibiotic use timeline).
[0068] As an option, structured fields are uniformly mapped to the ontology definition field structure, for example:
[0069] "Main diagnosis code" is mapped to the standard field diag_main_code;
[0070] "Surgical method" is mapped to surgery_main_code;
[0071] The medical expense breakdown fields are merged into independent fields such as fee_total, fee_drug, and fee_operation.
[0072] The unstructured data preprocessing process includes:
[0073] Medical entity extraction based on the BERT-BiLSTM-CRF model to identify diseases (D), symptoms (S), drugs (M), surgical procedures (OP), etc.
[0074] The terms "cesarean section" and "cesarean section" are synonymously normalized to "cesarean section", and the coding mapping is 75.0x00x;
[0075] The extraction results are stored in the standard triple format, for example: <appendicitis, disease, main diagnosis>, <laparoscopic resection, surgery, main surgical procedure>.
[0076] The time series data formatting process includes:
[0077] All time series are unified at hour-level granularity;
[0078] The SlidingWindow mechanism (window length 24 hours, sliding step length 1 hour) is used to generate the feature matrix;
[0079] The following matrix is generated for each patient:
[0080] Among them, X t is the time series feature vector at time t, Indicates the value of the i-th physical sign (such as body temperature).
[0081] When there are missing values in structured fields, a dynamic filling strategy is used to predict missing values based on the associated fields and historical case data. The filling function is as follows:
[0082] in, is the value of the field to be completed; f impute is the filling function, which can be mean filling, k-NN regression, or BERT-based field context reasoning; v1~v k Context fields for the current case, such as diagnosis, admission time, previous surgical history, etc.
[0083] In some embodiments, a knowledge graph-based path reasoning algorithm is used to generate field suggestion values:
[0084] If the main diagnosis is "gallstones", the recommended surgical code based on atlas reasoning is "cholecystectomy (51.22)";
[0085] If the past medical condition includes "hypertension", the comorbidity field will be automatically marked as "yes".
[0086] Data encryption transmission mechanism and desensitization rules:
[0087] The encryption method of the data transmission process is expressed by the following formula:
[0088] C=RSA pub (K)||AES K (D), where C is the final encrypted message; K is the symmetric encryption key (temporarily generated randomly); RSA pub Public key encryption algorithm configured for the hospital side; AES K The key K is used for AES symmetric encryption; D is the original data content of the medical record homepage, which has been desensitized.
[0089] Masked fields include:
[0090] Name (replaced with hash value);
[0091] Hospitalization number (partially masked, leaving only the first and second digits);
[0092] Medical insurance number, ID card number (using the full replacement strategy);
[0093] Physician signature field (converted to standard job identity, such as "Attending Physician A").
[0094] The structured data table, standardized text features, and time series feature matrix output by step S1 will be passed as input to step S2 to build a tensor model Where n represents the number of cases; m represents the number of features, including structured fields and entity labels; and k represents the time dimension or extraction dimension (optional).
[0095] In addition, the unstructured data is used as the entity matching basis for coding conflict detection in step S4, and the time series data is used to calculate the postoperative event response rate in step S6, affecting the conflict resolution score in the reward function.
[0096] S2. Classify the medical record homepage data according to the disease classification standard, divide the medical record homepage data into different disease categories, and construct a multi-dimensional tensor model related to each disease for subsequent data processing and feature extraction;
[0097] In the technical solution of the present invention, step S2 is mainly responsible for classifying the medical record homepage data into disease types, and constructing corresponding multi-dimensional tensor models for each disease based on these classifications for subsequent data processing and feature extraction. Specifically, through the disease classification module, the system divides the medical record homepage data into different categories according to the characteristics of the disease (such as main diagnosis, surgical method, number of days of hospitalization, etc.), and captures the multi-dimensional characteristics of each disease through multi-dimensional tensor modeling. On this basis, the rank of the tensor decomposition is dynamically selected for different disease categories in order to optimize the performance of the model and provide support for subsequent coding recommendations and anomaly detection.
[0098] In this embodiment, the disease classification of the medical record homepage data is based on the diagnosis information of the disease, specifically using the ICD-10 coding system or other standardized disease coding systems. By analyzing the primary diagnosis and secondary diagnosis on the medical record homepage, combined with the patient's medical history and other clinical information, the system divides the case into several disease categories. Common classifications include:
[0099] Tumor diseases (such as lung cancer, breast cancer, gastric cancer, etc.);
[0100] Cardiovascular and cerebrovascular diseases (such as hypertension, myocardial infarction, stroke, etc.);
[0101] Common internal medicine diseases (such as diabetes, chronic bronchitis, etc.);
[0102] Common surgical diseases (such as appendicitis, cholecystitis, etc.).
[0103] Alternatively, the system can further sub-classify the diagnosis based on the severity of the condition and treatment options. For example, for tumors, the system can sub-classify the disease based on different tumor types (such as primary tumors and metastatic tumors) to enable more accurate subsequent analysis.
[0104] After data classification is completed, the medical record homepage data will be mapped to a multi-dimensional tensor model. This tensor model contains at least the following four dimensions:
[0105] Age dimension: The age range of the patient. Categorize patients by age group (e.g., "0-18 years old," "19-40 years old," "41-60 years old," "over 60 years old").
[0106] Main diagnosis dimension: The main diagnostic information of the disease, expressed using ICD-10 codes or other standardized codes.
[0107] Surgical method dimension: the type of surgery the patient underwent, such as "laparoscopic surgery", "open surgery", "endoscopic surgery", etc.
[0108] Cost dimension: information related to treatment costs, including drug costs, surgical costs, hospitalization costs, etc.
[0109] Length of stay dimension: The number of days a patient is hospitalized, which serves as a key indicator of treatment and recovery time.
[0110] In some embodiments, the length of hospital stay may be further refined according to the type of disease. For example, the length of hospital stay for cancer patients may be significantly longer than that for patients with general internal medicine diseases.
[0111] These five dimensions together form a high-dimensional tensor Among them, I is the number of categories in the age dimension; J is the number of categories in the primary diagnosis dimension; K is the number of categories in the surgical method dimension; and L is the number of categories in other related fields such as cost and hospitalization days.
[0112] Specifically, each element T of the tensor model i,j,k,l Representing case data with specific age, diagnosis, surgical method, and cost. In this case, the tensor model can comprehensively represent the multidimensional information of each case, providing a basis for subsequent feature extraction and model training.
[0113] In tensor decomposition, choosing the appropriate rank is key. The dynamic selection of tensor rank is based on the complexity of different diseases to ensure that the model can effectively capture the characteristics of various diseases. The rank selection rules are as follows:
[0114] For tumor diseases, since tumor diseases usually have higher complexity in data and involve more features, the tensor rank is selected to be greater than or equal to 5.
[0115] For other common diseases, such as common internal medicine diseases and common surgical diseases, the disease features are relatively simple, and the tensor rank is set to 3, which helps reduce computational overhead while ensuring the efficiency of the model.
[0116] Specific implementation of rank selection:
[0117] Tumor diseases: The rank of the tensor is chosen to be R tumor ≥5, increasing the rank number allows tensor decomposition to better capture the high-order relationships and complexity between tumor cases.
[0118] Other diseases: For diseases such as internal medicine and surgery, the rank is set to R other =3 to simplify the model and reduce computational complexity.
[0119] By choosing an appropriate rank R, the tensor decomposition process can be expressed as:
[0120] in, is the tensor factor of the age dimension; is the tensor factor of the primary diagnosis dimension; is the tensor factor of the surgical approach dimension; is the tensor factor of the cost and hospitalization days dimensions; R is the rank of the tensor decomposition, which represents the number of decomposed components.
[0121] Specifically, the tensor factorization process reduces the dimensionality of each dimension and represents the high-dimensional tensor as the product of multiple low-rank tensor factors, thereby achieving the effect of data compression and feature extraction.
[0122] The optimization goal of tensor decomposition is to optimize the representation of tensor factors by minimizing the following loss function: in, is the predicted value obtained by tensor factorization; T i,j,k,l is the observed value in the original data;
[0123] The goal of the loss function is to minimize the error between the predicted value and the true data.
[0124] In some embodiments, in order to improve the decomposition accuracy, a regularization term (such as L2 regularization) may be added to prevent overfitting and enhance the generalization ability of the model.
[0125] Regularization loss term: Where ω is the regularization hyperparameter; ||U|| 2 ,||V|| 2 ,||W|| 2 ,||Z|| 2are the L2 norms of the tensor factors respectively.
[0126] The final loss function is: in, The total loss function, is the regularization loss function.
[0127] As an option, to improve the model's adaptability, the tensor rank can be dynamically adjusted based on the model's training results. Through methods such as cross-validation, the system can automatically select an appropriate rank based on the training error for each disease type, thereby improving the model's predictive ability.
[0128] In one possible implementation, the multi-dimensional tensor model constructed in step S2 can be combined with other types of deep learning models (e.g., neural networks and convolutional neural networks). By integrating multiple models, the system can fully leverage the strengths of different models to improve overall diagnostic accuracy and robustness.
[0129] The multidimensional tensor model T and tensor decomposition results from step S2 are further used to generate refined features in the subsequent step S3 (data preprocessing and feature extraction). These features are then fed into the reinforcement learning model in step S5 for conflict detection and reward function optimization. Furthermore, these features are passed as input to step S6 (coding recommendation), enabling more efficient disease coding.
[0130] S3. Applying a tensor decomposition method to the multi-dimensional tensor model constructed based on the medical record homepage data to decompose it into a core tensor and multiple factor matrices. The core tensor is used to represent the potential features of the medical record homepage data, and the factor matrix is used to describe the correlation between the dimensions in the medical record homepage data.
[0131] In step S2, the medical record homepage data is structured as a high-dimensional tensor T. This tensor, with dimensions such as age, primary diagnosis, surgery type, hospitalization expenses, and length of stay as distinct models, has strong multidimensional heterogeneous data representation capabilities. However, as the tensor dimension and data size increase, the original tensor not only contains a large amount of redundant information, but also significantly reduces subsequent processing efficiency.
[0132] To this end, in this step, the present invention proposes a tensor dimensionality reduction and feature extraction method based on Tucker decomposition, which uses core tensors and factor matrices to effectively compress the data structure while retaining the principal component information, and extracts potential interactive features for subsequent recommendation, prediction and other intelligent processing modules.
[0133] In this embodiment, the tensor constructed by the medical record homepage data is assumed to be: in, Represents the tensor of the original medical record homepage; N represents the order (number of dimensions) of the tensor, which is usually 5 in this example (corresponding to the five dimensions of age, diagnosis, surgical method, hospitalization cost, and hospitalization days); I n Indicates the dimension length of the tensor in the nth mode, such as I1 is the number of age groups, I2 is the number of diagnosis code categories, etc.
[0134] In order to achieve structure compression and latent feature acquisition, Applying Tucker decomposition, the expression is as follows:
[0135] in, Represents a low-dimensional core tensor; is the factor matrix of the nth mode, describing the linear transformation of the original nth mode space in the low-dimensional space; R n is the rank value on the nth mode, satisfying R n <<I n , used to control compression ratio and expression ability;× n Indicates the product operation on the nth mode (mode-nproduct).
[0136] In one possible implementation, the specific calculation process of Tucker decomposition is as follows:
[0137] Step 1: Initialize the rank parameter
[0138] Select rank R according to the complexity of the disease n Taking tumor cases as an example, we select a higher rank value R n ∈[6,10]; for surgical diseases, R n =3.
[0139] Step 2: Preprocess tensor data
[0140] Before decomposition, the tensor is normalized and missing values are filled. Mean imputation or K-nearest neighbor interpolation is used to fill missing units to ensure the convergence and stability of the decomposition.
[0141] Step 3: Use the ALS algorithm to solve the factor matrix
[0142] The alternating least squares (ALS) method is used for parameter estimation. The algorithm flow is as follows:
[0143] Randomly initialize all A (n) ;
[0144] For each iteration:
[0145] Fix all factor matrices except the nth dimension;
[0146] The solution minimizes the following objective function:
[0147] where ||·|| F represents the Frobenius norm, which is the sum of squared errors of the tensor.
[0148] Update core tensors Minimize the decomposition and reconstruction error;
[0149] Convergence criterion (such as error threshold ∈<10 -4 ), if satisfied, the iteration is terminated.
[0150] In some embodiments, the tensor structure can be expanded to a higher order, for example, by introducing "hospitalization department" or "medical insurance type" as additional dimensions to form a sixth-order or seventh-order tensor. In this case, Tucker decomposition is still applicable, and only the number of dimensions N and the rank parameter vector {R1,…,R N}.
[0151] Alternatively, to control the risk of overfitting, you can add sparsity constraints or regularization terms to the core tensors: Among them, δ is the regularization factor; Represents the L1 norm of the core tensor, used for sparse representation;
[0152] This optimization problem can be solved using the Hoop-Orthogonal Iteration (HOOI) method, which takes into account both compressibility and interpretability.
[0153] The core tensor obtained With each factor matrix A (n) This will serve as the feature combination input in step S4 to construct a cross-dimensional semantic representation. This structure can also be directly used in reward function modeling in step S5 or the encoding recommendation module in S6, providing the system with a multimodal input source and enabling dynamic understanding and context-aware processing of medical record features.
[0154] S4. Based on the core tensor and factor matrix obtained in step S3, the reconstruction error of the medical record homepage data in the tensor space is calculated, and the error is compared with a preset threshold. If the reconstruction error is greater than the threshold, the medical record homepage data is determined to be abnormal data, triggering subsequent abnormal data processing and verification optimization;
[0155] In the above step S3, the multidimensional data tensor of the medical record homepage has been decomposed by Tucker Represented as a core tensor With a set of factor matrices A (n) This structure realizes the compressed expression of data in tensor space and effectively extracts the latent semantic structure of medical record data.
[0156] Generally speaking, if the original tensor data can be well restored through the above-mentioned low-dimensional structure, it means that the sample belongs to the "normal" data within the training distribution; on the contrary, if the reconstruction error is large, it often means that the sample has significant deviations from historical laws at the structural level and needs to be included in the scope of anomaly detection.
[0157] Therefore, on this basis, the present invention proposes an anomaly detection mechanism based on tensor reconstruction error, combined with a dynamic threshold method of kernel density estimation, to achieve interpretable and highly robust automatic anomaly recognition in multidimensional medical record data scenarios.
[0158] Reconstruction error calculation:
[0159] For each medical record homepage data tensor sample Construct its estimated reconstruction value in the tensor decomposition structure The form is: in, Represents the tensor reconstruction value of the i-th case sample; is the core tensor; is the factor matrix on the nth mode; × n Represents a mode-n multiplication operation.
[0160] Reconstruction error ∈ (i) is defined as follows:
[0161] in, Represents the tensor sample of the i-th original medical record homepage; ||·|| F is the Frobenius norm, defined as the sum of the squared differences of all elements of the tensor:
[0162] in, is an N-order tensor, that is, a multidimensional array with N dimensions; is the square of the tensor Frobenius norm; I n is the size of the tensor in the nth dimension; j n The index variable on the nth dimension is used to enumerate the element positions in the dimension (from 1 to I n ).
[0163] In some embodiments, directly setting a fixed threshold may lead to misjudgment or missed judgment. Therefore, the present invention adopts a non-parametric density estimation method based on the historical sample error distribution and adaptively estimates the critical point through the kernel density function.
[0164] Construct the reconstruction error set: ∈ = {∈ (1) ,∈ (2) ,…,∈ (M)}, where M represents the number of historical samples.
[0165] The density function is estimated as follows:
[0166] Among them, K(·) is the kernel function, and the Gaussian kernel is commonly used. h is the kernel bandwidth parameter, which controls the smoothness of the estimated curve; f(∈) is the estimated probability density function of the reconstruction error.
[0167] Construct a dynamic decision threshold with significance level γ: Where,∈ is the reconstruction error; is the cumulative density;
[0168] Specifically, when the sample error is greater than the threshold: Mark as abnormal;
[0169] In this embodiment, the significance level is set to γ=0.05, that is, samples with errors in the top 95% are considered "normal", and samples with errors above this level are considered to have structural deviations.
[0170] In actual deployment, after detecting an abnormal sample, the system of the present invention will automatically call the abnormal data processing module, the functions of which include but are not limited to:
[0171] Trace back the original medical record input information;
[0172] Check the consistency between the main diagnosis and surgical method coding fields and the disease category;
[0173] Set automatic validation rules for key fields, such as "number of hospitalization days greater than 180 days" is an abnormal flag;
[0174] Initiate a review request to the manual review module.
[0175] In one possible implementation, the system sets a tag field for abnormal samples, and uses it as a filtering condition in subsequent coding recommendation strategies to prevent inconsistent data from polluting the model.
[0176] As an option, to improve the sensitivity of the system, a weighted reconstruction error mechanism can be introduced to give higher weights to certain key dimensions (such as diagnosis). Its expression is:
[0177] Among them, α n is the weight coefficient of the nth dimension, satisfying The vector expansion of the i-th sample in the n-th mode.
[0178] This mechanism is particularly suitable for practical scenarios where there is a higher semantic demand for specific dimensions.
[0179] S5. Using reinforcement learning to optimize the verification rules for abnormal data, dynamically adjust the weight distribution of each data feature in the verification rules according to the characteristics of the abnormal data in step S4, thereby achieving dynamic optimization of the data verification strategy for the medical record homepage;
[0180] In this embodiment, step S5 optimizes the validation rules for the medical record front page data using a reinforcement learning approach. This optimization process dynamically adjusts the weights assigned to each data feature in the validation rules based on the characteristics of the abnormal data in step S4, thereby optimizing the validation strategy for the medical record front page data and achieving efficient and accurate abnormal data detection and processing. Specifically, through the continuous learning and feedback mechanism of reinforcement learning, the system adaptively adjusts the weights of the validation rules, enabling the data validation strategy to be optimized based on actual conditions and improving the overall system performance.
[0181] In this embodiment, the design of the reinforcement learning reward function is the core part of system optimization. The reward function aims to balance the two objectives of conflict resolution rate and operation efficiency. In order to design a suitable reward function, considering that the system's goal is to optimize both the conflict resolution effect and the efficiency of the operation steps, the dual-objective reward function is defined as:
[0182] Among them, R(s t ,a t ) means in state s t Next, perform action a t The immediate reward obtained after acdlovd Indicates the number of conflicts that have been resolved in the current state, that is, the number of data conflicts that have been successfully corrected. The larger the value, the more effective the system's verification rules are; n conflict_total Indicates the total number of conflicts detected in the current state, that is, the number of data conflicts to be resolved; Indicates the conflict resolution rate, which is the ratio of the number of resolved conflicts to the total number of conflicts. The higher this value is, the better the effect of the verification rule is and the higher the efficiency of conflict resolution is. steps represents the number of steps required for the doctor to correct the conflict according to the system's suggestions; n max Indicates the maximum number of operation steps allowed by the system. This value sets the upper limit of the operation complexity and limits the maximum number of steps in each verification process; represents an operational efficiency indicator, indicating the ratio of the number of steps required to correct a conflict to the maximum number of steps allowed. Fewer steps indicate higher operational efficiency. α is the weighting coefficient for the conflict resolution rate, used to balance the impact of the conflict resolution rate against operational efficiency. β is the weighting coefficient for operational efficiency, used to balance the impact of operational efficiency against the conflict resolution rate.
[0183] In some embodiments, the values of α and β can be dynamically adjusted through a Nash equilibrium game model to ensure that the system can achieve an optimal balance between conflict resolution rate and operational efficiency during optimization.
[0184] In this embodiment, the problem of assigning validation rule weights is modeled as a Markov decision process (MDP), where the state space consists of eigenvectors related to the current number of conflicts and data characteristics. Each state reflects the system's current validation process and the anomaly characteristics of the current data, while the action space represents a set of weight adjustment strategies. The system needs to select the optimal weight adjustment strategy based on the current state.
[0185] In the reinforcement learning framework, the system uses the Q-learning algorithm to solve the optimal strategy. The Q-learning algorithm is a classic algorithm based on reinforcement learning. It optimizes the weight adjustment strategy by continuously learning the value function of each state-action pair. Its update formula is as follows:
[0186] Among them, Q(s t ,a t ) means in state s t Next, perform action a t The Q value represents the expected reward of the state-action pair; r t In state s t Next, perform action a t The immediate reward obtained after the reward is calculated using the previously defined reward function; μ is the discount factor, which indicates the importance of future rewards. It usually takes a value in the range of [0,1]. The larger the value, the more importance is attached to future rewards. Indicates the next state s t+1 The maximum Q value of all possible actions a′ represents the reward of the optimal strategy in the future state; ν is the learning rate, which represents the weight of the newly acquired information relative to the old information. A larger learning rate means that the model adapts to new environmental feedback faster.
[0187] Through the Q-learning algorithm, the system can continuously optimize the weight adjustment strategy so that the feature weights selected during each verification process are more in line with the actual needs of the data, thereby improving the verification effect of abnormal data.
[0188] In practical applications of reinforcement learning models, the system continuously adjusts and optimizes verification rules based on different types of abnormal data through reinforcement learning. For example, some abnormal data may be primarily due to errors or missing data in specific features. The system can dynamically adjust the weights of these features through reinforcement learning to maximize the impact of these data features.
[0189] As an option, the system can also introduce an exploration mechanism, which allows the system to randomly select strategies within a certain range to explore possible optimization solutions. The purpose of this is to prevent the system from falling into a local optimal solution and to encourage the system to explore more efficient verification strategies.
[0190] In this embodiment, the processing of abnormal data is closely linked to the verification optimization process. First, in step S4, the system identifies abnormal data and triggers subsequent abnormal data processing by judging the threshold based on the tensor reconstruction error. Once the system identifies abnormal data, the reinforcement learning algorithm in step S5 optimizes the weight distribution of the verification rules based on the characteristics of these abnormal data, thereby further improving the accuracy of data verification.
[0191] Specifically, in step S4, if the reconstruction error of the medical record homepage data exceeds a preset threshold, the data is identified as abnormal, triggering subsequent optimization of the verification rules. Based on this, step S5 continuously optimizes the data verification rules through reinforcement learning to cope with different types of abnormal data. Through repeated learning, the system can gradually adjust the weight of the verification rules based on the characteristics of the abnormal data, thereby improving data verification performance.
[0192] To further enhance the robustness and performance of the system, the system can introduce more data features into the reinforcement learning model, such as historical abnormal data trends and doctors' operating habits. This additional information helps the system consider more context when adjusting verification rules, thereby optimizing verification results.
[0193] Furthermore, the system can incorporate an adaptive exploration mechanism, automatically adjusting its exploration strategy based on its performance during training. For example, the system might initially rely more on random selection strategies, but as learning progresses, it will gradually shift towards selecting the optimal strategy and reduce ineffective exploration.
[0194] S6. Based on the verification rules optimized in step S5, a knowledge graph is constructed. The knowledge graph includes diseases, surgeries, complications, and their interrelationships. Based on the knowledge graph, coding reasoning is performed on the medical record homepage data. The reasoning model generates coding suggestions related to the disease type based on the analysis of the medical record homepage data, and applies them to the coding optimization of the medical record homepage data.
[0195] In this embodiment, the goal of step S6 is to construct a comprehensive knowledge graph based on the validation rules optimized in step S5. This graph covers diseases, surgeries, complications, and their interrelationships, and uses this knowledge graph to perform coding reasoning on the medical record front page data. The system generates coding suggestions related to the disease type through the reasoning model and optimizes the coding of the medical record front page data based on these suggestions. To ensure the accuracy of the reasoning and the operability of the results, this embodiment also pays special attention to the interpretability of the reasoning process, providing a detailed solution for anomaly location and correction path generation.
[0196] In this embodiment, the knowledge graph construction method includes two core links: the graph neural network (GNN) calculates the embedding vector of the node and the graph attention mechanism empowers the associated edges.
[0197] Specifically, the system uses graph neural networks to learn node embeddings of diseases, surgeries, and complications, thereby capturing the intrinsic connections between nodes, and assigns different weights to the relationships between nodes through the graph attention mechanism to improve the expressive power of the graph.
[0198] When constructing a knowledge graph, the system first maps elements such as diseases, surgeries, and complications into graph nodes and constructs edges between nodes to represent the relationships between these elements (such as the relationship between disease and surgery). The graph neural network updates the node embedding vector layer by layer by transferring local information between nodes and adjacent nodes. Each node's embedding vector captures its characteristics and its relationships with other nodes.
[0199] The update formula of the graph neural network is as follows:
[0200] in, Represents the embedding vector of node i at layer l; A represents the set of neighbor nodes of node i; ij is the adjacency matrix element between node i and node j, indicating the weight of the edge between nodes; W (l) is the weight matrix of the lth layer; σ(·) is the activation function, usually ReLU or Sigmoid, which is used to introduce nonlinear transformation.
[0201] Through multi-layer updates of the graph neural network, the embedding vector of each node is finally obtained These embedding vectors can capture the characteristics of nodes and the relationships between nodes.
[0202] To further enhance the expressive power of the knowledge graph, the system uses a graph attention mechanism (GAT) to weight the edges between nodes. The GAT adaptively assigns weights to edges, adjusting the intensity of information transfer based on the importance of the relationships between nodes. The system assigns weights to edges based on the relationships between nodes, enhancing the flow of upstream and downstream information related to specific nodes.
[0203] After completing the construction of the knowledge graph, the system can perform coding reasoning on the medical record homepage data based on the graph and generate relevant coding suggestions.
[0204] Specifically, the system combines the nodes and their relationships in the knowledge graph and generates coding suggestions related to the disease based on the data on the medical record homepage. These coding suggestions are based on the output of the inference model and applied to the optimization of the medical record homepage data.
[0205] To enhance the interpretability of inference results, the system provides a detailed feedback mechanism. Specifically, while generating coding suggestions, the system locates abnormal dimensions in the medical record homepage data based on tensor reconstruction errors, generates correction paths through the knowledge graph, and provides feedback. For example, when the system detects a conflict between the "primary diagnosis" field and the "type of surgery" field on the medical record homepage, the system generates the following feedback: "A conflict between [primary diagnosis] and [type of surgery] has been detected. It is recommended to verify the [primary diagnosis code]." This form of feedback helps users detect and correct data anomalies in a timely manner.
[0206] Through the tensor reconstruction error analysis in step S4, the system can accurately locate abnormal dimensions in the medical record homepage data. Assuming that the reconstruction error of a certain dimension (such as "disease code") is significantly larger, the system will determine that the dimension is an abnormal dimension through the tensor reconstruction error algorithm. At this time, based on the knowledge graph, the system will automatically generate a correction path and provide clear feedback to the user, such as: "A conflict between the [Disease Code] and [Surgery Type] fields has been detected. It is recommended to verify the [Surgery Code]."
[0207] To better handle these conflicts, the system leverages the relationships between disease types and surgeries in the knowledge graph to provide correction suggestions for conflicting fields during the reasoning process. This automatically generated correction path not only improves data verification efficiency but also provides users with direct and actionable suggestions.
[0208] To further enhance the performance of the knowledge graph, the system can incorporate more medical domain knowledge and data features. For example, in addition to diseases, surgeries, and complications, basic patient information (such as age, gender, and medical history) can be incorporated to further enhance the accuracy of the inference model. The system can also customize different inference paths and encoding rules based on the actual needs of different hospitals, thereby achieving targeted optimization.
[0209] In some embodiments, the system can also incorporate reinforcement learning methods to continuously optimize the knowledge graph. For example, by optimizing the model's reasoning path through reinforcement learning, the system can adjust coding suggestions based on historical feedback when faced with different types of medical record data, thereby improving coding efficiency and accuracy.
[0210] S7. Based on the coding suggestions generated in step S6, the knowledge graph is updated in combination with the real-time changes of the medical record homepage data, and the medical entities and their relationships in the knowledge graph are incrementally updated to ensure that the knowledge graph can be consistent with the latest medical record homepage data.
[0211] In this embodiment, the core goal of step S7 is to update and optimize the constructed knowledge graph based on the coding suggestions generated in step S6 and the real-time changes in the medical record homepage data. By incrementally updating the medical entities (such as diseases, surgeries, etc.) and their relationships in the knowledge graph, the knowledge graph is ensured to always be consistent with the latest medical record homepage data, thereby improving the accuracy of the inference model and data consistency. This incremental update method can not only respond to changes in medical record data, but also improve the adaptability of the knowledge graph to new data patterns.
[0212] In this embodiment, the knowledge graph is updated incrementally. This incremental update gradually adjusts edge weights in the graph based on newly added medical record data, ensuring that the graph can adapt to new medical record information in real time without being completely reconstructed. Incremental updates focus on the co-occurrence relationship between diseases and surgeries, and adjust edge weights in the graph to reflect new data features.
[0213] Edges in a knowledge graph represent the relationships between nodes. For example, in the relationship between disease and surgery, the weight of the edge can be dynamically adjusted based on the co-occurrence frequency of the disease and surgery.
[0214] Specifically, when the system receives new medical record data, it calculates the co-occurrence frequency between diseases and surgeries in the medical record data and updates the edge weights in the knowledge graph based on this information.
[0215] The updated edge weight calculation formula is as follows:
[0216] in, is the edge weight between node i and node j after update; is the edge weight between node i and node j before the update; λ is the attenuation factor, which is used to control the influence of new data on the edge weight update, and its value range is usually [0,1]; n co-occur (i, j) is the number of times node i and node j co-occur in the newly added data; n new_cases is the total number of new cases, indicating the number of medical record data.
[0217] Using this formula, the system dynamically adjusts edge weights based on the co-occurrence frequency of diseases and surgeries, ensuring that the knowledge graph can promptly reflect new patterns and relationships in medical record data. Specifically, if a disease and surgery co-occur frequently, the strength of the association (edge weight) between the two will increase accordingly; conversely, if the co-occurrence frequency is low, the strength of the association will decrease.
[0218] When new disease or surgery nodes appear in the newly added data, the system dynamically expands the knowledge graph using a graph embedding algorithm. This process involves calculating embedding vectors for the new nodes and connecting these new nodes to existing nodes in the graph. Specifically, when the system detects a new disease or surgery node, it uses the graph embedding algorithm to generate a high-dimensional vector for the node, representing its characteristics in the graph.
[0219] Common calculation methods for graph embedding algorithms include DeepWalk or Node2Vec, which can convert the structural information of the graph into vector representations of nodes.
[0220] Specifically, suppose a node v in the graph corresponds to a high-dimensional vector h v , then the graph embedding process minimizes the distance between nodes by optimizing the objective function, so that nodes with similar structural relationships are as close as possible in the embedding space.
[0221] In this way, the system can generate appropriate vector representations for newly added disease or surgery nodes and embed them into the existing knowledge graph. In this way, the new nodes can establish connections with related nodes in the graph and participate in subsequent reasoning and encoding processes.
[0222] The incremental update process of this embodiment specifically includes the following steps:
[0223] Receive new medical record data: The system extracts new entities such as diseases, surgeries, and their related relationships from the medical record homepage data.
[0224] Calculate the co-occurrence frequency of diseases and surgeries: The system calculates the co-occurrence frequency n between each pair of diseases and surgeries by counting the co-occurrence of diseases and surgeries in the newly added data. co-occur (i,j).
[0225] Update edge weights: Based on the co-occurrence frequency, the system uses the above update formula to adjust the edge weights between diseases and surgeries in the knowledge graph.
[0226] Processing new nodes: If the new data contains diseases or surgeries that do not appear in the knowledge graph, the system will use a graph embedding algorithm to generate an embedding vector for the new node and add it to the graph to ensure the integrity of the knowledge graph.
[0227] Output the updated knowledge graph: The system applies the updated knowledge graph to subsequent coding reasoning and anomaly detection to ensure that the reasoning process can reflect the latest medical record data.
[0228] As medical record data is continuously updated and the knowledge graph is incrementally updated, the system can dynamically optimize the encoding of the medical record homepage data based on the latest encoding suggestions and inference results. This optimization not only improves encoding accuracy but also enables the system to adapt to the ever-changing medical record data patterns and medical knowledge.
[0229] For example, if the co-occurrence frequency of a disease and surgery increases significantly in new medical records, the system will adjust the edge weights in the knowledge graph through incremental updates, which in turn affects subsequent coding inference results. In this way, the system can promptly capture new trends and patterns in the medical field, optimize coding strategies, and provide more accurate and personalized coding suggestions.
[0230] In some embodiments, to further improve the efficiency and accuracy of knowledge graph updates, the system can incorporate an adaptive decay factor, dynamically adjusting the decay factor based on different time periods and data sizes. Furthermore, the system can incorporate incremental learning methods, allowing the knowledge graph to be gradually optimized with each update without requiring retraining the entire model, thereby improving the efficiency of the update process.
[0231] The system can also take into account the user feedback mechanism. When medical data experts provide feedback on coding suggestions, the system automatically adjusts the weights of nodes and edges in the graph to further improve the quality and reasoning effect of the knowledge graph.
[0232] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
Claims
1. A multi-dimensional verification method for real-time quality control and coding of medical record homepage data, characterized in that: The following steps are involved: S1. Obtain medical record front page data from a hospital information system or an electronic medical record system, wherein the medical record front page data includes structured data, unstructured data, and time series data; S2. Classify the medical record homepage data according to the disease classification standard, divide the medical record homepage data into different disease categories, and construct a multi-dimensional tensor model related to each disease for subsequent data processing and feature extraction; S3. Applying a tensor decomposition method to the multi-dimensional tensor model constructed based on the medical record homepage data to decompose it into a core tensor and multiple factor matrices. The core tensor is used to represent the potential features of the medical record homepage data, and the factor matrix is used to describe the correlation between the dimensions in the medical record homepage data. S4. Based on the core tensor and factor matrix obtained in step S3, the reconstruction error of the medical record homepage data in the tensor space is calculated, and the error is compared with a preset threshold. If the reconstruction error is greater than the threshold, the medical record homepage data is determined to be abnormal data, triggering subsequent abnormal data processing and verification optimization; S5. Using reinforcement learning to optimize the verification rules for abnormal data, dynamically adjust the weight distribution of each data feature in the verification rules according to the characteristics of the abnormal data in step S4, thereby achieving dynamic optimization of the data verification strategy for the medical record homepage; S6. Based on the verification rules optimized in step S5, a knowledge graph is constructed. The knowledge graph includes diseases, surgeries, complications, and their interrelationships. Based on the knowledge graph, coding reasoning is performed on the medical record homepage data. The reasoning model generates coding suggestions related to the disease type based on the analysis of the medical record homepage data, and applies them to the coding optimization of the medical record homepage data. S7. Based on the coding suggestions generated in step S6, the knowledge graph is updated in combination with the real-time changes of the medical record homepage data, and the medical entities and their relationships in the knowledge graph are incrementally updated to ensure that the knowledge graph can be consistent with the latest medical record homepage data.
2. A multi-dimensional verification method for real-time quality control and coding of medical record front page data according to claim 1, characterized in that: The dimensions of the multidimensional tensor model in step S2 include at least four dimensions of age, primary diagnosis, surgical method, cost, and hospitalization days, and the rank of the tensor decomposition is dynamically selected for different disease categories. The specific rules are: If the disease type is tumor, the rank is ≥5; if it is other diseases, the rank is =3.
3. A multi-dimensional verification method for real-time quality control and coding of medical record front page data according to claim 1, characterized in that: The tensor decomposition method in step S3 is Tucker decomposition, and its expression is: Among them, the core tensor The dimension is consistent with the dynamically selected rank value, and the factor matrix A (i) Solve it iteratively using the alternating least squares (ALS) method. is the original multi-dimensional tensor, × N The product of tensors over their Nth mode.
4. A multi-dimensional verification method for real-time quality control and coding of medical record front page data according to claim 1, characterized in that: The threshold setting method in step S4 is: based on the distribution of the historical medical record data reconstruction error, the dynamic threshold θ is calculated using kernel density estimation. th , so that the abnormal data determination meets the significance level γ<0.
05.
5. The method for real-time quality control and coding of medical record front page data with multi-dimensional verification according to claim 1, characterized in that: The reinforcement learning method in step S5 includes: Design a dual-objective reward function: Among them, R(s t ,a t ) is in state s t Next, perform action a t Instant rewards received; acdlovd The number of resolved conflicts in the current state; n conflict_total is the total number of conflicts detected in the current state; is the conflict resolution rate; n steps The number of steps required for the doctor to correct the conflict according to the system's suggestions; n max The maximum number of operation steps allowed by the system; is the operational efficiency indicator; α and β are weight coefficients, which are dynamically adjusted through the Nash equilibrium game model to achieve a balance between the conflict resolution rate and operational efficiency.
6. A multi-dimensional verification method for real-time quality control and coding of medical record front page data according to claim 1, characterized in that: The method for constructing the knowledge graph in step S6 includes: Compute node embedding vectors for diseases, surgeries, and complications through graph neural networks; Assign weights to related edges based on the graph attention mechanism.
7. The method for real-time quality control and coding of medical record front page data with multi-dimensional verification according to claim 1, characterized in that: The knowledge graph update in step S7 is an incremental update, specifically including: According to the co-occurrence frequency of diseases and surgeries in the new medical record data, adjust the weight w of the edge in the knowledge graph ij , and its update formula is: in, is the edge weight between node i and node j after update; is the edge weight between node i and node j before updating; λ is the decay factor; n co-occur (i, j) is the number of times node i and node j co-occur in the newly added data; n new_cases is the total number of cases with newly added data; is the co-occurrence frequency of node i and node j in the newly added data; When new disease or surgery nodes are added, the knowledge graph is dynamically expanded through the graph embedding algorithm.
8. The method for real-time quality control and coding of medical record front page data with multi-dimensional verification according to claim 1, characterized in that: The specific method of dynamically adjusting the weight of the verification rule in step S5 is: The weight assignment of verification rules is modeled as a Markov decision process, where the state space includes the current number of conflicts and the eigenvector, and the action space is the weight adjustment strategy; The optimal strategy is solved through the Q-learning algorithm.
9. The method for real-time quality control and coding of medical record front page data with multi-dimensional verification according to claim 1, characterized in that: The coding suggestions generated in step S6 for the disease type must be accompanied by interpretability feedback, specifically including: Locating abnormal dimensions based on tensor reconstruction error; The correction path is generated through the knowledge graph, and the feedback is in the form of: "A conflict between [Dimension X] and [Dimension Y] is detected, and it is recommended to check [Field Z]."
Citation Information
Cited By
Medical record home page quality management and control system and method based on big data analysis
CN121565358A
Case history first page quality management and control system and method based on big data analysis
CN121565358B