Case history first page quality management and control system and method based on big data analysis
By combining big data analysis with information on source nodes and coders, the quality risks of medical record front pages are quantified, solving the problem of data interference and the lack of correlation between coding personnel workload in existing technologies. This enables precise control of medical record front page data and differentiated quality control strategies, thereby improving data quality and coding efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- THE FIRST AFFILIATED HOSPITAL OF ARMY MEDICAL UNIV
- Filing Date
- 2026-01-21
- Publication Date
- 2026-04-21
AI Technical Summary
Existing technologies fail to effectively link data interference from clinical sources and the workload of coding personnel in the quality analysis of medical record front pages, resulting in misjudgments in quality risk assessment and a lack of precision in quality control intervention strategies, making it difficult to fully cover complex risks.
By using big data analysis, combined with source node data and coding personnel information, we can quantify the quality, complexity, workload, and coding imbalance of data input, assess the quality risks of the medical record front page, and implement differentiated quality control interventions.
It enables precise quantitative control over medical record front page data, ensuring data integrity and standardization from the source, improving coding efficiency, reducing the risk of medical insurance refusal, and protecting the value of hospital data assets.
Smart Images

Figure CN121565358B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical record management technology, and in particular to a medical record front page quality control system and method based on big data analysis. Background Technology
[0002] As a crucial data carrier for recording core diagnostic and treatment information, supporting medical payment reform, and driving refined hospital management, the quality of the medical record cover sheet directly affects the operational efficiency of medical institutions, the rational use of medical insurance funds, and the objective evaluation of medical quality. During the generation and coding process, medical record cover sheet data is not only continuously affected by non-standard data entry at the clinical front end, leading to missing data or non-standard terminology at the source, thus creating potential quality risks; moreover, the high workload and complex cases faced in the coding stage exacerbate the cognitive pressure on coders, causing data quality problems to progress from initial errors in individual entries to inaccurate coding of specific diagnostic categories, and finally to a decline in overall data credibility. Failure to identify and intervene in these potential quality risks in a timely manner may lead to serious consequences such as medical insurance payment disputes, distorted hospital performance evaluations, and even biases in clinical decision support, causing not only economic losses to hospitals but also damaging their credibility in medical quality management. Traditional medical record cover sheet quality control relies heavily on manual review of samples after coding, which is not only inefficient but also fails to comprehensively cover and provide early warnings of complex risks hidden in massive amounts of data caused by multiple intertwined factors. Using big data analytics for quality control allows for in-depth analysis of the entire process from clinical data entry to coding review, leveraging its multi-source data fusion and intelligent algorithm modeling capabilities. This overcomes the limitations of traditional post-event sampling and significantly improves the foresight and efficiency of quality control.
[0003] However, current technologies for quality analysis of medical record front pages do not correlate data interference from clinical source nodes with the workload and skill limitations of the coders themselves. For example, when faced with a medical record with complex diagnostic descriptions and incomplete source data, it is difficult to accurately assess the risk of final coding failure of the front page if the coder is not handling a high-load, high-switch-frequency task, leading to misjudgments of the quality risk level. Furthermore, current technologies lack a comprehensive consideration of the dynamic correlation between data quality risk and personnel coding behavior, making it difficult to accurately pinpoint the root causes of quality problems in medical record front pages and failing to provide a comprehensive basis for developing differentiated and precise quality control intervention strategies.
[0004] To address these issues, this application presents a medical record front page quality control system and method based on big data analysis. Summary of the Invention
[0005] To overcome the defects and shortcomings of existing technologies, this invention provides a medical record front page quality control system and method based on big data analysis. By comprehensively analyzing the data coding interference at source nodes and the coding imbalance of coding personnel, the system assesses the quality risks of the medical record front page and implements differentiated quality control intervention plans based on the risk assessment results, thereby achieving precise control over the entire process from data source to personnel behavior.
[0006] To achieve the above objectives, the present invention adopts the following technical solution:
[0007] In a first aspect, embodiments of the present invention provide a method for quality control of medical record front pages based on big data analysis, comprising the following steps:
[0008] S1. Obtain the coding task list data and historical coded medical record data of the coding personnel of the medical record homepage data, and at the same time obtain the clinical work data and diagnostic record data of the source node of the medical record homepage data;
[0009] S2. Analyze the input quality and complexity of the medical record front page data by combining the clinical work data and diagnostic record data from the source nodes; and analyze the data encoding interference from the source nodes based on the analysis results of the input quality and complexity of the medical record front page data.
[0010] S3. Combining the coding task list data of the coding personnel on the front page of the medical record data with historical coding medical record data, analyze the coding task load status and coding level abnormalities of the coding personnel, and quantitatively analyze the coding imbalance of the coding personnel on the front page of the medical record data based on the coding task load status and coding level abnormalities.
[0011] S4. Based on the analysis results of source node data coding interference and the analysis results of coding imbalance of medical record front page data coding personnel, assess the quality risk of medical record front page;
[0012] S5. Based on the results of the quality risk assessment of the medical record front page, implement the quality control intervention plan for the medical record front page.
[0013] In one implementation of the present invention, step S2 combines clinical work data and diagnostic record data from the source node to analyze the input quality and complexity of the medical record front page data; and based on the analysis results of the input quality and complexity of the medical record front page data, analyzes the data encoding interference of the source node, including the following specific steps:
[0014] S21. Extract clinical work data and diagnostic record data from the source node;
[0015] S22. Import diagnostic record data into the data input quality risk analysis strategy for data input quality risk analysis;
[0016] S23. Import clinical work data and diagnostic record data into a data complexity analysis strategy for data complexity analysis;
[0017] S24. The data input quality risk analysis results and data complexity analysis results are weighted and summed to obtain the source node data encoding interference situation.
[0018] As one implementation of the present invention, the data input quality risk analysis strategy includes the following specific steps:
[0019] Extract the text content of all medical documents from the diagnostic record data; and obtain a predefined list of key medical elements;
[0020] A hybrid neural network model based on BERT-BiLSTM-CRF was used to perform medical entity recognition on the text content of all medical documents;
[0021] The model identifies medical entities and matches them with a predefined list of key medical elements. Based on the matching results, a data incompleteness index is calculated.
[0022] Extract the free text of all diagnostic descriptions and surgical procedure names from diagnostic record data; and construct a standard medical terminology knowledge graph containing hierarchical and synonym relationships between medical terms.
[0023] Based on the free text of all diagnostic descriptions and surgical procedure names in the diagnostic record data and the standard medical terminology knowledge graph, the ratio of non-standard terms was analyzed.
[0024] The data input quality risk analysis results are obtained by weighted summation of the data incompleteness index and the ratio of non-standard terms.
[0025] As one implementation of the present invention, the data complexity analysis strategy includes the following specific steps:
[0026] S231. Extract all diagnostic codes and surgical operation codes from the diagnostic record data; and extract historical diagnostic codes and historical surgical operation codes from the hospital's historical medical record database, and perform deduplication on the above diagnostic codes and surgical operation codes.
[0027] S232. Combining the diagnostic codes and surgical operation codes after deduplication, analyze the complexity of medical concepts in the medical records.
[0028] S233. Extract all diagnostic codes from the diagnostic record data; based on the hospital's historical medical record database, extract the frequency distribution of historical diagnostic codes.
[0029] S234. Calculate the rare diagnosis weight of the medical record based on all diagnostic codes in the diagnostic record data and the frequency distribution of historical diagnostic codes.
[0030] S235. Extract departmental consultation records from clinical work data and count the number of all clinical departments involved in diagnosis and treatment; simultaneously extract special diagnosis and treatment event markers from diagnostic record data; construct a diagnosis and treatment process relationship graph for medical records based on the above data, defining each clinical department and special diagnosis and treatment event as nodes in the graph, and defining the collaboration relationship between departments and the association of special diagnosis and treatment events as edges in the graph; count the number of nodes and edges in the diagnosis and treatment process relationship graph for medical records.
[0031] S236. From the hospital's historical medical record database, extract the maximum number of nodes and the maximum number of edges in the diagnosis and treatment process graph of all historical medical records; use the ratio of the number of nodes in the diagnosis and treatment process graph of the medical record to the maximum number of nodes as the node complexity, and the ratio of the number of edges in the diagnosis and treatment process graph of the medical record to the maximum number of edges as the edge complexity; perform a weighted sum of the node complexity and the edge complexity to obtain the process complexity of the medical record.
[0032] S237. The complexity of medical concepts, the weight of rare diagnoses, and the complexity of procedures in the medical records are weighted and summed to obtain the data complexity analysis results.
[0033] In one implementation of the present invention, step S3 combines the coding task list data of the coders on the front page of the medical record data with historical coded medical record data to analyze the coding task load status and coding level anomalies of the coders, and quantitatively analyzes the coding imbalance of the coders on the front page of the medical record data based on the coding task load status and coding level anomalies, including the following specific contents:
[0034] S31. Extract coding task list data and historical coding medical record data of coding personnel;
[0035] S32. Import the coding task list data into the coding task load status analysis strategy for coding task load status analysis.
[0036] S33. Import historical coded medical record data into the coding level anomaly analysis strategy to conduct coding level anomaly analysis;
[0037] S34. The results of the coding task load status analysis and the results of the coding level anomaly analysis are weighted and summed to obtain the coding imbalance of the coders of the medical record front page data.
[0038] As one implementation of the present invention, the coding task load state analysis strategy includes the following specific steps:
[0039] S321. Extract the main diagnostic categories of all medical records to be coded from the coding task list data; calculate the number of times each main diagnostic category appears; divide the number of times each main diagnostic category appears by the total number of coding tasks to obtain the frequency of occurrence of each main diagnostic category.
[0040] S322. Squaring the frequency of occurrence of each major diagnostic category, adding all the squaring results, and subtracting the addition result from the value 1 to obtain the frequency of cognitive switching of the coder.
[0041] S323. Analyze the task time dispersion of coders based on the coding task list data;
[0042] S324. The coding task load status is obtained by weighting and summing the frequency of cognitive switching of coders and the dispersion of task working hours.
[0043] As one implementation of the present invention, the coding level anomaly analysis strategy includes the following specific steps:
[0044] S331. Extract the coding quality review results of all completed coding medical records from the historical coding medical record data; classify and statistically analyze the coding error rate under each diagnostic category by diagnosis category;
[0045] S332. Using the proportion of coding volume of each diagnostic category as the weight, the error rate of each category is weighted and averaged to obtain the overall coding error rate;
[0046] S333. Extract all medical record records with coding errors from historical coded medical record data; classify and statistically analyze coding errors according to two dimensions: diagnosis category and error type;
[0047] S334. Identify combinations of diagnostic categories and error types with coding error rates greater than the average error rate, and define them as blind spots in the medical record knowledge of the coders; sum the ratios of coding error rates of all diagnostic categories in the blind spots to the average error rate and take the average value as the severity of the blind spots in the medical record knowledge.
[0048] S335. The overall coding error rate and the severity of blind spots in medical record knowledge are weighted and summed to obtain the coding level abnormalities of the coding personnel.
[0049] As one implementation of the present invention, based on the analysis results of source node data encoding interference and the analysis results of encoding imbalance by medical record front page data coders, the quality risk of the medical record front page is assessed, including the following specific contents:
[0050] S41. Extract the analysis results of data encoding interference from source nodes and the analysis results of coding imbalance of data coding personnel on the front page of medical records;
[0051] S42. The results of the source node data coding interference analysis and the results of the medical record front page data coding personnel coding imbalance analysis are weighted and summed to obtain the medical record front page quality risk assessment results.
[0052] As one implementation of the present invention, based on the quality risk assessment results of the medical record front page, a quality control intervention plan for the medical record front page is implemented, including the following specific contents:
[0053] Extract the quality risk assessment results of the medical record cover sheet, preset the quality risk threshold of the medical record cover sheet, and implement the quality control intervention plan for the medical record cover sheet when the quality risk assessment result of the medical record cover sheet is greater than or equal to the quality risk threshold of the medical record cover sheet; when the quality risk assessment result of the medical record cover sheet is less than the quality risk threshold of the medical record cover sheet, the data coding quality of the medical record cover sheet is deemed to be qualified.
[0054] Secondly, embodiments of the present invention also provide a medical record front page quality control system based on big data analysis, including:
[0055] The data acquisition module is used to acquire the coding task list data and historical coded medical record data of the coding personnel of the medical record front page data, and at the same time acquire the clinical work data and diagnostic record data of the source node of the medical record front page data;
[0056] The data interference detection module is used to analyze the input quality and complexity of the medical record front page data by combining the clinical work data and diagnostic record data of the source node; and to analyze the data encoding interference of the source node based on the analysis results of the input quality and complexity of the medical record front page data.
[0057] The coding personnel evaluation module is used to analyze the coding workload and coding level anomalies of coding personnel by combining the coding task list data of coding personnel on the medical record front page data and historical coding medical record data. Based on the coding workload and coding level anomalies of coding personnel, it quantifies the coding imbalance of coding personnel on the medical record front page data.
[0058] The risk quantification module is used to assess the quality risk of the medical record front page based on the analysis results of data coding interference at the source node and the analysis results of coding imbalance among coding personnel.
[0059] The risk intervention module is used to implement quality control intervention plans for medical record front pages based on the results of the quality risk assessment.
[0060] The control module is used to control the operation of the data acquisition module, data interference detection module, coder assessment module, risk quantification module, and risk intervention module.
[0061] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0062] 1. This invention analyzes the input quality and complexity of source node data, accurately quantifies the interference of the clinical data entry process on subsequent coding, ensures the integrity and standardization of the medical record front page data from the source, and improves the reliability of basic data.
[0063] 2. This invention quantifies coding imbalance by comprehensively analyzing the workload status and coding level anomalies of coders, thereby achieving accurate identification of human resource risks. This enables targeted task allocation and training arrangements, ensuring the quality and efficiency of the coding process.
[0064] 3. This invention provides quality risk warning and intervention based on a comprehensive assessment of data source interference and personnel imbalance. It enables early detection and precise control of quality problems in the medical record front page, reduces medical insurance reimbursement and operational management risks caused by poor data quality, and safeguards the long-term value and security of hospital data assets. Attached Figure Description
[0065] Other features, objects, and advantages of the invention will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings:
[0066] Figure 1 This is a schematic diagram of the overall process of the medical record front page quality control method based on big data analysis of the present invention;
[0067] Figure 2 This is a flowchart illustrating the data input quality risk analysis strategy in the big data analysis-based medical record front page quality control method of the present invention.
[0068] Figure 3 This is a flowchart illustrating the data complexity analysis strategy in the big data-based medical record front page quality control method of the present invention.
[0069] Figure 4 This is a schematic diagram of the structure of the medical record front page quality control system based on big data analysis of the present invention. Detailed Implementation
[0070] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the embodiments of the present invention and the specific features in the embodiments are detailed descriptions of the technical solution of the present invention, rather than limitations thereof. In the absence of conflict, the embodiments of the present invention and the technical features in the embodiments can be combined with each other.
[0071] Example 1
[0072] like Figure 1 As shown, this embodiment provides a method for quality control of medical record front pages based on big data analysis, specifically including the following steps:
[0073] S1. Obtain the coding task list data and historical coded medical record data of the coding personnel of the medical record homepage data, and at the same time obtain the clinical work data and diagnostic record data of the source node of the medical record homepage data;
[0074] S2. Analyze the input quality and complexity of the medical record front page data by combining the clinical work data and diagnostic record data from the source nodes; and analyze the data encoding interference from the source nodes based on the analysis results of the input quality and complexity of the medical record front page data.
[0075] S3. Combining the coding task list data of the coding personnel on the front page of the medical record data with historical coding medical record data, analyze the coding task load status and coding level abnormalities of the coding personnel, and quantitatively analyze the coding imbalance of the coding personnel on the front page of the medical record data based on the coding task load status and coding level abnormalities.
[0076] S4. Based on the analysis results of source node data coding interference and the analysis results of coding imbalance of medical record front page data coding personnel, assess the quality risk of medical record front page;
[0077] S5. Based on the results of the quality risk assessment of the medical record front page, implement the quality control intervention plan for the medical record front page.
[0078] In step S2 of this embodiment, the clinical work data and diagnostic record data of the source node are combined to analyze the input quality and complexity of the medical record front page data; and based on the analysis results of the input quality and complexity of the medical record front page data, the data encoding interference of the source node is analyzed, including the following specific steps:
[0079] S21. Extract clinical work data and diagnostic record data from the source node; This embodiment integrates clinical work data and diagnostic record data from the source node into the hospital information system to obtain comprehensive clinical work data, such as doctor's orders, nursing records, examination reports, and consultation information. Diagnostic record data covers both structured and unstructured content, including preliminary diagnosis, final diagnosis, surgical procedure records, and complication descriptions. In this embodiment, data extraction is achieved through the ETL process of the hospital data warehouse. First, raw data is extracted from the source system using SQL queries or API interfaces. Then, data cleaning and standardization are performed, such as removing duplicate records, handling missing values, and standardizing the time format to ensure data consistency and integrity. Clinical work data is obtained by capturing dynamic information during the diagnosis and treatment process, such as departmental collaboration records and special event markers. This data reflects the decision-making process and resource allocation of the medical team. The extraction of diagnostic record data focuses on disease codes and text descriptions, requiring preliminary verification according to the International Classification of Diseases (ICD-10) and surgical procedure codes (ICD-9-CM3) rules to avoid bias in subsequent analysis. Specifically, in the parameter acquisition process of this embodiment, clinical work data is analyzed by parsing hospital workflow logs to count the number of departments involved and the frequency of consultations, while diagnostic record data is extracted from free text using natural language processing technology, such as using regular expressions to match diagnostic terms and surgical names. This embodiment, by systematically integrating multi-source data, not only reduces the risk of errors from manual data collection but also enhances the coverage of big data analysis, helping to identify potential data inconsistencies and sources of coding interference, thus laying a solid foundation for quality control of medical record front pages.
[0080] S22. Import diagnostic record data into the data input quality risk analysis strategy for data input quality risk analysis;
[0081] S23. Import clinical work data and diagnostic record data into a data complexity analysis strategy for data complexity analysis;
[0082] S24. The results of the data input quality risk analysis and the data complexity analysis are weighted and summed to obtain the source node data coding interference. Specifically, this embodiment first extracts the data input quality risk analysis results and the data complexity analysis results from the previous analysis, and then performs a weighted summation using preset weight coefficients. In this embodiment, the weight coefficients are determined through regression analysis of historical medical record quality assessment data to ensure that the weights accurately reflect the contribution of each factor to coding interference. This embodiment quantifies the potential interference of source node data on the coding process. For example, high data input risk may stem from missing key diagnostic elements, causing coders to rely on incomplete information, while high data complexity may increase the probability of coding errors due to rare diseases or multi-departmental collaboration. Through weighted integration, this embodiment avoids the limitations of single analysis, provides a more comprehensive risk assessment, and thus provides a targeted basis for subsequent medical record front page quality control, helping hospitals prioritize high-risk data sources, optimize coding workflows, and improve overall data quality and the accuracy of medical decisions.
[0083] In this embodiment, as Figure 2 As shown, the data input quality risk analysis strategy includes the following specific steps:
[0084] S221. Extract the text content of all medical documents in the diagnostic record data; and obtain a predefined list of key medical elements;
[0085] S222. Use a hybrid neural network model based on BERT-BiLSTM-CRF to perform medical entity recognition on the text content of all medical documents. The BERT layer is responsible for deep semantic encoding of the text content, the BiLSTM layer captures the long-term dependencies of the text content sequence, and the CRF layer ensures the global optimality of the output labels.
[0086] S223. Match the medical entities identified by the model with a predefined list of key medical elements, count the number of key medical elements that failed to match in all documents as the numerator, and count the total number of key medical elements required to be included in the predefined list of key medical elements as the denominator; divide the two to calculate the data incompleteness index.
[0087] S224. Extract the free text of all diagnostic descriptions and surgical procedure names from the diagnostic record data; and construct a standard medical terminology knowledge graph containing hierarchical and synonymous relationships between medical terms. In this embodiment, the standard medical terminology knowledge graph is constructed as follows: using the National Health Commission's "Disease Classification and Code National Clinical Version," "Surgical Procedure Classification Code National Clinical Version," and "Common Clinical Medical Terms" as authoritative standard terminology sets, serving as the core nodes of the knowledge graph; using medical ontology construction tools, based on the hierarchical relationships (e.g., diabetes and type 2 diabetes), synonymous relationships (e.g., myocardial infarction and myocardial infarction), and correlation relationships between terms in the above standards, constructing a semantic relationship network between terms; this knowledge graph can be incrementally updated through newly added diagnostic and surgical procedure descriptions exported periodically from the hospital information system, using natural language processing technology to identify new terms and their relationships with existing terms, and incorporating them into the standard medical terminology knowledge graph after review; this standard medical terminology knowledge graph belongs to the existing medical terminology knowledge graph, and will not be elaborated further here.
[0088] S225. Use the TransE algorithm to learn the vector representation of standard medical terms in the standard medical terminology knowledge graph; based on the Sentence-BERT model, map the free text of all diagnostic descriptions and surgical procedure names to be matched to the standard medical terms in the standard medical terminology knowledge graph to the same vector space.
[0089] S226. Calculate the cosine similarity between the free text vector and the standard medical term vector, and preset a similarity threshold; count the number of terms in the free text whose similarity is less than the preset similarity threshold as the number of non-standard terms, and divide the number of non-standard terms by the total number of terms in the free text to obtain the non-standard term ratio; in this embodiment, the similarity threshold is obtained experimentally by those skilled in the art, specifically by: obtaining multiple historical standard medical records that have passed manual review, extracting the free text of the diagnostic descriptions and surgical operation names and the corresponding standard medical terms; substituting the free text and standard medical terms into steps S225 and S226 of this embodiment, calculating the cosine similarity of multiple historical text and term pairs; obtaining the judgment results of whether these terms are standardized in the historical review; importing the calculated multiple similarity values and the corresponding standardization judgment results into the fitting software, and outputting the value of the similarity threshold that meets the highest term standardization judgment accuracy; in this embodiment, the fitting software can use ROC curves for analysis.
[0090] S227. The data incompleteness index and non-standard terminology ratio are weighted and summed to obtain the data input quality risk analysis results. In this embodiment, the weights of the data incompleteness index and non-standard terminology ratio are obtained as follows: Multiple historical medical records that have completed coding and quality control are obtained, including their diagnostic record data; the diagnostic record data is substituted into steps S221 to S226 of this embodiment to obtain the data incompleteness index and non-standard terminology ratio for each medical record; the final results of historical quality control in determining whether each medical record has coding errors are obtained; the obtained historical multiple sets of data incompleteness index, non-standard terminology ratio, and corresponding judgment results on whether the medical record has coding errors are imported into the fitting software for training, and the values of the two weights that meet the highest coding error risk prediction accuracy are output. In this embodiment, the fitting software can use a logistic regression model for analysis.
[0091] It should be noted that this embodiment is based on a multi-stage model processing, aiming to identify incomplete and non-standardized data issues. First, it extracts the text content of all medical documents in the diagnostic record data, including progress notes, surgical reports, and discharge summaries, and compares them against a predefined list of key medical elements (such as essential diagnostic indicators, surgical details, and complication markers). In this embodiment, the core of the data input quality risk analysis strategy lies in using a hybrid neural network model based on BERT-BiLSTM-CRF for medical entity recognition. The BERT layer uses a pre-trained language model to perform deep semantic encoding on the text content, capturing the contextual meaning of words. The BiLSTM layer processes the text sequence through a bidirectional long short-term memory network, effectively identifying long-term dependencies, such as the coherence of diagnostic descriptions. The CRF layer, as the output layer, optimizes the label sequence using a conditional random field algorithm, ensuring global consistency and accuracy in entity recognition. Specifically, the steps for establishing a hybrid neural network model based on BERT-BiLSTM-CRF in this embodiment include: first, segmenting and vectorizing the medical document; generating word embeddings using BERT; then, inputting the BiLSTM layer to learn sequence features; and finally, outputting entity labels (e.g., disease names, surgical procedures) through the CRF layer, thereby identifying all key medical entities. In this embodiment, the data incompleteness index is calculated by statistically analyzing the ratio of the number of unmatched key medical elements to the total number of items in the list, reflecting the severity of data loss. The non-standard terminology ratio is calculated by learning the vector representation of the standard medical terminology knowledge graph using the TransE algorithm, and combining it with the Sentence-BERT model to map free text to the same vector space. After calculating the cosine similarity, the proportion of terms below a threshold is statistically analyzed to assess the standardization of medical terminology. This quantifies the risk of data input and helps identify interfering factors in the source nodes that may affect the accuracy of coding, such as missing key elements or the use of non-standard terms leading to coding errors or ambiguities. By using automated entity recognition and similarity calculation, this method improves analysis efficiency, reduces the bias of subjective judgment, and provides a reliable basis for subsequent data complexity analysis and overall quality assessment, thereby enhancing the overall consistency and traceability of medical record front page data.
[0092] In this embodiment, as Figure 3 As shown, the data complexity analysis strategy includes the following specific steps:
[0093] S231. Extract all diagnostic codes and surgical operation codes from the diagnostic record data; and extract historical diagnostic codes and historical surgical operation codes from the hospital's historical medical record database. Based on the ICD-10 and ICD-9-CM3 coding rules, use regular expressions to deduplicate the above diagnostic codes and surgical operation codes.
[0094] S232. Sum the number of all diagnostic codes and all surgical operation codes in the diagnostic record data after deduplication to obtain the number of medical record codes; sum the number of historical diagnostic codes and historical surgical operation codes in the hospital's historical medical record database after deduplication to obtain the number of historical codes; divide the number of medical record codes and the number of historical codes to obtain the medical concept complexity of the medical record.
[0095] S233. Extract all diagnostic codes from the diagnostic record data; based on the hospital's historical medical record database, extract the frequency distribution of historical diagnostic codes and count the number of times each diagnostic code appears in the hospital's historical medical record database; based on the frequency distribution of historical diagnostic codes, extract the frequency of each diagnostic code in the diagnostic record data in the historical diagnostic codes.
[0096] S234. Preset a rare diagnosis threshold, and use the diagnosis codes that appear less frequently than the rare diagnosis threshold in the statistical diagnosis record data as rare diagnosis codes; use the total frequency of rare diagnosis codes in the historical diagnosis codes as the rare diagnosis weight of the medical record.
[0097] S235. Extract departmental consultation records from clinical work data and count the number of all clinical departments involved in the diagnosis and treatment; simultaneously extract special diagnosis and treatment event markers from diagnostic record data; construct a diagnosis and treatment process relationship diagram for the medical records based on the above data, defining each clinical department and special diagnosis and treatment event as nodes in the diagram, and defining the collaborative relationships between departments and the associations of special diagnosis and treatment events as edges in the diagram; count the number of nodes and edges in the diagnosis and treatment process relationship diagram for the medical records; in this embodiment, special diagnosis and treatment event markers refer to key node events that occur during the diagnosis and treatment process and may increase the complexity of the condition or the difficulty of coding, including but not limited to: transfer to the intensive care unit (ICU), conducting multidisciplinary team (MDT) consultations, performing emergency surgery, using advanced life support technologies such as extracorporeal membrane oxygenation (ECMO), the occurrence of serious complications (such as postoperative massive bleeding, severe infection, etc.), or a major change in diagnosis. These event markers can be automatically extracted from specific fields of the hospital information system, surgical anesthesia system records, or intensive care unit records.
[0098] S236. From the hospital's historical medical record database, extract the maximum number of nodes and the maximum number of edges in the diagnosis and treatment process graph of all historical medical records; use the ratio of the number of nodes in the diagnosis and treatment process graph of the medical record to the maximum number of nodes as the node complexity, and the ratio of the number of edges in the diagnosis and treatment process graph of the medical record to the maximum number of edges as the edge complexity; perform a weighted sum of the node complexity and the edge complexity to obtain the process complexity of the medical record.
[0099] S237. The complexity of medical concepts, the weight of rare diagnoses, and the complexity of procedures in the medical records are weighted and summed to obtain the data complexity analysis results.
[0100] It should be noted that this embodiment calculates multi-dimensional indicators, first extracts all diagnostic codes and surgical operation codes from the diagnostic record data, and then uses regular expressions to perform deduplication according to ICD-10 and ICD-9-CM3 rules based on historical coding data in the hospital's historical medical record database, in order to eliminate duplicate entries and ensure data purity. In this embodiment, the medical concept complexity of a medical record is obtained by calculating the ratio of the number of current medical record codes to the number of historical codes, reflecting the diversity and rarity of the disease and surgery in the medical record. The calculation of the rare diagnosis weight is based on the frequency distribution of historical diagnosis codes, with a preset rare diagnosis threshold (in this embodiment, the rare diagnosis threshold is 1% by default, and codes with an occurrence frequency of less than 1% are considered rare diagnosis codes). Codes in the current medical record below this threshold are counted, and their total occurrence frequency is used as the weight to capture the impact of rare cases on coding complexity. Furthermore, in this embodiment, the process complexity is analyzed by constructing a diagnosis and treatment process relationship graph, where nodes represent clinical departments and special diagnosis and treatment events involved in the diagnosis and treatment, and edges represent collaborative relationships or event associations between departments. The index is obtained by weighted summing the ratios of the number of nodes and the number of edges to the historical maximum values, which can effectively measure the intensity of collaboration and resource input in the diagnosis and treatment process. Specifically, this embodiment uses graph theory to visualize the diagnosis and treatment process and quantifies the relationship complexity by counting the number of nodes and edges; at the same time, rare diagnosis patterns are identified based on historical data mining based on frequency analysis. This study reveals the inherent challenges of medical record data. For example, high complexity of medical concepts can increase the workload of coders, while high complexity of processes can introduce inconsistencies in multi-departmental collaboration, leading to coding interference. Based on the above, the data complexity analysis in this embodiment not only helps identify high-risk medical records but also provides data support for subsequent coding interference analysis, thereby optimizing resource allocation and quality control interventions, and improving the accuracy and completeness of the medical record front page data.
[0101] In step S3 of this embodiment, the coding task list data of the coders on the front page of the medical record data and the historical coded medical record data are combined to analyze the coding task load status and coding level anomalies of the coders. Based on the coding task load status and coding level anomalies of the coders, the coding imbalance of the coders on the front page of the medical record data is quantitatively analyzed, including the following specific contents:
[0102] S31. Extract coding task list data and historical coding medical record data from coding personnel. Specifically, this embodiment integrates real-time task data and historical archives through the hospital coding management system. The coding task list data includes information such as the main diagnostic category, estimated standard working hours, and task priority of the medical records to be coded, while the historical coding medical record data covers the review results, error records, and diagnostic category distribution of completed coding medical records. In this embodiment, data extraction utilizes database query and log analysis tools. First, the task queue of the current coder is extracted from the task management system, including the number of medical records, type, and deadline. Then, historical coding records are retrieved from the quality review database, focusing on collecting data related to coding error types, error frequencies, and diagnostic categories. Specifically, the coding task list data is analyzed by parsing task descriptions and classification tags to statistically determine the frequency of occurrence of main diagnostic categories, while the historical coding medical record data is analyzed by aggregating review reports to calculate the overall error rate and classification error rate.
[0103] S32. Import the coding task list data into the coding task load status analysis strategy for coding task load status analysis.
[0104] S33. Import historical coded medical record data into the coding level anomaly analysis strategy to conduct coding level anomaly analysis;
[0105] S34. The coding task load status analysis results and coding level anomaly analysis results are weighted and summed to obtain the coding imbalance of the coders on the front page of the medical record data. Specifically, this embodiment first extracts the coding task load status analysis results and coding level anomaly analysis results from the previous analysis, and then performs a weighted summation using preset weight coefficients. In this embodiment, the weight coefficients are determined through correlation analysis between historical coding quality data and performance indicators to ensure that the weights accurately reflect the impact of each factor on the imbalance. Based on the above, this embodiment quantifies the degree of imbalance in coders' workload and ability. For example, a high workload may lead to fatigue errors, while a high level of anomalies may amplify these errors. Through weighted integration, this method provides a comprehensive perspective, helping hospitals identify high-risk coders, optimize resource allocation and intervention strategies, thereby improving the overall quality of the front page of medical record data and the traceability of medical services.
[0106] In this embodiment, the coding task load status analysis strategy includes the following specific steps:
[0107] S321. Extract the main diagnostic categories of all medical records to be coded from the coding task list data; calculate the number of times each main diagnostic category appears; divide the number of times each main diagnostic category appears by the total number of coding tasks to obtain the frequency of occurrence of each main diagnostic category.
[0108] S322. Squaring the frequency of occurrence for each major diagnostic category, and then adding all the squaring results, subtracting the addition result from the value 1, yields the frequency of cognitive switching for the coder. This frequency of cognitive switching is based on the Simpson diversity index, effectively measuring the breadth and evenness of the coding task's distribution across disease diagnostic categories. When all cases to be coded belong to the same major diagnostic category, the frequency of cognitive switching for the coder is 0; the more diagnostic categories the task covers and the more evenly distributed they are, the closer the frequency of cognitive switching for the coder is to 1. This reflects the frequency with which the coder needs to switch between different disease knowledge domains.
[0109] S323. Extract the estimated standard working hours of all cases to be coded from the coding task list data; calculate the standard deviation of the estimated standard working hours of all cases to be coded, and divide the standard deviation of the estimated standard working hours of all cases to be coded by the average of the estimated standard working hours of all cases to be coded to obtain the task working hour dispersion of the coders; in this embodiment, the task working hour dispersion reflects the degree of fluctuation in the expected time consumption of the coding task. The higher the dispersion, the more it indicates that the task combination includes both simple and complex cases, which increases the difficulty of work rhythm management.
[0110] S324. The coding task load status is obtained by weighted summing of the frequency of cognitive switching and the task time dispersion of the coder. In this embodiment, the weights of the frequency of cognitive switching and the task time dispersion are obtained experimentally by those skilled in the art. The specific experimental method is as follows: obtain coding task list data of multiple coders in history; substitute the coding task list data into steps S321 to S323 of this embodiment to obtain the cognitive switching frequency value and task time dispersion value of each coder; obtain the average coding error rate of the coder after the corresponding task is completed as the actual load impact result; import the obtained historical multiple sets of cognitive switching frequency values, task time dispersion values and corresponding actual load impact results into fitting software for multiple regression analysis, and output the values of the two weights that meet the best fitting effect. In this embodiment, the higher the coding task load status, the greater the work challenge faced by the coder, who needs to frequently switch between different disease knowledge domains while also coping with the pressure brought by tight work time and uneven task difficulty.
[0111] It should be noted that this embodiment calculates based on task diversity and time volatility. First, it extracts the main diagnostic categories and estimated standard working hours for all cases to be coded from the coding task list data. Then, it calculates the frequency of cognitive switching using the Simpson diversity index. This parameter is obtained by squaring and summing the frequency of occurrence of each main diagnostic category, and then subtracting this value from 1. The principle is that when the task categories are evenly distributed, the index is close to 1, indicating that the coder needs to frequently switch between different disease knowledge domains, increasing cognitive load; conversely, the index is 0 for a single category, indicating a lower load. Specifically, the task time dispersion is obtained by calculating the ratio of the standard deviation to the mean of the estimated standard working hours for all cases to be coded. This reflects the uncertainty of task time consumption. High dispersion indicates that the task combination mixes simple and complex cases, making the work rhythm difficult to manage. This embodiment uses statistical methods to analyze the distribution of task categories and changes in working hours, and quantifies cognitive needs through the diversity index. It comprehensively assesses the work challenges faced by coders; for example, high cognitive switching frequency reduces coding efficiency, while high task time dispersion may lead to errors under time pressure. Analyzing the coding workload provides crucial input for subsequent coding imbalance assessments, helping hospitals monitor coders' status in real time, adjust task allocation, or provide support, thereby reducing coding errors and improving the accuracy and consistency of medical record front page data.
[0112] In this embodiment, the coding level anomaly analysis strategy includes the following specific steps:
[0113] S331. Extract the coding quality review results of all completed coding medical records from the historical coding medical record data; classify and calculate the coding error rate under each diagnostic category, wherein the coding error rate is calculated as the ratio of the number of incorrectly coded medical records under that diagnostic category to the total number of coded medical records under that diagnostic category;
[0114] S332. Using the proportion of coding volume in each diagnostic category as the weight, the error rate of each category is weighted and averaged to obtain the overall coding error rate. This embodiment reveals the strengths and weaknesses of the coder's professional knowledge structure by quantifying the coder's historical performance in each diagnostic category.
[0115] S333. Extract all medical record records with coding errors from historical coded medical record data; classify and statistically analyze coding errors according to two dimensions: diagnosis category and error type;
[0116] S334. Identify combinations of diagnostic categories and error types with coding error rates exceeding the average error rate, defining these as blind spots in the coder's medical record knowledge. Sum the ratios of the coding error rates of all diagnostic categories within the medical record knowledge blind spot to the average error rate, and take the average as the severity of the blind spot. The average error rate is calculated as follows: for each combination of diagnostic category and error type, compare the error rate of each diagnostic category with the overall coding error rate, and sum and average the comparison results for all diagnostic categories to obtain the average error rate. This embodiment, by systematically analyzing the coder's historical error patterns, accurately identifies their weaknesses in specific disease areas or specific coding rules. The severity of the knowledge blind spot quantifies the potential impact of these weaknesses on coding quality.
[0117] S335. The overall coding error rate and the severity of the medical record knowledge blind spot are weighted and summed to obtain the coding level anomalies of the coders. In this embodiment, the weights of the overall coding error rate and the severity of the medical record knowledge blind spot are obtained experimentally by those skilled in the art. The specific experimental method is as follows: obtain historical coding medical record data of multiple coders; substitute the historical coding medical record data of each coder into steps S331 to S334 of this embodiment to obtain their overall coding error rate and the severity of the medical record knowledge blind spot; obtain the frequency of coding quality problems of the coder in subsequent periods as the actual risk result; import the obtained historical multiple sets of overall coding error rates, the severity of medical record knowledge blind spots and the corresponding actual risk results into fitting software for regression analysis, and output the values of the two weights that meet the highest subsequent risk prediction accuracy.
[0118] It should be noted that the coding level anomaly analysis strategy in this embodiment is implemented through error rate statistics and knowledge blind spot identification. First, the coding quality review results of all completed coding cases in the historical coding medical record data are extracted. The coding error rate is statistically analyzed by diagnosis category and error type, where the error rate is calculated as the ratio of the number of erroneous coded medical records in that category to the total number of coded medical records. This embodiment uses a weighted average of the overall coding error rate based on the coding volume proportion of each diagnosis category, which balances the workload impact of different categories and more accurately reflects the overall performance. The severity of medical record knowledge blind spots is obtained by identifying combinations of diagnosis categories and error types with error rates higher than the average error rate, and calculating the average ratio of their error rates to the average error rate. The average error rate is determined based on a comparative analysis of error rates across all categories. Furthermore, this embodiment utilizes data mining techniques to aggregate historical error records, constructs an error pattern matrix, and identifies high-frequency error areas through threshold filtering. This can systematically reveal the coders' skill gaps; for example, a high overall error rate may indicate a general problem, while severe knowledge blind spots point to insufficient understanding of specific diseases or rules. By quantifying these anomalies, this embodiment provides a basis for targeted improvements, provides data support for subsequent coding imbalance analysis, and helps hospitals implement precise training and quality intervention to improve the professionalism and reliability of medical record front page coding.
[0119] In this embodiment, based on the analysis results of source node data encoding interference and the analysis results of coding imbalance among medical record front page data coders, the quality risk of the medical record front page is assessed, including the following specific contents:
[0120] S41. Extract the analysis results of data encoding interference from source nodes and the analysis results of coding imbalance of data coding personnel on the front page of medical records;
[0121] S42. The results of the source node data coding interference analysis and the results of the medical record front page data coding imbalance analysis are weighted and summed to obtain the medical record front page quality risk assessment result. Specifically, this embodiment adds the two analysis results based on preset weight coefficients. The weight coefficients are set by regression analysis of the correlation between historical medical record error events and various factors to ensure the predictive accuracy of the risk assessment. This embodiment can intuitively reflect the potential quality problems of medical record front page data. For example, a high score may indicate that the data source is unreliable and the coder's ability is insufficient, requiring immediate intervention. Through weighted integration, this method provides an actionable risk level, providing a basis for subsequent quality control intervention plans, thereby improving the overall quality, compliance, and medical data value of the medical record front page.
[0122] In this embodiment, based on the quality risk assessment results of the medical record cover sheet, a quality control intervention plan for the medical record cover sheet is implemented, including the following specific contents:
[0123] The system extracts the quality risk assessment results of the medical record cover page and presets a quality risk threshold. When the quality risk assessment result is greater than or equal to the threshold, a quality control intervention plan is implemented. When the result is less than the threshold, the data coding quality of the medical record cover page is deemed acceptable. The quality control intervention plan includes, but is not limited to: automatically assigning medical records to senior coding experts or quality control specialists for manual review, focusing on incomplete data, non-standard terminology, and coding entries related to complex diagnostic and treatment logic; simultaneously, the system automatically triggers alerts and notifies relevant clinical departments, requiring them to supplement or correct missing or non-standard key medical elements within a specified timeframe, ensuring data integrity and accuracy from the source. For low- to medium-risk medical records, a standardized intervention strategy is adopted, such as automatically pushing prompts through the quality control system to remind coders to pay attention to specific diagnostic categories or known knowledge gaps, and recommending relevant coding rules and standard terminology for reference. In this embodiment, the setting of the quality risk threshold for the medical record front page is optimized through ROC curve analysis. Specifically, this embodiment first constructs a classification model, using the comprehensive quality risk assessment score as the predictor variable and the coding errors actually found in historical medical records as the verification label. By adjusting a series of candidate thresholds, the ROC curves formed by the true positive rate (correctly identified high-risk medical records) and the false positive rate (misjudged low-risk medical records) corresponding to each threshold are calculated and plotted. The optimal balance point is found on this ROC curve by calculating the maximum Youden index. The maximum Youden index is the difference between the true positive rate and the false positive rate. The threshold corresponding to the maximum Youden index is selected as the trigger standard for quality control intervention, thereby maximizing resource utilization efficiency while ensuring the quality of the medical record front page data.
[0124] Example 2
[0125] like Figure 4 As shown, this embodiment provides a medical record front page quality control system based on big data analysis, including:
[0126] The data acquisition module is used to acquire the coding task list data and historical coded medical record data of the coding personnel of the medical record front page data, and at the same time acquire the clinical work data and diagnostic record data of the source node of the medical record front page data;
[0127] The data interference detection module is used to analyze the input quality and complexity of the medical record front page data by combining the clinical work data and diagnostic record data of the source node; and to analyze the data encoding interference of the source node based on the analysis results of the input quality and complexity of the medical record front page data.
[0128] The coding personnel evaluation module is used to analyze the coding workload and coding level anomalies of coding personnel by combining the coding task list data of coding personnel on the medical record front page data and historical coding medical record data. Based on the coding workload and coding level anomalies of coding personnel, it quantifies the coding imbalance of coding personnel on the medical record front page data.
[0129] The risk quantification module is used to assess the quality risk of the medical record front page based on the analysis results of data coding interference at the source node and the analysis results of coding imbalance among coding personnel.
[0130] The risk intervention module is used to implement quality control intervention plans for medical record front pages based on the results of the quality risk assessment.
[0131] The control module is used to control the operation of the data acquisition module, data interference detection module, coder assessment module, risk quantification module, and risk intervention module.
[0132] The steps for implementing the corresponding functions of each parameter and unit module in the medical record front page quality control system based on big data analysis of the present invention can be referred to the parameters and steps in the embodiments of the medical record front page quality control method based on big data analysis above, and will not be repeated here.
[0133] The various embodiments in this invention are described in a progressive manner. Similar or identical parts between embodiments can be referred to interchangeably. Each embodiment focuses on its differences from other embodiments. In particular, the embodiments for IoT devices and media are relatively simple in description because they are fundamentally similar to the method embodiments; relevant parts can be referred to the descriptions in the method embodiments.
[0134] The systems, media, and methods provided in the embodiments of the present invention are in one-to-one correspondence. Therefore, the systems and media also have similar beneficial technical effects as their corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the systems and media will not be repeated here.
[0135] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0136] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0137] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0138] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0139] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0140] Computer-readable media include both permanent and non-permanent, removable and non-removable media that can store information by any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0141] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0142] The above are merely embodiments of the present invention and are not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of the present invention should be included within the scope of the claims of the present invention.
Claims
1. A method for quality control of medical record front pages based on big data analysis, characterized in that, Includes the following steps: S1. Obtain the coding task list data and historical coded medical record data of the coding personnel of the medical record homepage data, and at the same time obtain the clinical work data and diagnostic record data of the source node of the medical record homepage data; S2. Analyze the input quality and complexity of the medical record front page data by combining the clinical work data and diagnostic record data from the source nodes; and analyze the data encoding interference from the source nodes based on the analysis results of the input quality and complexity of the medical record front page data. S3. Combining the coding task list data of the coding personnel on the front page of the medical record data with historical coding medical record data, analyze the coding task load status and coding level abnormalities of the coding personnel, and quantitatively analyze the coding imbalance of the coding personnel on the front page of the medical record data based on the coding task load status and coding level abnormalities. S4. Based on the analysis results of source node data coding interference and the analysis results of coding imbalance of medical record front page data coding personnel, assess the quality risk of medical record front page; S5. Based on the results of the medical record front page quality risk assessment, implement the medical record front page quality control intervention plan; In step S2, the clinical work data and diagnostic record data from the source node are combined to analyze the input quality and complexity of the medical record front page data; and based on the analysis results of the input quality and complexity of the medical record front page data, the data encoding interference from the source node is analyzed, including the following specific steps: S21. Extract clinical work data and diagnostic record data from the source node; S22. Import diagnostic record data into the data input quality risk analysis strategy for data input quality risk analysis; S23. Import clinical work data and diagnostic record data into a data complexity analysis strategy for data complexity analysis; S24. Weighted summation of the data input quality risk analysis results and data complexity analysis results to obtain the source node data encoding interference situation; In step S3, the coding task list data of the coders on the medical record front page data and the historical coded medical record data are combined to analyze the coding task load status and coding level anomalies of the coders. Based on the coding task load status and coding level anomalies of the coders, the coding imbalance of the coders on the medical record front page data is quantitatively analyzed, including the following specific contents: S31. Extract coding task list data and historical coding medical record data of coding personnel; S32. Import the coding task list data into the coding task load status analysis strategy for coding task load status analysis. S33. Import historical coded medical record data into the coding level anomaly analysis strategy to conduct coding level anomaly analysis; S34. The results of the coding task load status analysis and the results of the coding level anomaly analysis are weighted and summed to obtain the coding imbalance of the coders of the medical record front page data. The data input quality risk analysis strategy includes the following specific steps: Extract the text content of all medical documents from the diagnostic record data; and obtain a predefined list of key medical elements; A hybrid neural network model based on BERT-BiLSTM-CRF was used to perform medical entity recognition on the text content of all medical documents; The model identifies medical entities and matches them with a predefined list of key medical elements. Based on the matching results, a data incompleteness index is calculated. Extract the free text of all diagnostic descriptions and surgical procedure names from diagnostic record data; and construct a standard medical terminology knowledge graph containing hierarchical and synonym relationships between medical terms. Analyze the ratio of non-standard terms based on the free text of all diagnostic descriptions and surgical procedure names in the diagnostic record data and the standard medical terminology knowledge graph. The data input quality risk analysis results are obtained by weighted summation of the data incompleteness index and the non-standard terminology ratio. The data complexity analysis strategy includes the following specific steps: S231. Extract all diagnostic codes and surgical operation codes from the diagnostic record data; and extract historical diagnostic codes and historical surgical operation codes from the hospital's historical medical record database, and perform deduplication on the above diagnostic codes and surgical operation codes. S232. Combining the diagnostic codes and surgical operation codes after deduplication, analyze the complexity of medical concepts in the medical records. S233. Extract all diagnostic codes from the diagnostic record data; based on the hospital's historical medical record database, extract the frequency distribution of historical diagnostic codes. S234. Calculate the rare diagnosis weight of the medical record based on all diagnostic codes in the diagnostic record data and the frequency distribution of historical diagnostic codes. S235. Extract departmental consultation records from clinical work data and count the number of all clinical departments involved in diagnosis and treatment; simultaneously extract special diagnosis and treatment event markers from diagnostic record data; construct a diagnosis and treatment process relationship graph for medical records based on the above data, defining each clinical department and special diagnosis and treatment event as nodes in the graph, and defining the collaboration relationship between departments and the association of special diagnosis and treatment events as edges in the graph; count the number of nodes and edges in the diagnosis and treatment process relationship graph for medical records. S236. From the hospital's historical medical record database, extract the maximum number of nodes and the maximum number of edges in the diagnosis and treatment process graph of all historical medical records; use the ratio of the number of nodes in the diagnosis and treatment process graph of the medical record to the maximum number of nodes as the node complexity, and the ratio of the number of edges in the diagnosis and treatment process graph of the medical record to the maximum number of edges as the edge complexity; perform a weighted sum of the node complexity and the edge complexity to obtain the process complexity of the medical record. S237. The medical concept complexity, rare diagnosis weight, and process complexity of the medical records are weighted and summed to obtain the data complexity analysis results. The coding task load status analysis strategy includes the following specific steps: S321. Extract the main diagnostic categories of all medical records to be coded from the coding task list data; calculate the number of times each main diagnostic category appears; divide the number of times each main diagnostic category appears by the total number of coding tasks to obtain the frequency of occurrence of each main diagnostic category. S322. Squaring the frequency of occurrence of each major diagnostic category, adding all the squaring results, and subtracting the addition result from the value 1 to obtain the frequency of cognitive switching of the coder. S323. Analyze the task time dispersion of coders based on the coding task list data; S324. The coding task load status is obtained by weighting and summing the frequency of cognitive switching of the coder and the dispersion of task working hours. The coding level anomaly analysis strategy includes the following specific steps: S331. Extract the coding quality review results of all completed coding medical records from the historical coding medical record data; classify and statistically analyze the coding error rate under each diagnostic category by diagnosis category; S332. Using the proportion of coding volume of each diagnostic category as the weight, the error rate of each category is weighted and averaged to obtain the overall coding error rate; S333. Extract all medical record records with coding errors from historical coded medical record data; classify and statistically analyze coding errors according to two dimensions: diagnosis category and error type; S334. Identify combinations of diagnostic categories and error types with coding error rates greater than the average error rate, and define them as blind spots in the medical record knowledge of the coders; sum the ratios of coding error rates of all diagnostic categories in the blind spots to the average error rate and take the average value as the severity of the blind spots in the medical record knowledge. S335. The overall coding error rate and the severity of blind spots in medical record knowledge are weighted and summed to obtain the coding level abnormalities of the coding personnel.
2. The method for quality control of medical record front page based on big data analysis according to claim 1, characterized in that, The analysis results of source node data encoding interference and the analysis results of coding imbalance by coding personnel of medical record front page data are used to assess the quality risk of medical record front page data, including the following specific contents: S41. Extract the analysis results of data encoding interference from source nodes and the analysis results of coding imbalance of data coding personnel on the front page of medical records; S42. The results of the source node data coding interference analysis and the results of the medical record front page data coding personnel coding imbalance analysis are weighted and summed to obtain the medical record front page quality risk assessment results.
3. The method for quality control of medical record front page based on big data analysis according to claim 2, characterized in that, The implementation of the medical record cover page quality control intervention plan based on the medical record cover page quality risk assessment results includes the following specific contents: Extract the quality risk assessment results of the medical record cover page, preset the quality risk threshold of the medical record cover page, and implement the quality control intervention plan for the medical record cover page when the quality risk assessment result of the medical record cover page is greater than or equal to the quality risk threshold of the medical record cover page; When the quality risk assessment result of the medical record cover page is less than the quality risk threshold of the medical record cover page, the data coding quality of the medical record cover page is deemed to be qualified.
4. A medical record front page quality control system based on big data analysis, implemented based on the medical record front page quality control method based on big data analysis as described in any one of claims 1-3, characterized in that, The system includes: The data acquisition module is used to acquire the coding task list data and historical coded medical record data of the coding personnel of the medical record front page data, and at the same time acquire the clinical work data and diagnostic record data of the source node of the medical record front page data; The data interference detection module is used to analyze the input quality and complexity of the medical record front page data by combining the clinical work data and diagnostic record data of the source node; and to analyze the data encoding interference of the source node based on the analysis results of the input quality and complexity of the medical record front page data. The coding personnel evaluation module is used to analyze the coding workload and coding level anomalies of coding personnel by combining the coding task list data of coding personnel on the medical record front page data and historical coding medical record data. Based on the coding workload and coding level anomalies of coding personnel, it quantifies the coding imbalance of coding personnel on the medical record front page data. The risk quantification module is used to assess the quality risk of medical record front pages based on the analysis results of data coding interference at source nodes and the analysis results of coding imbalance among coding personnel. The risk intervention module is used to implement quality control intervention plans for medical record front pages based on the results of the quality risk assessment. The control module is used to control the operation of the data acquisition module, the data interference detection module, the coder evaluation module, the risk quantification module, and the risk intervention module.
Citation Information
Patent Citations
Quality analysis method and system based on medical insurance medical record home page data
CN117637086A
Real-time quality control and coding method for medical record home page data based on multi-dimensional verification
CN120636662A