Hospital big data laboratory data compliance dynamic auditing and error correcting method and system

By applying audit rule sets and blockchain technology to automate the processing of raw medical data in the hospital's big data laboratory, the problem of low data audit efficiency has been solved, and efficient data compliance management and enhanced security have been achieved.

CN121034574APending Publication Date: 2025-11-28KUNSHAN FIRST PEOPLES HOSPITAL
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511216456.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-28
Publication Date
2025-11-28

Smart Images

  • Figure CN121034574A_ABST
    Figure CN121034574A_ABST
Patent Text Reader

Abstract

The invention relates to a hospital big data laboratory data compliance dynamic auditing and error correction method and system, and relates to the technical field of data processing. The method mainly comprises the following steps: auditing original medical data through an auditing rule set to obtain an auditing label corresponding to each piece of original medical data; according to the audit tag, storing the index position, the audit tag, the data hash value, the timestamp and the node digital signature corresponding to the original medical data into a public chain or a private chain in the block chain; obtaining corresponding abnormal original medical data through an index position stored in a private chain in the block chain, and determining a data exception type of the abnormal original medical data corresponding to the index position; determining an error correction scheme according to the data exception type and the corresponding abnormal original medical data; and error correction is carried out on the abnormal original medical data through an error correction scheme.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and more specifically to a method and system for dynamic auditing and error correction of data compliance in a hospital big data laboratory. Background Technology

[0002] With the development of internet technology, internet-based businesses are becoming increasingly prevalent. During the execution of these businesses, malicious access may occur, potentially causing operational anomalies. Therefore, to ensure the normal operation of these businesses, it is necessary to audit the data accessed. The purpose of auditing is not only to identify problems, but more importantly, to promote rectification, standardize management, ensure security, improve quality, and ultimately safeguard the healthy operation of the hospital and the legitimate rights and interests of patients.

[0003] Currently, data auditing is mainly carried out manually by staff on remote interactive data. When the amount of remote interactive data is large, manual auditing is inefficient. Summary of the Invention

[0004] The present invention aims to provide a method and system for dynamic auditing and error correction of data compliance in hospital big data laboratories, so as to solve the shortcomings of the existing technology. The technical problem to be solved by the present invention is achieved through the following technical solution.

[0005] This invention provides a method for dynamic auditing and error correction of data compliance in a hospital big data laboratory, the method comprising:

[0006] Acquire raw medical data transmitted from hospital nodes; the raw medical data includes at least electronic medical records, test reports, and medication records;

[0007] The original medical data is audited using an audit rule set to obtain an audit label for each piece of original medical data. The audit labels include normal and abnormal.

[0008] Based on the audit tag, the index position, audit tag, data hash value, timestamp, and node digital signature of the original medical data are stored in the public or private chain of the blockchain; the public chain stores the original medical data with normal audit tags, and the private chain stores the original medical data with abnormal audit tags;

[0009] The corresponding abnormal original medical data is obtained by retrieving the index position stored in the private chain of the blockchain, and the data abnormality type of the abnormal original medical data corresponding to the index position is determined.

[0010] An error correction scheme is determined based on the data anomaly type and the corresponding original medical data; and the original medical data with the anomaly is corrected using the error correction scheme.

[0011] In an optional embodiment, storing the index position, audit tag, data hash value, timestamp, and node digital signature corresponding to the original medical data into a public or private blockchain based on the audit tag includes:

[0012] The index location, audit tag, data hash value, timestamp, and node digital signature corresponding to the original medical data with the audit tag being normal are stored in the public chain of the blockchain;

[0013] The index location, audit tag, data hash value, timestamp, and node digital signature corresponding to the original medical data with the audit tag being abnormal are stored in the private chain of the blockchain.

[0014] In an optional embodiment, the step of auditing the original medical data using an audit rule set to obtain audit tags corresponding to each piece of original medical data includes:

[0015] Calculate the rule violation severity value of the original medical data according to each audit rule in the audit rule set;

[0016] Obtain the maximum violation severity value among the violation severity values; and determine whether the maximum violation severity value is greater than an adaptive threshold.

[0017] If the value is greater than the adaptive threshold, the audit label of the original medical data is determined to be abnormal;

[0018] If the value is less than or equal to the adaptive threshold, the audit label of the original medical data is determined to be normal.

[0019] In an optional embodiment, calculating the violation severity value of the original medical data according to each audit rule in the audit rule set includes:

[0020] Obtain the historical medical data corresponding to the original medical data;

[0021] The historical standard deviation and historical mean were determined based on the historical medical data.

[0022] The degree of violation of the original medical data under each audit rule is calculated using the original medical data, historical standard deviation, and historical mean.

[0023] In an optional embodiment, the step of calculating the rule violation degree value of the original medical data under each audit rule using the original medical data, historical standard deviation, and historical mean includes:

[0024] Based on the original medical data, historical standard deviation, and historical mean, calculate the violation component corresponding to each condition under each audit rule;

[0025] The violation severity value of the original medical data under each audit rule is obtained by weighting the violation components corresponding to each condition under each audit rule.

[0026] In an optional embodiment, before calculating the degree of violation of the original medical data under each audit rule by weighting the violation components corresponding to each condition under each audit rule, the method further includes:

[0027] Obtain the gradient values ​​of all hospital nodes for each condition under each audit rule;

[0028] The final gradient value is obtained by averaging the gradient values ​​of each condition under each audit rule for all hospital nodes.

[0029] The weight value of each condition under each audit rule is determined based on the final gradient value.

[0030] In an optional embodiment, the data anomaly type of the original medical data corresponding to the index position includes:

[0031] Identify the abnormal data group corresponding to the abnormal original medical data at the index position. The abnormal data includes feature values, rule identifiers, rule violation degree values, operation user identifiers, timestamps, and device identifiers.

[0032] Based on a predefined medical error knowledge base, the abnormal data groups are identified to obtain the probability values ​​corresponding to each data abnormality type;

[0033] The data anomaly type of the original medical data is determined by the probability value corresponding to each data anomaly type.

[0034] In an optional embodiment, the data anomaly type of the original medical data corresponding to the index position includes:

[0035] Identify the abnormal data group corresponding to the abnormal original medical data at the index position. The abnormal data includes feature values, rule identifiers, rule violation degree values, operation user identifiers, timestamps, and device identifiers.

[0036] Get the abnormal data groups within a predetermined time range and arrange them in chronological order.

[0037] The abnormal data groups arranged in chronological order within a predetermined time range are converted into a feature vector matrix, where each row of the feature vector matrix represents an abnormal data group corresponding to a timestamp.

[0038] The data anomaly type of the original medical data corresponding to the index position is determined based on the feature vector matrix.

[0039] In an optional embodiment, determining the data anomaly type of the original medical data corresponding to the index position based on the feature vector matrix includes:

[0040] The feature vector matrix is ​​input into the data anomaly prediction model to obtain the data anomaly type of the original medical data corresponding to the index position.

[0041] This invention provides a dynamic audit and error correction system for data compliance in a hospital big data laboratory. The system includes:

[0042] The acquisition module is used to acquire raw medical data transmitted by the hospital node; the raw medical data includes at least electronic medical records, test reports, and medication records.

[0043] The audit module is used to audit the raw medical data using an audit rule set to obtain an audit tag corresponding to each piece of raw medical data. The audit tags include normal and abnormal.

[0044] The storage module is used to store the index position, audit tag, data hash value, timestamp, and node digital signature corresponding to the original medical data into a public or private chain in the blockchain according to the audit tag; the public chain stores the original medical data with normal audit tags, and the private chain stores the original medical data with abnormal audit tags;

[0045] The determination module is used to obtain the corresponding abnormal original medical data through the index position stored in the private chain of the blockchain, and determine the data abnormality type of the abnormal original medical data corresponding to the index position.

[0046] The error correction module is used to determine an error correction scheme based on the data anomaly type and the corresponding abnormal original medical data; and to correct the abnormal original medical data using the error correction scheme.

[0047] The embodiments of the present invention have the following advantages:

[0048] This invention provides a method and system for dynamic auditing and error correction of data compliance in a hospital big data laboratory. First, raw medical data transmitted from hospital nodes is acquired. This raw medical data includes at least electronic medical records, test reports, and medication records. Then, the raw medical data is audited using an audit rule set to obtain audit tags for each piece of raw medical data. Audit tags include normal and abnormal. Next, based on the audit tags, the index position, audit tag, data hash value, timestamp, and node digital signature of the raw medical data are stored in a public or private blockchain. The public blockchain stores raw medical data with normal audit tags, while the private blockchain stores raw medical data with abnormal audit tags. The corresponding abnormal raw medical data is obtained through the index position stored in the private blockchain, and the data anomaly type of the abnormal raw medical data corresponding to the index position is determined. Finally, an error correction scheme is determined based on the data anomaly type and the corresponding abnormal raw medical data, and the abnormal raw medical data is corrected using the error correction scheme. Compared to existing technologies where remote interactive data is manually audited by staff, this application allows hospital nodes to audit raw medical data according to a set of audit rules after uploading the data. The audit results are then stored on the blockchain, and the types of data anomalies are determined based on the private blockchain. This approach improves audit efficiency and enhances data security. Attached Figure Description

[0049] Figure 1 This is a flowchart of a method for dynamic auditing and error correction of data compliance in a hospital big data laboratory, provided by an embodiment of the present invention;

[0050] Figure 2 This is a schematic diagram of the structure of a hospital big data laboratory data compliance dynamic audit and error correction system provided in an embodiment of the present invention. Detailed Implementation

[0051] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. The present invention will now be described in detail with reference to the accompanying drawings and embodiments.

[0052] like Figure 1 As shown in this embodiment, a dynamic audit method for data compliance in a hospital big data laboratory is provided. This method is used to perform the following steps:

[0053] S101, Obtain the raw medical data transmitted by the hospital node.

[0054] The original medical data includes at least electronic medical records, laboratory reports, and medication records. In this embodiment, after obtaining the original medical data, relevant medical data can be extracted from the electronic medical records, laboratory reports, and medication records. This extraction can be done through text recognition or image recognition technology to extract key fields from the electronic medical records, laboratory reports, and medication records. These key fields may include: patient ID, timestamp, treatment items, measurement results (such as blood pressure, blood sugar, heart rate, etc.), etc., which are not specifically limited in this embodiment. After obtaining the key fields, they are standardized to conform to a unified medical data standard. Then, the key fields are normalized, outlier filtering is performed, and sensitive information is de-identified (e.g., using differential privacy technology to add noise to protect patient privacy).

[0055] S102, the original medical data is audited using an audit rule set to obtain an audit tag corresponding to each piece of original medical data.

[0056] The audit rules set includes multiple audit rules, which can be set according to actual medical conditions. For example, audit rules could include: hypertension medication rules: systolic blood pressure ≥140 or diastolic blood pressure ≥90; dosage between 0.1 mg / kg and 0.3 mg / kg; blood pressure safety rules for the elderly: systolic blood pressure ≤160 (safe upper limit for systolic blood pressure in the elderly), diastolic blood pressure ≤100 (safe upper limit for diastolic blood pressure in the elderly), etc. Another example is audit rule ID: HT-MED-001, whose rule logic is as follows:

[0057] When systolic blood pressure ≥ 140 mmHg or diastolic blood pressure ≥ 90 mmHg

[0058] Hypertension medication must be prescribed (drug code range: C02-C10).

[0059] The medication dosage must conform to the patient's weight standard (0.1mg / kg ≤ dosage ≤ 0.3mg / kg).

[0060] In this embodiment, the audit tags include "normal" and "abnormal." The audit tags serve as a key basis for data classification. When L=0 (normal data), the public blockchain storage process is triggered; when L=1 (abnormal data), the private blockchain storage process is triggered. An audit tag of "normal" indicates that the original medical data has no problems after auditing; an audit tag of "abnormal" indicates that the original medical data has been determined to be abnormal after auditing.

[0061] In one optional embodiment provided in this application, the step of auditing the original medical data using an audit rule set to obtain audit tags corresponding to each piece of original medical data includes:

[0062] S1021, calculate the rule violation degree value of the original medical data according to each audit rule in the audit rule set.

[0063] In this embodiment, the step of calculating the violation degree value of the original medical data according to each audit rule in the audit rule set includes: obtaining historical medical data corresponding to the original medical data; determining the historical standard deviation and historical mean based on the historical medical data; and calculating the violation degree value of the original medical data under each audit rule using the original medical data, historical standard deviation, and historical mean.

[0064] Specifically, the step of calculating the violation severity value of the original medical data under each audit rule using the original medical data, historical standard deviation, and historical mean includes: calculating the violation component corresponding to each condition under each audit rule based on the original medical data, historical standard deviation, and historical mean; and weighting the violation components corresponding to each condition under each audit rule to obtain the violation severity value of the original medical data under each audit rule. More specifically, in this embodiment, the absolute value of the difference between the actual value and the historical mean in the original medical data is first calculated, and then the absolute value of the difference is divided by the historical standard deviation to obtain the violation component.

[0065] For example, the patient's age is 65 years; systolic blood pressure (SBP) is 180 mmHg; diastolic blood pressure (DBP) is 110 mmHg; the medication record shows the drug code as "R06" (antihistamine), the dosage as 10 mg; and the patient's weight is 70 kg. The audit rule set contains two audit rules:

[0066] Audit Rule 1 (Hypertension Medication Rules):

[0067] Condition 1 (Continuous characteristic): Systolic blood pressure ≥140 (sub-condition 1) or diastolic blood pressure ≥90 (sub-condition 2)

[0068] Condition 2 (Discrete Feature): The drug code must be in the set of hypertension drugs (e.g., the set is {"C02","C03","C07","C08","C09"}).

[0069] Condition 3 (Continuous Characteristics): Dosage between 0.1 mg / kg and 0.3 mg / kg

[0070] Weights: w1 = [0.5, 0.3, 0.2] (corresponding to conditions 1, 2, and 3 respectively)

[0071] Calculate the violation component of rule 1

[0072] Condition 1 of Audit Rule 1 (abnormal blood pressure) is an OR relationship between two sub-conditions (systolic blood pressure ≥140 or diastolic blood pressure ≥90). In calculating the violation component, the violation components of both sub-conditions need to be calculated separately, and then the maximum value is taken (because exceeding either condition constitutes a violation).

[0073] Condition 1 (systolic blood pressure):

[0074] actual value Historical average Historical standard deviation .

[0075] Violation of quantity

[0076] It should be noted that sub-condition 1 requires a blood pressure of ≥140, while the actual value of 180 > 140, which is clearly a violation. However, this embodiment calculates the deviation from the historical mean. In reality, if the blood pressure value is higher than 140, it is considered a violation, but here we are calculating the standardized degree of the deviation. However, for sub-condition 1 (systolic blood pressure ≥140), this embodiment focuses on the portion exceeding 140. But the rule uses historical statistical parameters (mean and standard deviation) to calculate the violation component. Therefore, the deviation from the historical mean (standardized absolute deviation) calculated according to this embodiment is used as the violation component. In this way, even if the excess of 140 is not large, if it is far from the historical mean, the violation component will be large.

[0077] Subcondition 2 (diastolic blood pressure):

[0078] actual value Historical average Standard deviation .

[0079] Violation of quantity

[0080] Since condition 1 is an OR relationship between two sub-conditions, the maximum value of the two violated components is taken as the violated component of condition 1:

[0081] Condition 2 of Rule 1 (Drug Code):

[0082] The set of legal drugs is V = {"C02","C03","C07","C08","C09"}, but the actual drug code "R06" is not included.

[0083] Violation of quantity (Discrete features, those not in the valid set are anomalies 1)

[0084] Condition 3 of Rule 1 (dosage):

[0085] The required dosage is between 0.1 mg / kg and 0.3 mg / kg. For a patient weighing 70 kg, the reasonable dosage range is 7 mg to 21 mg. The actual dose of 10 mg is within this range, therefore it violates the dosage rule. .

[0086] If the dose exceeds the range, the proportion of the excess to the range size is calculated. For example, if the dose is 25mg, the excess of 21mg is 4mg, and the range width is 14mg (21-7), so the violation component can be calculated as 4 / 14≈0.286. However, in this case, there is no excess, so the violation component is 0.

[0087] Therefore, the penalty for violating audit rule 1 is... Calculation (weighted sum) of the degree of violation of audit rule 1:

[0088]

[0089] Audit Rule 2 (Blood Pressure Safety Rules for the Elderly):

[0090] Condition 1 (Continuous characteristic): Systolic blood pressure ≤ 160 (Safe upper limit of systolic blood pressure for the elderly)

[0091] Condition 2 (Continuous characteristic): Diastolic blood pressure ≤ 100 (Safe upper limit of diastolic blood pressure for the elderly)

[0092] Weights: w2 = [0.6, 0.4]

[0093] Calculate the violation component of rule 2

[0094] Condition 1 of Audit Rule 2 (Safe Upper Limit for Systolic Blood Pressure): Threshold 160, Reference Standard Deviation The actual systolic blood pressure is 180 > 160, therefore it violates the component law. .

[0095] Condition 2 of Audit Rule 2 (Safe Upper Limit for Diastolic Blood Pressure): Threshold 100, Reference Standard Deviation The actual diastolic blood pressure was 110 > 100, violating the component. .

[0096] The violation weight of audit rule 2 is

[0097] Calculation of violation of Rule 2 (weighted sum):

[0098]

[0099] Next, calculate the maximum rule violation severity value and compare the rule violation severity values ​​of the two rules: Rule 1 rule violation severity value: Rule 2 violation severity value: .

[0100] Maximum rule violation value .

[0101] In this embodiment, for continuous features, if the rule condition is exceeding a certain threshold T (or falling below a certain threshold), and a reference standard deviation σ is used, the violation component can be calculated using the following formula:

[0102] γ = max(0,(actual value-T) / σ) [When the condition exceeds the threshold, as in rule 2]

[0103] γ = max(0,(T-actual value) / σ) [When the condition is below the threshold]

[0104] It should be noted that Rule 1 uses historical mean and historical standard deviation to calculate absolute deviation. In fact, the condition for Rule 1 is a systolic blood pressure greater than or equal to 140 or a diastolic blood pressure greater than or equal to 90. However, when calculating the violation component, 140 and 90 are not used as thresholds; instead, the standardized distance is calculated centered on the historical mean. This is because the violation component calculation in Rule 1 is based on historical distribution, measuring the degree to which the current value deviates from the historical mean. Rule 2, on the other hand, is a safety rule and can directly use a fixed threshold. Therefore, the calculation method for the violation component depends on the definition of the rule. This embodiment can unify the calculation method for the violation component as follows:

[0105] If rule condition j requires feature k to be within a certain normal range (based on historical statistics), then the violation component... ,in, This refers to the actual or measured value. and These are the historical mean and historical standard deviation of feature k, respectively; if rule condition j requires feature k to not exceed (or not be lower than) threshold T, then the violation component... If the value is not lower than the threshold T, then the component violates the rule. , where σ is a reference standard deviation (which can be the historical standard deviation of the feature or a preset scaling factor).

[0106] For discrete features ( (If the condition is not met), that is, 1 indicates that the condition is not met, and 0 indicates that the condition is met.

[0107] In an optional embodiment, before calculating the degree of violation of the original medical data under each audit rule by weighting the violation components corresponding to each condition under each audit rule, the method further includes: obtaining the gradient values ​​of all hospital nodes for each condition under each audit rule; averaging the gradient values ​​of all hospital nodes for each condition under each audit rule to obtain the final gradient value; and determining the weight value of each condition under each audit rule based on the final gradient value.

[0108] The gradient value indicates the direction of adjustment for each weight (whether it should be increased or decreased) and approximately by how much. The goal of weight optimization is to minimize the overall violation error, and the gradient direction indicates which weight adjustment will most effectively reduce the error.

[0109] In this embodiment, for anomalous data, the rule with the highest violation degree is identified, and then the violation component (γ value) of which condition (such as "blood pressure value" or "drug code") contributes the most is determined. The greater the contribution of the condition, the stronger the instruction to increase its weight (positive gradient) is. For normal data, the rule with the lowest violation degree is identified, and then the condition that is actually correct but almost caused a misjudgment is determined. The weights of these conditions will receive an instruction to decrease (negative gradient).

[0110] It's important to note that the computation of a single hospital node may be biased or noisy. Therefore, all hospital nodes participating in federated learning only encrypt their calculated gradient values ​​(specifically directional gradients, not the original data) and upload them to a central coordinator. The central coordinator then aggregates the gradient values ​​from all hospitals. If most hospitals indicate "increase the weight of the blood pressure condition," the coordinator will generate a strong consensus instruction: "Increase the weight of the blood pressure condition overall." This consensus gradient value will then be distributed to the rule engines of all hospitals.

[0111] S1022, obtain the maximum violation degree value among the violation degree values; and determine whether the maximum violation degree value is greater than the adaptive threshold.

[0112] The adaptation threshold can be set according to actual needs, and can also be updated and iterated.

[0113] S1023, if the value is greater than the adaptive threshold, then the audit label of the original medical data is determined to be abnormal.

[0114] S1024, if it is less than or equal to the adaptive threshold, then the audit label of the original medical data is determined to be normal.

[0115] S103, based on the audit tag, store the index position, audit tag, data hash value, timestamp, and node digital signature of the original medical data into the public or private chain of the blockchain.

[0116] The public blockchain stores raw medical data with normal audit tags, while the private blockchain stores raw medical data with abnormal audit tags.

[0117] In one optional embodiment provided in this application, storing the index position, audit tag, data hash value, timestamp, and node digital signature corresponding to the original medical data in a public or private blockchain based on the audit tag includes: storing the index position, audit tag, data hash value, timestamp, and node digital signature corresponding to the original medical data with a normal audit tag in a public blockchain; and storing the index position, encrypted original medical data, audit tag, data hash value, timestamp, and node digital signature corresponding to the original medical data with an abnormal audit tag in a private blockchain. The index position indicates the origin of the original medical data, such as which hospital node it comes from.

[0118] Among them, the data hash value, as the unique digital fingerprint of the data, can be the SHA-256 digest of the original data, bound and stored with the audit tag to establish data identification; the timestamp serves as the time base for establishing the audit event, used to record the exact time the data was audited, providing a time dimension coordinate for subsequent anomaly pattern analysis; and the node digital signature serves as the identity signature of the hospital node.

[0119] S104: Obtain the corresponding abnormal original medical data through the index position stored in the private chain of the blockchain, and determine the data abnormality type of the abnormal original medical data corresponding to the index position.

[0120] It should be noted that the encrypted feature vectors or computable ciphertext stored on the private blockchain provide authorized nodes with the data needed for analysis, while cryptographic techniques ensure that the original data is not directly exposed. In contrast, the characteristic of a public blockchain is that data is visible to all participants. Therefore, this embodiment retrieves the corresponding abnormal original medical data through the index position stored in the private blockchain within the main blockchain, and determines the data anomaly type of the abnormal original medical data corresponding to the index position.

[0121] In one optional embodiment provided in this application, determining the data anomaly type of the original medical data corresponding to the index position includes: determining the abnormal data group corresponding to the abnormal original medical data corresponding to the index position, wherein the abnormal data includes feature values, rule identifiers, rule violation degree values, operation user identifiers, timestamps, and device identifiers; identifying the abnormal data group based on a predefined medical error knowledge base to obtain the probability values ​​corresponding to each data anomaly type; and determining the data anomaly type of the abnormal original medical data through the probability values ​​corresponding to each data anomaly type.

[0122] The abnormal data group consists of [feature value, rule identifier, rule violation degree value, operator identifier, timestamp, and device identifier]. This embodiment predefines a medical error knowledge base that can be divided into three dimensions to identify the types of data anomalies in the original medical data. These three dimensions are: planning dimension, numerical dimension, and environmental dimension. The planning dimension can involve resolving the clinical meaning of the rule identifier (e.g., R007 = severe hypertension threshold rule) and statistically analyzing the distribution of frequently triggered rules (e.g., 80% of anomalies trigger the same rule). The numerical dimension analyzes the distribution of feature values ​​(e.g., systolic blood pressure concentrated in the 190±5 range) and detects anomaly clusters (multiple similar anomalies within the same time period). The environmental dimension associates the device model (e.g., all anomalies originate from the same model of blood pressure monitor) and checks operation records (e.g., anomalies are concentrated in a specific operator shift).

[0123] In this embodiment, if a feature value is abnormal and remains consistently high, a device malfunction can be identified; if a feature value is abnormal and contains discrete abnormal values, an input error can be identified; if a single rule identifier is triggered frequently, an audit rule is identified as outdated; if multiple rules are triggered in association, a data logic error is identified; if the same device identifier exhibits multiple anomalies, a device malfunction is identified; if the same user identifier exhibits multiple anomalies, human error is identified.

[0124] For example, if multiple hospitals report abnormal blood pressure data, and analysis of the original medical data reveals that systolic blood pressure >180 and diastolic blood pressure >110 are concentrated between 9:00 and 10:00 every day, and all of them are using Type A blood pressure monitors, then the data anomaly type can be determined to be a device calibration failure (morning first use deviation) through a predefined medical error knowledge base.

[0125] In another optional embodiment provided in this application, the data anomaly type of the original medical data corresponding to the index position includes:

[0126] S1041, determine the abnormal data group corresponding to the original medical data of the abnormality at the index position.

[0127] The abnormal data includes feature values, rule identifiers, rule violation severity values, user identifiers, timestamps, and device identifiers.

[0128] S1042, Obtain abnormal data groups within a predetermined time range, and arrange the abnormal data groups within the predetermined time range in chronological order.

[0129] S1043, convert the abnormal data groups arranged in chronological order within a predetermined time range into a feature vector matrix, wherein each row of the feature vector matrix represents an abnormal data group corresponding to a timestamp.

[0130] S1044, Determine the data anomaly type of the original medical data corresponding to the index position based on the feature vector matrix.

[0131] The step of determining the data anomaly type of the original medical data corresponding to the index position based on the feature vector matrix includes: inputting the feature vector matrix into a data anomaly prediction model to obtain the data anomaly type of the original medical data corresponding to the index position.

[0132] S105, determine an error correction scheme based on the data anomaly type and the corresponding abnormal original medical data; and correct the abnormal original medical data using the error correction scheme.

[0133] In this embodiment, error correction schemes can be determined through an error correction scheme template knowledge base, in which each anomaly type corresponds to one or more standardized solution templates. For example, the equipment failure template library may contain: equipment immediate calibration template, equipment shutdown and replacement template, and equipment scheduling and usage template.

[0134] This embodiment uses the original abnormal data as parameters to populate a selected template, generating a specific and executable solution. Key parameters typically include: the faulty device number (extracted from the data), the erroneous measurement value (e.g., 170), a reasonable value range based on patient history or medical standards, and the time window for planned error correction operations. By matching the error correction solution template knowledge base, a complete, machine-readable error correction solution is generated.

[0135] This embodiment provides a method for dynamic auditing and error correction of data compliance in a hospital big data laboratory. First, it acquires raw medical data transmitted from hospital nodes; this raw medical data includes at least electronic medical records, test reports, and medication records. Then, it audits the raw medical data using an audit rule set to obtain audit tags for each piece of raw medical data, including normal and abnormal tags. Next, based on the audit tags, it stores the corresponding index position, audit tag, data hash value, timestamp, and node digital signature of the raw medical data in a public or private blockchain. The public blockchain stores raw medical data with normal audit tags, while the private blockchain stores raw medical data with abnormal audit tags. It retrieves the corresponding abnormal raw medical data based on the index position stored in the private blockchain and determines the data anomaly type of the abnormal raw medical data corresponding to the index position. Finally, it determines an error correction scheme based on the data anomaly type and the corresponding abnormal raw medical data, and corrects the abnormal raw medical data using the error correction scheme. Compared to existing technologies where remote interactive data is manually audited by staff, this application allows hospital nodes to audit raw medical data according to a set of audit rules after uploading the data. The audit results are then stored on the blockchain, and the types of data anomalies are determined based on the private blockchain. This approach improves audit efficiency and enhances data security.

[0136] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0137] In one embodiment, a dynamic audit and error correction system for data compliance in a hospital big data laboratory is provided. For example... Figure 2 As shown, the detailed descriptions of each functional module of the hospital's big data laboratory data compliance dynamic audit and error correction system are as follows:

[0138] The acquisition module 21 is used to acquire the raw medical data transmitted by the hospital node; the raw medical data includes at least electronic medical records, test reports, and drug records.

[0139] Audit module 22 is used to audit the original medical data through an audit rule set to obtain an audit label corresponding to each piece of original medical data, wherein the audit label includes normal and abnormal;

[0140] Storage module 23 is used to store the index position, audit tag, data hash value, timestamp, and node digital signature corresponding to the original medical data into a public or private chain in the blockchain according to the audit tag; the public chain stores the original medical data with normal audit tags, and the private chain stores the original medical data with abnormal audit tags;

[0141] The determination module 24 is used to obtain the corresponding abnormal original medical data through the index position stored in the private chain of the blockchain, and determine the data abnormality type of the abnormal original medical data corresponding to the index position.

[0142] The error correction module 25 is used to determine an error correction scheme based on the data anomaly type and the corresponding abnormal original medical data; and to correct the abnormal original medical data using the error correction scheme.

[0143] In an optional embodiment, the storage module 23 is specifically used for:

[0144] The index location, audit tag, data hash value, timestamp, and node digital signature corresponding to the original medical data with the audit tag being normal are stored in the public chain of the blockchain;

[0145] The index location, audit tag, data hash value, timestamp, and node digital signature corresponding to the original medical data with the audit tag being abnormal are stored in the private chain of the blockchain.

[0146] In an optional embodiment, audit module 22 is specifically used for:

[0147] Calculate the rule violation severity value of the original medical data according to each audit rule in the audit rule set;

[0148] Obtain the maximum violation severity value among the violation severity values; and determine whether the maximum violation severity value is greater than an adaptive threshold.

[0149] If the value is greater than the adaptive threshold, the audit label of the original medical data is determined to be abnormal;

[0150] If the value is less than or equal to the adaptive threshold, the audit label of the original medical data is determined to be normal.

[0151] In an optional embodiment, audit module 22 is specifically used for:

[0152] Obtain the historical medical data corresponding to the original medical data;

[0153] The historical standard deviation and historical mean were determined based on the historical medical data.

[0154] The degree of violation of the original medical data under each audit rule is calculated using the original medical data, historical standard deviation, and historical mean.

[0155] In an optional embodiment, audit module 22 is specifically used for:

[0156] Based on the original medical data, historical standard deviation, and historical mean, calculate the violation component corresponding to each condition under each audit rule;

[0157] The violation severity value of the original medical data under each audit rule is obtained by weighting the violation components corresponding to each condition under each audit rule.

[0158] In an optional embodiment, the acquisition module 21 is further configured to:

[0159] Obtain the gradient values ​​of all hospital nodes for each condition under each audit rule;

[0160] The final gradient value is obtained by averaging the gradient values ​​of each condition under each audit rule for all hospital nodes.

[0161] The weight value of each condition under each audit rule is determined based on the final gradient value.

[0162] In an optional embodiment, the determining module 24 is specifically used for:

[0163] Identify the abnormal data group corresponding to the abnormal original medical data at the index position. The abnormal data includes feature values, rule identifiers, rule violation degree values, operation user identifiers, timestamps, and device identifiers.

[0164] Based on a predefined medical error knowledge base, the abnormal data groups are identified to obtain the probability values ​​corresponding to each data abnormality type;

[0165] The data anomaly type of the original medical data is determined by the probability value corresponding to each data anomaly type.

[0166] In an optional embodiment, the determining module 24 is specifically used for:

[0167] Identify the abnormal data group corresponding to the abnormal original medical data at the index position. The abnormal data includes feature values, rule identifiers, rule violation degree values, operation user identifiers, timestamps, and device identifiers.

[0168] Get the abnormal data groups within a predetermined time range and arrange them in chronological order.

[0169] The abnormal data groups arranged in chronological order within a predetermined time range are converted into a feature vector matrix, where each row of the feature vector matrix represents an abnormal data group corresponding to a timestamp.

[0170] The data anomaly type of the original medical data corresponding to the index position is determined based on the feature vector matrix.

[0171] In an optional embodiment, the determining module 24 is specifically used for:

[0172] The feature vector matrix is ​​input into the data anomaly prediction model to obtain the data anomaly type of the original medical data corresponding to the index position.

[0173] It should be noted that the above detailed descriptions are exemplary and intended to provide further explanation of this application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains.

[0174] Specific limitations regarding the dynamic audit and correction system for data compliance in hospital big data laboratories can be found in the limitations of the dynamic audit and correction methods for data compliance in hospital big data laboratories mentioned above, and will not be repeated here. Each module in the aforementioned equipment can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in the computer device in hardware form, or stored in the memory of the computer device in software form, so that the processor can call and execute the corresponding operations of each module.

[0175] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the system can be divided into different functional units or modules to complete all or part of the functions described above.

[0176] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A method for dynamic auditing and error correction of data compliance in a hospital big data laboratory, characterized in that, The method includes: Acquire raw medical data transmitted from hospital nodes; the raw medical data includes at least electronic medical records, test reports, and medication records; The original medical data is audited using an audit rule set to obtain an audit label for each piece of original medical data. The audit labels include normal and abnormal. Based on the audit tag, the index position, audit tag, data hash value, timestamp, and node digital signature of the original medical data are stored in the public or private chain of the blockchain; the public chain stores the original medical data with normal audit tags, and the private chain stores the original medical data with abnormal audit tags; The corresponding abnormal original medical data is obtained by retrieving the index position stored in the private chain of the blockchain, and the data abnormality type of the abnormal original medical data corresponding to the index position is determined. An error correction scheme is determined based on the data anomaly type and the corresponding original medical data; and the original medical data with the anomaly is corrected using the error correction scheme.

2. The method according to claim 1, characterized in that, The step of storing the index position, audit tag, data hash value, timestamp, and node digital signature of the original medical data into a public or private blockchain based on the audit tag includes: The index location, audit tag, data hash value, timestamp, and node digital signature corresponding to the original medical data with the audit tag being normal are stored in the public chain of the blockchain; The index location, audit tag, data hash value, timestamp, and node digital signature corresponding to the original medical data with the audit tag being abnormal are stored in the private chain of the blockchain.

3. The method according to claim 1, characterized in that, The process of auditing the original medical data using an audit rule set to obtain audit tags corresponding to each piece of original medical data includes: Calculate the rule violation severity value of the original medical data according to each audit rule in the audit rule set; Obtain the maximum violation severity value among the violation severity values; and determine whether the maximum violation severity value is greater than an adaptive threshold. If the value is greater than the adaptive threshold, the audit label of the original medical data is determined to be abnormal; If the value is less than or equal to the adaptive threshold, the audit label of the original medical data is determined to be normal.

4. The method according to claim 3, characterized in that, The step of calculating the rule violation severity value of the original medical data according to each audit rule in the audit rule set includes: Obtain the historical medical data corresponding to the original medical data; The historical standard deviation and historical mean were determined based on the historical medical data. The degree of violation of the original medical data under each audit rule is calculated using the original medical data, historical standard deviation, and historical mean.

5. The method according to claim 4, characterized in that, The determination of the degree of violation of the original medical data under each audit rule, calculated using the original medical data, historical standard deviation, and historical mean, includes: Based on the original medical data, historical standard deviation, and historical mean, calculate the violation component corresponding to each condition under each audit rule; The violation severity value of the original medical data under each audit rule is obtained by weighting the violation components corresponding to each condition under each audit rule.

6. The method according to claim 4, characterized in that, Before calculating the degree of violation of the original medical data under each audit rule by weighting the violation components corresponding to each condition under each audit rule, the method further includes: Obtain the gradient values ​​of all hospital nodes for each condition under each audit rule; The final gradient value is obtained by averaging the gradient values ​​of each condition under each audit rule for all hospital nodes. The weight value of each condition under each audit rule is determined based on the final gradient value.

7. The method according to any one of claims 1-6, characterized in that, The data anomaly types of the original medical data corresponding to the determined index position include: Identify the abnormal data group corresponding to the abnormal original medical data at the index position. The abnormal data includes feature values, rule identifiers, rule violation degree values, operation user identifiers, timestamps, and device identifiers. Based on a predefined medical error knowledge base, the abnormal data groups are identified to obtain the probability values ​​corresponding to each data abnormality type; The data anomaly type of the original medical data is determined by the probability value corresponding to each data anomaly type.

8. The method according to any one of claims 1-6, characterized in that, The data anomaly types of the original medical data corresponding to the determined index position include: Identify the abnormal data group corresponding to the abnormal original medical data at the index position. The abnormal data includes feature values, rule identifiers, rule violation degree values, operation user identifiers, timestamps, and device identifiers. Get the abnormal data groups within a predetermined time range and arrange them in chronological order. The abnormal data groups arranged in chronological order within a predetermined time range are converted into a feature vector matrix, where each row of the feature vector matrix represents an abnormal data group corresponding to a timestamp. The data anomaly type of the original medical data corresponding to the index position is determined based on the feature vector matrix.

9. The method according to claim 8, characterized in that, The process of determining the data anomaly type of the original medical data corresponding to the index position based on the feature vector matrix includes: The feature vector matrix is ​​input into the data anomaly prediction model to obtain the data anomaly type of the original medical data corresponding to the index position.

10. A dynamic auditing and error correction system for data compliance in a hospital big data laboratory, characterized in that, The system includes: The acquisition module is used to acquire raw medical data transmitted by the hospital node; the raw medical data includes at least electronic medical records, test reports, and medication records. The audit module is used to audit the raw medical data using an audit rule set to obtain an audit tag corresponding to each piece of raw medical data. The audit tags include normal and abnormal. The storage module is used to store the index position, audit tag, data hash value, timestamp, and node digital signature corresponding to the original medical data into a public or private chain in the blockchain according to the audit tag; the public chain stores the original medical data with normal audit tags, and the private chain stores the original medical data with abnormal audit tags; The determination module is used to obtain the corresponding abnormal original medical data through the index position stored in the private chain of the blockchain, and determine the data abnormality type of the abnormal original medical data corresponding to the index position. The error correction module is used to determine an error correction scheme based on the data anomaly type and the corresponding abnormal original medical data; and to correct the abnormal original medical data using the error correction scheme.

Citation Information

Cited By

  • Medical data asset whole life cycle dynamic management and compliance auditing method and system

    CN122619399A