Medical data cleaning method based on AI large model

By building a medical data cleaning method based on AI large models, the problems of inconsistent medical data quality and format diversity have been solved, efficient and accurate data cleaning and format conversion have been achieved, and cross-institutional data consistency and emergency decision-making response speed have been improved.

CN120809046AActive Publication Date: 2025-10-17SICHUAN RUIFU ZHIJIAN MEDICAL TECH CO LTD

Patent Information

Application Number
CN202511250972.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-03
Publication Date
2025-10-17
Estimated Expiration
2045-09-03

AI Technical Summary

Technical Problem

Existing technologies in medical data processing have problems with uneven data quality and diverse formats, resulting in low data sharing and utilization efficiency. In addition, the application of large models in medical data cleaning and format conversion has not yet formed a mature solution.

Method used

By building a medical data cleaning method based on AI big models, including multimodal data acquisition, scenario analysis, time coordinate system calibration, broken node management and cross-institutional scenario migration, efficient and accurate data cleaning and format conversion can be achieved.

Benefits of technology

Significantly reduce the clinical misdiagnosis rate, improve cross-institutional terminology consistency, ensure the continuity and integrity of long-term health data, and improve the response speed of emergency decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120809046A_ABST
    Figure CN120809046A_ABST
Patent Text Reader

Abstract

The invention discloses a medical data cleaning method based on an AI large model, and relates to the technical field of medical data processing. Obtaining multi-modal data, obtaining source identification feature data through a source identification module, inputting the source identification feature data to a scene analysis module, and outputting a clinical feature set; constructing a time coordinate system through a scene-driven strategy, obtaining a continuity index based on the time coordinate system, executing dynamic calibration, and outputting first cleaning data containing a calibration data stream and a clinical feature set; when the continuity index is less than 95%, marking as a fracture node, performing segmented baseline management on the fracture node, generating clinical decision features, performing format conversion on the calibration data stream, and outputting second cleaning data; and inputting the second cleaning data into a clinical situation migration body, executing cross-mechanism situation migration and semantic reconstruction, and outputting third cleaning data.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of medical data processing, and in particular to a medical data cleaning method based on an AI large model. BACKGROUND

[0002] With the rapid development of medical informatization, medical data has shown explosive growth. However, these data come from different medical institutions and systems, have various formats, including electronic medical records, test reports, image data, etc., and contain a large amount of noise, redundancy and error information, and the data quality is uneven.

[0003] Traditional data cleaning methods mainly rely on manual rules and simple algorithms, and the processing effect is not good for complex medical data, and the efficiency is low. For example, when processing the free text in the electronic medical record, it is difficult for manual rules to cover all expression methods, resulting in many errors and noise data that cannot be identified and processed.

[0004] In terms of format conversion, the data formats between different systems differ greatly, and the traditional conversion method needs to write a large number of conversion rules for different formats, which has poor adaptability and high maintenance cost. When new data sources or data formats appear, conversion rules need to be rewritten, which seriously affects the sharing and utilization efficiency of medical data.

[0005] Although large models have achieved remarkable results in natural language processing and other fields, their application in medical data cleaning and format conversion gateway is still in the exploratory stage, and mature solutions have not yet been formed. The powerful learning and processing capabilities of large models provide a possibility for solving the problem of medical data processing, but how to effectively integrate them into the gateway system to achieve efficient and accurate data cleaning and format conversion is a problem that needs to be solved at present. SUMMARY

[0006] The present application provides a medical data cleaning method based on an AI large model, which solves the core problems of missing medical decision logic, cross-institutional data discontinuity and time continuity destruction in the prior art, and achieves the technical effects of significantly reducing clinical miswashing rate, greatly improving cross-institutional terminology consistency, and achieving industry-leading long-term health data continuity and completeness.

[0007] The present application provides a medical data cleaning method based on an AI large model, which includes: S1: Obtain multi-modal data, obtain source identification feature data through a source identification module, input the source identification feature data into a scene analysis module, and output a clinical feature set through dynamic expansion of feature dimension, symptom entity recognition and original feature retention; S2: constructing a time coordinate system through a scene-driven strategy, obtaining a continuity index based on the time coordinate system and performing dynamic calibration, and outputting first cleaning data containing calibration data stream and clinical feature set; S3: when the continuity index is less than 95%, marking as a broken node, performing segmented baseline management on the broken node, generating clinical decision features, and outputting second cleaning data after performing format conversion on the calibration data stream; S4: based on the scene-driven strategy, the time coordinate system, the clinical feature set and the clinical decision features, jointly constructing a clinical situation migration body, inputting the second cleaning data into the clinical situation migration body, performing cross-institutional situation migration and semantic reconstruction, and outputting third cleaning data.

[0008] Further, the multi-modal data includes: electronic medical record text data, vital sign monitoring time series data, image data and metadata, structured test data, device source data, clinical decision support data and spatiotemporal index data; The electronic medical record text data includes free text of patient complaints, medical history records and diagnosis conclusions; The vital sign monitoring time series data includes electrocardiogram waveform and blood oxygen saturation flow data; The image data and metadata include CT format files and their metadata; The structured test data is laboratory numerical values split by test items; The device source data is a device type identifier, which is used to call a device feature library to perform signal optimization; The clinical decision support data includes a standardized clinical terminology library, a hospital historical terminology frequency library, and a symptom test index logic chain library; The spatiotemporal index data includes hierarchical labels and audit tracking logs.

[0009] Further, the source identification feature data includes: scene attribute identification, professional attribute identification, device attribute identification and transmission protocol identification; The scene attribute identification is a Boolean emergency flag, which is used to mark whether the data belongs to an emergency scene; The professional attribute identification is a specialty code; The device attribute identification is a device type, which is used to associate signal optimization parameters of a device feature library; The transmission protocol identification is a data access protocol type, which is used to define data parsing rules.

[0010] Further, the analysis of clinical features includes: dynamic expansion of feature dimension, symptom entity recognition and original feature retention; The dynamically expanded feature dimension is a specialty code that adds a pathological grading offset feature; The symptom entity recognition is a keyword library scanning free text in emergency scenarios; The original feature retention is to reject early standardization and maintain device non-converted values.

[0011] Further, the scene-driven strategy includes processing channel strategy, calculation accuracy strategy, resource allocation strategy, and downstream control strategy. The processing channel strategy is to activate the timeout processing channel when in an emergency scenario and identify a symptom entity. The calculation accuracy strategy is to enable the accuracy processing mode of time series analysis when the specialty code is oncology and the research project number is detected. The resource allocation strategy generates a three-level scene priority label, including first priority, second priority, and third priority. The first priority triggers the calculation resource forced preemption mechanism in the emergency scenario. The second priority allocates a large-capacity memory cache area in the research scenario. The third priority enables the batch processing mechanism in the regular scenario. The downstream control strategy is to pass the priority label to the space-time coordination module to control the tolerance window range. The tolerance window is a time alignment fault tolerance interval in the time coordinate system.

[0012] Further, the construction of the time coordinate system includes: taking the R-wave peak as the electrocardiogram reference point, and combining the device time difference feature vector for dynamic delay compensation: , wherein, is the corrected timestamp, is the original timestamp, is the device inherent delay read from the device feature library, is the network transmission delay measured through the time synchronization protocol; based on the compensated time reference, the priority label passed by the downstream control strategy of the scene-driven strategy is received, the tolerance window range of the multi-modal synchronization window is set; a window feedback mechanism is established, when a single data exceeds the limit, a primary response is triggered to perform interpolation compensation and mark temporal anomalies; three consecutive overruns trigger a secondary response, recalibrate the device time difference vector, update the tolerance window range and generate a temporal anomaly report.

[0013] Further, the first cleaning data includes calibration data flow and clinical feature set; The calibration data flow is a multi-modal data set time-aligned by the tolerance window, including text symptom description, signal-optimized vital sign data, and abnormal marker data flow synchronized on the time axis; The clinical feature set is a set of unstandardized clinical features extracted from multi-modal data, including raw symptom descriptions, unconverted test values, and device native parameters.

[0014] Further, the second cleaning data includes a format conversion data stream and a standardized clinical feature set. The format conversion data stream includes a clinical pathway binding field, privacy level inheritance data, and a timestamp calibration stream. The standardized clinical feature set includes term standardized output, unit standardized numerical values, and contradiction resolution markers.

[0015] Further, the third cleaning data performs semantic reconstruction, context inheritance, and transmission optimization through a clinical context migration body. The semantic reconstruction includes: querying a receiver institution's term library, converting the term standardized output of the standardized clinical feature set into receiver standard terms, generating a term mapping table; calculating a time reference offset according to the time zone difference between institutions, updating all timestamps of the timestamp calibration stream; The context inheritance includes: analyzing the receiver's privacy compliance strategy, performing differential desensitization operations on the privacy level inheritance data: generating an audit-enhanced spatiotemporal index, recording term mapping tracks and privacy conversion logs; The transmission optimization includes: in an emergency scenario, prioritizing the transmission of symptom descriptions and vital sign abnormal values, and simplifying term mapping to emergency high-frequency abbreviations; in a research scenario, enabling a full data channel, preserving term history versions and device calibration certificates, and extending time axis accuracy to the microsecond level.

[0016] Further, the segmented baseline management includes: based on historical physiological data within the past six months, generating an exclusive reference curve through Gaussian process regression, marking circadian rhythm characteristics and fluctuation bandwidth; real-time monitoring of the deviation of the data from the reference curve, when the instantaneous deviation exceeds 3 times the standard deviation, determining the drift type through a LSTM model and a disease knowledge graph, including the gradual change of the disease development path of the disease course evolution type and the sudden unrelated fluctuation of the abnormal type.

[0017] Further, the determination of the drift type also includes a dynamic cleaning strategy: the disease course evolution type updates the bandwidth parameters of the reference curve and adds clinical event labels; the abnormal type performs wavelet denoising processing and generates a device quality control alarm; when a device calibration event occurs in a key interval of the reference curve, the parameter update is frozen and a calibration compensation note is added, and finally a cleaning data stream with medical markers is output; the key interval is a time period that meets the key period of disease diagnosis and treatment and the fluctuation rate of physiological parameters exceeds 1.5 times the historical fluctuation value.

[0018] One or more technical solutions provided in the present application have at least the following technical effects or advantages: By constructing a unified time coordinate system, the problem of time and space misalignment of multi-modal medical data is effectively solved, and the synchronization accuracy of key vital sign data such as electrocardiogram and blood oxygen is significantly improved; through the dynamic compensation mechanism driven by device calibration events, different types of calibration events are intelligently distinguished and differentiated compensation processing is performed, and a device traceability report is generated to ensure the continuity and integrity of long-term health monitoring data; through individualized baseline drift management technology, physiological parameter drift types are intelligently identified and targeted cleaning strategies are executed, significantly reducing the risk of data miswashing of chronic disease patients; finally, in the cross-institutional collaboration scene, relying on intelligent terminology mapping and decision chain verification technology, the semantic difference problem between medical institutions is effectively solved, the consistency of referral data and the emergency decision response speed are significantly improved, and a new standard for medical data cleaning covering clinical diagnosis and treatment, scientific research analysis and device management is formed. BRIEF DESCRIPTION OF DRAWINGS

[0019] Figure 1 A medical data cleaning method based on an AI large model in an embodiment of the present application is shown in the flowchart. DETAILED DESCRIPTION

[0020] In order to facilitate the understanding of the present application, the present application will be described more fully below with reference to the related drawings; the preferred embodiments of the present application are shown in the drawings, but the present application can be implemented in many different forms and is not limited to the embodiments described herein; on the contrary, the purpose of providing these embodiments is to make the disclosure of the present application more thorough and comprehensive.

[0021] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present application belongs; the terms used herein in the specification of the present application are only for the purpose of describing the specific embodiments and are not intended to limit the present application; the term "and / or" used herein includes any and all combinations of one or more related listed items.

[0022] Embodiment one: as shown, a medical data cleaning method based on an AI large model. Figure 1

[0023] S1: acquiring multi-modal data, obtaining source identification feature data through a source identification module, inputting the source identification feature data into a scene analysis module, and outputting a clinical feature set through dynamic expansion of feature dimension, symptom entity identification and original feature reservation; The multi-modal data includes electronic medical record text data, vital sign monitoring time series data, image data and metadata, structured test data, device source data, clinical decision support data and space index data; ​Specifically, the electronic medical record text data is obtained from the hospital EMR system in real time through the HL7 / FHIR interface, the patient's complaint, medical history record and the like free text are parsed, and the OCR scan is stripped; the vital sign monitoring time series data is collected from the bedside monitor, the electrocardiogram R wave peak and blood oxygen flow data are extracted, and the transmission delay is compensated; the image data and metadata are obtained from the PACS system, the device model and acquisition parameters in the CT file metadata are parsed; the structured test data is received from the LIS laboratory system, is split into atomic test items and is marked with the original unit value; the device source data is extracted from the device registration library to identify the device type; the clinical decision support data is loaded into the standardized clinical term library, the hospital historical term frequency library and the symptom test index logical chain library through the FHIR term service; and the space-time index data is obtained from the hospital audit system, is bound with a privacy classification label and generates an audit tracking log. All data is standardized through a dynamic protocol parser, and when the scene attribute identification is emergency, a 5-second timeout channel is activated for priority processing, and finally a complete multi-modal data stream supporting clinical decision embedding and cross-institution semantic alignment is formed.

[0024] The source identification feature data includes scene attribute identification, professional attribute identification, device attribute identification and transmission protocol identification; Specifically, the scene attribute identification is a Boolean (True / False) emergency flag, which is automatically set to True when the data contains an emergency triage label or continuously appears abnormal values of vital signs, and triggers a 5-second timeout processing channel and a computing resource forced preemption mechanism; the professional attribute identification maps the specialty coding system, matches the specialty type by parsing the electronic medical record diagnosis code or the symptom keyword, drives the high-precision processing mode and associates the symptom test index logical chain library; the device attribute identification parses the UDI code in the device source data, matches the device feature library to obtain signal optimization parameters; and the transmission protocol identification defines the protocol type enumeration value, automatically loads the parser or enables the protocol conversion layer to reconstruct the semantics according to the protocol header, and guarantees the integrity of the cross-institution context.

[0025] Extracting non-standardized clinical features from the multi-modal data stream, including original text of symptom description, original parameters of device and non-converted test values; dynamically expanding the feature dimension in combination with the professional attribute identification, such as adding pathological grading features in the oncology department scene; and adding the original measurement value of ST segment deviation in the electrocardiogram in the cardiovascular emergency.

[0026] Outputting the clinical feature set through dynamic expansion of the feature dimension, symptom entity recognition and original feature retention, including original symptom description, non-converted test values and device original parameters; When the scene attribute is identified as emergency (True) and the symptom entity is identified, the emergency channel is activated, the non-emergency task thread is forcibly interrupted, and the symptom entity keyword library is called to match the electronic medical record free text in real time; the symptom description free text in the electronic medical record is retained, and early term standardization is rejected to maintain the integrity of the original semantics; the native parameters of the equipment are collected; the original test values of the laboratory are recorded, and the original unit is retained to avoid pre-conversion distortion; and the feature dimensions are dynamically expanded based on the specialty scene, such as adding the pathological grading feature in the oncology scene, and adding the original measurement value of the ST segment offset of the electrocardiogram in the cardiovascular emergency scene, and the clinical feature set is complete to retain the native state of the medical data.

[0027] S2: Construct a time coordinate system through a scene-driven strategy, obtain a continuity index based on the time coordinate system, and perform dynamic calibration, outputting first cleaned data containing calibration data stream and a clinical feature set; The scene-driven strategy includes a processing channel strategy, a calculation accuracy strategy, a resource allocation strategy, and a downstream control strategy. Specifically, the processing channel strategy is divided into an emergency channel and a non-emergency channel. When the scene attribute is identified as emergency (True) and the symptom entity (such as chest pain, dyspnea) is identified, the emergency channel is activated, the non-emergency task thread is forcibly interrupted, and the symptom entity keyword library is called to match the electronic medical record free text in real time. The non-emergency channel enables a batch processing mechanism (60 seconds / batch) and is compatible with low-speed equipment.

[0028] The calculation accuracy strategy is divided into a high-precision mode and a regular mode. When the professional attribute is identified as oncology and the research plan number is detected, 0.1 ms level time sequence analysis is enabled and the tumor special logic chain library is loaded. The regular mode uses 10 ms level time sequence analysis to balance efficiency and accuracy.

[0029] The resource allocation strategy generates a three-level priority label. When the emergency flag is True and the vital signs are abnormal, the first priority is triggered, the CPU and memory of the system are preempted, and the regular task is suspended; in the research scene and there is a long-term tracking label, the second priority is triggered, the memory buffer area is allocated to store the original waveform data; in the regular scene and the device attribute is identified as a low-speed device, the third priority is triggered, batch compression processing (60 seconds / pack) is enabled, and the calculation load is reduced.

[0030] The downstream control strategy transmits the priority label to the space-time coordination module to dynamically adjust the tolerance window threshold; the emergency scene is accelerated to align, and the window is contracted to ±100 ms; the research scene prioritizes data integrity, and the window is expanded to ±500 ms. According to the device attribute identification, the signal optimization library is called, such as enabling airflow waveform smoothing algorithm for ventilator equipment and loading electromyographic interference suppression parameters for electrocardiogram equipment.

[0031] The first cleaned data includes a calibration data stream and a clinical feature set; The calibration data stream is a multimodal data set that is time-aligned with a tolerance window, including timeline-synchronized text symptom descriptions, signal-optimized vital sign data, and abnormality marker data streams; The clinical feature set is a set of unstandardized clinical features extracted from multimodal data, including original symptom descriptions, unconverted test values, and device native parameters.

[0032] Specifically, the data stream is calibrated through dynamic tolerance window alignment, device-specific signal optimization, abnormal marking and traceability output; the clinical feature set data is obtained by inheriting and analyzing clinical features.

[0033] The system receives priority tags from downstream control strategies and dynamically adjusts tolerance window thresholds. The first priority window is reduced to ±100ms to accelerate data alignment; the second priority window is reduced to ±500ms to ensure data integrity; and the third priority window is reduced to ±300ms to balance efficiency and accuracy. Multi-source data is aligned within the tolerance window, binding electronic medical record free text to the ECG R-wave peak timestamp; and ECG waveforms and blood oxygen flow data are synchronized to the millisecond level.

[0034] Every time data flows into the time coordinate system, the continuity index is calculated in real time: , in, is the continuity index, is the total amount of data to be processed in the time coordinate system, To tolerate the amount of valid data aligned within the window, potential break nodes are marked when the index decreases.

[0035] Using the ECG R-wave peak timestamp as the reference point, the electronic medical record's free text is strictly bound to the vital sign data for the corresponding time period within the dynamic tolerance window. The ECG waveform uses the device's signature library parameters and loads the device's proprietary signal optimization algorithm. A 50Hz notch filter is used to eliminate myoelectric interference; a 0.5Hz high-pass filter is used to correct for respiratory motion artifacts. A dynamic smoothing algorithm is used for blood oxygen flow data, with the filter coefficient adjusted in real time based on the intensity of the motion artifact.

[0036] Data outside the tolerance window is flagged as a temporal anomaly and a report is generated, including the device ID, the exceeded value, and recommended actions. When the ECG waveform signal-to-noise ratio falls below 20dB, a low-quality signal is flagged and a device quality control alarm is triggered. Linear interpolation is used to generate virtual data points for the exceeded data, which are then associated with the device ID to trace the cause of the anomaly. Finally, the first cleaned data is output.

[0037] S3: When the continuity index is less than 95%, it is marked as a broken node, segmented baseline management is performed on the broken node, clinical decision features are generated, and the calibration data stream is format converted to output the second cleaned data; Specifically, when the continuity index When the accuracy is <95%, it is automatically marked as a broken node, triggering segmented baseline management: the data stream is divided based on the broken node, the original baseline curve is retained in the front segment, and the Gaussian process regression curve is regenerated in the back segment and the offset is annotated. Decision transformation is performed on the clinical feature set, key clinical entities are extracted from the original symptom description, and the unconverted test values ​​are bound to the symptom entities to form a symptom test indicator logic chain library. The fluctuation bandwidth parameters of the back segment curve are extracted and bound to the symptom test indicator logic chain library to output clinical decision features with segmented labels. The 95% threshold is set based on the clinical data fault probability model: Analyzing the data of 100,000 patients, when the continuity index is ≥95%, the completeness of long-term health data reaches 99.3% (p<0.01); the emergency scenario is dynamically reduced to 90% to adapt to high-noise environments.

[0038] Key clinical entities are extracted from the original symptom descriptions of the clinical feature set, and non-clinical descriptions are stripped away using semantic parsing technology. The symptom test indicator logic chain library is called to match predefined clinical rule chains. If the device attribute is identified as an ECG device, device parameter constraints are added to ensure data validity. Based on the test values ​​and vital sign data, the rule chain trigger conditions are verified and decision-making action instructions are generated. If a broken node exists, a segmented baseline offset annotation is added. If a device calibration event occurs during the broken node period, the baseline curve update is frozen and a calibration compensation annotation is added. The output features are bound to the broken node.

[0039] When a conflicting instruction occurs, such as a blood glucose value >11mmol / L but no history of diabetes in the medical history feature marker, or ST-segment elevation detected when the ECG lead impedance is >20kΩ, a conflict resolution flag is added to the standardized clinical feature set, the current decision tree execution is paused, and the conflicting data packet is pushed to the manual review queue. After manual confirmation, the symptom test indicator logic chain library is updated.

[0040] The second cleaned data includes a format conversion data stream and a standardized clinical feature set; The format conversion data flow includes clinical pathway binding fields, privacy level inheritance data and timestamp calibration flow; The standardized clinical feature set includes standardized output of terminology, standardized values ​​of units, and conflict resolution markers.

[0041] Specifically, based on the decision label output by the clinical decision tree, the multi-modal data in the calibration data stream is bound to the standard clinical path; according to the privacy classification label in the space-time index, the identity information and the information of the medical image are desensitized, and the operation record in the audit log is inherited; all time stamps are converted into UTC standard format, the R-wave peak time of electrocardiogram is taken as the reference point, and the blood oxygen test data time stamp is synchronously offset compensated to realize cross-device alignment.

[0042] The standardized clinical feature set guarantees data consistency through three-level conversion: first, the standardized terminology library in the clinical decision support data is called, the original symptom description is converted into a standard terminology through an AI mapping engine, and the hospital commonly used terminology is preferentially matched according to the historical terminology frequency library. Second, the test values are uniformly converted into the international system of units, the original values are double-labeled, and the device parameters are synchronously standardized. Finally, for the logically conflicting data, a structured contradiction label is added to trigger the manual review process and freeze the related decision branch. For the period data of the broken node, special format conversion is performed: the time stamp calibration stream adds a broken interval label; the privacy level inheritance data forcibly enables the encryption level desensitization; the associated broken node has the highest priority and is upgraded to the highest review queue; and the second cleaned data is output.

[0043] The technical solutions in the embodiments of the application have at least the following technical effects or advantages: The application effectively solves the space-time misalignment problem of multi-modal medical data by constructing a unified time coordinate system, significantly improving the synchronization accuracy of key vital sign data such as electrocardiogram and blood oxygen; through a device calibration event-driven dynamic compensation mechanism, the application intelligently distinguishes abnormal types such as device delay and signal interference and performs differential compensation, generates a device traceability report, and guarantees the continuity and integrity of long-term health monitoring data; through individualized baseline drift management technology, the application generates a dedicated reference curve based on historical physiological data, intelligently identifies course evolution type and abnormal type drift, and performs targeted cleaning strategies to significantly reduce the risk of data miswashing for chronic disease patients; finally, in the cross-institutional collaboration scenario, relying on intelligent terminology mapping and decision chain verification technology, the application effectively solves the semantic difference problem between medical institutions, significantly improves the consistency of referral data and emergency decision response speed, and realizes full-scene medical data cleaning covering clinical diagnosis and treatment, scientific research analysis and device management.

[0044] Embodiment two: in embodiment one, a unified time coordinate system is constructed to effectively solve the space-time misalignment problem of multi-modal medical data, but it cannot solve the terminology difference between different institutions, there are time zone and clock reference differences between cross-institutional data, and it cannot adapt to cross-institutional diversified compliance requirements, this embodiment further optimizes embodiment one.

[0045] S4: Based on the scene-driven strategy, the time coordinate system, the clinical feature set and the clinical decision feature, a clinical context migration body is constructed, the second cleaning data is input into the clinical context migration body, cross-institutional context migration and semantic reconstruction are performed, and the third cleaning data is output.

[0046] Specifically, the processing channel strategy and the resource allocation strategy in the scene-driven strategy, the electrocardiogram R-wave reference time axis and the multi-modal synchronization window parameters constructed by the time coordinate system, the non-standardized features contained in the clinical feature set, and the standardized terms and conflict resolution markers output by the clinical decision feature are used as input elements to construct a clinical context migration body and output the third cleaning data.

[0047] During cross-institutional context migration, first, semantic reconstruction of terms is performed: the standardized clinical feature set is used as input data to query the receiving institution's term library, such as the hospital FHIR standard; if there is a term conflict, a term mapping table is added, the original term is retained and a standard term annotation is added; the semantic reconstruction features containing the term mapping table are output. Second, time reference compensation is performed: the time zone difference between institutions is calculated, the time reference offset is generated, and all timestamps are updated based on the time reference offset to ensure cross-institutional clock alignment. Finally, the privacy level is inherited: cross-institutional privacy policy conversion is performed, such as identity information, the sender's privacy level is community open level, and the receiver requires three-A encryption level, then it is converted to FHIR anonymous ID; the sender's privacy level is community encryption level, and the receiver requires three-A research level, then it is desensitized to ICD code B20; and an audit-enhanced spatiotemporal index containing record privacy conversion logs is output.

[0048] When the emergency flag in the scene strategy layer is True, the data is divided into ≤1MB / packet, and key fields such as symptom description and vital sign abnormal values are placed on top to ensure transmission rate; the term mapping is simplified and only basic desensitization is performed. In the second priority research scene, full data is retained, the original device parameter association calibration certificate and term mapping history version are stored, and during data transmission, data traceability is supported, including device serial number, operator ID field, and microsecond-level time axis accuracy. In other regular scenarios, balance efficiency and integrity are ensured, data is divided into ≤5MB / packet to reduce system load, term mapping uses basic rules and cache acceleration; the privacy layer adopts the default level, uses millisecond-level time axis accuracy, and the synchronization window tolerance is ±300ms.

[0049] The third cleaning data is a cross-institutional compatible data packet generated by the clinical context migration body, including semantic reconstruction features, context inheritance data, and transmission optimization structure. Specifically, the semantic reconstruction feature solves the semantic gap and time misalignment of cross-institutional data through dynamic term mapping table and precise time reference offset. The term mapping preserves the original term field and labels the confidence score while matching the recipient institution's standard term library, solving the expression difference of basic terms; the time reference offset realizes accurate alignment of cross-institutional clocks by calculating the time zone difference and inherent device delay, ensuring millisecond-level synchronization of ECG R-wave timestamps and CT image metadata.

[0050] The hierarchical privacy standardization engine dynamically loads the recipient institution's compliance strategy and performs differential de-identification on identity information: partial masking is used in community hospitals, FHIR anonymous ID is converted in first-class hospitals, and an approval number is added in research institutions; the emergency mode skips non-critical de-identification to ensure timeliness; the audit enhances the spatio-temporal index integration operation log, including term mapping trajectory, privacy conversion record, device traceability information and geographic spatial coordinates, forming a full-cycle tracking chain.

[0051] The emergency mode uses a streaming optimization architecture: the symptom description and abnormal values of vital signs are forced to the top of the transmission, and the term mapping is simplified to high-frequency emergency abbreviations to ensure response speed; the research mode uses a full-quantity data architecture: the term history version is preserved, the device calibration certificate is added, the time axis accuracy is expanded to the microsecond level, and the original waveform data access permission is opened.

[0052] The technical solutions in the embodiments of the application have at least the following technical effects or advantages: The application uses clinical context migration technology to achieve semantic lossless migration of cross-institutional medical data and spatio-temporal continuity guarantee through the synergistic effect of dynamic term mapping table and precise time reference offset; based on the hierarchical privacy standardization engine and the audit-enhanced spatio-temporal index, it dynamically adapts to the compliance requirements of different institutions, improves the privacy conversion efficiency while ensuring compliance; through the streaming transmission optimization of the emergency mode and the full-quantity data architecture of the research mode, combined with the cache acceleration mechanism of the conventional mode, a full-scene adaptive transmission system covering emergency response, research analysis and daily diagnosis and treatment is formed, ultimately realizing the effect of reducing misdiagnosis rate and shortening the access time of new institutions, providing core data support for the hierarchical diagnosis and treatment system.

[0053] In the embodiment one, semantic lossless migration of cross-institutional medical data and spatio-temporal continuity guarantee are achieved, but only clock alignment is realized, and the long-term data discontinuity caused by device calibration is not solved, so the embodiment further optimizes the embodiment two.

[0054] The construction of the time coordinate system includes: taking the R-wave peak as the ECG reference point, and combining the device time difference feature vector for dynamic delay compensation; Specifically, taking the R-wave peak of electrocardiogram as the core time reference, the R-wave position is accurately identified by wavelet transform algorithm, the time difference feature vector in the device source data is analyzed, dynamic compensation is performed, and the device calibration certificate containing device ID, delay value and calibration time is output: , Among them, is the corrected timestamp, is the original timestamp, is the device inherent delay read from the device feature library, is the network transmission delay measured by the time synchronization protocol.

[0055] Based on the compensated time reference, the priority label of the downstream control strategy of the scene driving strategy is received, the tolerance window range of the multi-modal synchronization window is set; when the single data is over-limit, the continuity index decreases by 5%, triggering the primary response: generating a virtual value based on the front and rear two effective data points to perform linear interpolation compensation, marking as temporal anomaly and generating quality control suggestion associated with device ID. When the continuity index is zero due to three consecutive over-limits, triggering the secondary response: calling the device feature library to update the delay parameter, recalibrating the device time difference vector; shrink the current tolerance window by 20%, output the temporal anomaly report containing over-limit value, device ID and repair record.

[0056] The system monitors the device log in real time, accurately identifies the zero calibration (such as blood pressure meter zeroing), sensitivity calibration (such as ECG gain adjustment) and composite calibration event, and generates an event fingerprint containing device ID, calibration type, parameter change value and timestamp. The calibration event triggers segmented baseline management: the data before calibration retains the original baseline curve for long-term trend analysis, and the data after calibration establishes a new baseline and marks the offset value. Finally, the calibration event ID is bound to the affected data segment, the calibration review label is inserted into the clinical decision tree, the device calibration history certificate is synchronized when migrating across institutions, so that the receiver can trace back to the original data before calibration, and eliminate the data gap caused by device calibration.

[0057] The technical solutions in the embodiments of the application have at least the following technical effects or advantages: The application solves the data discontinuity problem caused by medical equipment calibration by constructing a dynamic time benchmark calibration and calibration event continuity guarantee mechanism. Taking the R-wave peak of electrocardiogram as the core time benchmark, combining with the inherent delay of the equipment and the network transmission delay for dynamic compensation, a calibration certificate with equipment ID is generated, and the multi-modal data synchronization error is compressed to milliseconds. Through the tolerance window feedback mechanism, combined with intelligent identification and segmented baseline management of equipment calibration events, long-term data discontinuity caused by equipment calibration is eliminated, and the continuity and integrity of cross-year health data reaches 99.3%. At the same time, the calibration review mark is embedded in the clinical decision tree, and the equipment calibration history certificate is synchronized when migrating across institutions, ensuring that the receiver can trace back to the original data.

[0058] Embodiment four: Embodiment three solves the data discontinuity problem caused by equipment calibration, but does not distinguish the medical significance of physiological parameter changes and does not consider the conflict between calibration timing and critical periods of diagnosis and treatment. This embodiment further optimizes on the basis of embodiment three.

[0059] The segmented baseline management includes: based on the historical physiological data within half a year, generating an exclusive benchmark curve through Gaussian process regression, marking the characteristics of circadian rhythm and fluctuation bandwidth; real-time monitoring of the deviation of data from the benchmark curve, when the instantaneous deviation exceeds 3 times the standard deviation, the drift type is judged through the LSTM model and the disease knowledge graph, including the disease progression type of gradual change consistent with the disease development path and the abnormal type of sudden unrelated fluctuation.

[0060] Specifically, historical physiological data within half a year is collected, aligned by time axis and abnormal values are removed. Square exponential kernel (RBF) is used to capture long-term trends, and periodic kernel (period = 24 hours) is used to reflect circadian rhythm, to generate an exclusive benchmark curve, marking key features including circadian rhythm characteristics and fluctuation bandwidth.

[0061] When the instantaneous deviation of real-time data from the benchmark curve is greater than 3 times the standard deviation, LSTM time series analysis is triggered: input continuous 72-hour data, extract time series features such as slope change and fluctuation frequency, output drift pattern vector; call knowledge graph, use pre-defined medical logic chain for verification, if the drift features match the disease development path, mark as disease progression type representing true disease evolution; if the features are contradictory or trigger negative rules, mark as abnormal type pointing to equipment failure or noise interference.

[0062] The judgment of drift type also includes a dynamic cleaning strategy: the disease progression type updates the bandwidth parameters of the benchmark curve and adds clinical event labels; the abnormal type performs wavelet denoising and generates a device quality control alarm; when the equipment calibration event occurs in the key interval of the benchmark curve, the parameter update is frozen and a calibration compensation note is added, and finally the cleaned data stream with medical markers is output; the key interval is a time period that meets the disease diagnosis and treatment critical period and the physiological parameter fluctuation rate exceeds 1.5 times the historical fluctuation value.

[0063] Specifically, the disease progression type is high priority, the original data is reserved, and a clinical event label is added; the abnormal type is medium priority, wavelet denoising is triggered, and a device quality control work order is generated. When the data is in a critical period of diagnosis and treatment, such as 24 hours after surgery, and the fluctuation rate of physiological parameters exceeds 1.5 times the historical fluctuation value, even if the LSTM detects suspected abnormal type drift, the system freezes parameter update to prevent false washing of real disease signals, and outputs a double-path report for doctors to review.

[0064] When it is determined to be a disease progression type, the system immediately pushes the key label and associated treatment suggestions to the doctor workstation to drive clinical intervention; at the same time, the baseline curve parameters are automatically updated to dynamically adapt to the disease evolution trend. If it is determined to be an abnormal type, a device quality control alarm is automatically sent to the engineer end, a fault mode library is associated to generate a maintenance scheme; at the same time, wavelet denoising processing is performed on the original data, and double-track records of denoised data and repair labels are reserved to ensure that technical problems can be traced back and audited.

[0065] The technical solutions in the embodiments of the present application have at least the following technical effects or advantages: The present application realizes the dual improvement of clinical decision safety and data cleaning accuracy by constructing individualized baseline curve and intelligent drift judgment mechanism. Based on the historical physiological data within half a year, Gaussian process regression is used to generate an exclusive baseline curve, which marks the characteristics of circadian rhythm and fluctuation bandwidth, providing individualized reference standard for real-time monitoring. When the instantaneous deviation of real-time data and baseline curve exceeds 3 times the standard deviation, the system extracts time sequence features through the LSTM model, and cooperates with the disease knowledge graph for medical logic verification: if the features are consistent with the disease development path, it is determined to be a disease progression type; if the features are contradictory or have no pathological correlation, it is determined to be an abnormal type. For the disease progression type, the system updates the baseline curve bandwidth parameters and adds a clinical event label to drive clinical intervention; for the abnormal type, wavelet denoising processing is performed and a device quality control alarm is generated to promote technical optimization. When device calibration occurs in a critical period of diagnosis and treatment and the fluctuation rate of physiological parameters exceeds 1.5 times the historical mean value, the system freezes parameter update, adds calibration compensation annotations, avoids misjudgment of real disease, and outputs a double-baseline comparison report for doctors to review. Finally, the system outputs cleaned data stream with medical labels, which significantly reduces the false washing rate of chronic diseases by 37% and improves the safety of diagnosis and treatment of acute and severe diseases.

[0066] The above only describes the preferred embodiments of the present application and is not intended to limit the present application. Those skilled in the art can make various modifications and changes to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.

Claims

1. A medical data cleaning method based on AI big model, characterized by: include: S1: Acquire multimodal data, obtain source identification feature data through the source identification module, input the source identification feature data into the scenario analysis module, and output the clinical feature set by dynamically expanding feature dimensions, identifying symptom entities, and retaining original features; S2: Construct a time coordinate system through a scenario-driven strategy, obtain a continuity index based on the time coordinate system and perform dynamic calibration, and output the first cleaned data containing the calibration data stream and clinical feature set; S3: When the continuity index is less than 95%, it is marked as a broken node, segmented baseline management is performed on the broken node, clinical decision features are generated, and the calibration data stream is format converted to output the second cleaned data; S4: Based on the scenario-driven strategy, time coordinate system, clinical feature set and clinical decision-making features, a clinical scenario migration body is jointly constructed, the second cleaned data is input into the clinical scenario migration body, cross-institutional scenario migration and semantic reconstruction are performed, and the third cleaned data is output.

2. The medical data cleaning method based on AI big model according to claim 1, characterized in that: The multimodal data includes: electronic medical record text data, vital sign monitoring time series data, imaging data and metadata, structured test data, equipment source data, clinical decision support data and spatiotemporal index data; The electronic medical record text data includes free text of the patient's chief complaint, medical history record and diagnosis conclusion; The vital signs monitoring time series data includes electrocardiogram waveform and blood oxygen saturation streaming data; The image data and metadata include CT format files and metadata; The structured test data are laboratory values ​​split by test items; The device source data is a device type identifier, which is used to call the device feature library to perform signal optimization; The clinical decision support data includes a standardized clinical terminology database, a historical terminology frequency database of the hospital, and a symptom test indicator logic chain database; The spatiotemporal index data includes hierarchical tags and audit trail logs.

3. The medical data cleaning method based on AI big model according to claim 1, characterized in that: The source identification feature data includes: scene attribute identification, professional attribute identification, device attribute identification and transmission protocol identification; The scene attribute identifier is a Boolean emergency flag, which is used to mark whether the data belongs to an emergency scene; The professional attribute identifier is a specialist code; The device attribute identifier is a device type, which is used to associate the signal optimization parameters of the device feature library; The transmission protocol identifier is a data access protocol type, which is used to define data parsing rules.

4. The medical data cleaning method based on AI big model according to claim 1, characterized in that: The scenario-driven strategy includes processing channel strategy, calculation accuracy strategy, resource allocation strategy and downstream control strategy; The processing channel strategy is to activate the timeout processing channel when the patient is in an emergency scenario and a symptom entity is identified; The calculation accuracy strategy is to enable the accuracy processing mode of time series analysis when the specialty code is oncology and the study protocol number is detected; The resource allocation strategy generates three levels of scene priority tags, including first priority, second priority, and third priority; The downstream control strategy is to pass the priority mark to the spatiotemporal collaboration module to control the tolerance window range; The tolerance window is a time alignment error tolerance interval in a time coordinate system.

5. The medical data cleaning method based on AI big model according to claim 1, characterized in that: The construction of the time coordinate system includes: taking the peak value of the R wave as the ECG reference point and combining the device time difference characteristic vector to perform dynamic delay compensation: ,in, is the corrected timestamp, is the original timestamp, The device intrinsic delay is read from the device feature library. The network transmission delay is measured by the time synchronization protocol; based on the compensated time reference, the priority mark transmitted by the downstream control strategy of the scenario-driven strategy is received, and the tolerance window range of the multimodal synchronization window is set; a window feedback mechanism is established, and when a single data exceeds the limit, the primary response is triggered to perform interpolation compensation and mark the temporal anomaly; three consecutive exceeding the limit triggers the secondary response, recalibrates the device time difference vector, updates the tolerance window range and generates a temporal anomaly report.

6. The medical data cleaning method based on AI big model according to claim 1, characterized in that: The first cleaned data includes a calibration data stream and a clinical feature set; The calibration data stream is a multimodal data set that is time-aligned with a tolerance window, including timeline-synchronized text symptom descriptions, signal-optimized vital sign data, and abnormality marker data streams; The clinical feature set is a set of unstandardized clinical features extracted from multimodal data, including original symptom descriptions, unconverted test values, and device native parameters.

7. The medical data cleaning method based on AI big model according to claim 1, characterized in that: The second cleaned data includes a format conversion data stream and a standardized clinical feature set; The format conversion data flow includes clinical pathway binding fields, privacy level inheritance data and timestamp calibration flow; The standardized clinical feature set includes standardized output of terminology, standardized values ​​of units, and conflict resolution markers.

8. The medical data cleaning method based on AI big model according to claim 1, characterized in that: The third cleansed data is subjected to semantic reconstruction, context inheritance and transmission optimization through the clinical context migration body; The semantic reconstruction includes: querying the terminology library of the recipient institution, converting the standardized output of the terminology of the standardized clinical feature set into the recipient standard terminology, and generating a term mapping table; calculating the time base offset according to the time zone difference between the institutions, and updating all timestamps of the timestamp calibration stream; The context inheritance includes: parsing the privacy compliance policy of the recipient, performing differentiated desensitization operations on the privacy level inheritance data; generating audit-enhanced spatiotemporal indexes, and recording term mapping trajectories and privacy conversion logs; The transmission optimization includes: in an emergency scenario, symptom descriptions and abnormal vital signs are transmitted at the top, and terminology is simplified and mapped to high-frequency abbreviations for emergency cases; in a scientific research scenario, the full data channel is enabled, historical versions of terms and equipment calibration certificates are retained, and the timeline accuracy is extended to microseconds.

9. The medical data cleaning method based on AI big model according to claim 1, characterized in that: The segmented baseline management includes: generating a dedicated baseline curve based on historical physiological data within six months through Gaussian process regression, marking circadian rhythm characteristics and fluctuation bandwidth; real-time monitoring of the deviation between the data and the baseline curve. When the instantaneous deviation exceeds 3 times the standard deviation, the drift type is collaboratively judged through the LSTM model and the disease knowledge graph, including the gradual course evolution type that conforms to the disease development path and the abnormal type of sudden unrelated fluctuations.

10. The medical data cleaning method based on AI big model according to claim 9, characterized in that: The determination of drift type also includes a dynamic cleaning strategy: for disease progression type, the baseline curve bandwidth parameter is updated and clinical event labels are added; for abnormal type, wavelet denoising is performed and an equipment quality control alarm is generated; When a device calibration event occurs in the critical interval of the baseline curve, the parameter update is frozen and a calibration compensation annotation is added, and finally a clean data stream with medical tags is output; the critical interval is a period that simultaneously meets the critical period for disease diagnosis and treatment and the fluctuation rate of physiological parameters exceeds 1.5 times the historical fluctuation value.

Citation Information

Patent Citations

  • Medical data analysis method and device, equipment and storage medium

    CN116959733A

  • A system for implementing a method of matrix-digital transformation of a variable data set for generating a situational-strategic product program

    RU2020139765A

Cited By

  • Multi-source heterogeneous data management method and device, electronic equipment and storage medium

    CN121561379A