An AI large model-based medical data cleaning method
By constructing a unified time coordinate system and scenario-driven strategies, and combining AI large models for medical data cleaning, the problems of inconsistent medical data quality and diverse formats have been solved. This has enabled efficient cleaning and format conversion of cross-institutional data, improving the efficiency of data sharing and utilization.
Patent Information
- Application Number
- CN202511250972.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-03
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2045-09-03
AI Technical Summary
Existing technologies for medical data processing suffer from inconsistent data quality and diverse formats, leading to low efficiency in data sharing and utilization. Furthermore, there is no mature solution yet for the application of large models in medical data cleaning and format conversion.
By using an AI-based large model approach, multimodal medical data is acquired, a unified time coordinate system is constructed, dynamic calibration and format conversion are performed, and semantic reconstruction and transmission optimization of cross-institutional data are achieved using scenario-driven strategies and clinical context transfer models.
It significantly improves the semantic consistency of cross-institutional data and the response speed of referral data, reduces the false data cleaning rate, ensures the continuity and integrity of long-term health data, and forms a new standard for medical data cleaning across all scenarios.
Smart Images

Figure CN120809046B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of medical data processing, and in particular to a medical data cleaning method based on an AI large model. BACKGROUND
[0002] With the rapid development of medical informatization, medical data is growing explosively. However, these data come from different medical institutions and systems, have various formats, including electronic medical records, test reports, image data, etc., and contain a large amount of noise, redundancy and error information, and the data quality is uneven.
[0003] Traditional data cleaning methods mainly rely on manual rules and simple algorithms, which have poor processing effect and low efficiency for complex medical data. For example, when processing free text in electronic medical records, manual rules are difficult to cover all expression methods, resulting in many errors and noise data that cannot be identified and processed.
[0004] In terms of format conversion, there are large differences in data formats between different systems. Traditional conversion methods need to write a large number of conversion rules for different formats, which has poor adaptability and high maintenance cost. When new data sources or data formats appear, conversion rules need to be rewritten, which seriously affects the sharing and utilization efficiency of medical data.
[0005] Although large models have achieved remarkable results in natural language processing and other fields, their application in medical data cleaning and format conversion gateway is still in the exploratory stage, and mature solutions have not yet been formed. The powerful learning and processing capabilities of large models provide a possibility for solving the problem of medical data processing, but how to effectively integrate them into the gateway system to achieve efficient and accurate data cleaning and format conversion is a problem that needs to be solved at present. SUMMARY
[0006] The present application provides a medical data cleaning method based on an AI large model, which solves the core problems of missing medical decision logic, cross-institutional data discontinuity and time continuity destruction in the prior art, and achieves the technical effects of significantly reducing clinical miswashing rate, greatly improving cross-institutional terminology consistency, and achieving industry-leading level of long-term health data continuity and integrity.
[0007] The present application provides a medical data cleaning method based on an AI large model, which includes:
[0008] S1: Obtain multi-modal data, obtain source identification feature data through a source identification module, input the source identification feature data into a scene analysis module, and output a clinical feature set through dynamic expansion of feature dimension, symptom entity recognition and original feature retention;
[0009] S2: Constructing time coordinate system by scene-driven strategy, obtaining continuity index based on time coordinate system and performing dynamic calibration, outputting first cleaning data containing calibration data stream and clinical feature set;
[0010] S3: Marking as fracture node when continuity index is less than 95%, performing segmented baseline management on fracture node, generating clinical decision feature, performing format conversion on calibration data stream and outputting second cleaning data;
[0011] S4: Based on scene-driven strategy, time coordinate system, clinical feature set and clinical decision feature, constructing clinical situation migration body, inputting second cleaning data into clinical situation migration body, performing cross-institution situation migration and semantic reconstruction, outputting third cleaning data.
[0012] Further, the multi-modal data includes: electronic medical record text data, vital sign monitoring time series data, image data and metadata, structured test data, device source data, clinical decision support data and space-time index data;
[0013] The electronic medical record text data includes free text of patient complaints, medical history records and diagnosis conclusions;
[0014] The vital sign monitoring time series data includes electrocardiogram waveform and blood oxygen saturation flow data;
[0015] The image data and metadata include CT format files and their metadata;
[0016] The structured test data is laboratory numerical value split by test item;
[0017] The device source data is device type identification, which is used to call device feature library to perform signal optimization;
[0018] The clinical decision support data includes standardized clinical terminology library, hospital historical terminology frequency library and symptom test index logic chain library;
[0019] The space-time index data includes hierarchical labels and audit tracking logs.
[0020] Further, the source identification feature data includes: scene attribute identification, professional attribute identification, device attribute identification and transmission protocol identification;
[0021] The scene attribute identification is a Boolean emergency flag, which is used to mark whether the data belongs to an emergency scene;
[0022] The professional attribute identification is a specialty code;
[0023] The device attribute identification is a device type, which is used to associate the signal optimization parameters of the device feature library;
[0024] The transmission protocol identifier is of a data access protocol type, used to define data parsing rules.
[0025] Further, the analysis of clinical features includes dynamic expansion of feature dimensions, symptom entity recognition, and original feature retention.
[0026] The dynamically expanded feature dimensions are specialty code added pathological grading offset features.
[0027] The symptom entity recognition is a keyword library scanning free text in the emergency scenario.
[0028] The original feature retention is to reject early standardization and maintain device non-conversion numerical values.
[0029] Further, the scene-driven strategy includes processing channel strategy, calculation accuracy strategy, resource allocation strategy, and downstream control strategy.
[0030] The processing channel strategy is to activate the timeout processing channel when in the emergency scenario and identify the symptom entity.
[0031] The calculation accuracy strategy is to enable the accuracy processing mode of time series analysis when the specialty code is oncology and the research protocol number is detected.
[0032] The resource allocation strategy generates a three-level scene priority label, including first priority, second priority, and third priority.
[0033] The first priority triggers the calculation resource forced preemption mechanism in the emergency scenario.
[0034] The second priority allocates a large-capacity memory cache area in the research scenario.
[0035] The third priority enables the batch processing mechanism in the regular scenario.
[0036] The downstream control strategy is to pass the priority label to the space-time coordination module to control the tolerance window range.
[0037] The tolerance window is a time alignment fault tolerance interval in the time coordinate system.
[0038] Further, the construction of the time coordinate system includes taking the R-wave peak as the electrocardiogram reference point and combining the device time difference feature vector for dynamic delay compensation. wherein, is the corrected timestamp, is the original timestamp, is the device inherent delay read from the device feature library, network transmission delay measured by time synchronization protocol; based on the compensated time reference, receiving priority labels of downstream control strategy delivery of scene-driven strategy, setting the tolerance window range of multi-modal synchronization window; establishing a window feedback mechanism, triggering a primary response to perform interpolation compensation and mark temporal anomalies when a single data exceeds the limit; triggering a secondary response for three consecutive times, recalibrating the device time difference vector, updating the tolerance window range and generating a temporal anomaly report.
[0039] Further, the first cleaning data includes a calibration data stream and a clinical feature set;
[0040] The calibration data stream is a multi-modal data set time-aligned by the tolerance window, including text symptom description, signal-optimized vital sign data, and abnormal marker data stream synchronized on the time axis;
[0041] The clinical feature set is a non-standardized clinical feature set extracted from multi-modal data, including raw symptom description, non-converted test value, and device native parameter.
[0042] Further, the second cleaning data includes a format conversion data stream and a standardized clinical feature set;
[0043] The format conversion data stream includes a clinical path binding field, privacy level inheritance data, and a timestamp calibration stream;
[0044] The standardized clinical feature set includes term standardized output, unit standardized numerical value, and contradiction resolution marker.
[0045] Further, the third cleaning data performs semantic reconstruction, context inheritance, and transmission optimization by a clinical context migration body;
[0046] The semantic reconstruction includes: querying the recipient institution's term library, converting the term standardized output of the standardized clinical feature set into the recipient standard term, and generating a term mapping table; calculating the time reference offset according to the inter-institution time zone difference, updating all timestamps of the timestamp calibration stream;
[0047] The context inheritance includes: analyzing the recipient privacy compliance strategy, performing differential desensitization on the privacy level inheritance data: generating an audit-enhanced spatiotemporal index, recording the term mapping trajectory and privacy conversion log;
[0048] The transmission optimization includes: in the emergency scene, prioritizing the transmission of symptom description and vital sign abnormal values, and simplifying the term mapping to emergency high-frequency abbreviations; in the research scene, enable full data channel, retain term history version and device calibration certificate, and extend time axis accuracy to microsecond level.
[0049] Further, the segment baseline management includes: generating an exclusive baseline curve based on historical physiological data in the past six months through Gaussian process regression, marking the characteristics of circadian rhythm and fluctuation bandwidth; monitoring the deviation of real-time data from the baseline curve, and when the instantaneous deviation exceeds 3 times the standard deviation, judging the drift type through the LSTM model and the disease knowledge graph, including the disease progression evolution type of gradual change consistent with the disease development path and the abnormal type of sudden unrelated fluctuation.
[0050] Further, the judgment of the drift type further includes a dynamic cleaning strategy: the disease progression evolution type updates the bandwidth parameters of the baseline curve and adds a clinical event label; the abnormal type performs wavelet denoising processing and generates a device quality control alarm; when a device calibration event occurs in a key interval of the baseline curve, the parameter update is frozen and a calibration compensation note is added, and finally a cleaning data stream with medical markers is output; the key interval is a time period that meets the key period of disease diagnosis and treatment and the fluctuation rate of physiological parameters exceeds 1.5 times the historical fluctuation value.
[0051] One or more technical solutions provided in the present application have at least the following technical effects or advantages:
[0052] By constructing a unified time coordinate system, the problem of time and space misalignment of multi-modal medical data is effectively solved, and the synchronization accuracy of key vital sign data such as electrocardiogram and blood oxygen is significantly improved; through the dynamic compensation mechanism driven by device calibration events, different types of calibration events are intelligently distinguished and differential compensation processing is performed, a device traceability report is generated, and the continuity and integrity of long-term health monitoring data are ensured; through the individualized baseline drift management technology, the physiological parameter drift type is intelligently identified and the targeted cleaning strategy is executed, which greatly reduces the risk of data miswashing of chronic disease patients; finally, in the cross-institutional collaboration scene, relying on intelligent term mapping and decision chain verification technology, the semantic difference problem between medical institutions is effectively solved, the consistency of referral data and the emergency decision response speed are significantly improved, and a new standard for medical data cleaning covering clinical diagnosis and treatment, scientific research analysis and device management is formed. BRIEF DESCRIPTION OF DRAWINGS
[0053] Figure 1 A medical data cleaning method flowchart based on an AI large model in an embodiment of the present application; DETAILED DESCRIPTION
[0054] In order to facilitate the understanding of the present application, the present application will be described more fully below with reference to the related drawings; the preferred embodiments of the present application are shown in the drawings, but the present application can be realized in many different forms and is not limited to the embodiments described herein; on the contrary, the purpose of providing these embodiments is to make the disclosure of the present application more thorough and comprehensive.
[0055] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs; the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the present application; and the use herein of the terms "and / or" includes a set of one or more associated listed items.
[0056] Embodiment one: as shown, an AI large model-based medical data cleaning method. Figure 1
[0057] S1: Obtain multi-modal data, obtain source identification feature data through a source identification module, input the source identification feature data into a scene analysis module, output a clinical feature set through dynamic expansion of feature dimension, symptom entity identification and original feature retention;
[0058] The multi-modal data includes: electronic medical record text data, vital sign monitoring time series data, image data and metadata, structured test data, device source data, clinical decision support data and space-time index data;
[0059] Specifically, the electronic medical record text data is obtained in real time from the hospital EMR system through the HL7 / FHIR interface, the patient's chief complaint, medical history record and other free texts are parsed, and the OCR scanned pieces are stripped; the vital sign monitoring time series data is collected from the bedside monitor, the R wave peak value of electrocardiogram and blood oxygen flow data are extracted, and the transmission delay is compensated; the image data and metadata are obtained from the PACS system, the device model and acquisition parameters in the CT file metadata are parsed; the structured test data is received from the LIS laboratory system, which is split into atomic test items and labeled with original unit values; the device source data is extracted from the device registration library to identify the device type; the clinical decision support data loads the standardized clinical terminology library, the hospital historical terminology frequency library and the symptom test index logic chain library through the FHIR terminology service; the space-time index data is obtained from the hospital audit system, which is bound with privacy classification labels and generates audit tracking logs. All data is standardized by a dynamic protocol parser. When the scene attribute identification is emergency, a 5-second timeout channel is activated for priority processing, and finally a complete multi-modal data stream supporting clinical decision embedding and cross-institution semantic alignment is formed.
[0060] The source identification feature data includes: scene attribute identification, professional attribute identification, device attribute identification and transmission protocol identification;
[0061] Specifically, the scene attribute identifier is a Boolean type (True / False) emergency flag, which is automatically set to True when the data contains an emergency triage label or continuously appears abnormal values of vital signs, and triggers a 5-second timeout processing channel and a computing resource forced preemption mechanism; the professional attribute identifier maps the specialty coding system, matches the specialty type by analyzing the electronic medical record diagnosis code or symptom keyword, drives high-precision processing mode and associates the symptom test index logical chain library; the device attribute identifier analyzes the UDI code in the device source data, matches the device feature library to obtain signal optimization parameters; the transmission protocol identifier defines the protocol type enumeration value, automatically loads the parser or enables the protocol conversion layer to reconstruct the semantics according to the protocol header mark, and ensures the integrity of the cross-institutional context.
[0062] Unstandardized clinical features are extracted from the multi-modal data stream, including raw text of symptom description, device native parameters, and non-converted test values; combined with the professional attribute identifier, the feature dimension is dynamically expanded, such as adding pathological grading features in the oncology department; and adding ECG ST segment offset raw measurement values in cardiovascular emergency.
[0063] The clinical feature set is output by dynamically expanding the feature dimension, symptom entity recognition, and preserving the original features, including raw symptom description, non-converted test values, and device native parameters;
[0064] When the scene attribute identifier is emergency (True) and the symptom entity is recognized, the emergency channel is activated, the non-emergency task thread is forcibly interrupted, and the symptom entity keyword library is called to match the electronic medical record free text in real time; the symptom description free text in the electronic medical record is preserved, and early term standardization is rejected to maintain the integrity of the original semantics; the native parameters of the device are collected; the original test values of the laboratory are recorded, and the original unit is preserved to avoid pre-conversion distortion; and the feature dimension is dynamically expanded based on the specialty scene, such as adding pathological grading features in the oncology department, and adding ECG ST segment offset raw measurement values in cardiovascular emergency, and the clinical feature set is complete and retains the native state of medical data.
[0065] S2: Construct a time coordinate system through a scene-driven strategy, obtain a continuity index based on the time coordinate system, and perform dynamic calibration, output the first cleaned data containing the calibrated data stream and the clinical feature set;
[0066] The scene-driven strategy includes processing channel strategy, computing precision strategy, resource allocation strategy, and downstream control strategy;
[0067] Specifically, the processing channel strategy is divided into emergency channel and non-emergency channel. When the scene attribute is identified as emergency (True) and the symptom entity (such as chest pain, dyspnea) is identified, the emergency channel is activated, the non-emergency task thread is forcibly interrupted, and the symptom entity keyword library is called to match the electronic medical record free text in real time. The non-emergency channel enables the batch processing mechanism (60 seconds / batch), which is compatible with low-speed devices.
[0068] The calculation accuracy strategy is divided into high-precision mode and normal mode. When the professional attribute is identified as oncology and the research plan number is detected, 0.1 ms level time sequence analysis is enabled and tumor special logic chain library is loaded. The normal mode uses 10 ms level time sequence analysis to balance efficiency and accuracy.
[0069] The resource allocation strategy generates a three-level priority label. When the emergency flag is True and the vital signs are abnormal, the first priority is triggered, the CPU and memory of the system are preempted, and the regular tasks are suspended. In the research scene and there is a long-term tracking mark, the second priority is triggered, and the memory buffer area is allocated to store the original waveform data. In the normal scene and the device attribute is identified as a low-speed device, the third priority is triggered, and batch compression processing (60 seconds / pack) is enabled to reduce the calculation load.
[0070] The downstream control strategy passes the priority label to the space-time coordination module to dynamically adjust the tolerance window threshold. In the emergency scene, the alignment is accelerated, and the window is contracted to ±100 ms. In the research scene, data integrity is prioritized, and the window is expanded to ±500 ms. According to the device attribute identification, the signal optimization library is called, such as enabling airflow waveform smoothing algorithm for ventilator devices and loading electromyographic interference suppression parameters for electrocardiogram devices.
[0071] The first cleaning data includes a calibration data stream and a clinical feature set;
[0072] The calibration data stream is a multi-modal data set for time alignment of the tolerance window, including text symptom description, signal optimized vital sign data, and abnormal marker data stream synchronized on the time axis;
[0073] The clinical feature set is a non-standardized clinical feature set extracted from multi-modal data, including original symptom description, non-converted test values, and device native parameters.
[0074] Specifically, the output calibration data stream is aligned by dynamic tolerance window, optimized by device-specific signal, and output by abnormal marker and traceability. The clinical feature set inherits the clinical feature set data obtained by analyzing the clinical features.
[0075] Receiving priority labels passed by downstream control strategy, dynamically adjusting tolerance window threshold; the first priority window shrinks to ±100ms, used to accelerate data alignment; the second priority window shrinks to ±500ms, used to guarantee data integrity; the third priority window shrinks to ±300ms, used to balance efficiency and accuracy. Aligning multi-source data within the tolerance window, binding electronic medical record free text with ECG R-wave peak timestamp; ECG waveform and blood oxygen streaming data are synchronized at millisecond level.
[0076] Each time data flows into the time coordinate system, real-time calculation of continuity index:
[0077] ,
[0078] Wherein, is the continuity index, is the total amount of data to be processed in the time coordinate system, is the amount of effective data aligned within the tolerance window, when the index decreases, mark the potential fracture node.
[0079] Taking ECG R-wave peak timestamp as the reference point, strictly binding electronic medical record free text with corresponding period of vital sign data within the dynamic tolerance window. ECG waveform calls device feature library parameters, loads device-specific signal optimization algorithm. For electromyographic interference, enable 50Hz notch filter to eliminate noise; for baseline drift, use 0.5Hz high-pass filter to correct respiratory motion artifacts; blood oxygen streaming data enables dynamic smoothing algorithm, and the filter coefficient is adjusted in real time according to the strength of motion artifacts.
[0080] For data beyond the tolerance window, mark it as a temporal anomaly and generate a report containing device ID, over-limit value and recommended operation. When the ECG waveform signal-to-noise ratio is less than 20dB, mark the low-quality signal and trigger the device quality control alarm. Over-limit data generates virtual data points through linear interpolation, while associating device ID to trace the abnormal cause. Finally, output the first cleaned data.
[0081] S3: When the continuity index is less than 95%, mark it as a fracture node, perform segmented baseline management on the fracture node, generate clinical decision features, and output the second cleaned data after format conversion on the calibration data stream;
[0082] Specifically, when the continuity index <95% is automatically marked as a broken node, triggering segmented baseline management: the data stream is divided by the broken node, the original baseline curve is retained in the front segment, and the Gaussian process regression curve is regenerated in the rear segment. Perform decision transformation on the clinical feature set, extract key clinical entities from the original symptom description, and bind the unconverted test values to the symptom entities to form a symptom test index logical chain library. Extract the fluctuation bandwidth parameters of the rear curve and bind them to the symptom test index logical chain library to output clinical decision features containing segment markers. The 95% threshold is based on the clinical data fault probability model: analysis of 100,000 patient data shows that when the continuity index is ≥95%, the long-term health data integrity reaches 99.3% (p<0.01); in emergency scenarios, it is dynamically adjusted to 90% to adapt to high-noise environments.
[0083] Extract key clinical entities from the original symptom description of the clinical feature set, and use semantic analysis technology to strip non-clinical descriptions; call the symptom test index logical chain library to match the pre-defined clinical rule chain; if the device attribute identifier is an electrocardiogram device, append the device parameter constraint to ensure data validity. According to the test value and vital sign data, verify the rule chain trigger condition to generate decision action instructions. If the broken node exists, append the segment baseline offset annotation; when the device calibration event occurs in the broken node period, freeze the baseline curve update and add calibration compensation annotations; output features bound to the broken node.
[0084] When the instructions conflict, such as blood glucose test value > 11 mmol / L, but no history of diabetes in the medical history feature marker; the impedance of the electrocardiogram lead > 20kΩ, ST segment elevation is detected. Add a conflict resolution marker to the standardized clinical feature set, pause the current decision tree execution, and push the conflicting data package to the manual review queue. After manual confirmation, update the symptom test index logical chain library.
[0085] The second cleaning data includes format conversion data stream and standardized clinical feature set;
[0086] The format conversion data stream includes clinical path binding field, privacy level inheritance data and timestamp calibration stream;
[0087] The standardized clinical feature set includes term standardization output, unit standardized numerical value and conflict resolution marker.
[0088] Specifically, based on the decision label output by the clinical decision tree, the multi-modal data in the calibration data stream is bound to the standard clinical path; according to the privacy classification label in the space-time index, the identity information and medical image information are desensitized, and the operation records in the audit log are inherited; all timestamps are converted to UTC standard format, and the R wave peak time of electrocardiogram is taken as the reference point, and the blood oxygen test data timestamp is synchronously offset compensated to realize cross-device alignment.
[0089] The standardized clinical feature set guarantees data consistency through three levels of conversion: first, the standardized terminology library in the clinical decision support data is called to convert the original symptom description into standard terminology through the AI mapping engine, and the hospital's commonly used terminology is preferentially matched according to the historical terminology frequency library. Second, the test values are converted into the international system of units, the original values are double-labeled, and the device parameters are standardized. Finally, for the logically conflicting data, structured conflict labels are added to trigger the manual review process and freeze the related decision branches. Special format conversion is performed on the data of the fracture node period: the timestamp calibration flow adds a fracture interval label; the privacy level inheritance data is forced to enable encryption level desensitization; the associated fracture nodes are prioritized to the highest level of review queue; and the second cleaned data is output.
[0090] The technical solutions in the embodiments of the application have at least the following technical effects or advantages:
[0091] The application effectively solves the time and space misalignment problem of multi-modal medical data by constructing a unified time coordinate system, significantly improving the synchronization accuracy of key vital sign data such as electrocardiogram and blood oxygen; through a device calibration event-driven dynamic compensation mechanism, it intelligently distinguishes abnormal types such as device delay and signal interference and performs differentiated compensation, generates a device traceability report, and guarantees the continuity and integrity of long-term health monitoring data; through individualized baseline drift management technology, it generates a dedicated reference curve based on historical physiological data, intelligently identifies course evolution type and abnormal type drift, and executes targeted cleaning strategies to significantly reduce the risk of data miswashing for chronic disease patients; finally, in a cross-institutional collaboration scenario, relying on intelligent terminology mapping and decision chain verification technology, it effectively solves the semantic difference problem between medical institutions, significantly improves the consistency of referral data and emergency decision response speed, and realizes full-scene medical data cleaning covering clinical diagnosis and treatment, scientific research analysis and device management.
[0092] Embodiment Two: In Embodiment One, a unified time coordinate system is constructed to effectively solve the time and space misalignment problem of multi-modal medical data, but it cannot solve the terminology difference between different institutions, there are time zone and clock reference differences between cross-institutional data, and it cannot adapt to the diversified compliance requirements of cross-institutions. This embodiment makes further optimization based on Embodiment One.
[0093] S4: Based on the scene-driven strategy, time coordinate system, clinical feature set and clinical decision feature, a clinical context migration body is constructed, the second cleaned data is input into the clinical context migration body, cross-institutional context migration and semantic reconstruction are performed, and the third cleaned data is output.
[0094] Specifically, the processing channel strategy and the resource allocation strategy in the scene-driven strategy; the electrocardiogram R-wave benchmark time axis constructed by the time coordinate system and the multi-modal synchronization window parameters; the non-standardized features contained in the clinical feature set; and the standardized terms and conflict resolution markers output by the clinical decision feature are taken as input elements to construct a clinical context migration body and output third cleaning data.
[0095] When migrating across institutions, first, the semantic reconstruction of terms is performed: the standardized clinical feature set is taken as input data to query the receiving institution's term library, such as the hospital FHIR standard; if there is a term conflict, a term mapping table is added, the original term is retained and a standard term note is attached; and the semantic reconstruction features containing the term mapping table are output. Second, the time benchmark compensation is performed: the time zone difference between institutions is calculated to generate a time benchmark offset, and all timestamps are updated based on the time benchmark offset to ensure cross-institution clock alignment. Finally, the privacy level is inherited: cross-institution privacy policy conversion is performed, such as identity information, the sender's privacy level is community open level, and the receiver requires three-A encryption level, which is converted to FHIR anonymous ID; the sender's privacy level is community encryption level, and the receiver requires three-A research level, which is desensitized to ICD code B20; and an audit-enhanced spatiotemporal index containing record privacy conversion logs is output.
[0096] When the emergency flag of the scene strategy layer is True, the data is divided into ≤1 MB / packet, and key fields such as symptom description and vital sign abnormal values are placed on top to ensure transmission rate; the term mapping is simplified, and only basic desensitization is performed. In the second priority research scene, full-amount data is retained, the original device parameter association calibration certificate and term mapping history version are stored, data traceability is supported during data transmission, including device serial number and operator ID fields, and microsecond-level time axis accuracy is used. In other regular scenarios, balance efficiency and integrity are ensured, data is divided into ≤5 MB / packet to reduce system load, term mapping uses basic rules and is cached for acceleration; the privacy layer adopts the default level, uses millisecond-level time axis accuracy, and the synchronization window tolerance is ±300 ms.
[0097] The third cleaning data is a cross-institution compatible data packet generated by the clinical context migration body, including semantic reconstruction features, context inheritance data, and transmission optimization structure.
[0098] Specifically, the semantic reconstruction features solve the semantic gap and time misalignment problems of cross-institution data through dynamic term mapping tables and accurate time benchmark offsets. Term mapping retains the original term field and labels the confidence score while matching the receiving institution's standard term library, solving the expression differences of basic terms; the time benchmark offset realizes accurate cross-institution clock alignment by calculating the time zone difference and the inherent delay of the device, ensuring millisecond-level synchronization of electrocardiogram R-wave timestamps and CT image metadata.
[0099] The hierarchical privacy standardization engine dynamically loads the compliance strategy of the receiving agency, and performs differential desensitization on identity information: partial masking is used in community hospitals, FHIR anonymous ID is converted in first-class hospitals, and approval numbers are added in research institutions; the emergency mode skips non-key desensitization to ensure timeliness; the audit enhances the spatiotemporal index integration operation log, including the term mapping track, privacy conversion record, device traceability information and geographic space coordinates, to form a full-cycle tracking chain.
[0100] The emergency mode adopts a streaming optimization architecture: the symptom description and abnormal value of vital signs are forced to the top of the transmission, and the term mapping is simplified to high-frequency emergency abbreviations to ensure response speed; the research mode enables a full-quantity data architecture: the term history version is retained, the device calibration certificate is added, the time axis accuracy is expanded to the microsecond level, and the original waveform data access permission is opened.
[0101] The technical solutions in the embodiments of the application have at least the following technical effects or advantages:
[0102] The application adopts a clinical context migration body technology, and through the synergistic action of a dynamic term mapping table and an accurate time reference offset, semantic lossless migration and spatiotemporal continuity guarantee of cross-institutional medical data are realized; based on the hierarchical privacy standardization engine and the audit-enhanced spatiotemporal index, the compliance requirements of different institutions are dynamically adapted, the privacy conversion efficiency is improved while the compliance is guaranteed; through the streaming transmission optimization of the emergency mode and the full-quantity data architecture of the research mode, combined with the cache acceleration mechanism of the conventional mode, a full-scene adaptive transmission system covering emergency response, research analysis and daily diagnosis and treatment is formed, and finally the effects of reducing misdiagnosis rate in referral and shortening access time to new institutions are realized, providing core data support for the hierarchical diagnosis and treatment system.
[0103] Embodiment three: in embodiment one, semantic lossless migration and spatiotemporal continuity guarantee of cross-institutional medical data are realized, but only clock alignment is realized, and the long-term data discontinuity caused by device calibration is not solved, and this embodiment further optimizes embodiment two.
[0104] The construction of the time coordinate system includes: taking the R wave peak as the electrocardiogram reference point, and combining the device time difference characteristic vector for dynamic delay compensation;
[0105] Specifically, the R wave peak of electrocardiogram is taken as the core time reference, the R wave position is accurately identified through the wavelet transform algorithm, the time difference characteristic vector in the device source data is analyzed, dynamic compensation is performed, and a device calibration certificate containing device ID, delay value and calibration time is output:
[0106] ,
[0107] wherein, is the corrected timestamp, is the original timestamp, a device-specific delay read from a device feature library, a network transmission delay measured by a time synchronization protocol.
[0108] Based on the compensated time reference, receive the priority label of the downstream control strategy delivery of the scene-driven strategy, set the tolerance window range of the multi-modal synchronization window; when the single data exceeds the limit, the continuity index decreases by 5%, triggering the primary response: generating a virtual value based on the two effective data points before and after to perform linear interpolation compensation, marking as a temporal anomaly and generating a quality control suggestion associated with the device ID. When the continuity index is zeroed for three consecutive times due to continuous over-limit, triggering the secondary response: calling the device feature library to update the delay parameters, recalibrating the device time difference vector; shrink the current tolerance window by 20%, outputting a temporal anomaly report containing over-limit values, device ID, and repair records.
[0109] The system monitors device logs in real time, accurately identifies zero-point calibration (such as blood pressure meter zeroing), sensitivity calibration (such as ECG gain adjustment), and composite calibration events, and generates event fingerprints containing device ID, calibration type, parameter change value, and timestamp. Calibration event triggers segmented baseline management: data before calibration retains the original baseline curve for long-term trend analysis, and data after calibration establishes a new baseline and marks the offset value. Finally, bind the calibration event ID to the affected data segment, insert the calibration review label in the clinical decision tree, synchronize the device calibration history certificate when migrating across institutions, so that the receiving party can trace back to the original data before calibration, and eliminate the data gap caused by device calibration.
[0110] The technical solutions in the embodiments of the application have at least the following technical effects or advantages:
[0111] The application solves the problem of data gap caused by medical device calibration by constructing a dynamic time reference calibration and calibration event continuity guarantee mechanism. Taking the ECG R-wave peak value as the core time reference, combining device-specific delay and network transmission delay for dynamic compensation, generating a calibration certificate with device ID, and compressing multi-modal data synchronization error to the millisecond level. Through the tolerance window feedback mechanism, combined with intelligent identification of device calibration events and segmented baseline management, the long-term data gap caused by device calibration is eliminated, and the continuity and integrity of cross-year health data reaches 99.3%. At the same time, the calibration review label is embedded in the clinical decision tree, and the device calibration history certificate is synchronized when migrating across institutions, ensuring that the receiving party can trace back to the original data.
[0112] Embodiment Four: Embodiment Three solves the problem of data gap caused by device calibration, but does not distinguish the medical significance of physiological parameter changes and does not consider the conflict between calibration timing and critical periods of diagnosis and treatment. This embodiment further optimizes based on Embodiment Three.
[0113] The segmented baseline management includes: based on the historical physiological data in the past six months, generating a dedicated baseline curve through Gaussian process regression, marking the characteristics of circadian rhythm and fluctuation bandwidth; real-time monitoring of the deviation of the data from the baseline curve, when the instantaneous deviation exceeds 3 times the standard deviation, the drift type is judged through the LSTM model and the disease knowledge graph, including the disease development path consistent gradual change of the disease course evolution type and the sudden unrelated fluctuation of the abnormal type.
[0114] Specifically, the historical physiological data in the past six months is collected, aligned by time axis and abnormal values are removed. The square exponential kernel (RBF) is used to capture long-term trends, and the periodic kernel (period = 24 hours) is used to reflect the circadian rhythm, to generate a dedicated baseline curve, marking the key features containing circadian rhythm characteristics and fluctuation bandwidth.
[0115] When the instantaneous deviation of real-time data from the baseline curve is greater than 3 times the standard deviation, LSTM time series analysis is triggered: input continuous 72-hour data, extract time series features such as slope change and fluctuation frequency, output drift mode vector; call knowledge graph, use pre-defined medical logic chain for verification, if the drift feature matches the disease development path, it is marked as the disease course evolution type representing the true disease evolution; if the feature is contradictory or triggers a negative rule, it is marked as the abnormal type pointing to device failure or noise interference.
[0116] The judgment of the drift type also includes a dynamic cleaning strategy: the disease course evolution type updates the bandwidth parameters of the baseline curve and adds a clinical event label; the abnormal type performs wavelet denoising and generates a device quality control alarm; when the device calibration event occurs in the key interval of the baseline curve, the parameter update is frozen and a calibration compensation note is added, and finally the cleaned data stream with medical labels is output; the key interval is the time period that meets the key period of disease diagnosis and treatment and the physiological parameter fluctuation rate exceeds 1.5 times the historical fluctuation value.
[0117] Specifically, the disease course evolution type is high priority, the original data is retained, and a clinical event label is added; the abnormal type is medium priority, wavelet denoising is triggered, and a device quality control work order is generated. When the data is in the key period of diagnosis and treatment, such as 24 hours after surgery, and the physiological parameter fluctuation rate exceeds 1.5 times the historical fluctuation value, even if the LSTM detects a suspected abnormal type drift, the system freezes the parameter update to prevent false cleaning of the true disease signal, and outputs a double-path report for the doctor to review.
[0118] When it is determined to be the disease course evolution type, the system immediately pushes the key label and associated treatment suggestions to the doctor workstation to drive clinical intervention; at the same time, the baseline curve parameters are automatically updated to dynamically adapt to the disease evolution trend. If it is determined to be the abnormal type, a device quality control alarm is automatically sent to the engineer end, a fault mode library is associated to generate a maintenance plan; at the same time, wavelet denoising is performed on the original data, and the denoised data and repair labels are recorded on the double-track record to ensure that technical problems can be traced back and audited.
[0119] The technical solutions in the embodiments of the present application have at least the following technical effects or advantages:
[0120] The present application realizes the dual improvement of clinical decision safety and data cleaning accuracy by constructing individualized reference curve and intelligent drift judgment mechanism. Based on the historical physiological data within half a year, the Gaussian process regression is used to generate the exclusive reference curve, mark the characteristics of circadian rhythm and fluctuation bandwidth, and provide individualized reference standard for real-time monitoring. When the instantaneous deviation of real-time data and reference curve exceeds 3 times of standard deviation, the system extracts time sequence features through LSTM model and cooperates with disease knowledge graph for medical logic verification: if the features are consistent with the disease development path, it is judged as the disease evolution type; if the features are contradictory or have no pathological correlation, it is judged as the abnormal type. For the disease evolution type, the system updates the bandwidth parameters of the reference curve and adds clinical event labels to drive clinical intervention; for the abnormal type, wavelet denoising processing is performed and device quality control alarm is generated to promote technology optimization. When the device calibration occurs in the key period of diagnosis and treatment and the fluctuation rate of physiological parameters exceeds 1.5 times of the historical mean value, the system freezes the parameter update, adds calibration compensation annotation to avoid misjudgment of the real disease, and outputs the double reference comparison report for the doctor to review. Finally, the system outputs the cleaning data stream with medical labels, which significantly reduces the false cleaning rate of chronic diseases by 37% and improves the safety of diagnosis and treatment of acute and severe diseases.
[0121] The above only describes the preferred embodiments of the present application and is not used to limit the present application. For those skilled in the art, the present application can have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A medical data cleaning method based on an AI large-scale model, characterized in that, include: S1: Acquire multimodal data, obtain source identification feature data through the source identification module, input the source identification feature data into the scene analysis module, and output the clinical feature set by dynamically expanding feature dimensions, identifying symptom entities and preserving original features; S2: Construct a time coordinate system through a scenario-driven strategy, obtain a continuity index based on the time coordinate system and perform dynamic calibration, and output the first cleaned data containing the calibration data stream and clinical feature set; The scenario-driven strategy includes processing channel strategy, computational accuracy strategy, resource allocation strategy and downstream control strategy; The processing channel strategy is to activate the timeout processing channel when the emergency scene is in progress and the symptom entity is identified. The accuracy calculation strategy is to enable the accuracy processing mode of time series analysis when the specialty code is oncology and the research protocol number is detected. The resource allocation strategy generates three levels of scene priority markers, including first priority, second priority and third priority; The downstream control strategy is to pass the priority flag to the spatiotemporal coordination module to control the tolerance window range; The tolerance window is the time alignment tolerance interval in the time coordinate system; The construction of the time coordinate system includes: using the R-wave peak value as the ECG reference point, and performing dynamic delay compensation by combining the device time difference feature vector. ,in, For the corrected timestamp, This is the original timestamp. The inherent latency of the device read from the device feature library, To measure network transmission delay via time synchronization protocol; based on the compensated time base, receive priority flags from downstream control strategies driven by scenario, set the tolerance window range of multimodal synchronization window; establish a window feedback mechanism, triggering a primary response to perform interpolation compensation and mark temporal anomalies when a single data exceedance occurs; triggering a secondary response after three consecutive exceedances, recalibrating the device time difference vector, updating the tolerance window range, and generating a temporal anomaly report; S3: When the continuity index is less than 95%, it is marked as a breakpoint. Segmented baseline management is performed on the breakpoint, clinical decision features are generated, and the second cleaned data is output after format conversion of the calibration data stream. The segmented baseline management includes: generating a dedicated baseline curve based on six months of historical physiological data through Gaussian process regression, marking diurnal rhythm characteristics and fluctuation bandwidth; real-time monitoring of the deviation between the data and the baseline curve, and when the instantaneous deviation exceeds 3 times the standard deviation, using an LSTM model and a disease knowledge graph to collaboratively determine the drift type, including the progressive disease progression type that conforms to the disease development path and the abnormal type with sudden unrelated fluctuations. S4: Based on the scenario-driven strategy, time coordinate system, clinical feature set and clinical decision features, a clinical scenario transfer body is constructed. The second cleaned data is input into the clinical scenario transfer body, cross-institutional scenario transfer and semantic reconstruction are performed, and the third cleaned data is output. The third cleaned data undergoes semantic reconstruction, context inheritance, and transmission optimization through a clinical context transfer model. The semantic reconstruction includes: querying the terminology database of the receiving institution, converting the standardized output of the terminology of the standardized clinical feature set into the standard terminology of the receiving institution, and generating a terminology mapping table; calculating the time base offset based on the time zone difference between institutions, and updating all timestamps in the timestamp calibration stream; The context inheritance includes: parsing the recipient's privacy compliance policy, performing differentiated desensitization operations on the privacy-level inherited data, generating an audit-enhanced spatiotemporal index, and recording terminology mapping trajectories and privacy conversion logs; The transmission optimizations include: in emergency scenarios, transmitting symptom descriptions and abnormal vital signs at the top, and simplifying terminology mapping to high-frequency emergency abbreviations; in research scenarios, enabling the full data channel, retaining historical versions of terminology and equipment calibration certificates, and extending the time axis accuracy to the microsecond level.
2. The medical data cleaning method based on an AI large model as described in claim 1, characterized in that, The multimodal data includes: electronic medical record text data, vital sign monitoring time series data, image data and metadata, structured test data, equipment source data, clinical decision support data and spatiotemporal index data; The electronic medical record text data includes free text of the patient's chief complaint, medical history, and diagnostic conclusions; The time-series data for vital sign monitoring includes electrocardiogram waveforms and blood oxygen saturation flow cytometry data. The image data and metadata include CT format files and their metadata; The structured test data consists of laboratory values broken down by test item; The device source data is a device type identifier, used to call the device feature library to perform signal optimization; The clinical decision support data includes a standardized clinical terminology database, a historical terminology frequency database of the hospital, and a logical chain database of symptom test indicators. The spatiotemporal index data includes hierarchical labels and audit trail logs.
3. The medical data cleaning method based on an AI large model as described in claim 1, characterized in that, The source identification feature data includes: scene attribute identifier, professional attribute identifier, device attribute identifier, and transmission protocol identifier; The scenario attribute identifier is a Boolean emergency flag, used to mark whether the data belongs to an emergency scenario; The professional attribute identifier is the associate degree code; The device attribute identifier is the device type, which is used to associate signal optimization parameters with the device feature library; The transmission protocol identifier is the data access protocol type, used to define data parsing rules.
4. The medical data cleaning method based on an AI large model as described in claim 1, characterized in that, The first cleaning data includes calibration data streams and clinical feature sets; The calibration data stream is a multimodal data set with time alignment within a tolerance window, including time-axis synchronized textual symptom descriptions, signal-optimized vital sign data, and abnormality marker data streams. The clinical feature set is a collection of unstandardized clinical features extracted from multimodal data, including original symptom descriptions, unconverted test values, and native device parameters.
5. The medical data cleaning method based on an AI large model as described in claim 1, characterized in that, The second cleaned data includes format-converted data streams and standardized clinical feature sets; The format conversion data stream includes clinical pathway-bound fields, privacy-level inherited data, and timestamp calibration stream; The standardized clinical feature set includes standardized terminology output, standardized unit values, and conflict resolution markers.
6. The medical data cleaning method based on an AI large model as described in claim 1, characterized in that, The determination of drift type also includes a dynamic cleaning strategy: for disease progression type, update the baseline curve bandwidth parameter and add clinical event labels; for abnormal type, perform wavelet denoising and generate equipment quality control alarms. When a device calibration event occurs in the critical interval of the baseline curve, the parameters are frozen, updated, and calibration compensation annotations are added. The final output is a clean data stream with medical markers. The critical interval is a period that simultaneously meets the criteria of critical period for disease diagnosis and treatment and physiological parameter volatility exceeding 1.5 times the historical volatility value.
Citation Information
Patent Citations
Medical data analysis method and device, equipment and storage medium
CN116959733A
Medical data multi-source fusion verification and correction method and system
CN120180014A