A medical data full-link blood relationship and version management method based on lake and warehouse integration
Patent Information
- Application Number
- CN202610989491.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-03
- Publication Date
- 2026-09-25
AI Technical Summary
在此背景下,基于传统数据湖架构的医疗数据管理体系面临多重技术瓶颈,已难以满足精细化治理需求:
[0018]本发明实施例提供的技术方案带来的有益效果至少包括:
Smart Images

Figure CN122817166A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical data governance technology, and in particular to a lake-warehouse integrated medical data lineage and version management platform, a lake-warehouse integrated medical data lineage and version management method and its application, electronic devices and computer-readable storage media. Background Technology
[0002] As the digital transformation of the healthcare industry deepens, medical data is exhibiting characteristics of multimodalization (electronic medical records, medical images, laboratory reports, gene sequencing data, etc.) and petabyte-scale growth, supporting an ever-expanding range of business scenarios such as clinical decision support, medical AI model training, and medical compliance auditing. Against this backdrop, medical data management systems based on traditional data lake architectures face multiple technical bottlenecks and are no longer sufficient to meet the demands of refined governance. Data dependencies are not transparent: The medical data processing chain involves dozens of collection, cleaning, transformation and analysis nodes. When it is necessary to modify or take a data table offline, it is not possible to quickly locate the upstream and downstream related tasks, reports, indicators, datasets and AI models. This can easily lead to business failures due to the omission of dependencies, affecting clinical decision-making or the accuracy of model training.
[0003] Inefficient problem localization: Frequent issues such as task failures, process reruns, and data anomalies occur throughout the entire data processing chain. The lack of end-to-end lineage tracing capabilities makes it impossible to quickly identify data dependencies and assess the scope of impact, resulting in data fault troubleshooting time of several hours or even days, which is difficult to meet the high availability requirements of medical services.
[0004] Insufficient data backtracking capability: When erroneous data is generated due to operational errors, processing logic errors, changes in business requirements, etc., the existing data lake architecture lacks unified version management and rapid rollback capability, and cannot restore multimodal data in batches to a consistent state at a specified historical point in time. This not only affects data consistency, but may also violate the compliance requirements for traceability of medical data in the "Administrative Measures for Cybersecurity of Medical and Health Institutions".
[0005] Therefore, there is an urgent need to establish a full-chain lineage and version management mechanism for medical multimodal data that is adapted to the lakeware architecture. This mechanism would enable full-chain lineage tracking, automatic impact analysis, version snapshot management, and rapid rollback of multimodal data under a unified storage platform, thereby addressing the core pain points in medical data governance. Summary of the Invention
[0006] To address the technical problems existing in the prior art, the present invention provides the following technical solution: On the one hand, a medical multimodal data lineage and version management platform based on a lake-warehouse integrated architecture is provided, including: The lake warehouse storage layer is used to uniformly store structured, semi-structured, and unstructured data in the medical field using an architecture that integrates data lake and data warehouse, and to achieve data version snapshot management through the native time travel capability of the underlying storage engine. The fusion acquisition layer is used to synchronously and automatically collect data metadata and flow dependencies in each processing stage of the data lifecycle based on the extended metadata management plugin to form lineage information, and trigger the lake warehouse storage layer to generate corresponding data version snapshots, establishing a mapping relationship between lineage information and version snapshots; The service layer is used to provide functions such as bloodline retrieval and impact analysis, data version management and batch rollback, and audit report generation that meet medical compliance requirements, based on the bloodline and version mapping relationship map constructed by the fusion acquisition layer.
[0007] On the other hand, a method for end-to-end lineage and version management of medical multimodal data based on a lake-warehouse integrated architecture is also provided, applied to the aforementioned platform, including the following steps: During the data access phase, metadata and data source dependency information of the accessed multimodal medical data are automatically collected as initial lineage nodes. The accessed structured data is written into a tabular storage that supports transactions and time travel capabilities and an initial version snapshot is generated. The accessed unstructured data is stored in object storage and its file identifier and metadata information are associated with it and written into the tabular storage to generate an initial version snapshot. Establish a mapping relationship between the initial lineage node and the corresponding initial version snapshot and store it in the lineage graph library.
[0008] Preferably, the data preprocessing stage further includes the following steps: Automatically collect input data, processing logic, and output data dependencies of pre-processing tasks, and report them as newly added lineage information; After the preprocessing task is completed, the output structured data is written to the tabular storage to automatically generate a new version snapshot, or the output unstructured data processing results are updated to the object storage and the associated tabular storage to generate a new version snapshot; The pre-processing task is constructed as a new lineage link segment, which includes an input node, a processing logic node, and an output node. The output node is then bound to the newly generated snapshot and updated to the lineage graph library.
[0009] Preferably, the data processing and dataset generation stage further includes the following steps: Automatically collect all upstream input data sources, processing logic, and downstream output dataset dependencies involved in business-oriented data processing tasks, as well as related project and business scenario attribute information; After the processing task is completed, a fixed version snapshot with business tags is generated in the tabular storage for the output structured and unstructured data; Based on the collected full-link dependency relationship, the current processing node is dynamically associated with its upstream and downstream lineage nodes through dependency relationship resolution and automatic identification of context nodes, forming a complete lineage link from the original data access to the final business dataset output. The output node of the final business dataset is then bound to the fixed business version snapshot, completing the closed loop of lineage link and version mapping.
[0010] Preferably, the "automatic collection of dependencies between all upstream input data sources, processing logic, and downstream output datasets involved in business-oriented data processing tasks" specifically includes: The metadata management framework's acquisition plugin automatically parses the input source table or file, the definition of the processing logic, and the output target table or file of the data processing task, extracting the complete dependency chain of "input-processing-output".
[0011] Preferably, the "dynamically associating the current processing node with its upstream and downstream lineage nodes" specifically includes: The upstream parent node and downstream child node of the current processing node are automatically identified through metadata matching rules; The bloodline service center associates the upstream parent node, the current processing node, and the downstream child node through relationship edges of the type "input dependency" and "output generation".
[0012] Preferably, the method further includes: constructing a bidirectional mapping graph of lineage and version based on the lineage graph library, wherein each lineage node in the graph is associated with one or more historical version snapshots.
[0013] On the other hand, a data lineage impact analysis method based on the platform or method described above is provided, including: In response to a user's operation on a specified node in the lineage graph, based on the bidirectional mapping graph of lineage and version, the system automatically traverses and identifies all downstream dependent nodes of the specified node. The downstream dependent nodes include data tables, data processing tasks, reports, indicators, datasets, and AI models that directly or indirectly depend on the specified node. An impact scope analysis report is generated and output, which clearly lists all affected downstream dependent nodes and related tasks that may need to be re-executed.
[0014] On the other hand, a multimodal data linkage rollback method based on the platform or method described above is provided, including: In response to an instruction to roll back a target node in the lineage graph to a specified historical version, based on the bidirectional mapping graph of lineage and version, the system automatically retrieves and locates the specified historical version snapshot corresponding to the target node, as well as the historical version snapshots of all upstream and downstream nodes that are related to the target node at the same time or logically corresponding to the target node; triggers the storage engine's time travel function to batch roll back the structured data tables and unstructured data files involved in the target node and all its associated upstream and downstream nodes to the data consistency state corresponding to the specified historical version snapshot.
[0015] On the other hand, a method for supporting medical data compliance auditing and scientific research reproduction based on the aforementioned platform or method is provided, including: Automatically records and stores operation logs, processing entities, and data change content of each lineage node in the entire data chain during data access, processing, and version generation; in response to audit or reproduction requests, for the specified final business data output node and its associated business version tag, based on the bidirectional mapping map of lineage and version, traces back and outputs the complete processing chain from the original data source to the output node, as well as the version snapshot information and operation records corresponding to each node in the chain, in order to generate a compliance audit report or provide a precise scientific research data reproduction path.
[0016] On the other hand, an electronic device is provided, comprising: a processor; and a memory storing computer-readable instructions, which, when executed by the processor, implement the method described above.
[0017] On the other hand, a computer-readable storage medium is provided, wherein at least one instruction is stored therein, the at least one instruction being loaded and executed by a processor to implement the above method.
[0018] The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following: 1. Enhanced multimodal adaptability: For the first time, it achieves unified management of lineage and version of structured, semi-structured, and unstructured medical data under the lake warehouse integrated architecture, supporting the traceability needs of all types of medical data such as electronic medical records, medical images, and gene sequencing data.
[0019] 2. Improved problem handling efficiency: The ability to automatically collect and analyze the lineage throughout the entire data chain reduces the average time for troubleshooting data faults from 4 hours to less than 10 minutes; the ability to batch roll back multimodal data reduces the data recovery time from several hours to minutes.
[0020] 3. Reduced compliance costs: It natively meets the regulatory requirements for traceability of medical data, eliminating the need to build an additional independent audit system, thus reducing governance costs by more than 40%.
[0021] 4. Enhanced research value: By linking lineage and version management, the reproducibility of the medical research dataset generation process is ensured, thereby improving the credibility of research results. Attached Figure Description
[0022] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0023] Figure 1 This is a flowchart of a medical data lineage and version management process provided in an embodiment of the present invention. Detailed Implementation
[0024] The technical solution of the present invention will now be described with reference to the accompanying drawings.
[0025] In embodiments of the present invention, words such as "exemplarily," "for example," etc., are used to indicate that something is an example, illustration, or description. Any embodiment or design described as "exemplary" in the present invention should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of the word "exemplary" is intended to present the concept in a concrete manner. Furthermore, in embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one.
[0026] In the embodiments of this invention, the terms "image" and "picture" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning. Similarly, the terms "of," "corresponding (relevant)," and "corresponding" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning.
[0027] In this embodiment of the invention, sometimes a subscript such as W1 may be mistakenly written as a non-subscript form such as W1. When the difference is not emphasized, the meaning they express is the same.
[0028] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.
[0029] Terminology Explanation: Lake-warehouse integration is a new type of data architecture that combines the core advantages of data lakes and data warehouses. It combines the low cost and native multimodal data storage capabilities of data lakes with the strong transaction consistency and high-performance query and analysis capabilities of data warehouses. It supports the unified storage, processing and analysis of structured, semi-structured and unstructured data (including multimodal data such as medical images, electronic medical records, and audio and video) on the same storage platform. It is the mainstream underlying architecture for current medical big data platforms and multimodal data governance systems.
[0030] Data lineage: A structured graph used to characterize the relationships in the entire data flow chain. It records core information such as the source, flow direction, transformation rules, and processing entities of data throughout the entire process from collection, processing, cleaning, fusion to consumption. It can support the needs of scenarios such as data traceability, impact analysis, and data compliance auditing.
[0031] Data version: This refers to the status record of raw data, intermediate processing results, or final datasets in the lake warehouse storage system at different time points, including metadata snapshots and data content operation logs at the corresponding time points. It is used to solve problems such as historical backtracking, consistency verification, and version conflict resolution during dynamic data changes.
[0032] Time travel: This is one of the core capabilities provided by the underlying storage engine of Lakeware. By automatically recording version snapshots with corresponding timestamps when data is written / modified, it allows users to quickly query the data status at historical time points based on timestamps or version IDs, or roll back data to a specified historical version. It is also known as data snapshot rollback capability.
[0033] Atlas is an open-source metadata management and governance framework under the Apache Software Foundation. It provides centralized metadata storage, classification management, automatic lineage tracing, data access control, and security policy management capabilities, and is one of the mainstream metadata governance components in current data lake warehouse architectures.
[0034] Iceberg is an open table format designed for big data scenarios. It supports transactional consistency, schema evolution, hidden partitions, and native time travel capabilities. It is compatible with multiple computing engines and is the mainstream underlying table format implementation solution in the current lakeware architecture.
[0035] Metadata: Data that describes data, mainly including core content such as data structure definition, business meaning, storage location, access permissions, and change history. It is the core foundation of data governance, data lineage, and data versioning.
[0036] To address the technical problems in existing medical data governance, such as opaque lineage relationships, low efficiency in problem localization, and insufficient multimodal data backtracking capabilities, this invention proposes a full-link lineage and version management method for medical multimodal data based on a lake-warehouse integrated architecture. The core innovation lies in constructing a "lineage-version bidirectional mapping graph" to achieve linked management of data flow links and historical versions, while simultaneously meeting the requirements of traceability, consistency, and compliance of medical data.
[0037] I. Platform Architecture On the one hand, a medical multimodal data lineage and version management platform based on a lake-warehouse integrated architecture is provided, including: The lake warehouse storage layer is used to uniformly store structured, semi-structured, and unstructured data in the medical field using an architecture that integrates data lake and data warehouse, and to achieve data version snapshot management through the native time travel capability of the underlying storage engine. The fusion acquisition layer is used to synchronously and automatically collect data metadata and flow dependencies in each processing stage of the data lifecycle based on the extended metadata management plugin to form lineage information, and trigger the lake warehouse storage layer to generate corresponding data version snapshots, establishing a mapping relationship between lineage information and version snapshots; The service layer is used to provide functions such as bloodline retrieval and impact analysis, data version management and batch rollback, and audit report generation that meet medical compliance requirements, based on the bloodline and version mapping relationship map constructed by the fusion acquisition layer.
[0038] To achieve the governance goals of data version traceability, lineage association, and operation rollback, this application constructs a core application mechanism of "lineage and version linkage management." This mechanism runs through the entire data lifecycle, triggering version snapshot generation at each stage of data processing (access, pre-processing, business processing, and service output), and dynamically binding it with lineage nodes collected in real time by Atlas to build a "lineage-version" mapping relationship chain. The specific implementation of this mechanism relies on the three-layer platform architecture shown in Table 1 below: Table 1 The platform's overall logic is as follows: Based on the lake warehouse integrated storage base, it synchronously collects lineage data and version snapshot information throughout the entire lifecycle of data access, pre-processing, processing, and application. It associates lineage nodes with corresponding version data through a globally unique entity ID, which not only supports the tracing of upstream and downstream dependencies across the entire chain, but also supports the historical version backtracking and rollback of any lineage node, solving the problems of fragmented lineage and version management and poor multimodal data adaptability in traditional solutions.
[0039] Based on this mechanism and architecture, the management steps and implementation methods at each stage of the data lifecycle will be detailed below.
[0040] II. Methods and Steps On the other hand, a method for end-to-end lineage and version management of medical multimodal data based on a lake-warehouse integrated architecture is also provided, applied to the aforementioned platform, including the following steps: During the data access phase, metadata and data source dependency information of the accessed multimodal medical data are automatically collected as initial lineage nodes. The accessed structured data is written into a tabular storage that supports transactions and time travel capabilities and an initial version snapshot is generated. The accessed unstructured data is stored in object storage and its file identifier and metadata information are associated with it and written into the tabular storage to generate an initial version snapshot. Establish a mapping relationship between the initial lineage node and the corresponding initial version snapshot and store it in the lineage graph library.
[0041] The specific end-to-end lineage and version fusion management process is as follows: This solution simultaneously achieves lineage collection and version association at all stages of the data lifecycle. The specific steps are as follows: Step S1: Data Access Phase - Lineage and Version Initialization S11 Metadata and Lineage Collection: Through the Atlas extended medical multimodal collection plugin, it automatically collects metadata information of the access data, including the table structure and field information of structured data, and the file type and business attributes of unstructured data. At the same time, it records the access task name, execution subject, collection time, data source and other dependency information, and reports it to the lineage service center.
[0042] S12 Initial Version Generation: For accessed structured data, it is directly written to the Iceberg table and an initial version snapshot is automatically generated, and a globally unique version ID is assigned; for unstructured data, it is written to object storage and a unique file identifier is generated. The identifier, index information and metadata are associated and written to the Iceberg table, an initial version snapshot is generated synchronously, and a mapping relationship between the initial version and the current lineage node is established.
[0043] After the S13 data is written, the lineage node information, associated version ID, and storage path are stored in the lineage graph library.
[0044] Step S2: Data Preprocessing Stage - Lineage Link Expansion and Version Update S21 Processing Dependency Collection: For structured data preprocessing tasks, Atlas automatically collects the SQL execution logic and the dependency relationships of input / output tables / fields, and reports them to the lineage service center; for unstructured data preprocessing tasks (such as image format conversion and desensitization processing), the collection plugin automatically records the relationship between the input files, processing logic, and output files of the processing task, and reports them synchronously to the lineage service center.
[0045] S22 Processing Version Synchronization Generation: After the pre-processing task is completed, the output results of structured data are written to the Iceberg table, and a new version snapshot is automatically generated through Iceberg time travel; after the processing results of unstructured data are written to object storage, the file identifier and index information in the Iceberg table are updated, and the corresponding version snapshot is generated synchronously.
[0046] The S23 lineage service center constructs the input nodes, output nodes, and processing logic of the current processing task into a new lineage link segment. At the same time, it binds the output node with the version ID generated this time and updates it to the lineage graph library, forming a traceable processing link evidence chain.
[0047] Step S3: Data Processing and Dataset Generation Stage - End-to-End Lineage Closure and Version Consolidation S31 Full-Link Dependency Collection: For business-oriented processing tasks (such as dataset construction, AI training data generation, and report calculation), Atlas automatically collects all input data sources (including structured tables and unstructured file directories), processing logic, and dependencies of the output dataset. At the same time, it collects attribute information such as project information, user entity, and business scenario, and reports it to the lineage service center.
[0048] S32 Business Version Fixation: After the processing task is completed, the output structured and unstructured data are uniformly generated into a version snapshot through Iceberg, which supports setting version tags (such as "Clinical Research V1.0" and "AI Training Dataset 202601") to achieve the fixation and retention of business versions.
[0049] Based on the full-link dependency relationships collected by Atlas (including metadata and logical dependencies of upstream input, current processing, and downstream output), the S33 lineage service center dynamically associates the current processing node with upstream and downstream nodes through three stages: dependency relationship resolution, automatic identification of context nodes, and lineage edge association. Atlas automatically parses the task's input (source table / file), processing logic (SQL / operator), and output (target table / file), extracting the dependency chain of "input → processing → output". By matching metadata rules (consistency of table name, path, and task ID), the upstream parent node and downstream child node of the current node are automatically identified. The lineage service center associates the upstream and downstream nodes with the current node through lineage edges such as "input dependency" and "output generation", forming a complete lineage link from the original data access to the generation of the business dataset. At the same time, the final output node is bound to the fixed business version, completing the full-link lineage and version mapping closed loop.
[0050] III. Application of Methods and Mechanisms 1. Analysis of the Influence of Bloodline on Data In response to a user's operation on a specified node in the lineage graph, based on the bidirectional mapping graph of lineage and version, the system automatically traverses and identifies all downstream dependent nodes of the specified node. The downstream dependent nodes include data tables, data processing tasks, reports, indicators, datasets, and AI models that directly or indirectly depend on the specified node. An impact scope analysis report is generated and output, which clearly lists all affected downstream dependent nodes and related tasks that may need to be re-executed.
[0051] Supports full-chain lineage retrieval and impact analysis: Supports filtering of lineage links by data type (data table, field, file / directory), business domain (data access, pre-processing, project management), lineage depth, etc., and visualizes the entire flow path from raw data to business applications; When the data of a specified node changes or an anomaly occurs, it automatically traverses the lineage graph, assesses the impact of the node change on all downstream tasks, reports, datasets, and AI models, and outputs an impact analysis report to avoid business failures.
[0052] 2. Multimodal data linkage rollback In response to an instruction to roll back a target node in the lineage graph to a specified historical version, based on the bidirectional mapping graph of lineage and version, the system automatically retrieves and locates the specified historical version snapshot corresponding to the target node, as well as the historical version snapshots of all upstream and downstream nodes that are related to the target node at the same time or logically corresponding to the target node; triggers the storage engine's time travel function to batch roll back the structured data tables and unstructured data files involved in the target node and all its associated upstream and downstream nodes to the data consistency state corresponding to the specified historical version snapshot.
[0053] The platform also supports multimodal data version management and linked rollback: it allows for quick querying of all historical versions through lineage nodes, providing version difference comparison capabilities (including metadata changes and data content comparison); when data errors occur and rollback is required, any historical version of any lineage node can be specified, and the system automatically retrieves all upstream and downstream dependent data associated with that version, and rolls back in batches to a consistent state at the same point in time based on Iceberg time travel, solving the problem of poor consistency in multimodal data rollback in traditional solutions.
[0054] 3. Support for medical data compliance auditing and scientific research replication Automatically records and stores operation logs, processing entities, and data change content of each lineage node in the entire data chain during data access, processing, and version generation; in response to audit or reproduction requests, for the specified final business data output node and its associated business version tag, based on the bidirectional mapping map of lineage and version, traces back and outputs the complete processing chain from the original data source to the output node, as well as the version snapshot information and operation records corresponding to each node in the chain, in order to generate a compliance audit report or provide a precise scientific research data reproduction path.
[0055] It also provides compliance audit and data reproducibility support: automatically records operation logs, processing entities, and changes for all versions across the entire chain, and outputs compliance audit reports that meet the requirements of the "Administrative Measures for Cybersecurity of Medical and Health Institutions"; for medical research scenarios, it can completely reproduce the generation process of datasets through bloodline links and corresponding versions, ensuring the traceability and reproducibility of research results.
[0056] Therefore, this invention can be applied to products such as medical data lake warehouse systems, medical big data platforms, and multimodal medical data management platforms.
[0057] To more clearly illustrate the technical solution of the present invention, the following describes the implementation process and technical effects of the present invention in detail, taking into account a specific medical multimodal data application scenario.
[0058] Example 1 Scenario Setting: Suppose a hospital's data platform needs to provide data support for a research project titled "Training an AI Model for Pneumonia Auxiliary Diagnosis Based on CT Images." The data processing chain for this project is as follows: Data access: Access the "Patient Basic Information Form" and a batch of original chest CT image files from the hospital information system.
[0059] Data preprocessing: Desensitization and cleaning of the "Patient Basic Information Form"; format standardization and anonymization of the original CT images.
[0060] Data processing and dataset generation: The processed structured information is correlated and filtered with unstructured images to generate the final "Pneumonia Research Dataset V1.0" that can be used for model training.
[0061] 1. Implementation process of the example Step S1: Lineage and Version Initialization in the Data Access Phase S11 Metadata and Lineage Collection: The Atlas acquisition plugin automatically captures the access task task_collect_patient.
[0062] The structured data source collected is the ods_patient_info table, which contains fields such as patient_id, name, age, and gender.
[0063] Unstructured data source acquired: object storage directory ofs: / / server / hot / imaging / ct_scans / 2026 / 01 / , containing 1000 raw CT files in .dcm format.
[0064] The bloodline service center records the initial bloodline node: Hospital system -> task_collect_patient -> Output node ods_patient_info table and directory ofs: / / server / hot / imaging / ct_scans / 2026 / 01 / .
[0065] S12 initial version generated: The ods_patient_info table is written in Iceberg format, automatically generating the initial version V1.0.
[0066] The CT file is stored in object storage, and its path, MD5 value and other index information are written as a record into the Iceberg table imaging_file_index, and the corresponding initial version V1.0 is generated.
[0067] Establish mapping: bloodline node ods_patient_info bound to version V1.0; bloodline node directory ofs: / / server / hot / imaging / ct_scans / 2026 / 01 / bound to version V1.0.
[0068] Step S2: Lineage Link Expansion and Version Update in the Data Preprocessing Stage S21 Processing Depends on Acquisition: Task 1 (Structured Data Masking): Execute the SQL task task_clean_patient, read data from ods_patient_info, mask the name field, and output it to the dwd_patient_info_clean table. Atlas collects the input, output, and processing logic dependencies.
[0069] Task 2 (Image Standardization): Execute the standardization task `task_convert_dcm`, read files in the `ofs: / / server / hot / imaging / ct_scans / 2026 / 01 / ` directory, convert them to .png format, remove privacy information, and output them to the `ofs: / / server / hot / imaging / standardized / ct_scans / 2026 / 01 / ` directory. The Atlas acquisition plugin records this file-level processing dependency.
[0070] S22 processing version generated synchronously: The dwd_patient_info_clean table is written to Iceberg, generating a new version V1.2 (version ID: ver_patient_clean).
[0071] The standardized PNG files are written to a new directory, and the imaging_file_index table is updated, generating a new version V1.2.
[0072] S23 Bloodline Update: A new link has been added to the bloodline service center: ods_patient_info -> task_clean_patient -> dwd.patient_info_clean.
[0073] New link: ofs: / / server / hot / imaging / ct_scans / 2026 / 01 / -> task_convert_dcm-> ofs: / / server / hot / imaging / standardized / ct_scans / 2026 / 01 / .
[0074] Bind to the new version: bind the node dwd_patient_info_clean to V1.2; bind the node ofs: / / server / hot / imaging / standardized / ct_scans / 2026 / 01 / to V1.2.
[0075] Step S3: End-to-end lineage closed loop and version solidification in the data processing and dataset generation stages S31 end-to-end dependency collection: The core processing task `task_build_dataset` is executed. Its logic is to associate the `dwd_patient_info_clean` table with the `ofs: / / server / hot / imaging / standardized / ct_scans / 2026 / 01 / ` directory, filter out cases "diagnosed as suspected pneumonia", and generate the final `ads_pneumonia_dataset` (a structured table containing patient IDs and image path associations) and the corresponding image file subset directory `ofs: / / server / hot / imaging / training / pneumonia / 202601 / `.
[0076] Atlas automatically collects all upstream dependencies for this task: the dwd_patient_info_clean table, the ofs: / / server / hot / imaging / standardized / ct_scans / 2026 / 01 / directory, and the processing logic (SQL code).
[0077] S32 service version fixation: After the task is successful, the ads_pneumonia_dataset table generates an Iceberg version snapshot and is labeled with the business tag "Pneumonia Research Dataset V1.0".
[0078] Information from a subset of the output image files is synchronously recorded in the imaging_file_index, generating associated versions.
[0079] S33 Closed-loop bloodline: The bloodline service center added task_build_dataset as a new node to the graph. Its upstream nodes are dwd_patient_info_clean and ofs: / / server / hot / imaging / standardized / ct_scans / 2026 / 01 / , and its downstream nodes are ads_pneumonia_dataset and ofs: / / server / hot / imaging / training / pneumonia / 202601 / .
[0080] The final output node ads_pneumonia_dataset is bound to the fixed business version V1.0.
[0081] This completes the "lineage-version bidirectional mapping map" from raw data to business datasets.
[0082] 2. Demonstration of Technical Effects and Data Examples of the Implementation Examples Based on the constructed graph above, the core service capabilities and quantification effects of the present invention are demonstrated: Scenario A: Impact Analysis (When Upstream Data Needs to Be Changed) Scenario: The hospital discovers a large number of data entry errors in the age field of the ods_patient_info table, requiring correction and rerun of the process.
[0083] Traditional approach: Operations personnel need to manually compile all tasks, scripts, and reports that may use this table, which is prone to omissions. Impact assessment relies on personnel experience, takes approximately 2-4 hours, and carries the risk of incomplete assessments.
[0084] Applying this invention: Select the lineage node ods_patient_info in the service layer interface and execute "Impact Analysis".
[0085] Demonstration: The system automatically traverses the pedigree graph and generates a report within one minute, accurately listing all affected downstream nodes: task_clean_patient -> dwd_patient_info_clean -> task_build_dataset -> ads_pneumonia_dataset, as well as the final AI training file directory. The report clearly indicates that three tasks need to be re-executed.
[0086] Quantitative results: The impact analysis time has been reduced from hours to minutes (<2 minutes), with an accuracy of 100%, avoiding the risk of business interruption due to dependency omissions.
[0087] Scenario B: Multimodal data linkage rollback (when processing logic fails) Scenario: An error was found in the filtering logic of the task_build_dataset, which caused the data in ads_pneumonia_dataset (structured table) and ofs: / / server / hot / imaging / training / pneumonia / 202601 / (image file set) to be incorrect, and the entire system needs to be rolled back to yesterday's state.
[0088] Traditional approach: This requires separately locating historical snapshots of the structured table and historical backups in the object storage directory, manually coordinating rollback points, which is complex and prone to causing inconsistencies between table and file states. The rollback process typically takes several hours.
[0089] Applying this invention: In the service layer, query the historical versions of the node ads_pneumonia_dataset, select the correct version snapshot V1.0 from yesterday, and execute "version rollback".
[0090] Demonstration: The system automatically locates the following through "lineage-version mapping": The target table ads_pneumonia_dataset needs to be rolled back to Iceberg version V1.0.
[0091] The associated unstructured file set directory needs to be rolled back to the file snapshot state mapped to the same version ID.
[0092] Automatically trigger linked rollback operations to ensure that structured and unstructured data are restored to the same consistent point in time in terms of business logic.
[0093] Quantitative effect: The complex multimodal data consistency rollback operation is simplified to a single click, the rollback time is reduced from several hours to minutes (about 5 minutes), and the data inconsistency problem caused by rollback asynchrony is completely eliminated.
[0094] Scenario C: Compliance Audit and Scientific Research Reproduction Scenario: One year later, the auditing department examines the entire data flow of the pneumonia research project, or another research team needs to reproduce the dataset.
[0095] By applying this invention: auditors or researchers only need to locate the final dataset ads_pneumonia_dataset and its version, pneumonia research dataset V1.0.
[0096] Demonstration of effects: Audit: The system can automatically output the entire chain from the original data in the hospital system to the final dataset generation, including operation logs at each node, processing entities, and data version change records, forming a complete chain of compliance evidence.
[0097] Reproduction: Researchers can use the complete lineage diagram to roll back (or re-execute) each node to the exact version (V1.0, V1.2, etc.) that was used to generate the V1.0 dataset, and completely reproduce the dataset generation process.
[0098] Quantifiable results: The time required for manually compiling audit materials is reduced from several person-days to instant generation. It provides automated, version snapshot-based, and precise support for the reproducibility of scientific research, going beyond simply enabling repeatable code execution.
[0099] 3. Summary of Implementation Examples This embodiment fully demonstrates the three core steps of this invention—"synchronous generation of bloodline data and version snapshots," "bidirectional mapping map construction," and "map-based intelligent services"—through a real-world medical multimodal data processing pipeline. Data shows that this invention effectively addresses the three major pain points mentioned in the background section: "opaque dependency relationships," "low efficiency in problem localization," and "insufficient backtracking capabilities." It achieves a significant improvement in governance efficiency and a marked reduction in compliance costs, fully demonstrating the inventiveness, practicality, and technical advantages of this invention.
[0100] The above embodiments can be implemented, in whole or in part, by software, hardware (such as circuits), firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. A semiconductor medium can be a solid-state drive.
[0101] It should be understood that the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. A and B can be singular or plural. Additionally, the character " / " in this article generally indicates an "or" relationship between the preceding and following related objects, but it can also represent an "and / or" relationship. Please refer to the context for a more accurate understanding.
[0102] In this invention, "at least one" means one or more, and "more than one" means two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of a single item or a plurality of items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be a single item or multiple items.
[0103] It should be understood that, in various embodiments of the present invention, the order of the above-mentioned process numbers does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0104] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0105] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the devices, apparatuses, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0106] In the several embodiments provided by this invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0107] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0108] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0109] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0110] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A medical multimodal data end-to-end lineage and version management platform based on a lake-warehouse integrated architecture, characterized in that, include: The lake warehouse storage layer is used to uniformly store structured, semi-structured, and unstructured data in the medical field using an architecture that integrates data lake and data warehouse, and to achieve data version snapshot management through the native time travel capability of the underlying storage engine. The fusion acquisition layer is used to synchronously and automatically collect data metadata and flow dependencies in each processing stage of the data lifecycle based on the extended metadata management plugin to form lineage information, and trigger the lake warehouse storage layer to generate corresponding data version snapshots, establishing a mapping relationship between lineage information and version snapshots; The service layer is used to provide functions such as bloodline retrieval and impact analysis, data version management and batch rollback, and audit report generation that meet medical compliance requirements, based on the bloodline and version mapping relationship map constructed by the fusion acquisition layer.
2. A method for end-to-end lineage and version management of multimodal medical data based on a lake-warehouse integrated architecture, characterized in that, Applied to the platform of claim 1, comprising the steps of: During the data access phase, metadata and data source dependency information of the accessed multimodal medical data are automatically collected as initial lineage nodes. The accessed structured data is written into a tabular storage that supports transactions and time travel capabilities and an initial version snapshot is generated. The accessed unstructured data is stored in object storage and its file identifier and metadata information are associated with it and written into the tabular storage to generate an initial version snapshot. Establish a mapping relationship between the initial lineage node and the corresponding initial version snapshot and store it in the lineage graph library.
3. The method according to claim 2, characterized in that, The data preprocessing stage also includes the following steps: Automatically collect input data, processing logic, and output data dependencies of pre-processing tasks, and report them as newly added lineage information; After the preprocessing task is completed, the output structured data is written to the tabular storage to automatically generate a new version snapshot, or the output unstructured data processing results are updated to the object storage and the associated tabular storage to generate a new version snapshot; The pre-processing task is constructed as a new lineage link segment, which includes an input node, a processing logic node, and an output node. The output node is then bound to the newly generated snapshot and updated to the lineage graph library.
4. The method according to claim 3, characterized in that, The data processing and dataset generation stage also includes the following steps: Automatically collect all upstream input data sources, processing logic, and downstream output dataset dependencies involved in business-oriented data processing tasks, as well as related project and business scenario attribute information; After the processing task is completed, a fixed version snapshot with business tags is generated in the tabular storage for the output structured and unstructured data; Based on the collected full-link dependency relationship, the current processing node is dynamically associated with its upstream and downstream lineage nodes through dependency relationship resolution and automatic identification of context nodes, forming a complete lineage link from the original data access to the final business dataset output. The output node of the final business dataset is then bound to the fixed business version snapshot, completing the closed loop of lineage link and version mapping.
5. The method according to claim 4, characterized in that, The phrase "automatically collecting the dependencies between all upstream input data sources, processing logic, and downstream output datasets involved in business-oriented data processing tasks" specifically includes: The metadata management framework's acquisition plugin automatically parses the input source table or file, the definition of the processing logic, and the output target table or file of the data processing task, extracting the complete dependency chain of "input-processing-output".
6. The method according to claim 4, characterized in that, The phrase "dynamically associating the current processing node with its upstream and downstream lineage nodes" specifically includes: The upstream parent node and downstream child node of the current processing node are automatically identified through metadata matching rules; The bloodline service center associates the upstream parent node, the current processing node, and the downstream child node through relationship edges of the type "input dependency" and "output generation".
7. The method according to any one of claims 2 to 6, characterized in that, The method further includes: constructing a bidirectional mapping graph of lineage and version based on the lineage graph library, wherein each lineage node in the graph is associated with one or more historical version snapshots.
8. A method for analyzing the impact of data lineage on data based on the platform of claim 1 or the method of any one of claims 2 to 7, characterized in that, include: In response to a user's operation on a specified node in the lineage graph, based on the bidirectional mapping graph of lineage and version, the system automatically traverses and identifies all downstream dependent nodes of the specified node. The downstream dependent nodes include data tables, data processing tasks, reports, indicators, datasets, and AI models that directly or indirectly depend on the specified node. An impact scope analysis report is generated and output, which clearly lists all affected downstream dependent nodes and related tasks that may need to be re-executed.
9. A multimodal data linkage rollback method based on the platform of claim 1 or the method of any one of claims 2 to 7, characterized in that, include: In response to an instruction to roll back a target node in the lineage graph to a specified historical version, based on the bidirectional mapping graph of lineage and version, the system automatically retrieves and locates the specified historical version snapshot corresponding to the target node, as well as the historical version snapshots of all upstream and downstream nodes that are related to the target node at the same time or logically corresponding to the target node; triggers the storage engine's time travel function to batch roll back the structured data tables and unstructured data files involved in the target node and all its associated upstream and downstream nodes to the data consistency state corresponding to the specified historical version snapshot.
10. A method for supporting medical data compliance auditing and scientific research reproduction based on the platform of claim 1 or the method of any one of claims 2 to 7, characterized in that, include: Automatically record and store the operation logs, processing entities, and data change content of each lineage node in the entire data chain during data access, processing, and version generation; In response to audit or reproduction requests, for the specified final business data output node and its associated business version tag, based on the bidirectional mapping graph of lineage and version, the complete processing link from the original data source to the output node is traced back and output, as well as the version snapshot information and operation records corresponding to each node on the link, in order to generate a compliance audit report or provide a precise scientific research data reproduction path.