Medical big data processing method, device and equipment based on lake and warehouse integration and medium

By unifying the access and processing of medical data through a lake-warehouse integrated architecture, the problems of data dispersion and inefficiency in traditional architectures are solved, enabling efficient and real-time data processing and analysis, and supporting the development of smart healthcare.

CN120977466APending Publication Date: 2025-11-18山东浪潮智慧医疗科技有限公司
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510834871.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Traditional medical big data processing architecture suffers from data fragmentation, multimodal collaboration failure, and low research efficiency, failing to meet the needs of medical business and management decision-making. Furthermore, it is difficult to guarantee throughput efficiency, real-time performance, and consistency when processing large-scale data.

Method used

Adopting a lake-warehouse integrated architecture, medical data is uniformly accessed through full synchronization, incremental synchronization, or data virtualization. The data is then cleaned, standardized, and structurally transformed to build thematic databases for specific business topics, and query interfaces and analysis services are provided.

Benefits of technology

It improves data utilization efficiency, meets the data needs of different business scenarios, ensures data integrity and real-time performance, reduces system complexity and operation and maintenance costs, and supports automated operation and maintenance and intelligent monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120977466A_ABST
    Figure CN120977466A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of data processing, and particularly relates to a medical big data processing method, device and equipment based on lake and warehouse integration and a medium, the method comprises the following steps: collecting original medical data, and uniformly accessing the original medical data to a data access layer of a lake and warehouse integrated architecture in an original format and full-amount information to form a source pasting library; preprocessing the original medical data in the source library, performing unified coding and format conversion according to the medical data standard set, generating a standardized data set, and storing the standardized data set in a normalization library; on the basis of data in the normalization library and the source pasting library, multi-modal medical data are associated and integrated through cross-source federated query and computational flow technologies according to business scene requirements of clinical diagnosis and treatment, scientific research analysis or management decision, a special topic library oriented to a specific business topic is constructed, and a special topic library view is generated; and providing a query interface and an analysis service for the data application platform through the thematic database. The data processing speed and efficiency are improved, delay is reduced, and the data throughput efficiency is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of data processing technology, specifically relating to a method, apparatus, equipment, and medium for medical big data processing based on a lake-warehouse integration. Background Technology

[0002] In the medical field, medical big data is complex and diverse, including structured electronic medical records, semi-structured test reports, and unstructured medical images. At the same time, it has extremely high requirements for the real-time performance, security, and long-term storage of data, which poses challenges to traditional data processing architectures.

[0003] In traditional architectures, data is often scattered across multiple systems such as HIS, EMR, and PACS, leading to data fragmentation, failure of multimodal collaboration, low research efficiency, and accumulated technical debt, failing to meet the needs of medical business and management decision-making. Traditional solutions typically employ a data lake + data warehouse architecture, first aggregating data into a data lake for unified storage, and then improving data quality through a data warehouse. To enhance the governance capabilities of the data lake, a layered data warehouse design is usually adopted. For example, heterogeneous data is uniformly collected at the ODS layer (source layer), real-time and offline data are integrated into wide tables at the DWD layer (detail layer) to build a streaming ETL data processing link, and finally, common indicators are formed and data interfaces are exposed at the DWS layer (summary layer) to provide services.

[0004] However, this method consumes a lot of system resources, and each ETL process requires manual design. When processing large-scale data, it is difficult to guarantee data throughput efficiency, real-time performance, and consistency. Summary of the Invention

[0005] In view of the above-mentioned shortcomings of the prior art, the present invention provides a medical big data processing method, device, equipment and medium based on lake-warehouse integration to solve the above-mentioned technical problems.

[0006] In a first aspect, the present invention provides a medical big data processing method based on a lake-warehouse integration, comprising: S1. Collect raw medical data from various medical business systems, and integrate the raw medical data into the data access layer of the lake warehouse integrated architecture in its original format and with all information to form a source library; S2. Clean, standardize, and structure the original medical data in the source database, and perform unified encoding and format conversion according to the medical data standard set to generate a standardized dataset and store it in the unified database. S3. Based on the data in the unified database and the source database, through cross-source federated query and computation flow technology, multimodal medical data is associated and integrated according to the business scenario needs of clinical diagnosis and treatment, scientific research analysis or management decision-making, to build a thematic database for specific business themes and generate thematic database views. S4. Provide query interfaces and analysis services for data application platforms through thematic databases.

[0007] Based on data from the unified and source-attached databases, cross-source federated queries and computational flow technologies enable the construction of thematic databases tailored to specific business scenarios, generating thematic database views. This provides flexible data support for clinical diagnosis and treatment, scientific research analysis, and management decision-making. By providing query interfaces and analysis services to data application platforms through these thematic databases, business personnel can quickly obtain the data they need, improving data utilization efficiency and providing strong support for the development of smart healthcare.

[0008] As a further limitation of the technical solution of the present invention, step S1 specifically includes: S11. Identify the medical data sources to be collected from hospital information systems, electronic medical record systems, medical image archiving and communication systems, medical data platforms, and regional medical shared data platforms; S12. Medical data from medical data sources is uniformly connected to the data access layer of the lake warehouse integrated architecture through full synchronization, incremental synchronization, or data virtualization to form a source database; among which, medical data is stored in the source database in its original format, retaining all fields and metadata; S13. Add unified metadata tags to the data in the post source library, including data source, collection time, data format and business meaning description; and establish a full-link lineage relationship for the data in the post source library, recording the data flow path from the source system to the post source library.

[0009] By unifying medical data from various data sources through full synchronization, incremental synchronization, or data virtualization into the data access layer of the lakeware architecture, a source database is formed. This meets the access needs of different data sources, improving the flexibility and efficiency of data access. Medical data is stored in the source database in its original format, retaining all fields and metadata. A unified metadata tag is added to the data in the source database, establishing a complete lineage relationship, facilitating data traceability and management, and ensuring data integrity and traceability.

[0010] As a further limitation of the technical solution of the present invention, in S12, the step of uniformly connecting medical data from the medical data source to the data access layer of the lake warehouse integrated architecture through full synchronization, incremental synchronization, or data virtualization to form a source library includes: S121. Perform full synchronization of initial data or periodically updated data in medical data sources through batch processing to completely extract the data into the post source library. S122. For real-time changes in medical data sources, monitor the database logs of medical data sources using change data capture technology, and capture incremental data in real time using stream processing and synchronize it to the post source database. S123. For external data in medical data sources that cannot be directly physically integrated, establish logical mapping relationships through data virtualization technology to achieve dynamic access to cross-system data; In S121, batch processing full synchronization is triggered according to a preset cycle, and the data range is marked based on timestamp or version number during extraction. In S122, the stream processing incrementally parses the database transaction log in real time to capture insert, update, and delete operation events.

[0011] By combining batch full synchronization, stream incremental synchronization, and data virtualization technologies, data from medical data sources can be efficiently synchronized to the patch database, ensuring data integrity while improving data real-time performance and update efficiency. Batch full synchronization is triggered at preset intervals, and data ranges are marked based on timestamps or version numbers during extraction, enabling precise extraction of initial data or periodically updated data. Stream incremental synchronization parses the database transaction log in real time, capturing insert, update, and delete operation events, and can capture real-time changes in the data source to ensure data timeliness.

[0012] As a further limitation of the technical solution of the present invention, in S123, the logical mapping relationship is implemented in the following way: S1231. For external API data, encapsulate REST / SOAP interface calls into virtual tables; specifically including: Register the access endpoints, authentication methods, and data formats of external APIs in the lakewarehousing integrated architecture; Map the query parameters required by the API to queryable fields of a virtual table; Parse the JSON / XML data returned by the API, extract structured fields, and define the column names and data types of the virtual table; Based on the parsed data model, a virtual table is created in the lake warehouse integrated architecture. This table only stores metadata, and the actual data is dynamically obtained by calling the API in real time. When the application layer queries the virtual table, it automatically converts the SQL or analysis request into an API call, retrieves data in real time, and returns the results. S1232. Establish a streaming virtual view for message queue data; specifically including: Configure message queue connection parameters in a lakeware architecture, including server address, topic name, consumer group, and deserialization format; Parse the data structure in the message queue, define the field names, data types, and timestamp fields of the virtual view, and register them with the metadata management system; Based on real-time data streams from message queues, a virtual view is created in the lakeware architecture that dynamically consumes the latest messages on demand. Micro-batch processing of message streams by time window or event window generates queryable temporary datasets; Perform real-time queries or subscribe to data changes on virtual views using a streaming SQL engine or API interface.

[0013] By encapsulating external API data into virtual tables and establishing streaming virtual views with message queues, flexible access and dynamic access to external data that cannot be directly physically integrated are achieved. This expands data sources, enriches data types, and provides support for the fusion analysis of multimodal medical data. Creating virtual tables and views based on the parsed data schema enables real-time acquisition of data from external APIs and message queues. Real-time queries or subscriptions to data changes can be executed through a streaming SQL engine or API interface, improving data real-time performance and availability, and meeting the needs of real-time business scenarios.

[0014] As a further limitation of the technical solution of the present invention, step S2 specifically includes: S21. Perform integrity checks, validity verifications, and consistency checks on the original medical data in the patch source database; S22. Based on a predefined set of medical data standards, map non-standard terms in the original medical data to standard codes, and perform unified conversion on data units and formats; S23. Perform structured processing on the data in the source library after format conversion; including text recognition of unstructured data and extraction of key information through regular expressions, converting the extracted key information into structured data; parsing semi-structured data and extracting the structured data from it; S24. Store the processed structured data in a unified database according to subject domains, and establish a lineage tracing relationship with the source database data.

[0015] Based on a predefined set of medical data standards, non-standard terms in the raw medical data are mapped to standard codes, and data units and formats are uniformly converted, achieving data standardization and normalization. This provides a unified data format and standard for subsequent data processing and analysis. The converted data undergoes structuring processing, including text recognition and key information extraction from unstructured data and parsing of semi-structured data. This process converts multimodal medical data into structured data, improving its usability and analyzability.

[0016] As a further limitation of the technical solution of the present invention, step S3 specifically includes: S31. Based on the business scenario needs of clinical diagnosis and treatment, scientific research analysis, or management decision-making, determine the thematic direction of the topic database, including the clinical diagnosis and treatment topic database, the disease research topic database, and the hospital operation topic database; S32. Using a unified query engine, perform a joint query on the structured data in the unified database, the original data in the source database, and the external data source to extract multimodal medical data related to the topic, including structured data, semi-structured data, and unstructured data. S33. Dynamically scheduled computing tasks process the extracted multimodal medical data, including direct computation of raw data at the storage layer; accelerated processing of structured data using the MPP engine; and collaborative analysis of cross-lake warehouse data. S34. Integrate and organize the processed data according to business needs, and build a thematic library for specific business topics; S35. Generate SQL views, REST APIs, or OLAP cubes based on thematic database tables to display the data in thematic databases; among them, the clinical diagnosis and treatment thematic database contains a view of the entire patient's diagnosis and treatment cycle; the scientific research and analysis thematic database contains disease-specific research datasets; and the management decision-making thematic database contains a cube of hospital operation indicators.

[0017] By executing joint queries through a unified query engine and dynamically scheduling computational tasks to process the extracted multimodal medical data, the system can efficiently integrate and organize data, improving the efficiency and flexibility of data processing and meeting the data processing needs of different business scenarios. SQL views, REST APIs, or OLAP cubes are generated based on thematic database tables to display the data in the thematic database, meeting the data display needs of different users and business scenarios, improving data visualization and usability, and providing strong support for clinical diagnosis and treatment, scientific research analysis, and management decision-making.

[0018] As a further limitation of the technical solution of the present invention, in S32, performing a joint query includes: Transform business logic into a cross-source query plan; Push the filtering conditions to the data source for execution; Merge query results from heterogeneous data sources.

[0019] Transforming business logic into a cross-source query plan, pushing filtering conditions to the data source for execution, and merging query results from heterogeneous data sources can optimize the performance of cross-source queries, reduce data transmission volume, improve query efficiency, and ensure the efficiency and accuracy of joint queries.

[0020] Secondly, the present invention also provides a medical big data processing device based on a lake-warehouse integration, comprising: The unified data integration module is used to collect raw medical data from various medical business systems and integrate the raw medical data into the data access layer of the lake warehouse integrated architecture in its original format and with full information to form a source library. The data governance and storage management module is used to clean, standardize, and structure the raw medical data in the source database, and to perform unified encoding and format conversion according to the medical data standard set, generating a standardized dataset and storing it in the unified database. The Lake Warehouse Collaborative Processing Module is used to integrate multimodal medical data based on data in the unified database and the source database, through cross-source federated query and computation flow technology, according to the business scenario needs of clinical diagnosis and treatment, scientific research analysis or management decision-making, to build thematic databases for specific business themes and generate thematic database views; The unified access interface module is used to provide query interfaces and analysis services to the data application platform through thematic databases.

[0021] As a further limitation of the technical solution of the present invention, the unified data integration module includes a data acquisition unit, a data access unit, and a data integration unit; The data acquisition unit is used to identify the medical data sources to be collected from hospital information systems, electronic medical record systems, medical image archiving and communication systems, medical data platforms, and regional medical shared data platforms. The data access unit is used to uniformly access medical data from medical data sources into the data access layer of the lake warehouse integrated architecture through full synchronization, incremental synchronization, or data virtualization, forming a source database; among which, medical data is stored in the source database in its original format, retaining all fields and metadata; The data integration unit is used to add unified metadata tags to the data in the post source library, including data source, collection time, data format and business meaning description; and to establish a full-link lineage relationship for the data in the post source library, recording the flow path of data from the source system to the post source library.

[0022] Thirdly, the present invention also provides an electronic device, the electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; the memory storing computer program instructions executable by the at least one processor, the computer program instructions being executed by the at least one processor to enable the at least one processor to execute the lake-warehouse integrated medical big data processing method as described in the first aspect.

[0023] Fourthly, the present invention also provides a non-transitory computer-readable storage medium that stores computer instructions that cause the computer to execute the lake-warehouse integrated medical big data processing method as described in the first aspect.

[0024] The beneficial effects of this invention are that clinical decision support systems need to extract data from multiple data sources such as EMR, LIS, and PACS, and then load it into a data warehouse for analysis after cleaning and transformation. In traditional hybrid architectures, this process involves data interaction across multiple systems, making it complex and time-consuming. The lakeware warehouse integrated architecture of this application simplifies the ETL process. Through a unified data access and processing platform, it can extract, clean, and transform data in real-time or in batches. For example, it can extract the latest patient medical records from the EMR system in real-time, process them, and directly load them into the data warehouse for use by the clinical decision support system, improving data processing speed and efficiency, reducing latency, and ensuring data throughput efficiency and real-time performance. By unifying data standards and specifications, data is cleaned and transformed according to unified standards during the data acquisition phase, ensuring consistency in format, content, and quality across different data sources.

[0025] This application's lake-warehouse integrated architecture combines the functions of a data lake and a data warehouse. Medical institutions only need to deploy and maintain a unified data management platform to comprehensively manage medical data, reduce system complexity, reduce maintenance workload and costs, and support automated maintenance and intelligent monitoring, thereby improving system stability and reliability. Attached Figure Description

[0026] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0027] Figure 1 This is a schematic flowchart illustrating a method according to an embodiment of the present invention.

[0028] Figure 2 This is a schematic block diagram of an apparatus according to an embodiment of the present invention.

[0029] Figure 3 This is a diagram illustrating the processing application environment in an embodiment of the present invention. Detailed Implementation

[0030] To make the objectives, features, and advantages of this invention more apparent and understandable, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings of the specific embodiments. Obviously, the embodiments described below are only some embodiments of this invention, and not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0031] like Figure 1 As shown, this embodiment of the invention provides a medical big data processing method based on a lake-warehouse integration, including: S1. Collect raw medical data from various medical business systems, and integrate the raw medical data into the data access layer of the lake warehouse integrated architecture in its original format and with all information to form a source library; S2. Clean, standardize, and structure the original medical data in the source database, and perform unified encoding and format conversion according to the medical data standard set to generate a standardized dataset and store it in the unified database. S3. Based on the data in the unified database and the source database, through cross-source federated query and computation flow technology, multimodal medical data is associated and integrated according to the business scenario needs of clinical diagnosis and treatment, scientific research analysis or management decision-making, to build a thematic database for specific business themes and generate thematic database views. S4. Provide query interfaces and analysis services for data application platforms through thematic databases.

[0032] This application, based on a lake-warehouse integrated architecture, collects raw medical data from various business systems and aggregates it into a source database, preserving the original characteristics and full information of the data. The data in the source database is cleaned, standardized, and structurally transformed to form a standardized dataset, which is then stored in a unified database to ensure data quality and consistency. Based on the unified database data, further integration, correlation, and mining are performed to construct specialized databases for clinical diagnosis and treatment, scientific research analysis, and management decision-making scenarios, realizing the extraction of data value. These specialized databases provide efficient query interfaces and analysis services for the data application platform, supporting multi-dimensional data exploration and intelligent decision-making. The entire processing flow relies on the lake-warehouse integrated architecture to achieve the collection, hierarchical storage, flexible scheduling, and fusion analysis of multi-source heterogeneous medical data, ultimately forming a full-chain medical big data processing system covering data collection, governance, modeling, and application, contributing to the construction of smart healthcare.

[0033] In healthcare scenarios, various business systems continuously generate massive amounts of diverse data, such as patient basic information, medical records, and examination and test reports. This data is scattered across different business systems, forming data silos. The data access layer of the lakehouse architecture acts as a powerful data aggregation center, unifying the access of this data from different business systems. Regardless of which business system the data was originally stored in, or its format, the data access layer can integrate the data through physical or logical methods such as full synchronization, incremental synchronization, and data virtualization to form a healthcare data lake. With its flexibility and large capacity, it provides a free space for the aggregation of various types of raw data, ensuring that all data is completely preserved. Step S1 specifically includes: S11. Identify the medical data sources to be collected from hospital information systems, electronic medical record systems, medical image archiving and communication systems, medical data platforms, and regional medical shared data platforms; S12. Medical data from medical data sources is uniformly connected to the data access layer of the lake warehouse integrated architecture through full synchronization, incremental synchronization, or data virtualization to form a source database; among which, medical data is stored in the source database in its original format, retaining all fields and metadata; S13. Add unified metadata tags to the data in the post source library, including data source, collection time, data format and business meaning description; and establish a full-link lineage relationship for the data in the post source library, recording the data flow path from the source system to the post source library.

[0034] It should be noted that in S12, the steps to uniformly connect medical data from medical data sources to the data access layer of the lakewarehouse integrated architecture, forming the source library, through full synchronization, incremental synchronization, or data virtualization, include: S121. Perform full synchronization of initial data or periodically updated data in medical data sources through batch processing to completely extract the data into the post source library. S122. For real-time changes in medical data sources, monitor the database logs of medical data sources using change data capture technology, and capture incremental data in real time using stream processing and synchronize it to the post source database. S123. For external data in medical data sources that cannot be directly physically integrated, establish logical mapping relationships through data virtualization technology to achieve dynamic access to cross-system data; In S121, batch processing full synchronization is triggered according to a preset cycle, and the data range is marked based on timestamp or version number during extraction. In S122, the stream processing incrementally parses the database transaction log in real time to capture insert, update, and delete operation events.

[0035] It should be further explained that in S123, the logical mapping relationship is implemented in the following way: S1231. For external API data, encapsulate REST / SOAP interface calls into virtual tables; specifically including: Register the access endpoints, authentication methods, and data formats of external APIs in the lakewarehousing integrated architecture; Map the query parameters required by the API to queryable fields of a virtual table; Parse the JSON / XML data returned by the API, extract structured fields, and define the column names and data types of the virtual table; Based on the parsed data model, a virtual table is created in the lake warehouse integrated architecture. This table only stores metadata, and the actual data is dynamically obtained by calling the API in real time. When the application layer queries the virtual table, it automatically converts the SQL or analysis request into an API call, retrieves data in real time, and returns the results. S1232. Establish a streaming virtual view for message queue data; specifically including: Configure message queue connection parameters in a lakeware architecture, including server address, topic name, consumer group, and deserialization format; Parse the data structure in the message queue, define the field names, data types, and timestamp fields of the virtual view, and register them with the metadata management system; Based on real-time data streams from message queues, a virtual view is created in the lakeware architecture that dynamically consumes the latest messages on demand. Micro-batch processing of message streams by time window or event window generates queryable temporary datasets; Perform real-time queries or subscribe to data changes on virtual views using a streaming SQL engine or API interface.

[0036] Raw medical data accessed into a data lake often presents numerous problems, such as inconsistent data formats, duplicate or erroneous information, etc. The data governance layer of the lakehouse architecture then comes into play, acting like a meticulous data steward, carefully cleaning and standardizing the data in the data lake. Through a series of business rules and data quality verification mechanisms, noise and redundancy are removed from the data, transforming it into a unified format and standard. The processed data is then stored in the more structured part of the lakehouse architecture; step S2 specifically includes: S21. Perform integrity checks, validity verifications, and consistency checks on the original medical data in the patch source database; S22. Based on a predefined set of medical data standards, map non-standard terms in the original medical data to standard codes, and perform unified conversion on data units and formats; S23. Perform structured processing on the data in the source library after format conversion; including text recognition of unstructured data and extraction of key information through regular expressions, converting the extracted key information into structured data; parsing semi-structured data and extracting the structured data from it; S24. Store the processed structured data in a unified database according to subject domains, and establish a lineage tracing relationship with the source database data.

[0037] In healthcare, different business scenarios have varying data needs. For example, clinical diagnosis and treatment require information such as patient medical records, diagnostic results, and treatment plans; scientific research and analysis focus on data such as disease pathogenesis and treatment efficacy comparisons; and management decision-making requires understanding hospital operational indicators and resource utilization. The lakeware architecture's business theme layer extracts relevant data from the governed data warehouse area based on these different business needs, constructing thematic repositories tailored to specific business themes. Each thematic repository closely revolves around a specific business theme, further integrating and organizing data to form targeted datasets. These thematic repositories act like tailor-made data treasure troves for different business scenarios, fully preparing for subsequent in-depth analysis and enabling business personnel to more easily access data that meets their needs. Specifically, step S3 includes: S31. Based on the business scenario needs of clinical diagnosis and treatment, scientific research analysis, or management decision-making, determine the thematic direction of the topic database, including the clinical diagnosis and treatment topic database, the disease research topic database, and the hospital operation topic database; S32. Using a unified query engine, perform a joint query on the structured data in the unified database, the original data in the source database, and the external data source to extract multimodal medical data related to the topic, including structured data, semi-structured data, and unstructured data. S33. Dynamically scheduled computing tasks process the extracted multimodal medical data, including direct computation of raw data at the storage layer; accelerated processing of structured data using the MPP engine; and collaborative analysis of cross-lake warehouse data. S34. Integrate and organize the processed data according to business needs, and build a thematic library for specific business topics; S35. Generate SQL views, REST APIs, or OLAP cubes based on thematic database tables to display the data in thematic databases; among them, the clinical diagnosis and treatment thematic database contains a view of the entire patient's diagnosis and treatment cycle; the scientific research and analysis thematic database contains disease-specific research datasets; and the management decision-making thematic database contains a cube of hospital operation indicators.

[0038] It should be noted here that in S32, executing a union query includes: converting business logic into a cross-source query plan; pushing filter conditions to the data source for execution; and merging query results from heterogeneous data sources.

[0039] Well-constructed thematic databases provide abundant resources for data applications in healthcare. The data application layer of the lakeware architecture presents the data from these databases to business personnel in an intuitive and convenient way. Whether it's doctors needing to quickly retrieve relevant patient information during clinical decision-making, researchers needing to deeply analyze data on specific diseases during research projects, or managers needing comprehensive operational data to formulate hospital development strategies, all can be easily achieved through the data application layer. The data application layer provides various data query, analysis, and visualization tools, allowing business personnel to explore and analyze data from multiple dimensions according to their business needs without requiring complex technical knowledge, thereby gaining valuable insights and providing strong support for clinical decision-making, scientific research, and healthcare management.

[0040] like Figure 2 As shown, this embodiment of the invention also provides a medical big data processing device based on a lake-warehouse integration, comprising: The unified data integration module is used to collect raw medical data from various medical business systems and integrate the raw medical data into the data access layer of the lake warehouse integrated architecture in its original format and with full information to form a source library. The data governance and storage management module is used to clean, standardize, and structure the raw medical data in the source database, and to perform unified encoding and format conversion according to the medical data standard set, generating a standardized dataset and storing it in the unified database. The Lakewarehouse collaborative processing module is used to integrate multimodal medical data based on data from the unified database and the data source database. Through cross-source federated queries and computational flow technology, it connects and integrates data according to business scenarios such as clinical diagnosis and treatment, scientific research analysis, or management decision-making. This allows for the construction of thematic databases for specific business themes and the generation of thematic database views. This module enables efficient collaborative processing of medical big data. The cross-source federated query function breaks down data silos within the Lakewarehouse. For example, when diagnosing complex diseases, doctors need to simultaneously refer to the patient's imaging data (stored in the data lake) and historical medical records (stored in the data warehouse). Cross-source federated queries can quickly and jointly analyze these data, providing comprehensive support for diagnosis. Lakewarehouse query acceleration achieves rapid data retrieval through technologies such as caching, index optimization, and materialized views. For instance, in medical emergency scenarios, medical personnel need to quickly retrieve patients' past medical history and allergy information; query acceleration technology ensures accurate data is obtained in a short time. The computation flow dynamically schedules computational tasks between the data lake and the data warehouse. For example, the raw medical image data in the data lake is preprocessed (such as image enhancement and feature extraction), and then the processed data is flowed into the data warehouse for further analysis and mining to assist in the formulation of disease diagnosis and treatment plans. This ensures efficient collaboration of data in a heterogeneous environment and improves the efficiency of medical big data processing and analysis.

[0041] The unified access interface module provides query interfaces and analysis services to the data application platform through specialized databases. This module offers unified services externally via standardized interfaces (such as JDBC and REST API) to meet the diverse application needs of healthcare operations. In terms of report analysis, hospital management can obtain departmental operational data reports through the interface to understand departmental workload, revenue, and other information. The ad-hoc query function allows doctors to query relevant patient information at any time according to their needs, such as querying the treatment status of patients with specific diseases within a certain period. The user profiling function helps healthcare institutions understand patients' health status and medical habits, supporting personalized healthcare services. Through these standardized interfaces, different healthcare business systems and users can easily access and use medical data in the lake repository, maximizing data value while improving system compatibility and scalability, facilitating integration with other healthcare systems.

[0042] In some embodiments, the unified data integration module includes a data acquisition unit, a data access unit, and a data integration unit; The data acquisition unit is used to identify the medical data sources to be collected from hospital information systems, electronic medical record systems, medical image archiving and communication systems, medical data platforms, and regional medical shared data platforms. The data access unit is used to uniformly access medical data from medical data sources into the data access layer of the lake warehouse integrated architecture through full synchronization, incremental synchronization, or data virtualization, forming a source database; among which, medical data is stored in the source database in its original format, retaining all fields and metadata; The data integration unit is used to add unified metadata tags to the data in the post source library, including data source, collection time, data format and business meaning description; and to establish a full-link lineage relationship for the data in the post source library, recording the flow path of data from the source system to the post source library.

[0043] In some embodiments, the data access unit includes a full synchronization access submodule, an incremental synchronization access submodule, and a virtualization access submodule; The full synchronization access submodule is used to perform full synchronization of initial data or periodically updated data in medical data sources through batch processing, and completely extract the data into the post source library. The incremental synchronization access submodule is used to monitor the database logs of medical data sources in real time by using change data capture technology, and to capture incremental data in real time and synchronize it to the post source library in a stream processing manner. The virtualization access submodule is used to establish logical mapping relationships for external data in medical data sources that cannot be directly physically integrated, thereby enabling dynamic cross-system data access. Among them, the full synchronization access submodule batch processing full synchronization is triggered according to a preset cycle, and the data range is marked based on timestamp or version number during extraction. The incremental synchronization access submodule stream processing incremental synchronization parses the database transaction log in real time, capturing insert, update, and delete operation events.

[0044] In some embodiments, the virtualization access submodule establishes logical mapping relationships through the following units: The first processing subunit is used to encapsulate external API data, REST / SOAP interface calls, into virtual tables; specifically, it is used for: Register the access endpoints, authentication methods, and data formats of external APIs in the lakewarehousing integrated architecture; Map the query parameters required by the API to queryable fields of a virtual table; Parse the JSON / XML data returned by the API, extract structured fields, and define the column names and data types of the virtual table; Based on the parsed data model, a virtual table is created in the lake warehouse integrated architecture. This table only stores metadata, and the actual data is dynamically obtained by calling the API in real time. When the application layer queries the virtual table, it automatically converts the SQL or analysis request into an API call, retrieves data in real time, and returns the results. The second processing subunit is used to create a streaming virtual view of the message queue data; specifically, it is used for: Configure message queue connection parameters in a lakeware architecture, including server address, topic name, consumer group, and deserialization format; Parse the data structure in the message queue, define the field names, data types, and timestamp fields of the virtual view, and register them with the metadata management system; Based on real-time data streams from message queues, a virtual view is created in the lakeware architecture that dynamically consumes the latest messages on demand. Micro-batch processing of message streams by time window or event window generates queryable temporary datasets; Perform real-time queries or subscribe to data changes on virtual views using a streaming SQL engine or API interface.

[0045] In some embodiments, the data governance and storage management module includes an inspection and verification unit, a format conversion unit, a structured processing unit, and a unified library storage unit; The inspection and verification unit is used to perform integrity checks, validity verifications, and consistency checks on the original medical data in the patch source library. The format conversion unit is used to map non-standard terms in the original medical data to standard codes according to a predefined set of medical data standards, and to perform unified conversion of data units and formats. The structured processing unit is used to perform structured processing on the data in the source library after format conversion; this includes text recognition of unstructured data and extraction of key information through regular expressions, converting the extracted key information into structured data; and parsing semi-structured data to extract the structured data from it. The unified library storage unit is used to classify and store the processed structured data into the unified library according to subject domains, and to establish a lineage tracking relationship with the source library data.

[0046] In some embodiments, the lake warehouse collaborative processing module includes a business scenario topic determination unit, a cross-source federated query processing unit, a dynamic scheduling processing unit, a topic library construction unit, and a view service unit; The business scenario theme determination unit is used to determine the theme direction of the topic library based on the business scenario needs of clinical diagnosis and treatment, scientific research analysis or management decision-making, including clinical diagnosis and treatment topic library, disease research topic library, and hospital operation topic library; The cross-source federated query processing unit is used to perform joint queries on structured data in the unified database, raw data in the source database, and external data sources through a unified query engine to extract multimodal medical data related to the topic direction, including structured data, semi-structured data, and unstructured data. The dynamic scheduling processing unit is used to dynamically schedule computing tasks to process the extracted multimodal medical data, including direct computation of raw data at the storage layer; accelerated processing of structured data using the MPP engine; and collaborative analysis processing of cross-lake warehouse data. Thematic library construction unit is used to integrate and organize processed data according to business needs, and build thematic libraries oriented towards specific business themes; The view service unit is used to generate SQL views, REST APIs, or OLAP cubes based on thematic database tables to display the data in thematic databases; among them, the clinical diagnosis and treatment thematic database contains a view of the entire patient's diagnosis and treatment cycle; the scientific research and analysis thematic database contains disease-specific research datasets; and the management decision-making thematic database contains a cube of hospital operation indicators.

[0047] In some embodiments, the cross-source federated query processing unit performs federated queries by: converting business logic into a cross-source query plan; pushing filter conditions to the data source for execution; and merging query results from heterogeneous data sources.

[0048] It should be noted that in the embodiments of the present invention, such as Figure 3As shown, this embodiment of the invention also provides a lake-warehouse integrated medical big data processing application environment, including: a rich data source for the medical big data processing system. Within the hospital, an Oracle database stores detailed patient medical records (including symptoms, diagnoses, medications, etc.) and complex medical financial information (such as costs and medical insurance reimbursements); a SQL Server database is used to store results data for specific examination items in certain departments (such as specialist test reports and imaging results). External systems such as medical insurance data platforms and regional medical shared data platforms provide patient cross-institutional medical data (such as medical insurance reimbursement records and diagnostic results from other hospitals). These collectively constitute the raw data pool of medical big data.

[0049] Database read / write separation clusters are crucial in healthcare big data processing architectures. Data is collected from various data sources, such as timestamps and CDC (Data Criteria for Disease Control and Prevention), via a data acquisition module. During write operations, the cluster centrally processes write requests, ensuring data consistency and integrity. For example, when entering new patient information or updating diagnostic records, write nodes write precisely. For read operations, multiple read nodes handle a large number of query requests. In healthcare, queries are frequent, such as doctors accessing historical medical records and researchers accessing disease research data. The read / write separation cluster distributes read requests across different read nodes, significantly improving query efficiency and reducing response time. Furthermore, this cluster possesses high availability and fault tolerance; if one node fails, other nodes can automatically take over, ensuring the normal operation of healthcare services.

[0050] MPP clusters play a central role in medical big data processing. In practice, new patient registration, medical records, electronic medical records, and medical imaging data generated by hospital business systems (such as Oracle and SQL Server) are first received by the MPP cluster in the form of source database external tables, and then quickly imported into a unified database via INSERT INTO operations. The unified database performs preliminary data integration and standardization, including cleaning, transformation, and format unification. Different unified databases achieve real-time data synchronization through the CCR synchronization mechanism, ensuring data consistency and timeliness, which is crucial for medical big data processing because data timeliness and accuracy are related to patient treatment outcomes and medical quality. The data processed by the unified database is then made available externally through thematic database views. Thematic database views associate and summarize unified database data according to medical business needs. For example, in medical quality monitoring scenarios, patient, examination, and treatment data from different departments can be integrated to form a comprehensive view of medical information. The MPP cluster analyzes and integrates data in real time, enabling the timely detection of potential medical quality problems and providing decision support for medical management and research. Based on the processing flow of MPP clusters, hospitals can efficiently manage and utilize massive amounts of medical data, improve the quality and efficiency of medical services, and provide patients with precise and personalized diagnosis and treatment services.

[0051] The application system embodies the results of medical big data processing. As a decision support system for hospital administrators, it connects to the MPP cluster via interfaces such as JDBC, providing operational data reports, such as departmental workload, revenue, and patient satisfaction, to assist administrators in optimizing resource allocation. For medical staff, the clinical support system utilizes MPP cluster data to provide real-time patient information queries, diagnostic suggestions, and treatment plan recommendations, assisting doctors in diagnosis and treatment. This invention also provides a computer device for integrated lakehouse-based medical big data processing, comprising: a memory and a processor, wherein the memory stores a specific computer program. When the processor executes the program, it first collects heterogeneous medical data from various business systems in batches or in real time, and aggregates this data into a source database. The source database stores the original collected data, retaining its original characteristics and full information. Subsequently, the device cleans, standardizes, and structures the data in the source database, forming a standard dataset, which is then stored in a unified database. The data in the unified database, after processing, has high quality and consistency, providing a good foundation for subsequent data processing. Based on the data in the unified database, the device further integrates, correlates, and mines the data to construct thematic databases for different scenarios such as clinical diagnosis and treatment, scientific research analysis, and management decision-making, thereby extracting the value of the data. Finally, through a computing engine, the device can quickly query the processed data from the real-time lakehouse (integrating the source database, unified database, and thematic database) and provide data services for related applications in the medical field, contributing to the effective utilization of medical big data and the development of smart healthcare.

[0052] This invention provides a non-transitory computer-readable storage medium containing a computer program. When executed by a processor, the program can collect heterogeneous medical data, such as medical records, examination reports, and vital sign data, in batches or in real time from various business systems and equipment in the hospital. The data is then cleaned according to its data type to remove duplicates, errors, and missing values. The cleaned data is stored in a real-time lake warehouse that can efficiently store and manage large-scale multimodal medical data. Next, a processing chain is constructed for the data in the real-time lake warehouse to extract, transform, and load data to extract its value. Finally, a computing engine queries the processed data from the real-time lake warehouse, providing data services for clinical diagnosis and treatment, scientific research analysis, and management decision-making, thus contributing to the development of smart healthcare.

[0053] This invention further constructs a comprehensive and detailed set of medical data standards. This data standards set acts as a precise medical data navigation map, comprehensively covering multiple key categories in the medical field, providing solid support for the in-depth application and value mining of medical data.

[0054] The outpatient standard results set records detailed patient information, including appointment time, department, and attending physician. It also comprehensively records patient symptom descriptions, accurately documenting both subjective self-reported discomfort and objective signs observed by the physician. Regarding diagnostic results, it clearly identifies the preliminary diagnosis, suspected diagnosis, and final confirmed diagnosis, coded according to the International Classification of Diseases (ICD) to ensure standardized and accurate diagnosis. Treatment plan information includes the names, dosages, usage, and duration of prescribed medications, as well as recommended examinations and treatments, providing clear guidance for subsequent treatment and recovery. Furthermore, it records the patient's allergy history and past medical history, enabling physicians to comprehensively understand the patient's health status and make more accurate treatment decisions.

[0055] The inpatient standard outcome set focuses on the entire process of patient data during hospitalization. Admission information details the patient's admission time, admission route (e.g., emergency admission, outpatient transfer), and admission condition assessment. The inpatient progress notes comprehensively present changes in the patient's condition during hospitalization, the implementation of various treatment measures, and the evaluation of treatment effects. Surgical information includes the surgery name, surgery time, surgeon, anesthesia method, and surgical procedure record, providing important evidence for surgical quality assessment and medical research. Discharge information includes discharge time, discharge diagnosis, and discharge orders (e.g., follow-up medication, follow-up appointments), ensuring patients receive continuous medical services after discharge. Additionally, nursing records are included, including vital sign monitoring data, nursing operation records, and the patient's daily care, comprehensively reflecting the patient's nursing condition during hospitalization.

[0056] This standardized results set for laboratory tests comprehensively integrates various medical laboratory test items. For blood tests, it records detailed results for complete blood counts and biochemical indicators, including specific values ​​for red blood cell count, white blood cell count, blood glucose, and blood lipids, along with normal reference ranges, facilitating rapid assessment of abnormal results by physicians. Urine tests cover all indicators from routine urinalysis, such as urine protein, urine glucose, and urine red blood cells. Microbiological tests record results from bacterial culture and drug sensitivity testing, providing crucial information for the diagnosis and treatment of infectious diseases. Furthermore, it includes results from other specialized tests, such as tumor marker detection and gene testing, providing strong support for early disease diagnosis and personalized treatment. Information such as sample collection time, collection site, and delivery time is also recorded to ensure the accuracy and traceability of test results.

[0057] The standardized results set for medical examinations includes data from various medical examinations. In imaging, it details the imaging findings and diagnostic conclusions of X-rays, CT scans, and MRIs, such as the size, location, and shape of lung nodules, and the extent and nature of brain lesions. Ultrasound examinations cover abdominal ultrasound, cardiac ultrasound, and obstetric / gynecological ultrasound, clearly displaying the morphology, structure, and blood flow of organs. Endoscopic examinations record endoscopic findings and pathological diagnoses from gastroscopy and colonoscopy, which are crucial for the diagnosis and treatment of digestive tract diseases. It also includes results from functional examinations such as electrocardiograms and electroencephalograms (EEGs), providing important evidence for assessing cardiac and nervous system function. Furthermore, detailed information on the model of the examination equipment and examination parameters is recorded, facilitating accurate interpretation and comparative analysis of the results.

[0058] The drug management standards set records detailed information about hospital drugs, including basic attributes such as drug name, specifications, dosage form, manufacturer, and approval number. It also tracks and records drug procurement information, such as procurement time, quantity, and price, to facilitate cost control and inventory management. In the drug dispensing process, it records the dispensing time, department, quantity, and corresponding prescription information to ensure accuracy and traceability. Furthermore, it monitors adverse drug reactions, recording symptoms, timing, and treatment measures for patients experiencing adverse reactions, thus ensuring the safe use of drugs.

[0059] The medical quality management outcome set revolves around medical quality assessment indicators, covering key metrics such as surgical success rate, infection rate, and complication rate. It records quality control data throughout the medical process, such as surgical safety checklists and medical record quality inspection results. Simultaneously, it meticulously records medical disputes and complaints, including the time, cause, and outcome of each dispute, to facilitate analysis of existing problems and the implementation of targeted improvement measures to enhance medical quality and service levels.

[0060] The research and teaching outcome set integrates clinical research data, including basic information about research projects, inclusion and exclusion criteria for research subjects, observation indicators and data records during the research process. It also collects medical literature search and citation information to provide reference for researchers. In terms of teaching, it records information such as the time, location, participants, and content of teaching activities, as well as student assessment scores and feedback, which helps improve teaching quality and cultivate medical talent.

[0061] This complete set of medical big data processing standards provides unified standards and norms for the integration, analysis and application of medical data by comprehensively covering and recording multiple categories such as outpatient, inpatient, laboratory, examination, drug management, medical quality management, scientific research and teaching. It helps to improve medical quality, promote the development of medical research and optimize medical management decisions.

[0062] This invention provides a non-transitory computer-readable storage medium that stores computer instructions that cause a computer to execute the methods provided in the above-described method embodiments. These instructions include, for example,: establishing a rule configuration table for different types of shared documents, comprising a document encoding field, an element path field, and an attribute validation field group; converting the constraints of the shared document specification into structured configuration rule data and storing it in the corresponding rule configuration table, wherein the structured configuration rule data includes a mapping relationship between element paths and constraints; parsing the shared document XML file to be validated, locating the XML document node through the element path, and obtaining the document type encoding; loading configuration rule data from the rule configuration database according to the document type encoding, and establishing a validation queue according to the hierarchical order of the element paths; traversing the validation queue and validating the XML node by comparing its attribute values ​​and text content with the corresponding constraints in the rule configuration table; and recording validation error information, including the element path and error details, for nodes that do not meet the constraints.

[0063] Although the present invention has been described in detail with reference to the accompanying drawings and preferred embodiments, the present invention is not limited thereto. Various equivalent modifications or substitutions can be made to the embodiments of the present invention by those skilled in the art without departing from the spirit and essence of the invention, and such modifications or substitutions should all be within the scope of the present invention. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should also be covered within the protection scope of the present invention.

Claims

1. A method for processing medical big data based on a lake-warehouse integration, characterized in that, include: S1. Collect raw medical data from various medical business systems, and integrate the raw medical data into the data access layer of the lake warehouse integrated architecture in its original format and with all information to form a source library; S2. Clean, standardize, and structure the original medical data in the source database, and perform unified encoding and format conversion according to the medical data standard set to generate a standardized dataset and store it in the unified database. S3. Based on the data in the unified database and the source database, through cross-source federated query and computation flow technology, multimodal medical data is associated and integrated according to the business scenario needs of clinical diagnosis and treatment, scientific research analysis or management decision-making, to build a thematic database for specific business themes and generate thematic database views. S4. Provide query interfaces and analysis services for data application platforms through thematic databases.

2. The medical big data processing method based on lake-warehouse integration according to claim 1, characterized in that, Step S1 specifically includes: S11. Identify the medical data sources to be collected from hospital information systems, electronic medical record systems, medical image archiving and communication systems, medical data platforms, and regional medical shared data platforms; S12. Medical data from medical data sources is uniformly connected to the data access layer of the lake warehouse integrated architecture through full synchronization, incremental synchronization, or data virtualization to form a source database; among which, medical data is stored in the source database in its original format, retaining all fields and metadata; S13. Add unified metadata tags to the data in the post source library, including data source, collection time, data format and business meaning description; and establish a full-link lineage relationship for the data in the post source library, recording the data flow path from the source system to the post source library.

3. The medical big data processing method based on lake-warehouse integration according to claim 2, characterized in that, In S12, the steps to unify the medical data from medical data sources into the data access layer of the lake warehouse integrated architecture, forming the source library, include: (Full synchronization, incremental synchronization, or data virtualization methods) S121. Perform full synchronization of initial data or periodically updated data in medical data sources through batch processing to completely extract the data into the post source library. S122. For real-time changes in medical data sources, monitor the database logs of medical data sources using change data capture technology, and capture incremental data in real time using stream processing and synchronize it to the post source database. S123. For external data in medical data sources that cannot be directly physically integrated, establish logical mapping relationships through data virtualization technology to achieve dynamic access to cross-system data; In S121, batch processing full synchronization is triggered according to a preset cycle, and the data range is marked based on timestamp or version number during extraction. In S122, the stream processing incrementally parses the database transaction log in real time to capture insert, update, and delete operation events.

4. The medical big data processing method based on lake-warehouse integration according to claim 3, characterized in that, In S123, the logical mapping relationship is implemented in the following way: S1231. For external API data, encapsulate REST / SOAP interface calls into virtual tables; specifically including: Register the access endpoints, authentication methods, and data formats of external APIs in the lakewarehousing integrated architecture; Map the query parameters required by the API to queryable fields of a virtual table; Parse the JSON / XML data returned by the API, extract structured fields, and define the column names and data types of the virtual table; Based on the parsed data model, a virtual table is created in the lake warehouse integrated architecture. This table only stores metadata, and the actual data is dynamically obtained by calling the API in real time. When the application layer queries the virtual table, it automatically converts the SQL or analysis request into an API call, retrieves data in real time, and returns the results. S1232. Establish a streaming virtual view for message queue data; specifically including: Configure message queue connection parameters in a lakeware architecture, including server address, topic name, consumer group, and deserialization format; Parse the data structure in the message queue, define the field names, data types, and timestamp fields of the virtual view, and register them with the metadata management system; Based on real-time data streams from message queues, a virtual view is created in the lakeware architecture that dynamically consumes the latest messages on demand. Micro-batch processing of message streams by time window or event window generates queryable temporary datasets; Perform real-time queries or subscribe to data changes on virtual views using a streaming SQL engine or API interface.

5. The medical big data processing method based on lake-warehouse integration according to claim 1, characterized in that, Step S2 specifically includes: S21. Perform integrity checks, validity verifications, and consistency checks on the original medical data in the patch source database; S22. Based on a predefined set of medical data standards, map non-standard terms in the original medical data to standard codes, and perform unified conversion on data units and formats; S23. Perform structured processing on the data in the source library after format conversion; including text recognition of unstructured data and extraction of key information through regular expressions, converting the extracted key information into structured data; parsing semi-structured data and extracting the structured data from it; S24. Store the processed structured data in a unified database according to subject domains, and establish a lineage tracing relationship with the source database data.

6. The medical big data processing method based on lake-warehouse integration according to claim 1, characterized in that, Step S3 specifically includes: S31. Based on the business scenario needs of clinical diagnosis and treatment, scientific research analysis, or management decision-making, determine the thematic direction of the topic database, including the clinical diagnosis and treatment topic database, the disease research topic database, and the hospital operation topic database; S32. Using a unified query engine, perform a joint query on the structured data in the unified database, the original data in the source database, and the external data source to extract multimodal medical data related to the topic, including structured data, semi-structured data, and unstructured data. S33. Dynamically scheduled computing tasks process the extracted multimodal medical data, including direct computation of raw data at the storage layer; accelerated processing of structured data using the MPP engine; and collaborative analysis of cross-lake warehouse data. S34. Integrate and organize the processed data according to business needs, and build a thematic library for specific business topics; S35. Generate SQL views, REST APIs, or OLAP cubes based on thematic database tables to display the data in thematic databases; among them, the clinical diagnosis and treatment thematic database contains a view of the entire patient's diagnosis and treatment cycle; the scientific research and analysis thematic database contains disease-specific research datasets; and the management decision-making thematic database contains a cube of hospital operation indicators.

7. The medical big data processing method based on lake-warehouse integration according to claim 6, characterized in that, In S32, executing a join query includes: Transform business logic into a cross-source query plan; Push the filtering conditions to the data source for execution; Merge query results from heterogeneous data sources.

8. A medical big data processing device based on a lake-warehouse integration, characterized in that, include: The unified data integration module is used to collect raw medical data from various medical business systems and integrate the raw medical data into the data access layer of the lake warehouse integrated architecture in its original format and with full information to form a source library. The data governance and storage management module is used to clean, standardize, and structure the raw medical data in the source database, and to perform unified encoding and format conversion according to the medical data standard set, generating a standardized dataset and storing it in the unified database. The Lake Warehouse Collaborative Processing Module is used to integrate multimodal medical data based on data in the unified database and the source database, through cross-source federated query and computation flow technology, according to the business scenario needs of clinical diagnosis and treatment, scientific research analysis or management decision-making, to build thematic databases for specific business themes and generate thematic database views; The unified access interface module is used to provide query interfaces and analysis services to the data application platform through thematic databases.

9. An electronic device, characterized in that, The electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; the memory stores computer program instructions executable by the at least one processor, the computer program instructions being executed by the at least one processor to enable the at least one processor to perform the lake-warehouse integrated medical big data processing method as described in any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium, characterized in that, The non-transitory computer-readable storage medium stores computer instructions that cause the computer to execute the lake-warehouse integrated medical big data processing method as described in any one of claims 1 to 7.

Citation Information

Cited By

  • Data processing method and device, equipment, storage medium and program product

    CN121188012A

  • Digital base system

    CN121689559A

  • A digital submount system

    CN121689559B